Ulysses sequence-parallel all-to-all as a torch custom op, moved by the GPU copy engines into torch symmetric memory. Zero SM usage; 1.66-2.17x over torch.distributed on NVLink.
cuda ulysses pytorch nvlink diffusion-models collective-communication sequence-parallelism custom-operator
-
Updated
Aug 25, 2026 - Python