Skip to content

refactor: split PD disaggregation into separate EngineGroups - #1609

Merged
zhuzilin merged 3 commits into
mainfrom
zhuzilin/pd-split-engine-groups
Feb 21, 2026
Merged

zhuzilin merged 3 commits into
mainfrom
zhuzilin/pd-split-engine-groups

Conversation

@zhuzilin

Copy link
Copy Markdown
Contributor

No description provided.

- Add role and rank_offset fields to EngineGroup so each group knows
  its worker type and global rank starting point
- Refactor init_rollout_engines to accept worker_type/rank_offset
  instead of computing prefill/decode internally
- Update _allocate_rollout_engine_addr_and_ports_normal to use
  worker_type parameter instead of prefill_limit calculation
- Split start_rollout_server to create separate prefill and decode
  EngineGroups when prefill_num_servers is set
- Change health monitors from single to per-group list so each
  EngineGroup is monitored independently
Move the engine initialization logic from the standalone
init_rollout_engines() function into EngineGroup.init_engines().
This eliminates the duplication of role/rank_offset between
the function call and the EngineGroup constructor.

start_rollout_server now creates EngineGroup with [None] engines
and calls group.init_engines() to fill them in.

reinit_engines is kept as an alias of init_engines for clarity
in fault-recovery paths.
When debug_train_only is set, self.server is None. Add None
checks to offload() and onload() to prevent AttributeError.
@zhuzilin
zhuzilin merged commit 029bbed into main Feb 21, 2026
15 of 16 checks passed
@zhuzilin
zhuzilin deleted the zhuzilin/pd-split-engine-groups branch February 21, 2026 15:05
gxlvera pushed a commit to gxlvera/slime that referenced this pull request Feb 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant