-
Notifications
You must be signed in to change notification settings - Fork 1k
Pull requests: OpenRLHF/OpenRLHF
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
fix(dpo): normalize IPO log probabilities by response length
#1362
opened Sep 20, 2026 by
Excelius-Wang
Contributor
Loading…
fix(ppo): set dynamic gradient accumulation boundaries
#1361
opened Sep 20, 2026 by
Excelius-Wang
Contributor
Loading…
fix(ppo): train only complete gradient accumulation windows
#1360
opened Sep 19, 2026 by
Excelius-Wang
Contributor
Loading…
fix(ppo): stop co-located sync barriers from blocking unrelated forward dispatches
#1359
opened Sep 18, 2026 by
RainieLLM
Loading…
fix: persist EMA updates to ZeRO-3 parameter partitions
#1357
opened Sep 15, 2026 by
Excelius-Wang
Contributor
Loading…
fix(ppo): avoid buffering actor experiences during critic warmup
#1356
opened Sep 15, 2026 by
Excelius-Wang
Contributor
Loading…
fix(ppo): seed prompt shuffling with the training seed
#1355
opened Sep 15, 2026 by
Excelius-Wang
Contributor
Loading…
fix(ppo): prevent binary-KL roundoff from rejecting valid samples
#1354
opened Sep 14, 2026 by
Excelius-Wang
Contributor
Loading…
fix(eval): handle integer rewards in PPO metrics
#1353
opened Sep 14, 2026 by
Excelius-Wang
Contributor
Loading…
fix(agent): mark context-exhausted trajectories as truncated
#1352
opened Sep 14, 2026 by
Excelius-Wang
Contributor
Loading…
fix: count DPO and RM scheduler steps across epochs
#1351
opened Sep 14, 2026 by
Excelius-Wang
Contributor
Loading…
fix: align DPO and RM checkpoint resume with loader epochs
#1350
opened Sep 14, 2026 by
Excelius-Wang
Contributor
Loading…
fix(rl): cast synced params to bf16 before vLLM weight update
#1348
opened Sep 14, 2026 by
Tonystarkw12
Loading…
fix: restore text rollouts in the OpenAI-compatible agent example
#1347
opened Sep 14, 2026 by
Excelius-Wang
Contributor
Loading…
fix: handle split message columns in multi-turn SFT
#1346
opened Sep 14, 2026 by
Excelius-Wang
Contributor
Loading…
fix(launcher): apply all custom resources to placement-group bundles
#1345
opened Sep 14, 2026 by
aha-jansen
Loading…
fix(reward): pass pad_sequence through to model instead of hardcoding…
#1344
opened Sep 14, 2026 by
aha-jansen
Loading…
fix: restore reward normalization buffers after model loading
#1343
opened Sep 14, 2026 by
Excelius-Wang
Contributor
Loading…
fix: preserve reward model heads through LoRA save and merge
#1342
opened Sep 14, 2026 by
Excelius-Wang
Contributor
Loading…
fix: save async best checkpoints from the evaluated training step
#1341
opened Sep 14, 2026 by
Excelius-Wang
Contributor
Loading…
feat: AMD ROCm (MI300X) support — validated container, training scripts, quick-start guide
#1331
opened Sep 11, 2026 by
jun-amd
Loading…
fix: authenticate remote reward HTTP and bind loopback by default
#1326
opened Sep 9, 2026 by
apurv-1
Loading…
2 tasks
fix(data): retain trainable turns before a truncated prompt
#1324
opened Sep 6, 2026 by
feng1201
Loading…
fix(async): persist oversampling buffer state across checkpoints
#1323
opened Sep 4, 2026 by
ai-yang
Contributor
Loading…
5 tasks done
fix(data): filter preference pairs collapsed by truncation
#1313
opened Aug 21, 2026 by
ai-yang
Contributor
Loading…
Previous Next
ProTip!
What’s not been updated in a month: updated:<2026-08-21.