Skip to content

Pull requests: OpenRLHF/OpenRLHF

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

fix(dpo): normalize IPO log probabilities by response length
#1362 opened Sep 20, 2026 by Excelius-Wang Contributor Loading…
fix(ppo): set dynamic gradient accumulation boundaries
#1361 opened Sep 20, 2026 by Excelius-Wang Contributor Loading…
fix(ppo): train only complete gradient accumulation windows
#1360 opened Sep 19, 2026 by Excelius-Wang Contributor Loading…
fix: persist EMA updates to ZeRO-3 parameter partitions
#1357 opened Sep 15, 2026 by Excelius-Wang Contributor Loading…
fix(ppo): avoid buffering actor experiences during critic warmup
#1356 opened Sep 15, 2026 by Excelius-Wang Contributor Loading…
fix(ppo): seed prompt shuffling with the training seed
#1355 opened Sep 15, 2026 by Excelius-Wang Contributor Loading…
fix(ppo): prevent binary-KL roundoff from rejecting valid samples
#1354 opened Sep 14, 2026 by Excelius-Wang Contributor Loading…
fix(eval): handle integer rewards in PPO metrics
#1353 opened Sep 14, 2026 by Excelius-Wang Contributor Loading…
fix(agent): mark context-exhausted trajectories as truncated
#1352 opened Sep 14, 2026 by Excelius-Wang Contributor Loading…
fix: count DPO and RM scheduler steps across epochs
#1351 opened Sep 14, 2026 by Excelius-Wang Contributor Loading…
fix: align DPO and RM checkpoint resume with loader epochs
#1350 opened Sep 14, 2026 by Excelius-Wang Contributor Loading…
fix: restore text rollouts in the OpenAI-compatible agent example
#1347 opened Sep 14, 2026 by Excelius-Wang Contributor Loading…
fix: handle split message columns in multi-turn SFT
#1346 opened Sep 14, 2026 by Excelius-Wang Contributor Loading…
fix: restore reward normalization buffers after model loading
#1343 opened Sep 14, 2026 by Excelius-Wang Contributor Loading…
fix: preserve reward model heads through LoRA save and merge
#1342 opened Sep 14, 2026 by Excelius-Wang Contributor Loading…
fix: save async best checkpoints from the evaluated training step
#1341 opened Sep 14, 2026 by Excelius-Wang Contributor Loading…
fix(async): persist oversampling buffer state across checkpoints
#1323 opened Sep 4, 2026 by ai-yang Contributor Loading…
5 tasks done
fix(data): filter preference pairs collapsed by truncation
#1313 opened Aug 21, 2026 by ai-yang Contributor Loading…
ProTip! What’s not been updated in a month: updated:<2026-08-21.