RL post-training for code reasoning
DeepCoder—14B
A fully open-source coding model trained with distributed reinforcement learning. I worked on the RLLM pipeline, asynchronous infrastructure, execution-grounded evaluation, and the reward loop that makes “good code” testable.
Read the research ↗Pass@1
training run
pipeline