Highlights
- Pro
Stars
slime is an LLM post-training framework for RL Scaling.
J-tau: A Japanese tau-bench for Benchmarking Tool-Agent-User Interaction in Real-World Domains
A fast and soft pattern search for trillion-scale corpora.
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems (Submitted on 26 Sep 2025)
[EMNLP2025] Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling
Muon is an optimizer for hidden layers in neural networks
Helpful kernel tutorials, examples and SKILLs for tile-based GPU programming
We introduce COMMUNITYNOTES dataset for predicting note helpfulness and its reasons, propose an automatic reason-optimization framework, and show its benefits for evidence sufficiency and fact-chec…
Dataset for evaluating the knowledge of Yokai in language models.
Code of "Evaluation of Best-of-N Sampling Strategies for Language Model Alignment"
Code of "Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment" (2025).
Evaluate your LLM's response with Prometheus and GPT4 đŸ’¯
AirLLM 70B inference with single 4GB GPU
Code of Annotation-Efficient Language Model Alignment via Diverse and Representative Response Texts (EMNLP Findings 2025)
A curated list of research papers and resources on Cultural LLM.
Code of "Model-Based Minimum Bayes Risk Decoding for Text Generation" 2024
Flexible evaluation tool for language models
Robust recipes to align language models with human and AI preferences
[EMNLP 2024] Introducing Filtered Direct Preference Optimization (fDPO) that enhances language model alignment with human preferences by discarding lower-quality samples compared to those generated…