Stars
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
The simplest, fastest repository for training/finetuning medium-sized GPTs.
18 Lessons to Get Started Building AI Agents
A PyTorch native platform for training generative AI models
A version of verl to support diverse tool use [TMLR 2026]
SkyRL: A Modular Full-stack RL Library for LLMs
This repo contains the Hugging Face Deep Reinforcement Learning Course.
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
This is the homepage of a new book entitled "Mathematical Foundations of Reinforcement Learning."
Companion code for FanOutQA: Multi-Hop, Multi-Document Question Answering for Large Language Models (ACL 2024)
yeayee / joyful-pandas
Forked from datawhalechina/joyful-pandasPandas中文教程
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
A collection of LLM papers, blogs, and projects, with a focus on OpenAI o1 🍓 and reasoning techniques.
A bibliography and survey of the papers surrounding o1
Robust recipes to align language models with human and AI preferences
Implementation for "Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs"
Llama中文社区,实时汇总最新Llama学习资料,构建最好的中文Llama大模型开源生态,完全开源可商用
北京航空航天大学大数据高精尖中心自然语言处理研究团队开展了智能问答的研究与应用总结。包括基于知识图谱的问答(KBQA),基于文本的问答系统(TextQA),基于表格的问答系统(TableQA)、基于视觉的问答系统(VisualQA)和机器阅读理解(MRC)等,每类任务分别对学术界和工业界进行了相关总结。
大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP
Faker is a Python package that generates fake data for you.
A recipe for online RLHF and online iterative DPO.
Scalable toolkit for efficient model alignment
JS tokenizer for LLaMA 3 and LLaMA 3.1