Skip to content
View hitwsl's full-sized avatar

Block or report hitwsl

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Rust 195,068 109,182 Updated Aug 6, 2026

The simplest, fastest repository for training/finetuning medium-sized GPTs.

Python 62,059 10,697 Updated Nov 12, 2025

18 Lessons to Get Started Building AI Agents

Jupyter Notebook 72,013 23,861 Updated Jul 29, 2026

📚 从零开始构建大模型

Jupyter Notebook 32,909 3,123 Updated Aug 8, 2026

A PyTorch native platform for training generative AI models

Python 5,620 946 Updated Aug 13, 2026

A version of verl to support diverse tool use [TMLR 2026]

Python 1,031 88 Updated Jul 15, 2026

SkyRL: A Modular Full-stack RL Library for LLMs

Python 2,147 403 Updated Aug 11, 2026

This repo contains the Hugging Face Deep Reinforcement Learning Course.

MDX 4,982 809 Updated May 26, 2026

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,933 4,389 Updated Aug 12, 2026

verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"

Python 2,214 212 Updated Jun 9, 2026

This is the homepage of a new book entitled "Mathematical Foundations of Reinforcement Learning."

MATLAB 17,447 1,659 Updated Aug 10, 2026

Reproduce R1 Zero on Logic Puzzle

Python 2,450 163 Updated Mar 20, 2025

Companion code for FanOutQA: Multi-Hop, Multi-Document Question Answering for Large Language Models (ACL 2024)

Python 63 7 Updated Jul 30, 2026

Pandas中文教程

Jupyter Notebook 997 491 Updated Nov 21, 2024

An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)

Python 9,906 998 Updated Jul 14, 2026

A collection of LLM papers, blogs, and projects, with a focus on OpenAI o1 🍓 and reasoning techniques.

6,899 369 Updated Dec 17, 2025

O1 Replication Journey

2,001 61 Updated Jan 14, 2025

A bibliography and survey of the papers surrounding o1

TeX 1,215 50 Updated Jul 7, 2026

Robust recipes to align language models with human and AI preferences

Python 5,659 489 Updated May 26, 2026

Implementation for "Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs"

Python 397 16 Updated Jan 19, 2025

Llama3-中文后训练版

Python 4,148 332 Updated Feb 21, 2026

Llama中文社区,实时汇总最新Llama学习资料,构建最好的中文Llama大模型开源生态,完全开源可商用

Python 14,747 1,296 Updated Apr 6, 2025

北京航空航天大学大数据高精尖中心自然语言处理研究团队开展了智能问答的研究与应用总结。包括基于知识图谱的问答(KBQA),基于文本的问答系统(TextQA),基于表格的问答系统(TableQA)、基于视觉的问答系统(VisualQA)和机器阅读理解(MRC)等,每类任务分别对学术界和工业界进行了相关总结。

1,816 260 Updated Apr 6, 2023

大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP

9,907 1,554 Updated Feb 6, 2026

Faker is a Python package that generates fake data for you.

Python 19,370 2,108 Updated Aug 3, 2026

搜索所有中文NLP数据集,附常用英文NLP数据集

Python 4,460 627 Updated Nov 21, 2022

A recipe for online RLHF and online iterative DPO.

Python 544 48 Updated Dec 28, 2024
Python 4,598 503 Updated Apr 22, 2026

Scalable toolkit for efficient model alignment

Python 853 109 Updated Oct 6, 2025

JS tokenizer for LLaMA 3 and LLaMA 3.1

JavaScript 118 5 Updated Jul 28, 2025
Next