-
Zhejiang University
- Beijing
- https://www.zhihu.com/people/hai-tan-shang-chong-hua-47
Stars
A summary of technical reports for various large language models (LLMs).
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
OpenClaw-RL: Train any agent simply by talking
Curated, opinionated index of post-R1 LLM × Reinforcement Learning. Many deep-dive blog posts cross-linked to many papers — GRPO, DAPO, DPO, PPO, RLHF, GSPO, CISPO, VAPO, Reward Modeling, MoE RL st…
📖 This is a repository for organizing papers, codes, and other resources related to Latent Reasoning.
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
slime is an LLM post-training framework for RL Scaling.
SGLang is a high-performance serving framework for large language models and multimodal models.
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
deepspeedai / Megatron-DeepSpeed
Forked from NVIDIA/Megatron-LMOngoing research training transformer language models at scale, including: BERT & GPT-2
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
Ongoing research training transformer models at scale
Lime: Explaining the predictions of any machine learning classifier
AISystem 主要是指AI系统,包括AI芯片、AI编译器、AI推理和训练框架等AI全栈底层技术
[ICML 2024] LESS: Selecting Influential Data for Targeted Instruction Tuning
Multipack distributed sampler for fast padding-free training of LLMs
Code for evaluating with Flow-Judge-v0.1 - an open-source, lightweight (3.8B) language model optimized for LLM system evaluations. Crafted for accuracy, speed, and customization.
Summarize existing representative LLMs text datasets.
This is the first released survey paper on hallucinations of large vision-language models (LVLMs). To keep track of this field and continuously update our survey, we maintain this repository of rel…
A Survey of LLM Alignment (SFT & RLHF), and A Survey of RLHF methods (2023~2024)
Unsloth is a local UI for training and running Gemma 4, Qwen3.6, DeepSeek, Kimi, GLM and other models.
Similarities: a toolkit for similarity calculation and semantic search. 相似度计算、匹配搜索工具包,支持亿级数据文搜文、文搜图、图搜图,python3开发,开箱即用。
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama mode…
Evaluate your LLM's response with Prometheus and GPT4 💯