Lists (13)
Sort Name ascending (A-Z)
Stars
A Business-Driven Real-World Financial Benchmark for Evaluating LLMs
a toolkit on knowledge distillation for large language models
Official code for "Self-Distilled Agentic Reinforcement Learning"
Post-training with Tinker
Claw Code No Rust No TypeScript Only Python. Easy to work with. Fast to iterate. 🔥 Zero external dependencies 🔥
Token efficient Claude Code full Python rebuild. AI Coding Agent in 270K LoC pure Python. Up to 200X Cost Saving!
Uni-Agent is a framework for training long-horizon agents.
与其蒸馏别人,不如蒸馏自己。欢迎加入数字永生!Inspired by colleague-skill(同事skill)。
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
AgentFlow: In-the-Flow Agentic System Optimization
将冰冷的离别化为温暖的 Skill,欢迎加入数字生命1.0!Transforming cold farewells into warm skills? It's giving rebirth era. Welcome to Digital Life 1.0. 🫶
你想蒸馏的下一个员工,何必是同事。蒸馏任何人的思维方式——心智模型、决策启发式、表达DNA。Distill how anyone thinks.
A project implementing various agentic RL based on the Slime post-training framework
OpenSeeker: A search agent with open-source data and models
OpenClaw-RL: Train any agent simply by talking
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
A modern GUI client based on Tauri, designed to run in Windows, macOS and Linux for tailored proxy experience
Agent harness to publish your agent chat history as Huggingface datasets.
bilibili video course src code
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
An agentic skills framework & software development methodology that works.
slime is an LLM post-training framework for RL Scaling.
Scalable toolkit for efficient model reinforcement
SkyRL: A Modular Full-stack RL Library for LLMs
A curated list of reinforcement learning (RL) for agents.
全员 AI 的公司,只有老板是人。7×24 为你做调研、发媒体、写代码、管投资。基于 Claude Code 构建。