Starred repositories
One local control plane for every AI agent: route across models, fuse new capabilities, orchestrate tools, and stay fully in control.
A programmable Mixture-of-Models router for heterogeneous LLM inference
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
PyTorch version of Stable Baselines, reliable implementations of reinforcement learning algorithms.
Train transformer language models with reinforcement learning.
TensorDict is a pytorch dedicated tensor container.
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
深度学习500问,以问答形式对常用的概率知识、线性代数、机器学习、深度学习、计算机视觉等热点问题进行阐述,以帮助自己及有需要的读者。 全书分为18个章节,50余万字。由于水平有限,书中不妥之处恳请广大读者批评指正。 未完待续............ 如有意合作,联系scutjy2015@163.com 版权所有,违权必究 Tan 2018.06
ViewTurbo | A Blazing Fast VPN, Beautifully Designed, Stable & Easy-to-Use | 极速 VPN、精美设计、超20G顶级带宽、不限速度、不限流量、AI智能分流、更适合 ChatGPT、Claude、GitHub、Gemini 等场景、绝对的爽感体验
Structured deep research skill for Claude Code/Open Code/Codex with human-in-the-loop control
深度调研报告生成 Skill — 一条命令,十分钟出券商级深度调研报告 / Professional deep research report generation Skill · Supports 19 languages
🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman
⚡️SwanLab - an open-source, modern-design AI training tracking and visualization tool. Supports Cloud / Self-hosted use. Integrated with PyTorch / Transformers / verl / LLaMA Factory / ms-swift / U…
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Byted PyTorch Distributed for Hyperscale Training of LLMs and RLs
Set up your own IPsec VPN server in just a few minutes, with IPsec/L2TP, Cisco IPsec and IKEv2. Supports Ubuntu, Debian, CentOS/RHEL, Alpine Linux and Raspberry Pi OS. Includes client config and ma…
A kernel library written in tilelang
A bidirectional pipeline parallelism algorithm for computation-communication overlap in DeepSeek V3/R1 training.
Assign issues to Claude Code, Codex, Cursor, and 17 more coding agents like teammates — open-source and self-hostable.
Self-referential self-improving agents that can optimize for any computable task
A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
The agent that grows with you
🔥 A collection of the Claude Code open source
AI agents running research on single-GPU nanochat training automatically
Muon is an optimizer for hidden layers in neural networks