-
Ant Financial
- Bay area, USA
- http://merlintang.github.io/
Stars
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
My learning notes for ML SYS.
Accelerating MoE with IO and Tile-aware Optimizations
FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Trae, Traycer AI…
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
slime is an LLM post-training framework for RL Scaling.
The Memory Layer for AI Agents - Drop-in memory infrastructure for AI agents and apps. Context that persists. Built for production.
TextPy: Collaborative Agent Workflow through Programming and Prompting
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
[EMNLP2025] LightRAG: Simple and Fast Retrieval-Augmented Generation
End-to-end Generative Optimization for AI Agents
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthr…
AG2 (formerly AutoGen): The Open-Source AgentOS.Join us at: https://discord.gg/sNGSwQME3x
A collection of notebooks/recipes showcasing some fun and effective ways of using Claude.
SGLang is a high-performance serving framework for large language models and multimodal models.
Minimalistic 4D-parallelism distributed training framework for education purpose
Scalable RL solution for advanced reasoning of language models
Ongoing research training transformer models at scale
Recipes to scale inference-time compute of open models
Supercharge Your LLM Application Evaluations 🚀
🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
"BadPart: Unified Black-box Adversarial Patch Attacks against Pixel-wise Regression Tasks"
An Efficient LLM Fine-Tuning Factory Optimized for MoE PEFT
An Efficient "Factory" to Build Multiple LoRA Adapters
A generative speech model for daily dialogue.