Stars
[arxiv'26] GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation
The agent that grows with you
OpenAPI 3.x code generator for deterministic, type-safe SDKs and API tooling across languages and runtimes.
SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
FlashRec is a CUDA-graph engine for generative recommendation: wide beam search (3–5 SID steps, n=50–512+) over a trie-constrained catalog, in-process FP8 serving, and ranked beams on /v1/chat/comp…
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…
SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
An end-to-end agent project for GPU kernel implementation, analysis, profiling, and iterative optimization. It helps an agent turn PyTorch logic or an existing kernel into a high-performance GPU ke…
Deep Code 是专为 deepseek-v4 模型优化的终端 AI 编码助手,支持深度思考、推理强度控制以及 Agent Skills。
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
Modular SenseNova skills for building AI-powered office assistants and productivity workflows
A unified inference and post-training framework for accelerated video generation.
🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"
🚀 Efficient implementations for emerging model architectures
Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps
Fast, Sharp & Reliable Agentic Intelligence
The AI that really does things. Any OS. Any Platform. The lobster way. 🦞
FlagGems is an operator library for large language models implemented in the Triton Language.
LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.
NEO Series: Native Vision-Language Models from First Principles
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
A framework for efficient model inference with omni-modality models
cuTile is a programming model for writing parallel kernels for NVIDIA GPUs
LightTTS is a lightweight TTS inference framework optimized for CosyVoice2 and CosyVoice3, enabling fast and scalable speech synthesis in Python and supports stream and bistream modes.