Stars
A markdown native slides tool for academics building with agents.
40+ tips for getting the most out of Claude Code, from basics to advanced - includes a custom status line script and Claude Code running itself in a container. Also includes the dx plugin: skills f…
slime is an LLM post-training framework for RL Scaling.
A comprehensive collection of Agent Skills for context engineering, multi-agent architectures, and production agent systems. Use when building, optimizing, or debugging agent systems that require e…
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
UEval: A Benchmark for Unified Multimodal Generation
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
Fully Open Framework for Democratized Multimodal Training
StreamingVLM: Real-Time Understanding for Infinite Video Streams
Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"
Official codebase used to develop Vision Transformer, SigLIP, MLP-Mixer, LiT and more.
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
Minimalistic large language model 3D-parallelism training
Training Large Language Model to Reason in a Continuous Latent Space
Official codebase for the paper Latent Visual Reasoning
SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
[NeurIPS 2024] Visual Perception by Large Language Model’s Weights
Tongyi Deep Research, the Leading Open-source Deep Research Agent
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
A toolkit for developing and comparing reinforcement learning algorithms.
SGLang is a high-performance serving framework for large language models and multimodal models.
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)