Highlights
- Pro
Stars
Evaluation code for "Benchmarking Visual State Tracking in Multimodal Video Understanding"
Cambrian-P: Pose-Grounded Video Understanding
Official Implemenation for RAEv2: Improved Baselines with Representation Autoencoders
[ICML 2026] ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
⚡️SwanLab - an open-source, modern-design AI training tracking and visualization tool. Supports Cloud / Self-hosted use. Integrated with PyTorch / Transformers / verl / LLaMA Factory / ms-swift / U…
A linear estimator on top of clip to predict the aesthetic quality of pictures
The open agent skills tool - npx skills
A lightweight, AI-native training framework for large language models. Designed for fast iteration, reproducible experiments, and modular configuration across SFT, RLVR, and evaluation workflows.
The first multiplayer video world model in Minecraft
MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
ViPE: Video Pose Engine for Geometric 3D Perception
Cambrian-S: Towards Spatial Supersensing in Video
Native Multimodal Models are World Learners
[NeurIPS'25] Official repository of Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"
[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.
A simple, performant, and scalable Jax LLM!
🧑🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), ga…
[ICLR 2026] On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification.