Stars
🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"
📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥
🚀 Efficient implementations for emerging model architectures
Native Multimodal Models are World Learners
The official GitHub repo for the survey paper "A Survey on Diffusion Language Models".
[TMLR 2026] Multimodal Large Language Models for Code Generation under Multimodal Scenarios
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
[NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team.
The Open Cookbook for Top-Tier Code Large Language Model
Refine high-quality datasets and visual AI models
[CVPR'24] Interactive3D: Create What You Want by Interactive 3D Generation
CustomDiffusion360: Customizing Text-to-Image Diffusion with Camera Viewpoint Control
[ECCV 2024 Oral] LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation.
[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)
[CVPR 2024] Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. Foundation Model for Monocular Depth Estimation
This repository provides the code and model checkpoints for AIMv1 and AIMv2 research projects.
Code for MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World
[ECCV 2024] Street Gaussians: Modeling Dynamic Urban Scenes with Gaussian Splatting
Implementation of MeshGPT, SOTA Mesh generation using Attention, in Pytorch
[CVPR'24] Text-to-3D Generation with Bidirectional Diffusion using both 2D and 3D priors
MeshGPT: Generating Triangle Meshes with Decoder-Only Transformers
Single Image to 3D using Cross-Domain Diffusion for 3D Generation
Curated list of papers and resources focused on 3D Gaussian Splatting, intended to keep pace with the anticipated surge of research in the coming months.
VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
🪐 Objaverse-XL is a Universe of 10M+ 3D Objects. Contains API Scripts for Downloading and Processing!