Highlights
- Pro
Lists (1)
Sort Name ascending (A-Z)
Stars
Official Repo For PerceptionDLM Codebase
[ICML 2026] The official implementation of paper "Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer is Key to Unification"
ARM: An AutoRegressive Large Multimodal Model with Discrete Representations
Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
[ICLR'26] Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
Official Codebase For paper "One-step Language Modeling via Continuous Denoising"
Skill package for ML/CV/NLP paper writing, curated and adapted from Prof. Peng Sida's open notes for Codex, Claude Code, and Gemini.
[🚀 ICLR 2026 Oral] NextStep-1: SOTA Autogressive Image Generation with Continuous Tokens. A research project developed by the StepFun’s Multimodal Intelligence team.
Personal PyTorch implementation of "Generative Modeling via Drifting" with Claude
Elevate your AI research writing, no more tedious polishing ✨
Official code implementation of Context Cascade Compression: Exploring the Upper Limits of Text Compression
PyTorch implementation of JiT https://arxiv.org/abs/2511.13720
Official Implementation of "MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation"
[ICLR 2026 Oral & ICML 2026] Generative Universal Verifier as Multimodal Meta-Reasoner
Official implementation of "Continuous Autoregressive Language Models"
GPU-optimized framework for training diffusion language models at any scale. The backend of Quokka, Super Data Learners, and OpenMoE 2 training.
A Curated List of Awesome Works in World Modeling, Aiming to Serve as a One-stop Resource for Researchers, Practitioners, and Enthusiasts Interested in World Modeling.
[NeurIPS 2025] Encoder-Decoder Diffusion Language Models for Efficient Training and Inference
[ICLR 2026] 🐻 Uniform Discrete Diffusion with Metric Path for Video Generation
[ICLR'26] Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs