-
Stanford University
- Stanford, California
- https://bchao1.github.io
- @BrianCChao
- in/brian-chao-85425415a
Starred repositories
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
Solve puzzles. Improve your pytorch.
A skill to stop your coding agent from burying the answer. ADHD-friendly output.
Official implementation of MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models
Official codebase for Fast-WAM: Do World Action Models Need Test-time Future Imagination?
Official PyTorch Implementation of Unified Video Action Model (RSS 2025)
[ICML 2026] Official implementation for "DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion".
[to appear at NeurIPS 2026] Official implementation of "MilliVid: Adaptive Latents for Long-Range Consistency in Video Generation"
Sekai2: From World Exploration to Interactive World Modeling
[ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.
An open-source AI agent that brings the power of Gemini directly into your terminal.
[ICLR 2026] Official Implementation of Muddit [Meissonic II]: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model.
MMaDA - Open-Sourced Multimodal Large Diffusion Language Models (dLLMs with block diffusion, mixed-CoT, unified RL)
NVIDIA FastGen: Fast Generation from Diffusion Models
🪨 why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
Official implementation of "Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer" (ECCV 2026)
The official code for NeurIPS 2025 "MagCache: Fast Video Generation with Magnitude-Aware Cache"
LLM Wiki is a cross-platform desktop application that turns your documents into an organized, interlinked knowledge base — automatically. Instead of traditional RAG (retrieve-and-answer from scratc…
Florence-2 is a novel vision foundation model with a unified, prompt-based representation for a variety of computer vision and vision-language tasks.
Gemma open-weight LLM library, from Google DeepMind
Code repository for "Spectral Progressive Diffusion for Efficient Image and Video Generation"
Code release for "Foveated Diffusion: Efficient Spatially Adaptive Image and Video Generation"
bchao1 / sglang
Forked from sgl-project/sglangSGLang is a high-performance serving framework for large language models and multimodal models.
An agentic skills framework & software development methodology that works.
Ideogram 4: Open image model at the forefront of design
Wrapper of 50+ image matching models with a unified interface
Implementation of Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players