Skip to content
View sebgao's full-sized avatar
🌮
🌮

Block or report sebgao

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

[ICLR 2026] PixNerd: Pixel Neural Field Diffusion

Python 1 Updated Jun 29, 2026

Official implementation of AsymFlow, pi-Flow, GMFlow

Python 461 26 Updated Jul 14, 2026

Flow Map OPD for AnyStep Video Diffusion

Python 411 11 Updated Aug 14, 2026

Awesome Visual Tokenizers/Autoencoders

20 Updated Nov 19, 2025

[ICLR 2026] PixNerd: Pixel Neural Field Diffusion

Python 185 7 Updated Dec 10, 2025

[CVPR 2025 Highlight] Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency

Python 81 9 Updated Jul 16, 2025

Matrix imputation leveraging external covariance structure information

Python 2 Updated Jun 5, 2025

Official implementation of BLIP3o-Series

Python 1,665 79 Updated Nov 29, 2025

the official repo for "D-AR: Diffusion via Autoregressive Models"

Python 138 3 Updated Jan 29, 2026

[CVPR 2026] DDT: Decoupled Diffusion Transformer

Python 416 22 Updated May 22, 2026

Code for: "Long-Context Autoregressive Video Modeling with Next-Frame Prediction"

Python 312 15 Updated Apr 23, 2025

Helpful tools and examples for working with flex-attention

Python 1,225 77 Updated Aug 15, 2026

[CVPR 2025] Open-source, End-to-end, Vision-Language-Action model for GUI Agent & Computer Use.

Python 1,891 142 Updated Apr 24, 2026

[NeurIPS 2024] Exploring DCN-like Architectures for Fast Image Generation with Arbitrary Resolution

Python 38 1 Updated Dec 23, 2024

Code for [CVPR 2025] ROICtrl: Boosting Instance Control for Visual Generation

Python 110 Updated Apr 16, 2025

FQGAN: Factorized Visual Tokenization and Generation

Python 59 3 Updated Mar 29, 2025

[ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.

Python 1,971 93 Updated Jan 8, 2026

[NeurlPS 2024] One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

Python 150 6 Updated Dec 26, 2024

[ICLR 2025] Binary Spherical Quantization + [CVPR 2026] Leech Spherical Quantization

Python 223 7 Updated Aug 5, 2026

Reference implementation for DPO (Direct Preference Optimization)

Python 2,905 237 Updated Aug 11, 2024

VideoLLM-online: Online Video Large Language Model for Streaming Video (CVPR 2024)

Python 680 73 Updated Nov 26, 2025

[ECCV 2024] DragAnything: Motion Control for Anything using Entity Representation

Python 506 17 Updated Jul 2, 2024

[CVPR 2024] Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Python 702 29 Updated Oct 25, 2024

Strong and Open Vision Language Assistant for Mobile Devices

Python 1,370 89 Updated Apr 15, 2024

A family of lightweight multimodal models.

Python 1,052 75 Updated Nov 18, 2024

TensorDict is a pytorch dedicated tensor container.

Python 1,034 116 Updated Aug 14, 2026
Python 75 3 Updated May 10, 2024

HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video

Python 69 6 Updated Dec 12, 2023

MLX: An array framework for Apple silicon

C++ 27,954 2,137 Updated Aug 14, 2026
Next