Stars
Bili-Sakura / PixNerd-diffusers
Forked from MCG-NJU/PixNerd[ICLR 2026] PixNerd: Pixel Neural Field Diffusion
Official implementation of AsymFlow, pi-Flow, GMFlow
[ICLR 2026] PixNerd: Pixel Neural Field Diffusion
[CVPR 2025 Highlight] Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
Matrix imputation leveraging external covariance structure information
the official repo for "D-AR: Diffusion via Autoregressive Models"
Code for: "Long-Context Autoregressive Video Modeling with Next-Frame Prediction"
Helpful tools and examples for working with flex-attention
[CVPR 2025] Open-source, End-to-end, Vision-Language-Action model for GUI Agent & Computer Use.
[NeurIPS 2024] Exploring DCN-like Architectures for Fast Image Generation with Arbitrary Resolution
Code for [CVPR 2025] ROICtrl: Boosting Instance Control for Visual Generation
FQGAN: Factorized Visual Tokenization and Generation
[ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.
[NeurlPS 2024] One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
[ICLR 2025] Binary Spherical Quantization + [CVPR 2026] Leech Spherical Quantization
Reference implementation for DPO (Direct Preference Optimization)
VideoLLM-online: Online Video Large Language Model for Streaming Video (CVPR 2024)
[ECCV 2024] DragAnything: Motion Control for Anything using Entity Representation
[CVPR 2024] Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
Strong and Open Vision Language Assistant for Mobile Devices
TensorDict is a pytorch dedicated tensor container.
HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video