Stars
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
[Arxiv 2025] ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions
A unified diffusers implementation for MVDream and ImageDream
Hunyuan-DiT : A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
A unified framework for 3D content generation.
[NeurIPS 2023] This repo contains the code for our paper Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP
MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion, NeurIPS 2023 (spotlight)
Direct voxel grid optimization for fast radiance field reconstruction.
Code for robust monocular depth estimation described in "Ranftl et. al., Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer, TPAMI 2022"
VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis [ICLR 2023]
算法面试必备,推荐刷题网站www.lintcode.com。北大学霸的《LeetCode刷题模板》+V领取: jiuzhangfeifei
Contains system design materials to prepare for system design interviews 🚩👨💻👨💻👨💻
OpenMMLab Detection Toolbox and Benchmark
Productive, portable, and performant GPU programming in Python.
10 differentiable physical simulators built with Taichi differentiable programming (DiffTaichi, ICLR 2020)
Real-Time SLAM for Monocular, Stereo and RGB-D Cameras, with Loop Detection and Relocalization Capabilities
Differentiable architecture search for convolutional and recurrent networks
LEGO: Learning Edge with Geometry all at Once by Watching Videos
Train the HRNet model on ImageNet