Stars
A Roadmap to Build World Models for Robot Policy Evaluation
ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works…
GigaWorld-Policy: An Efficient Action-Centered World–Action Model
A comprehensive list of papers for the definition of World Models and using World Models for General Video Generation, Embodied AI, and Autonomous Driving, including papers, codes, and related webs…
We release Evo-RL, the opensource real-world offline RL on So-101 and AgileX PiPER for easier reproduction.
GigaBrain-0: A World Model-Powered Vision-Language-Action Model
GigaWorld-0: World Models as Data Engine to Empower Embodied AI
GigaTrain: An Efficient and Scalable Training Framework for AI Models
Simulation platform for general-purpose robotics & embodied AI learning.
[CSUR] A Survey on Video Diffusion Models
[ECCV`24&ICLR`25] CityGaussian Series for High-quality Large-Scale Scene Reconstruction with Gaussians
[CVPR'25 Highlight] You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
Official Pytorch Implementation of Synthesizing Coherent Story with Auto-Regressive Latent Diffusion Models
[CVPR 2025] Official implementation of "AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models"
Official implementation of "MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling"
Official implementation of `Splatter Image: Ultra-Fast Single-View 3D Reconstruction' CVPR 2024
[ICCV'25]DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion
3D高斯论文,持续更新,欢迎交流讨论。
Open-Sora: Democratizing Efficient Video Production for All
The best OSS video generation models, created by Genmo
This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
deep learning for image processing including classification and object-detection etc.
Large World Model -- Modeling Text and Video with Millions Context
Multi-Target Multi-Camera Human Tracking (Non-overlapping camera system)
深蓝学院 多传感器定位融合第四期 学习笔记