-
The University of Hong Kong
- Hong Kong SAR
-
17:34
(UTC +08:00) - jinghuahou7@gmail.com
Highlights
- Pro
Stars
A latent text-to-image diffusion model
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
Reference PyTorch implementation and models for DINOv3
Taming Transformers for High-Resolution Image Synthesis
Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation
A data generation pipeline for creating semi-realistic synthetic multi-object videos with rich annotations such as instance segmentation masks, depth maps, and optical flow.
Official repository of 'Visual-RFT: Visual Reinforcement Fine-Tuning' & 'Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning'’
A suite of image and video neural tokenizers
[CVPR 2024] LMDrive: Closed-Loop End-to-End Driving with Large Language Models
[AAAI2024] Far3D: Expanding the Horizon for Surround-view 3D Object Detection
[NeurIPS 2023] Dynamo-Depth: Fixing Unsupervised Depth Estimation for Dynamical Scenes