Skip to content
View jihanyang's full-sized avatar
🐢
I may be slow to respond.
🐢
I may be slow to respond.

Highlights

  • Pro

Organizations

@vision-x-nyu @cambrian-mllm

Block or report jihanyang

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

maximal update parametrization (µP)

Jupyter Notebook 1,749 105 Updated Jul 17, 2024

Evaluation code for "Benchmarking Visual State Tracking in Multimodal Video Understanding"

Python 40 1 Updated Jun 3, 2026

Cambrian-P: Pose-Grounded Video Understanding

Python 104 3 Updated Jul 28, 2026

Official Implemenation for RAEv2: Improved Baselines with Representation Autoencoders

Python 315 13 Updated May 21, 2026

[CVPR 2026 Oral] VGGT Omega

Python 3,915 263 Updated Jul 15, 2026

[ICML 2026] ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning

Python 83 1 Updated Jul 8, 2026

⚡️SwanLab - an open-source, modern-design AI training tracking and visualization tool. Supports Cloud / Self-hosted use. Integrated with PyTorch / Transformers / verl / LLaMA Factory / ms-swift / U…

Python 4,129 214 Updated Aug 9, 2026

A linear estimator on top of clip to predict the aesthetic quality of pictures

Jupyter Notebook 729 28 Updated Aug 15, 2022

The open agent skills tool - npx skills

TypeScript 28,484 2,418 Updated Aug 9, 2026

A lightweight, AI-native training framework for large language models. Designed for fast iteration, reproducible experiments, and modular configuration across SFT, RLVR, and evaluation workflows.

Python 582 43 Updated May 18, 2026

The first multiplayer video world model in Minecraft

Python 221 9 Updated Mar 3, 2026

MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence

Python 62 1 Updated Mar 11, 2026

Web-based 3D visualization in Python

Python 2,725 211 Updated Aug 9, 2026

Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders

Python 255 4 Updated Feb 13, 2026

ViPE: Video Pose Engine for Geometric 3D Perception

Python 2,065 168 Updated Jun 9, 2026

SIMS-V: simulated instruction-tuning data generation for spatial video understanding

Python 11 Updated Aug 9, 2026

Cambrian-S: Towards Spatial Supersensing in Video

Python 564 20 Updated Apr 3, 2026
Python 715 71 Updated Apr 12, 2025

The best ChatGPT that $100 can buy.

Python 57,077 7,911 Updated Aug 2, 2026

Native Multimodal Models are World Learners

Python 1,543 69 Updated Dec 30, 2025

[NeurIPS'25] Official repository of Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations

Python 533 29 Updated Apr 7, 2026

Contexts Optical Compression

Python 23,758 2,195 Updated Jan 27, 2026

Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"

Python 1,988 87 Updated Feb 25, 2026

[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.

Python 512 52 Updated Mar 30, 2026

A simple, performant, and scalable Jax LLM!

Python 2,382 582 Updated Aug 9, 2026

Long Video Gen Infrastructure

Python 2,525 242 Updated Aug 7, 2026

🧑‍🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), ga…

Python 67,301 6,748 Updated Jan 22, 2026

[ICLR 2026] On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification.

Python 1,095 26 Updated Aug 1, 2026
Next