Skip to content
View jihanyang's full-sized avatar
🐢
I may be slow to respond.
🐢
I may be slow to respond.

Highlights

  • Pro

Organizations

@vision-x-nyu @cambrian-mllm

Block or report jihanyang

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

maximal update parametrization (µP)

Jupyter Notebook 1,746 104 Updated Jul 17, 2024

Evaluation code for "Benchmarking Visual State Tracking in Multimodal Video Understanding"

Python 38 1 Updated Jun 3, 2026

Cambrian-P: Pose-Grounded Video Understanding

Python 104 3 Updated Jul 28, 2026

Official Implemenation for RAEv2: Improved Baselines with Representation Autoencoders

Python 312 13 Updated May 21, 2026

[CVPR 2026 Oral] VGGT Omega

Python 3,801 245 Updated Jul 15, 2026

[ICML 2026] ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning

Python 82 1 Updated Jul 8, 2026

⚡️SwanLab - an open-source, modern-design AI training tracking and visualization tool. Supports Cloud / Self-hosted use. Integrated with PyTorch / Transformers / verl / LLaMA Factory / ms-swift / U…

Python 4,108 214 Updated Aug 2, 2026

A linear estimator on top of clip to predict the aesthetic quality of pictures

Jupyter Notebook 728 28 Updated Aug 15, 2022

The open agent skills tool - npx skills

TypeScript 27,832 2,348 Updated Jul 31, 2026

A lightweight, AI-native training framework for large language models. Designed for fast iteration, reproducible experiments, and modular configuration across SFT, RLVR, and evaluation workflows.

Python 581 43 Updated May 18, 2026

The first multiplayer video world model in Minecraft

Python 221 9 Updated Mar 3, 2026

MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence

Python 61 1 Updated Mar 11, 2026

Web-based 3D visualization in Python

Python 2,718 210 Updated Aug 2, 2026

Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders

Python 255 4 Updated Feb 13, 2026

ViPE: Video Pose Engine for Geometric 3D Perception

Python 2,059 166 Updated Jun 9, 2026
11 Updated Nov 7, 2025

Cambrian-S: Towards Spatial Supersensing in Video

Python 564 20 Updated Apr 3, 2026
Python 715 71 Updated Apr 12, 2025

The best ChatGPT that $100 can buy.

Python 56,874 7,871 Updated Aug 1, 2026

Native Multimodal Models are World Learners

Python 1,543 69 Updated Dec 30, 2025

[NeurIPS'25] Official repository of Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations

Python 531 29 Updated Apr 7, 2026

Contexts Optical Compression

Python 23,716 2,188 Updated Jan 27, 2026

Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"

Python 1,983 87 Updated Feb 25, 2026

[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.

Python 512 52 Updated Mar 30, 2026

A simple, performant, and scalable Jax LLM!

Python 2,379 578 Updated Aug 2, 2026

Long Video Gen Infrastructure

Python 2,507 240 Updated Jul 31, 2026

🧑‍🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), ga…

Python 67,255 6,748 Updated Jan 22, 2026

[ICLR 2026] On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification.

Python 1,084 26 Updated Aug 1, 2026
Next