Skip to content
View JamesHujy's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report JamesHujy

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Long Video Gen Infrastructure

Python 2,541 242 Updated Aug 7, 2026
Jupyter Notebook 220 3 Updated Dec 19, 2025
Python 2 1 Updated Oct 16, 2025
Python 42 6 Updated Dec 26, 2025

Official implementation of "Figure It Out: Improve the Frontier of Reasoning with Active Visual Thinking"

Python 17 Updated Jan 13, 2026
Python 631 59 Updated Feb 26, 2026

Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"

Python 1,991 87 Updated Feb 25, 2026

A General, Accurate, Long-Horizon, and Efficient Mobile Agent driven by Multimodal Foundation Models

Python 246 13 Updated Nov 18, 2025

LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence https://arxiv.org/abs/2509.03505

Python 3,961 306 Updated Jun 16, 2026
Python 890 52 Updated Sep 15, 2025

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,971 4,410 Updated Aug 15, 2026

Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team.

Python 16,788 1,238 Updated Mar 24, 2026
Python 83 1 Updated Oct 18, 2025

FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Trae, Traycer AI…

142,848 34,843 Updated Aug 11, 2026

Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.

1,503 47 Updated Mar 9, 2026

OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871

Jupyter Notebook 4,113 31 Updated Mar 20, 2026

MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model.

Python 3,169 283 Updated Jul 7, 2025

Official implementation of BLIP3o-Series

Python 1,665 79 Updated Nov 29, 2025

[ICLR 2025 Oral] Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models

Python 1,027 79 Updated Jul 10, 2025

UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Python 888 30 Updated Dec 23, 2025

The official repo of One RL to See Them All: Visual Triple Unified Reinforcement Learning

Python 332 17 Updated May 31, 2025

Open-source unified multimodal model

Python 6,146 547 Updated May 4, 2026

[AAAI 2026] - Official repo for paper: "Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't"

Python 292 30 Updated Mar 11, 2026

Scalable RL solution for advanced reasoning of language models

Python 1,867 116 Updated Mar 18, 2025

PoC for "SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning" [NeurIPS '25]

Python 75 9 Updated Oct 2, 2025

Video-R1: Reinforcing Video Reasoning in MLLMs [🔥the first paper to explore R1 for video]

Python 885 46 Updated Dec 14, 2025

Cosmos-Reason1 models understand the physical common sense and generate appropriate embodied decisions in natural language through long chain-of-thought reasoning processes.

Python 952 84 Updated Jun 7, 2026

This package contains the original 2012 AlexNet code.

Cuda 2,965 392 Updated Mar 12, 2025
Next