Skip to content
View JamesHujy's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report JamesHujy

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Long Video Gen Infrastructure

Python 2,532 242 Updated Aug 7, 2026
Jupyter Notebook 220 3 Updated Dec 19, 2025
Python 2 1 Updated Oct 16, 2025
Python 42 6 Updated Dec 26, 2025

Official implementation of "Figure It Out: Improve the Frontier of Reasoning with Active Visual Thinking"

Python 17 Updated Jan 13, 2026
Python 630 59 Updated Feb 26, 2026

Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"

Python 1,988 87 Updated Feb 25, 2026

A General, Accurate, Long-Horizon, and Efficient Mobile Agent driven by Multimodal Foundation Models

Python 246 13 Updated Nov 18, 2025

LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence https://arxiv.org/abs/2509.03505

Python 3,932 306 Updated Jun 16, 2026
Python 890 52 Updated Sep 15, 2025

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,918 4,382 Updated Aug 11, 2026

Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team.

Python 16,778 1,237 Updated Mar 24, 2026
Python 83 1 Updated Oct 18, 2025

FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Trae, Traycer AI…

142,732 34,839 Updated Aug 11, 2026

Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.

1,503 47 Updated Mar 9, 2026

OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871

Jupyter Notebook 4,112 31 Updated Mar 20, 2026

MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model.

Python 3,166 283 Updated Jul 7, 2025

Official implementation of BLIP3o-Series

Python 1,664 79 Updated Nov 29, 2025

[ICLR 2025 Oral] Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models

Python 1,027 79 Updated Jul 10, 2025

UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Python 888 30 Updated Dec 23, 2025

The official repo of One RL to See Them All: Visual Triple Unified Reinforcement Learning

Python 331 17 Updated May 31, 2025

Open-source unified multimodal model

Python 6,144 546 Updated May 4, 2026

[AAAI 2026] - Official repo for paper: "Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't"

Python 292 30 Updated Mar 11, 2026

Scalable RL solution for advanced reasoning of language models

Python 1,869 116 Updated Mar 18, 2025

PoC for "SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning" [NeurIPS '25]

Python 75 9 Updated Oct 2, 2025

Video-R1: Reinforcing Video Reasoning in MLLMs [🔥the first paper to explore R1 for video]

Python 886 46 Updated Dec 14, 2025

Cosmos-Reason1 models understand the physical common sense and generate appropriate embodied decisions in natural language through long chain-of-thought reasoning processes.

Python 951 84 Updated Jun 7, 2026

This package contains the original 2012 AlexNet code.

Cuda 2,963 392 Updated Mar 12, 2025
Next