-
Shanghai Jiao Tong University
- https://sotamak1r.github.io/
- @SOTAMak1r
- https://scholar.google.com/citations?user=BXL9nMgAAAAJ&hl=zh-CN
Stars
MAGI-2-preview: Scaling Video Generation Models Efficiently
Open-source Dreamer world-model implementation in JAX
Open-source agentic video editing skills — an AI-powered alternative to Opus Clip and CapCut
Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
Training library for Megatron-based models with bidirectional Hugging Face conversion capability
Official Code of NAVA: Native Audio-Visual Alignment for Generation.
Kandinsky 5.0: A family of diffusion models for Video & Image generation
A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training
The first MoE-based framework for unified and scalable world modeling.
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Infinite Worlds with Versatile Interactions
Open-source native multimodal pretraining — without catastrophic forgetting.
Code for MIRA: Multiplayer Interactive World Models with Representation Autoencoders
Claude Code plugin: automated code review loop with Codex
KVAE-Audio: a continuous full-band audio waveform autoencoder
Humanizer 的汉化版本,Claude Code Skills,旨在消除文本中 AI 生成的痕迹。
Agent skill that removes signs of AI-generated writing from text
[NeurIPS 2025] PyTorch implementation of [ThinkSound], a unified framework for generating audio from any modality, guided by Chain-of-Thought (CoT) reasoning.
"ViMax: Agentic Video Generation (Director, Screenwriter, Producer, and Video Generator All-in-One)"
Boogu-Image-0.1 is an Apache-2.0 open-source image generation and editing model family that delivers near-closed-source performance with an order of magnitude less data.
Code release for "i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models"
AI turns documents or topics into real, native PowerPoint decks—with native shapes, transitions and animations, data-backed charts and tables on demand, audio narration from speaker notes, and supp…
World Model Self-Distillation project website
Ideogram 4: Open image model at the forefront of design