Skip to content
View xinli2008's full-sized avatar
🌴
On vacation
🌴
On vacation

Block or report xinli2008

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!

Jupyter Notebook 2,329 157 Updated Apr 13, 2026
Python 10 Updated Jun 11, 2026

[ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing

Python 3,882 267 Updated Oct 17, 2025

Official PyTorch re-implementation of MiniT2I.

Python 288 12 Updated Jun 24, 2026

Official implementation for "Multimodal Chain-of-Thought Reasoning in Language Models" (stay tuned and more will be updated)

Python 3,985 332 Updated Jun 12, 2024

A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.

Python 1,291 91 Updated Jul 14, 2026

Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.

Python 1,176 91 Updated Jul 13, 2026

Official repo for "Let ViT Speak: Generative Language-Image Pre-training"

Python 133 4 Updated Jun 10, 2026

Presentation Slides for Developers

TypeScript 47,840 2,127 Updated Jul 22, 2026

🎙️ 「大模型」从0训练0.1B能听能说能看的全模态Omni模型!A 0.1B Omni model trained from scratch, capable of listening, speaking, and seeing!

Python 2,182 253 Updated Jun 28, 2026

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 30,713 7,385 Updated Jul 25, 2026

Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation

Python 740 30 Updated Jul 22, 2026

SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles

Python 4,359 370 Updated Jul 16, 2026

LLaDA2.0-Uni: Understanding and Generation the World.

Python 769 49 Updated May 29, 2026

1K resolution vision transformers pretrained on 1B human images.

Python 880 60 Updated May 24, 2026

An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of…

Python 77,791 10,611 Updated Jul 24, 2026

This is an ultra-simple, single-file PyTorch implementation of MoonViT, the native-resolution vision encoder from Kimi-VL.

Python 28 5 Updated Apr 25, 2026

🔍大模型应用开发实战一:RAG 技术全栈指南,在线阅读地址:https://datawhalechina.github.io/all-in-rag/

Python 9,765 4,886 Updated Jul 20, 2026

An open source implementation of CLIP.

Python 14,020 1,296 Updated Jul 17, 2026

📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程

Python 68,436 8,517 Updated Jul 17, 2026

18 Lessons to Get Started Building AI Agents

Jupyter Notebook 70,275 23,273 Updated Jul 22, 2026

Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1

Python 72,153 11,698 Updated Jun 26, 2026

This repository collects Visual Autoregressive (VAR) modeling papers from 2024 to 2026 published at top-tier conferences, as well as relevant works available on arXiv.

21 1 Updated Mar 13, 2026

AIInfra(AI 基础设施)指AI系统从底层芯片等硬件,到上层软件栈支持AI大模型训练和推理。

Jupyter Notebook 7,706 992 Updated Dec 22, 2025

Ongoing research training transformer models at scale

Python 17,204 4,282 Updated Jul 25, 2026

large scale pre-training VLMs

Python 25 1 Updated Jul 6, 2026

Awesome Multimodal Modeling [Covers MLLM, UMM, and NMM]

505 25 Updated Jul 19, 2026

JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.

Python 2,241 158 Updated Jul 17, 2026

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,650 4,275 Updated Jul 24, 2026

Based on Nano-vLLM, a simple replication of vLLM with self-contained paged attention and flash attention implementation

Python 937 149 Updated Jul 22, 2026
Next