Stars
(NIPS 2025) OpenOmni: Official implementation of Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis
OCR, layout analysis, reading order, table recognition in 90+ languages
This is for ACL 2025 Findings Paper: From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalitiesModels
[ICCV 2025] MM-IFEngine: Towards Multimodal Instruction Following
[ICLR 2025 Spotlight] OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
A controllable image composition model which could be used for image blending, image harmonization, view synthesis.
"Pexel Downloader: Python-based web scraper for effortlessly downloading high-quality photos and videos from Pexels.com, open-source with MIT License."
Official implementation of LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment.
VideoGen-Eval: Agent-based System for Video Generation Evaluation
[ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling
【ArXiv】PDF-Wukong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
Mobile-Agent: The Powerful GUI Agent Family
MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment
Super-Efficient RLHF Training of LLMs with Parameter Reallocation
Generate Color Palette from your images using Kmeans and DBSCAN
Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
NeurIPS 2024 Paper: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
Paper list about multimodal and large language models, only used to record papers I read in the daily arxiv for personal needs.
Dino V2 for Classification, PCA Visualization, Instance Retrival: https://arxiv.org/abs/2304.07193
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
通用版面分析 | 中文文档解析 |Document Layout Analysis | layout paser