Skip to content
View DWCTOD's full-sized avatar

Block or report DWCTOD

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

(NIPS 2025) OpenOmni: Official implementation of Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis

Python 142 7 Updated May 9, 2026

CD-Reasoning code

Python 10 Updated Jul 14, 2025
Python 44 2 Updated Jul 9, 2025

OCR, layout analysis, reading order, table recognition in 90+ languages

Python 21,202 1,524 Updated Jul 23, 2026

This is for ACL 2025 Findings Paper: From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalitiesModels

103 3 Updated Mar 22, 2026

[ICCV 2025] MM-IFEngine: Towards Multimodal Instruction Following

Python 126 Updated Feb 13, 2026

[ICLR 2025 Spotlight] OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Python 425 8 Updated May 5, 2025

An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.

Python 1,945 220 Updated Jul 25, 2026

A controllable image composition model which could be used for image blending, image harmonization, view synthesis.

Python 187 12 Updated Jun 28, 2026

Official repository for LTX-Video

Python 10,797 1,107 Updated Jan 5, 2026

"Pexel Downloader: Python-based web scraper for effortlessly downloading high-quality photos and videos from Pexels.com, open-source with MIT License."

Python 22 3 Updated Mar 16, 2026

Official implementation of LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment.

Python 85 4 Updated May 4, 2025

VideoGen-Eval: Agent-based System for Video Generation Evaluation

268 14 Updated Dec 16, 2025

[ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling

Python 3,203 302 Updated Dec 21, 2024

【ArXiv】PDF-Wukong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling

132 4 Updated Jun 4, 2025

Mobile-Agent: The Powerful GUI Agent Family

Python 9,024 903 Updated Jul 7, 2026

MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment

Python 35 1 Updated Jul 1, 2024

Super-Efficient RLHF Training of LLMs with Parameter Reallocation

Python 336 22 Updated Apr 24, 2025

Generate Color Palette from your images using Kmeans and DBSCAN

Python 1 Updated Feb 16, 2021

Cambrian-1 is a family of multimodal LLMs with a vision-centric design.

Python 2,013 139 Updated Nov 7, 2025

NeurIPS 2024 Paper: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing

Python 576 34 Updated Oct 20, 2024

InstructionGPT-4

Python 42 2 Updated Dec 29, 2023

Paper list about multimodal and large language models, only used to record papers I read in the daily arxiv for personal needs.

760 43 Updated May 21, 2026

Dino V2 for Classification, PCA Visualization, Instance Retrival: https://arxiv.org/abs/2304.07193

Jupyter Notebook 206 12 Updated Jul 5, 2023

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

Python 7,986 723 Updated Aug 3, 2026
Python 4,711 471 Updated Jun 15, 2026

✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

788 30 Updated Dec 8, 2025

Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …

Python 15,026 1,569 Updated Aug 3, 2026

通用版面分析 | 中文文档解析 |Document Layout Analysis | layout paser

Python 47 8 Updated Jun 13, 2024
Next