Skip to content
View qyc-98's full-sized avatar
  • Tsinghua University
  • Beijing

Block or report qyc-98

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

Reading notes about Multimodal Large Language Models, Large Language Models, and Diffusion Models

1,188 50 Updated Jul 14, 2026

[ICLR 2026] Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and Vision

Python 234 7 Updated May 31, 2026

Official repository of "GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing"

Jupyter Notebook 317 12 Updated Sep 28, 2025
Python 83 1 Updated Oct 18, 2025

Just for linear system control course ENGG5403

MATLAB 4 Updated May 11, 2020

This is the official implementation of our Señorita-2M [Weights and Dataset] : A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists

Python 112 1 Updated Apr 9, 2025

The official VOT Challenge evaluation and analysis toolkit

Python 205 53 Updated Jun 9, 2026

Official repository of 'Visual-RFT: Visual Reinforcement Fine-Tuning' & 'Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning'’

Jupyter Notebook 2,264 111 Updated Oct 29, 2025

Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without …

TypeScript 151,340 23,889 Updated Aug 4, 2026

📖 This is a repository for organizing papers, codes and other resources related to unified multimodal models.

828 41 Updated Oct 10, 2025

Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.

Jupyter Notebook 19,724 1,831 Updated Jan 30, 2026

[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型

Python 10,117 790 Updated Sep 22, 2025

The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use th…

Jupyter Notebook 19,648 2,522 Updated May 30, 2026

[ECCV 2022] XMem: Long-Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model

Python 1,983 211 Updated Nov 15, 2024

[NeurIPS 2021] Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object Segmentation

Python 568 73 Updated Mar 15, 2024

Video-Inpaint-Anything: This is the inference code for our paper CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility.

Python 325 11 Updated Sep 24, 2024

[ECCV 2024] Elysium: Exploring Object-level Perception in Videos via MLLM

Python 89 5 Updated Oct 25, 2024

Tarsier -- a family of large-scale video-language models, which is designed to generate high-quality video descriptions , together with good capability of general video understanding.

Python 548 32 Updated Aug 14, 2025

Official Repo For OMG-LLaVA and OMG-Seg codebase [CVPR-24 and NeurIPS-24]

Python 1,353 55 Updated Oct 15, 2025

[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with object segmentation masks.

Python 965 55 Updated Aug 5, 2025

The HD-VG-130M Dataset

126 2 Updated Apr 8, 2024

[TPAMI2024] Codes and Models for VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Python 311 18 Updated Dec 25, 2024

Long Context Transfer from Language to Vision

Python 408 18 Updated Mar 18, 2025

A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone

Python 26,095 2,043 Updated Aug 4, 2026

This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.

Python 12,159 1,061 Updated Mar 8, 2026

[CVPR 2024] Code release for "InstanceDiffusion: Instance-level Control for Image Generation"

Python 614 32 Updated Jun 17, 2025

A collection of awesome video generation studies.

TeX 779 43 Updated Mar 31, 2026

[ECCV 2024] DragAnything: Motion Control for Anything using Entity Representation

Python 506 17 Updated Jul 2, 2024

Official repo for VGen: a holistic video generation ecosystem for video generation building on diffusion models

Python 3,156 274 Updated Jan 10, 2025

[ICLR'24 spotlight] Chinese and English Multimodal Large Model Series (Chat and Paint) | 基于CPM基础模型的中英双语多模态大模型系列

Python 1,063 88 Updated Jun 13, 2024
Next