Skip to content
View luyao-cv's full-sized avatar

Organizations

@CLFClub

Block or report luyao-cv

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!

1,440 68 Updated Aug 2, 2026

Witness the aha moment of VLM with less than $3.

Python 4,061 282 Updated May 19, 2025

利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.

Python 103,724 15,719 Updated Aug 13, 2026

Next-Token Prediction is All You Need

Python 2,437 100 Updated Jan 12, 2026

Codebase for Aria - an Open Multimodal Native MoE

Jupyter Notebook 1,087 89 Updated Jan 22, 2025

[ICML'25][TPAMI'26] Official implementation of paper "SparseVLM" and "SparseVLM+".

Python 272 22 Updated Jul 30, 2026

INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model

Python 42 Updated Aug 4, 2024

🔍 An LLM-based Multi-agent Framework of Web Search Engine (like Perplexity.ai Pro and SearchGPT)

JavaScript 6,913 689 Updated Jul 4, 2025

[ACL2024 Findings] Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

361 10 Updated Mar 22, 2024

Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audi…

Python 10,231 847 Updated Mar 25, 2026

VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior

801 49 Updated Dec 5, 2023
Python 230 270 Updated Nov 25, 2023

A simple and open-source analogue of the HeyGen system

Python 1,024 211 Updated Aug 1, 2024
Python 429 38 Updated Nov 1, 2023

Instant voice cloning by MIT and MyShell. Audio foundation model.

Python 37,148 4,143 Updated Apr 19, 2025

[ACL2024] T-Eval: Evaluating Tool Utilization Capability of Large Language Models Step by Step

Python 312 16 Updated Apr 3, 2024

Some meaningless nscripter tools.

Python 677 1,084 Updated Jul 8, 2020

Text-to-Audio/Music Generation

Python 2,637 210 Updated Sep 29, 2024

飞桨大模型开发套件,提供大语言模型、跨模态大模型、生物计算大模型等领域的全流程开发工具链。

Python 481 165 Updated May 24, 2024

Stable Diffusion web UI

Python 164,498 30,554 Updated Mar 2, 2026

The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.

Python 127,677 15,029 Updated Aug 15, 2026

OpenMMLab Pre-training Toolbox and Benchmark

Python 3,849 1,109 Updated Nov 1, 2024

AudioLDM: Generate speech, sound effects, music and beyond, with text.

Python 2,905 271 Updated Jun 25, 2025

AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

Python 10,172 850 Updated Jul 6, 2024

骆驼(Luotuo): Open Sourced Chinese Language Models. Developed by 陈启源 @ 华中师范大学 & 李鲁鲁 @ 商汤科技 & 冷子昂 @ 商汤科技

Jupyter Notebook 3,589 241 Updated Sep 3, 2023

[TPAMI2024] Codes and Models for VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Python 311 18 Updated Dec 25, 2024

Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high …

Python 724 223 Updated Mar 6, 2026

Code release for paper "You Only Segment Once: Towards Real-Time Panoptic Segmentation" [CVPR 2023]

Python 292 24 Updated Jul 21, 2023

用来进行简单查重以及基于OpenAI的GPT3接口进行文章润色的小程序,使用pyqt作为GUI框架

Python 6 1 Updated Mar 15, 2023

Object Detection toolkit based on PaddlePaddle. It supports object detection, instance segmentation, multiple object tracking and real-time multi-person keypoint detection.

Python 14,374 3,025 Updated May 28, 2026
Next