Stars
This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!
Witness the aha moment of VLM with less than $3.
利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.
Codebase for Aria - an Open Multimodal Native MoE
[ICML'25][TPAMI'26] Official implementation of paper "SparseVLM" and "SparseVLM+".
INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model
🔍 An LLM-based Multi-agent Framework of Web Search Engine (like Perplexity.ai Pro and SearchGPT)
[ACL2024 Findings] Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audi…
VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior
A simple and open-source analogue of the HeyGen system
Instant voice cloning by MIT and MyShell. Audio foundation model.
[ACL2024] T-Eval: Evaluating Tool Utilization Capability of Large Language Models Step by Step
飞桨大模型开发套件,提供大语言模型、跨模态大模型、生物计算大模型等领域的全流程开发工具链。
Stable Diffusion web UI
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
OpenMMLab Pre-training Toolbox and Benchmark
AudioLDM: Generate speech, sound effects, music and beyond, with text.
AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
骆驼(Luotuo): Open Sourced Chinese Language Models. Developed by 陈启源 @ 华中师范大学 & 李鲁鲁 @ 商汤科技 & 冷子昂 @ 商汤科技
[TPAMI2024] Codes and Models for VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high …
Code release for paper "You Only Segment Once: Towards Real-Time Panoptic Segmentation" [CVPR 2023]
用来进行简单查重以及基于OpenAI的GPT3接口进行文章润色的小程序,使用pyqt作为GUI框架
Object Detection toolkit based on PaddlePaddle. It supports object detection, instance segmentation, multiple object tracking and real-time multi-person keypoint detection.