Highlights
- Pro
Stars
Automatically Update Text-to-speech (TTS) Papers Daily using Github Actions (Update Every 12th hours)
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singi…
Generative models for conditional audio generation
Provide with pre-build flash-attention 2 and 3 package wheels on Linux and Windows using GitHub Actions
A list of tools, papers and code related to Fake Audio Detection.
We Speech Toolkit, LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
The project is associated with the recently-launched INTERSPEECH 2025 Workshop on Multilingual Conversational Speech Language Model (MLC-SLM) to provide participants with baseline systems for speec…
OSUM & OSUM-EChat, open speech understanding model and empathetic spoken chatbot based on it, open-sourced by ASLP@NPU.
A comprehensive mapping database of English to Chinese technical vocabulary in the artificial intelligence domain
Awesome speech/audio LLMs, representation learning, and codec models
主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题
Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.
《代码随想录》LeetCode 刷题攻略:200道经典题目刷题顺序,共60w字的详细图解,视频难点剖析,50余张思维导图,支持C++,Java,Python,Go,JavaScript等多语言版本,从此算法学习不再迷茫!🔥🔥 来看看,你会发现相见恨晚!🚀
Build local voice agents with open-source models
The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
A python package to analyze and compare voices with deep learning
🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.
Oh my tmux! My self-contained, pretty & versatile tmux configuration made with 💛🩷💙🖤❤️🤍
🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language Model(VLM), Vision Language Action(VLA), AI Generated Content(AIGC), the related Datasets and Applic…
LLMs interview notes and answers:该仓库主要记录大模型(LLMs)算法工程师相关的面试题和参考答案
Crack LeetCode, not only how, but also why.
Minimal extension of OpenAI's Whisper adding speaker diarization with special tokens
Evaluate your speech-to-text system with similarity measures such as word error rate (WER)
AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
西北工业大学硕博学位论文模版 | Yet Another Thesis Template for Northwestern Polytechnical University