Stars
Human + AI music production workflow for Suno - skills, templates, and tools
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
AKShare is an elegant and simple financial data interface library for Python, built for human beings! 开源财经数据接口库
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
Tools to download and cleanup Common Crawl data
Alibaba Java Diagnostic Tool Arthas/Alibaba Java诊断利器Arthas
Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.
LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.
Phrase-Based & Neural Unsupervised Machine Translation
formiel / fairseq
Forked from facebookresearch/fairseqFacebook AI Research Sequence-to-Sequence Toolkit written in Python.
code for paper "Cross-modal Contrastive Learning for Speech Translation" (NAACL 2022)
MooER: Moore-threads Open Omni model for speech-to-speech intERaction. MooER-omni includes a series of end-to-end speech interaction models along with training and inference code, covering but not …
Code for our INTERSPEECH paper Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
Whisper realtime streaming for long speech-to-text transcription and translation
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
A Framework for Speech, Language, Audio, Music Processing with Large Language Model
《SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts》
Foundational Models for State-of-the-Art Speech and Text Translation
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)