Skip to content
View Mickey-Stone's full-sized avatar

Block or report Mickey-Stone

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Human + AI music production workflow for Suno - skills, templates, and tools

Python 413 101 Updated Aug 5, 2026

PersonaPlex code.

Python 10,338 1,444 Updated Mar 2, 2026

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

TypeScript 385,919 81,098 Updated Aug 11, 2026

Spark-TTS Inference Code

Python 11,004 1,160 Updated Apr 9, 2025

Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Python 222 15 Updated Feb 28, 2025

AKShare is an elegant and simple financial data interface library for Python, built for human beings! 开源财经数据接口库

Python 21,956 3,435 Updated Aug 10, 2026

An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)

Python 9,901 998 Updated Jul 14, 2026

Use for saving demo of web-api

C 11 13 Updated May 20, 2022

Dataset

2 Updated Jan 28, 2026

Tools to download and cleanup Common Crawl data

Python 1,045 154 Updated Apr 25, 2023

Alibaba Java Diagnostic Tool Arthas/Alibaba Java诊断利器Arthas

Java 37,483 7,635 Updated Aug 11, 2026

Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.

Python 10,842 1,004 Updated May 16, 2026

LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.

Python 3,144 225 Updated May 19, 2025

Phrase-Based & Neural Unsupervised Machine Translation

Python 1,499 262 Updated Sep 15, 2021

Facebook AI Research Sequence-to-Sequence Toolkit written in Python.

Python 20 1 Updated May 13, 2026

code for paper "Cross-modal Contrastive Learning for Speech Translation" (NAACL 2022)

Python 64 5 Updated May 25, 2022

MooER: Moore-threads Open Omni model for speech-to-speech intERaction. MooER-omni includes a series of end-to-end speech interaction models along with training and inference code, covering but not …

Python 219 17 Updated Jan 8, 2025

Code for our INTERSPEECH paper Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection

Python 114 10 Updated Mar 30, 2025

Whisper realtime streaming for long speech-to-text transcription and translation

Python 3,662 409 Updated Nov 12, 2025

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

Python 22,699 2,620 Updated May 25, 2026

Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.

C 9,055 805 Updated Jul 27, 2026
Python 857 80 Updated Jun 7, 2024

A Framework for Speech, Language, Audio, Music Processing with Large Language Model

Python 1,053 117 Updated Jan 15, 2026

《SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts》

77 5 Updated Jun 9, 2023

Foundational Models for State-of-the-Art Speech and Text Translation

Jupyter Notebook 11,839 1,176 Updated Jul 28, 2026

ASR text preprocessing utility

Python 21 5 Updated Aug 5, 2024

Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

Python 73,993 9,052 Updated Aug 10, 2026

Go ahead and axolotl questions

Python 12,339 1,401 Updated Aug 11, 2026
Next