Skip to content
View Mickey-Stone's full-sized avatar

Block or report Mickey-Stone

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Human + AI music production workflow for Suno - skills, templates, and tools

Python 423 103 Updated Aug 15, 2026

PersonaPlex code.

Python 10,361 1,443 Updated Mar 2, 2026

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

TypeScript 386,371 81,206 Updated Aug 15, 2026

Spark-TTS Inference Code

Python 11,003 1,159 Updated Apr 9, 2025

Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Python 222 15 Updated Feb 28, 2025

AKShare is an elegant and simple financial data interface library for Python, built for human beings! 开源财经数据接口库

Python 22,047 3,449 Updated Aug 13, 2026

An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)

Python 9,914 1,001 Updated Aug 13, 2026

Use for saving demo of web-api

C 11 13 Updated May 20, 2022

Dataset

2 Updated Jan 28, 2026

Tools to download and cleanup Common Crawl data

Python 1,045 154 Updated Apr 25, 2023

Alibaba Java Diagnostic Tool Arthas/Alibaba Java诊断利器Arthas

Java 37,486 7,636 Updated Aug 14, 2026

Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.

Python 10,868 1,006 Updated May 16, 2026

LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.

Python 3,145 225 Updated May 19, 2025

Phrase-Based & Neural Unsupervised Machine Translation

Python 1,499 261 Updated Sep 15, 2021

Facebook AI Research Sequence-to-Sequence Toolkit written in Python.

Python 20 1 Updated May 13, 2026

code for paper "Cross-modal Contrastive Learning for Speech Translation" (NAACL 2022)

Python 64 5 Updated May 25, 2022

MooER: Moore-threads Open Omni model for speech-to-speech intERaction. MooER-omni includes a series of end-to-end speech interaction models along with training and inference code, covering but not …

Python 219 17 Updated Jan 8, 2025

Code for our INTERSPEECH paper Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection

Python 114 10 Updated Mar 30, 2025

Whisper realtime streaming for long speech-to-text transcription and translation

Python 3,662 409 Updated Nov 12, 2025

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

Python 22,767 2,628 Updated May 25, 2026

Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.

C 9,078 807 Updated Aug 12, 2026
Python 857 79 Updated Jun 7, 2024

A Framework for Speech, Language, Audio, Music Processing with Large Language Model

Python 1,054 117 Updated Jan 15, 2026

《SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts》

77 5 Updated Jun 9, 2023

Foundational Models for State-of-the-Art Speech and Text Translation

Jupyter Notebook 11,839 1,175 Updated Jul 28, 2026

ASR text preprocessing utility

Python 21 5 Updated Aug 5, 2024

Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

Python 74,125 9,069 Updated Aug 13, 2026

Go ahead and axolotl questions

Python 12,360 1,406 Updated Aug 13, 2026
Next