Skip to content
View xjchenGit's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report xjchenGit

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech

Jupyter Notebook 372 41 Updated Aug 14, 2025

WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

Python 23,592 2,383 Updated Jul 13, 2026

Train transformer language models with reinforcement learning.

Python 19,082 2,910 Updated Aug 16, 2026

NeurIPS 2025 Spotlight; ICLR2024 Spotlight; CVPR 2024; EMNLP 2024

Python 1,851 80 Updated Aug 11, 2026

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,974 4,410 Updated Aug 15, 2026

Covo-Audio is a 7B-parameter end-to-end large audio language model that directly processes continuous audio inputs and generates audio outputs within a single unified architecture.

Python 177 17 Updated Mar 17, 2026

蒸餾李宏毅老師的skill,結合Karpathy的LLM,Fable 5加持 以及 本人親自訪談

HTML 964 100 Updated Aug 11, 2026

Expert code review skill: SOLID, security, performance, error handling, boundary conditions

Python 3,853 340 Updated May 11, 2026

From Automated Idea Factory to Realization

Shell 1,391 124 Updated Jul 18, 2026

Official Repository of UltraVoice

JavaScript 64 2 Updated Oct 28, 2025

Plug-and-play streaming semantic VAD for real-time full-duplex spoken dialogue systems.

Python 295 39 Updated Jul 17, 2026

Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Python 1,037 142 Updated Dec 2, 2025

High-Quality Voice Cloning TTS for 600+ Languages

Python 9,149 1,503 Updated Aug 10, 2026

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

Python 35,710 4,087 Updated Aug 12, 2026

LLM-Codec: Neural Audio Codec Meets Language Model Objectives

Python 23 1 Updated May 3, 2026

Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.

Python 1,495 112 Updated Mar 16, 2026

MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech

Python 205 12 Updated Jun 20, 2026

Liquid Audio - Speech-to-Speech audio models by Liquid AI

Python 561 96 Updated Jun 5, 2026

Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding

Jupyter Notebook 10,423 1,094 Updated Aug 4, 2026

verl: Volcano Engine Reinforcement Learning for LLMs

Python 42 5 Updated Jun 23, 2025

Generative World Renderer: an AI-native Renderer for Games and Virtual Worlds.

Python 711 19 Updated May 5, 2026

Reverse Engineering of Supervised Semantic Speech Tokenizer (S3Tokenizer) proposed in CosyVoice

Python 527 69 Updated Dec 22, 2025

A survey of spoken dialogue models (SDMs) with speech input and speech output. Focus on their Intermediate Representation and Generation Pattern

32 Updated Mar 24, 2026

Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …

Python 15,209 1,599 Updated Aug 16, 2026

Align Anything: Training All-modality Model with Feedback

Python 4,666 505 Updated Nov 27, 2025

open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.

Python 3,570 312 Updated Nov 5, 2024

Simple text to phones converter for multiple languages

Python 1,567 201 Updated Aug 4, 2026

Wrap Antigravity, ChatGPT Codex, Claude Code, Grok Build as an OpenAI/Gemini/Claude/Codex compatible API service, allowing you to enjoy the free Gemini 3.1 Pro, GPT 5.6 Series, Grok 4.5, Claude mod…

Go 47,389 7,336 Updated Aug 15, 2026

Modeling, training, eval, and inference code for OLMo

Python 6,628 792 Updated Nov 24, 2025

Paper list for Efficient Reasoning.

901 47 Updated May 29, 2026
Next