-
National Taiwan University
- Taipei, Taiwan
-
14:37
(UTC +08:00) - xjchen.tech
- @xjchen_ntu
- in/jun-ntu
Lists (1)
Sort Name ascending (A-Z)
Starred repositories
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Train transformer language models with reinforcement learning.
NeurIPS 2025 Spotlight; ICLR2024 Spotlight; CVPR 2024; EMNLP 2024
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
Covo-Audio is a 7B-parameter end-to-end large audio language model that directly processes continuous audio inputs and generates audio outputs within a single unified architecture.
蒸餾李宏毅老師的skill,結合Karpathy的LLM,Fable 5加持 以及 本人親自訪談
Expert code review skill: SOLID, security, performance, error handling, boundary conditions
Plug-and-play streaming semantic VAD for real-time full-duplex spoken dialogue systems.
Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
High-Quality Voice Cloning TTS for 600+ Languages
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
LLM-Codec: Neural Audio Codec Meets Language Model Objectives
Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
Liquid Audio - Speech-to-Speech audio models by Liquid AI
Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding
hyc2026 / sft-qwen2.5-omni-thinker
Forked from verl-project/verlverl: Volcano Engine Reinforcement Learning for LLMs
Generative World Renderer: an AI-native Renderer for Games and Virtual Worlds.
Reverse Engineering of Supervised Semantic Speech Tokenizer (S3Tokenizer) proposed in CosyVoice
A survey of spoken dialogue models (SDMs) with speech input and speech output. Focus on their Intermediate Representation and Generation Pattern
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
Align Anything: Training All-modality Model with Feedback
open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.
Simple text to phones converter for multiple languages
Wrap Antigravity, ChatGPT Codex, Claude Code, Grok Build as an OpenAI/Gemini/Claude/Codex compatible API service, allowing you to enjoy the free Gemini 3.1 Pro, GPT 5.6 Series, Grok 4.5, Claude mod…
Modeling, training, eval, and inference code for OLMo
Paper list for Efficient Reasoning.