- Japan
-
05:57
(UTC +09:00) - https://yousan.notion.site/
- @ayousanz
- https://ayousanz.hatenadiary.jp/archive
- https://zenn.dev/ayousanz
Lists (5)
Sort Name ascending (A-Z)
Stars
Implementation of RIFT-SVC, a singing voice conversion model based on Rectified Flow Transformer.
We propose LeapTalk, a novel framework that achieves stable and real-time talking-head generation with a single forward step, scaling to arbitrarily long videos.
A Robust Low-Latency Streaming Zero-Shot Voice Conversion
Japanese Braille translator originally developed for NVDAJP
The ConvFill Repository provides training and inference code for the Conversational Infill task proposed and implemented in Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive…
[NAACL 2025] WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching
Haqumei (薄明/拍命) is a Japanese Grapheme-to-Phoneme (G2P) library.
Training code and dataset cleasing with Sidon
ITAコーパス emotion 50名分の読み上げ音声データ ( CC BY 4.0 )
[Under Review] A Vibrato Controlling Method by Predicting High-frequency F0 contour for Singing Voice Conversion
Causal zero-shot voice conversion at 44.1kHz.
Miso TTS is an 8 billion, highly emotive text-to-speech model
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action
First foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every other model falls apart. **You'll come …
Even Realities Hub Simulator - multi application test environment
Teaching Speech Enhancement Models to Sing: Domain Adaptation from Speech Enhancement to Singing Voice Separation
🎙️ 「大模型」从0训练0.1B能听能说能看的全模态Omni模型!A 0.1B Omni model trained from scratch, capable of listening, speaking, and seeing!
X-VC: Zero-shot Streaming Voice Conversion in Codec Space
Open-source AI sandbox infrastructure with unified API for VMMs -- Firecracker, QEMU and libkrun.
[ACL 2026] VoxMind: An End-to-End Agentic Spoken Dialogue System
[MLsys2026]: RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate, and 100% private RAG application on your personal device.
浏览器端虚拟歌姬工作站 — 导入 MIDI,填写歌词,驱动 UTAU 声库演唱,支持 SeedVC 音色转换与多轨伴奏混音。