Skip to content
View ayutaz's full-sized avatar

Block or report ayutaz

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Implementation of RIFT-SVC, a singing voice conversion model based on Rectified Flow Transformer.

Python 71 18 Updated Nov 10, 2025

We propose LeapTalk, a novel framework that achieves stable and real-time talking-head generation with a single forward step, scaling to arbitrarily long videos.

Python 162 15 Updated Aug 11, 2026

A Robust Low-Latency Streaming Zero-Shot Voice Conversion

Python 126 22 Updated Aug 10, 2026
Python 17 1 Updated Jul 11, 2026

Japanese Braille translator originally developed for NVDAJP

Python 7 1 Updated Jul 18, 2026

A shader combining PBR and NPR

HLSL 193 5 Updated Jul 14, 2026

The ConvFill Repository provides training and inference code for the Conversational Infill task proposed and implemented in Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive…

Python 19 2 Updated Jun 26, 2026

[NAACL 2025] WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching

Python 133 12 Updated Apr 8, 2026

Haqumei (薄明/拍命) is a Japanese Grapheme-to-Phoneme (G2P) library.

Rust 7 Updated Aug 8, 2026

Training code and dataset cleasing with Sidon

Jupyter Notebook 180 18 Updated Apr 24, 2026

ITAコーパス emotion 50名分の読み上げ音声データ ( CC BY 4.0 )

HTML 14 Updated May 28, 2026

[Under Review] A Vibrato Controlling Method by Predicting High-frequency F0 contour for Singing Voice Conversion

Python 16 5 Updated Jun 15, 2026

Causal zero-shot voice conversion at 44.1kHz.

Python 29 2 Updated Jul 3, 2026
Python 1,160 121 Updated Aug 12, 2026

Miso TTS is an 8 billion, highly emotive text-to-speech model

Python 3,209 328 Updated Jun 9, 2026

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling

Python 210 6 Updated Jun 6, 2026

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

122 Updated May 20, 2026

First foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every other model falls apart. **You'll come …

Python 1,127 72 Updated Jun 2, 2026
Python 416 47 Updated Jul 23, 2026

Even Realities Hub Simulator - multi application test environment

TypeScript 83 15 Updated Mar 23, 2026

Teaching Speech Enhancement Models to Sing: Domain Adaptation from Speech Enhancement to Singing Voice Separation

Python 9 1 Updated Jul 14, 2026

🎙️ 「大模型」从0训练0.1B能听能说能看的全模态Omni模型!A 0.1B Omni model trained from scratch, capable of listening, speaking, and seeing!

Python 2,325 273 Updated Aug 6, 2026

talkie is a vintage language model from 1930

Python 977 60 Updated May 19, 2026
Jupyter Notebook 8 Updated Jun 12, 2026

X-VC: Zero-shot Streaming Voice Conversion in Codec Space

Python 74 8 Updated May 6, 2026

Open-source AI sandbox infrastructure with unified API for VMMs -- Firecracker, QEMU and libkrun.

Python 767 59 Updated Aug 4, 2026

[ACL 2026] VoxMind: An End-to-End Agentic Spoken Dialogue System

Python 40 4 Updated Jul 8, 2026

PersonaPlex code.

Python 10,363 1,443 Updated Mar 2, 2026

[MLsys2026]: RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate, and 100% private RAG application on your personal device.

Python 12,785 1,144 Updated Jul 31, 2026

浏览器端虚拟歌姬工作站 — 导入 MIDI,填写歌词,驱动 UTAU 声库演唱,支持 SeedVC 音色转换与多轨伴奏混音。

C# 51 4 Updated May 24, 2026
Next