-
Tokyo Metropolitan University
- Tokyo
-
19:27
(UTC +09:00) - https://portfolio.ayutaso.com
- @aya172957
Highlights
- Pro
Lists (1)
Sort Name ascending (A-Z)
Stars
Benchmarking STT service TTFB and semantic WER for real-time AI applications
Coco-Nut (Corpus of connecting NIHONGO utterance and text) corpus
Post-training with Tinker
Powerful system-level package manager for Linux, macOS and Windows written in Rust – building on top of the Conda ecosystem.
A fast, cross-platform build tool inspired by Make, designed for modern workflows.
AIST Toolkit for Accelerating Machine Learning Research
agent multiplexer that lives in your terminal.
A toolkit for speaker diarization.
MoshiRAG is a compact full-duplex speech language model augmented with asynchronous knowledge retrieval to improve factuality without sacrificing real-time interactivity.
FlashCosyVoice: A lightweight vLLM implementation built from scratch for CosyVoice.
The agent that grows with you
Large-scale, Informative, and Diverse Multi-round Chat Data (and Models)
Reference implementation of an end-to-end voice agent built using the NVIDIA Nemotron models
[EMNLP 2025 Findings] Code for "Distilling Many-Shot In-Context Learning into a Cheat Sheet"
List of open-source TTS, voice cloning, and music generation models
Erasing concepts from neural representations with provable guarantees
A package for NeuCodec: a 50hz, 0.8kbps, 24kHz audio codec.
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenario…
Whisper-Flow is a framework designed to enable real-time transcription of audio content using OpenAI’s Whisper model. Rather than processing entire files after upload (“batch mode”), Whisper-Flow a…