- Hyogo, Japan
- https://kun432.github.io/
- @kun432
- https://zenn.dev/kun432
Highlights
- Pro
Stars
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action
⚖️ HAKARI-Bench is a lightweight IR benchmark that rebuilds retrieval tasks as small Nano-sets, making model selection, regression checks, quantization, truncation, and reranking comparisons practi…
Self-hosted native-Rust runtime for real-time voice agents. Own the stack: one binary in your own VPC or air-gapped, no hosted control plane. pipecat-compatible pipeline, in-process SIP/RTP, single…
A curated solutions to building a self-evolving second brain that helps AI agents understand your personal and team context.
Preference-supervised naturalness scorer for modern neural TTS . best way to measure naturalness
A Repository for Single- and Multi-modal Speaker Verification, Speaker Recognition and Speaker Diarization
A self-hosted household AI chat app built on OpenClaw — multiple conversations, multi-user family life, and long-term memory on your hardware.
Real-time voice assistant — WebRTC streaming, faster-whisper ASR, local LLM, Vui Nano (300M) TTS. OpenAI Realtime API compatible. Voice cloning, barge-in, ~9× realtime on a 4090. Apache 2.0.
LLM speculative inference server for consumer & heterogeneous hardware
The context API to search, scrape, and interact with the web at scale. 🔥
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
An open-source wake word library for creating voice-enabled applications.
REAM: Merging Improves Pruning of Experts in LLMs
Open diffusion language model for code generation — releasing pretraining, evaluation, inference, and checkpoints.
Wyoming server for using Pocket-TTS with Home Assistant or other Wyoming aware services.
Demo of publishing audio from local microphone device and playing audio on local audio output device
Deploying LLMs offline on the NVIDIA Jetson platform marks the dawn of a new era in embodied intelligence, where devices can function independently without continuous internet access.
Soprano: Instant, Ultra-Realistic Text-to-Speech
VyvoTTS: LLM-Based Text-to-Speech Training Framework
Nornicdb is a distributed low-latency, Graph+Vector, Temporal MVCC with all sub-ms HNSW search, graph traversal, and writes. Using Neo4j Bolt/Cypher and qdrant's gRPC means you can switch with no c…
FlashCosyVoice: A lightweight vLLM implementation built from scratch for CosyVoice.
Streamlit Component to quickly create Interactive Flow Diagrams using React Flow
Learn OpenCV : C++ and Python Examples
Automatically detects and fixes syntax errors in Mermaid diagrams within Markdown files