15 unstable releases (3 breaking)
| 0.4.2 | Mar 22, 2026 |
|---|---|
| 0.4.0 | Mar 21, 2026 |
| 0.3.3 | Mar 20, 2026 |
| 0.2.3 | Mar 19, 2026 |
| 0.1.4 | Nov 9, 2016 |
#107 in Audio
7MB
9K
SLoC
voice
Like say, but with Kokoro TTS and Moonshine STT. A command-line speech tool for macOS, powered by MLX on Apple Silicon.
Install
Pre-built binary (recommended)
Install with cargo-binstall to get a pre-built binary — no compilation required:
# Install cargo-binstall if you don't have it
cargo install cargo-binstall
# Install voice
cargo binstall voice
Build from source
Requires Git LFS for embedded voice/model data:
# Install git-lfs if you don't have it
brew install git-lfs
git lfs install
# Clone and build
git clone https://github.com/rgbkrk/voice.git
cd voice
cargo install --path crates/voice-cli
Note:
cargo install voicecompiles from source on crates.io, but the Metal shader library path can break — see the main README for details. Usecargo binstall voiceor build from a local clone instead.
Usage
# Just talk (backward compatible — no subcommand needed)
voice Hello world
# Text-to-speech with the say subcommand
voice say -v am_michael "How are you today?"
voice say -f script.txt -o output.wav
echo "Hello" | voice say
voice say --markdown -f post.mdx
# Speech-to-text from microphone
voice listen
voice listen --continuous
# Transcribe an audio file
voice transcribe recording.wav
# JSON-RPC 2.0 server on stdin/stdout
voice serve -v am_michael
Options
Top-level
Usage: voice [OPTIONS] [COMMAND] [TEXT]...
Commands:
say Speak text aloud (default when no subcommand given)
listen Record from microphone and transcribe (speech-to-text)
transcribe Transcribe a WAV audio file
serve Run as a JSON-RPC 2.0 server on stdin/stdout
Arguments:
[TEXT]... Text to speak (shorthand for `voice say <text>`)
Options:
-q, --quiet Suppress progress output
-h, --help Print help
voice say
Usage: voice say [OPTIONS] [TEXT]...
Options:
-f, --input-file <FILE> Read text from a file (use - for stdin)
--phonemes <IPA> Raw phoneme string (IPA)
-v, --voice <VOICE> Voice name [default: af_heart]
-o, --output <PATH> Write WAV to file instead of playing
-s, --speed <SPEED> Speech speed factor [default: 1.0]
--markdown Strip markdown/MDX formatting before speaking
--sub <WORD=REPLACEMENT> Word substitution (repeatable)
--sub-file <PATH> Load substitutions from a file
voice listen
Usage: voice listen [OPTIONS]
Options:
--continuous Record and transcribe segments continuously
voice transcribe
Usage: voice transcribe <FILE>
Voices
American: af_heart, af_bella, af_nicole, af_sarah, af_sky, am_adam, am_michael
British: bf_emma, bf_isabella, bm_george, bm_lewis
See the full list in the main README.
License
MIT
Dependencies
~41–79MB
~1.5M SLoC