VibeVoice AI: Turn Text Into 90‑Minute Multi‑Speaker Podcasts
VibeVoice AI is Microsoft's open-source framework for long-form, multi-speaker text-to-speech. Generate hours of natural dialogue, up to four speakers, in English or Chinese, with full local control.
Hear VibeVoice AI in Action
Context-Aware Expression
Spontaneous Emotion & Singing
Natural emotional responses and spontaneous singing integration
Multi-Speaker Podcasts
3-4 Speaker Conversations
Professional podcast-style discussions with background music
Cross-Lingual Speech
English ↔ Chinese
Seamless language switching within single conversations
Long-Form Conversations
45-90 Minute Sessions
Extended dialogue with consistent speaker identity throughout
Natural Turn-Taking
Conversational Flow
Realistic pauses, interruptions, and dialogue pacing
Expressive Prosody
Emotional Nuance
Rich intonation patterns and contextual expression
Live Audio Examples
Real VibeVoice outputs from the Microsoft Research demo
🎭Spontaneous Emotion & Singing
Spontaneous Argument Scene
Natural emotional escalation between two speakers with realistic dialogue dynamics
📝Transcript
"See You Again" - Spontaneous Singing
Conversation naturally flowing into singing, showcasing VibeVoice's musical capabilities
📝Transcript
🌐Cross-Lingual Examples
Chinese to English Transition
Seamless language switching demonstrating VibeVoice's bilingual capabilities
📝Transcript
All audio examples generated on consumer GPUs, completely unedited.
Timestamps and transcripts extracted from generated audio may contain minor errors.
VibeVoice AI Core Features
Discover what makes VibeVoice AI the most advanced open-source TTS framework for long-form content
Long-Form Conversational Synthesis
Multi-Speaker Dialogue Support
Next-Token Diffusion Framework
Ultra-Low Frame Rate Tokenizer
Hybrid Audio Representations
Superior Quality Despite Compression
Industry-Leading Performance
Flexible Length Adaptation
Scalable Model Variants
Open Source & Research-Ready
How VibeVoice AI Works
The next-token diffusion pipeline that makes VibeVoice AI's 90-minute multi-speaker synthesis possible
Next-token diffusion pipeline with ultra-compressed speech tokens
Input Preparation
Multi-speaker script processing
Users provide text scripts with speaker role identifiers (Speaker A, Speaker B) and optional voice prompts. The system maintains consistent voice characteristics across long dialogues.
Role-based speaker identification • Voice prompt conditioning • Long-context preparation
Hybrid Representation Encoding
Dual tokenizer architecture
Two parallel tokenizers work together: Acoustic tokenizer compresses audio to 7.5 Hz (capturing timbre and prosody), while semantic tokenizer learns content-level features aligned with text meaning.
σ-VAE acoustic tokenizer • ASR-trained semantic tokenizer • 3200× compression ratio
Context Modeling with LLM
Qwen2.5 language understanding
Large Language Model (1.5B or 7B parameters) interprets conversational flow, predicts hidden states, and ensures coherence across up to 90 minutes of speech generation.
64K token context window • Multi-speaker conversation modeling • Long-form narrative coherence
Token-Level Diffusion Refinement
Next-token diffusion process
For each sequence step, diffusion head iteratively refines Gaussian noise into clean acoustic features using classifier-free guidance and fast samplers like DPM-Solver++.
Classifier-free guidance • DPM-Solver++ sampling • Token-by-token refinement
Decoding to Waveform
High-quality audio reconstruction
Predicted acoustic features pass through the decoder, reconstructing natural-sounding waveform audio with smooth multi-speaker dialogue flow without stitched-together artifacts.
Neural vocoder reconstruction • Continuous waveform generation • Multi-speaker voice consistency
Quality Assurance
Benchmark-beating output
Final audio achieves top scores on PESQ, STOI, and UTMOS metrics while maintaining natural prosody, speaker consistency, and dialogue flow across the entire generation.
PESQ/STOI/UTMOS optimization • Speaker similarity preservation • Prosody naturalness
VibeVoice Key Innovation: Next-Token Diffusion
Unlike traditional TTS that separates text and audio modeling, VibeVoice uses a unified next-token diffusion approach. This allows it to scale to long conversations, multiple speakers, and bilingual outputs while remaining efficient thanks to ultra-compressed tokens (7.5 Hz vs typical 40-50 Hz).
VibeVoice AI Real-World Applications
Discover how VibeVoice AI transforms content creation across industries
Podcast Prototyping
Creators can quickly turn written scripts into 90-minute multi-speaker podcast drafts without booking studios or hiring voice actors. Perfect for experimenting with episode formats, dialogue pacing, and guest interactions before final production.
Key Benefits
Draft episodes • Test dialogue flow • Experiment with speaker dynamics
Audiobook Narration
Authors and publishers can generate multi-character audiobook recordings with up to four distinct voices per story. Unlike traditional single-speaker narration, each character gets their own consistent voice throughout the entire book.
Key Benefits
Character dialogue • Story narration • Voice consistency across chapters
Educational Content & Training
Teachers and course designers transform text lessons into engaging spoken dialogues between professors and students. Makes e-learning materials more dynamic and accessible, especially beneficial for auditory learners.
Key Benefits
Lecture dialogues • Q&A sessions • Interactive learning scenarios
Language Learning & Bilingual Content
With native English and Chinese support, create bilingual dialogues for language practice, listening comprehension, and immersive learning. Generate roleplay conversations between teachers and students directly from text scripts.
Key Benefits
Language practice • Listening comprehension • Cultural dialogue exchange
Game Development & Interactive Stories
Game designers prototype in-game dialogue between multiple characters during early narrative design. Test pacing, tone, and emotional delivery without professional voice actors, accelerating the creative process.
Key Benefits
Character dialogue • Narrative pacing • Emotional tone testing
Accessibility & Assistive Technology
Convert long documents, articles, or reports into natural, conversational audio. Makes content easier to consume for visually impaired users or anyone who prefers listening over reading lengthy materials.
Key Benefits
Document conversion • Article narration • Report accessibility
VibeVoice Research-First, Responsible Usage
VibeVoice is designed primarily for research and creative experimentation. While technically MIT licensed, we recommend limiting use to non-commercial prototyping, academic research, and clearly disclosed AI-generated content.
VibeVoice AI Limitations: Know the Boundaries
Understanding VibeVoice AI's current limitations helps set appropriate expectations
Language Scope
medium impactVibeVoice is currently optimized for English and Chinese. While it may produce outputs in other languages, results can be unstable or unintelligible. Cross-lingual generalization shows promise but remains experimental.
Impact
Non-English/Chinese applications should be treated as experimental
EN/ZH training focus • Limited multilingual data • Experimental cross-lingual transfer
Overlapping Speech
high impactThe system does not support simultaneous speakers. All conversations assume turn-taking, which limits realism in debates, panel discussions, or natural group conversations where interruptions are common.
Impact
Cannot model realistic multi-party conversations with interruptions
Sequential generation only • No simultaneous speech modeling • Turn-based architecture
Background Sound & Non-Speech Audio
medium impactDesigned strictly for speech synthesis without background noise, music, or sound effects. Occasional artifacts like faint 'background music' may appear due to training data but are uncontrollable noise, not features.
Impact
Cannot add atmosphere or soundscape to generated audio
Speech-only training • Uncontrolled background artifacts • No environmental audio modeling
Computational Cost
high impactDespite ultra-low-rate tokenizer efficiency, long-form synthesis remains computationally heavy. Generating tens of minutes requires powerful GPUs (7–24GB VRAM) and significant runtime, making real-time deployment impractical.
Impact
High-end hardware required for practical use
7-24GB VRAM requirement • Non-real-time generation • GPU-intensive processing
Risk of Misuse
critical impactLike all high-quality voice synthesis, VibeVoice could enable deepfakes, impersonation, fraud, or disinformation. The research team explicitly warns against deceptive or harmful applications and requires disclosure of AI-generated content.
Impact
Potential for harmful misuse requires responsible deployment
High-fidelity synthesis • Identity spoofing risk • Requires ethical safeguards
Research-Only Recommendation
critical impactOfficial guidance emphasizes research and development use only, not direct commercial deployment. Further testing, safeguards, and development are needed before production applications like voice assistants or commercial audiobooks.
Impact
Not ready for commercial production without additional safeguards
Research-stage technology • Requires safety testing • Needs production hardening
🔧VibeVoice Additional Technical Considerations
VibeVoice Responsible Use Disclaimer
This technology is intended for research purposes only. Users must disclose AI-generated content, avoid deceptive applications, and implement appropriate safeguards. Commercial deployment requires additional testing, safety measures, and ethical review. The creators are not responsible for misuse.
Run VibeVoice AI Yourself
VibeVoice AI hardware guidance and quick commands
- 1.5B → ~7–10GB VRAM (≈90 min)
- Large → ~18–24GB VRAM (≈45 min, more natural voices)
docker run --gpus all --rm -it nvcr.io/nvidia/pytorch:24.07-py3 git clone https://github.com/microsoft/VibeVoice.git cd VibeVoice && pip install -e . python demo/gradio_demo.py --model_path microsoft/VibeVoice-1.5B --share
Models: VibeVoice-1.5B, VibeVoice-Large (Hugging Face)
- Streaming 0.5B low-latency variant
- Improved multilingual stability
- Emotion & prosody control
- VibePod end‑to‑end podcast tool
- Safety, watermarking, responsible usage guides
VibeVoice AI FAQ: Frequently Asked Questions
Everything you need to know about using VibeVoice AI effectively