VibeVoice LogoVibeVoice

Try Online Free • No Sign-In

VibeVoice AI: Turn Text Into 90‑Minute Multi‑Speaker Podcasts

VibeVoice AI is Microsoft's open-source framework for long-form, multi-speaker text-to-speech. Generate hours of natural dialogue, up to four speakers, in English or Chinese, with full local control.

Hear VibeVoice AI in Action

🎭

Context-Aware Expression

Spontaneous Emotion & Singing

Natural emotional responses and spontaneous singing integration

🎙️

Multi-Speaker Podcasts

3-4 Speaker Conversations

Professional podcast-style discussions with background music

🌐

Cross-Lingual Speech

English ↔ Chinese

Seamless language switching within single conversations

Long-Form Conversations

45-90 Minute Sessions

Extended dialogue with consistent speaker identity throughout

🔄

Natural Turn-Taking

Conversational Flow

Realistic pauses, interruptions, and dialogue pacing

🎵

Expressive Prosody

Emotional Nuance

Rich intonation patterns and contextual expression

🔴

Live Audio Examples

Real VibeVoice outputs from the Microsoft Research demo

🎭Spontaneous Emotion & Singing

Spontaneous Argument Scene

Natural emotional escalation between two speakers with realistic dialogue dynamics

00:00.000
00:00.000

📝Transcript

No transcript available

"See You Again" - Spontaneous Singing

Conversation naturally flowing into singing, showcasing VibeVoice's musical capabilities

00:00.000
00:00.000

📝Transcript

No transcript available

🌐Cross-Lingual Examples

Chinese to English Transition

Seamless language switching demonstrating VibeVoice's bilingual capabilities

00:00.000
00:00.000

📝Transcript

No transcript available

All audio examples generated on consumer GPUs, completely unedited.
Timestamps and transcripts extracted from generated audio may contain minor errors.

VibeVoice AI Core Features

Discover what makes VibeVoice AI the most advanced open-source TTS framework for long-form content

Long-Form Conversational Synthesis

Generate up to 90 minutes of continuous audio within a 64K token context. Maintains coherent dialogue flow and natural turn-taking across long spans, perfect for podcasts and audiobooks.
🎭

Multi-Speaker Dialogue Support

Native support for up to 4 distinct speakers in one conversation. Role identifiers ensure consistent timbre and speaker-specific characteristics throughout.
🧠

Next-Token Diffusion Framework

Unified approach where Large Language Models predict hidden states and diffusion head refines them into acoustic features. Improves speech realism and long-form stability.
🗜️

Ultra-Low Frame Rate Tokenizer

Revolutionary 7.5 Hz speech tokenizer compresses audio by up to 3200×. Drastically reduces compute costs while preserving perceptual fidelity.
🔀

Hybrid Audio Representations

Parallel acoustic and semantic tokenizers balance timbre preservation with linguistic meaning. σ-VAE for prosody, ASR objectives for content accuracy.
🎯

Superior Quality Despite Compression

Top scores on PESQ, STOI, and UTMOS benchmarks. Retains natural timbre and intelligibility while outperforming Encodec, DAC, and other baselines.
🏆

Industry-Leading Performance

Higher human evaluation scores in realism and richness than ElevenLabs v3 Alpha and Google Gemini 2.5 Pro. Proven reliability for English and Chinese.
📏

Flexible Length Adaptation

Optimized for long-form but generalizes well to short utterances. Low word error rates and high speaker similarity across all content lengths.
⚖️

Scalable Model Variants

VibeVoice-1.5B for efficiency (90 min), VibeVoice-7B for quality (45 min). Choose the perfect balance between performance and resource usage.
🔓

Open Source & Research-Ready

MIT licensed with pretrained weights on GitHub and Hugging Face. Full local control with extensive documentation and research-oriented design.

How VibeVoice AI Works

The next-token diffusion pipeline that makes VibeVoice AI's 90-minute multi-speaker synthesis possible

TextTokenizerLLMDiffusionDecoderAudio

Next-token diffusion pipeline with ultra-compressed speech tokens

1
📝

Input Preparation

Multi-speaker script processing

Users provide text scripts with speaker role identifiers (Speaker A, Speaker B) and optional voice prompts. The system maintains consistent voice characteristics across long dialogues.

Role-based speaker identification • Voice prompt conditioning • Long-context preparation

2
🔀

Hybrid Representation Encoding

Dual tokenizer architecture

Two parallel tokenizers work together: Acoustic tokenizer compresses audio to 7.5 Hz (capturing timbre and prosody), while semantic tokenizer learns content-level features aligned with text meaning.

σ-VAE acoustic tokenizer • ASR-trained semantic tokenizer • 3200× compression ratio

3
🧠

Context Modeling with LLM

Qwen2.5 language understanding

Large Language Model (1.5B or 7B parameters) interprets conversational flow, predicts hidden states, and ensures coherence across up to 90 minutes of speech generation.

64K token context window • Multi-speaker conversation modeling • Long-form narrative coherence

4
🌊

Token-Level Diffusion Refinement

Next-token diffusion process

For each sequence step, diffusion head iteratively refines Gaussian noise into clean acoustic features using classifier-free guidance and fast samplers like DPM-Solver++.

Classifier-free guidance • DPM-Solver++ sampling • Token-by-token refinement

5
🎵

Decoding to Waveform

High-quality audio reconstruction

Predicted acoustic features pass through the decoder, reconstructing natural-sounding waveform audio with smooth multi-speaker dialogue flow without stitched-together artifacts.

Neural vocoder reconstruction • Continuous waveform generation • Multi-speaker voice consistency

6
🎧

Quality Assurance

Benchmark-beating output

Final audio achieves top scores on PESQ, STOI, and UTMOS metrics while maintaining natural prosody, speaker consistency, and dialogue flow across the entire generation.

PESQ/STOI/UTMOS optimization • Speaker similarity preservation • Prosody naturalness

💡

VibeVoice Key Innovation: Next-Token Diffusion

Unlike traditional TTS that separates text and audio modeling, VibeVoice uses a unified next-token diffusion approach. This allows it to scale to long conversations, multiple speakers, and bilingual outputs while remaining efficient thanks to ultra-compressed tokens (7.5 Hz vs typical 40-50 Hz).

VibeVoice AI Real-World Applications

Discover how VibeVoice AI transforms content creation across industries

🎙️

Podcast Prototyping

Content Creation

Creators can quickly turn written scripts into 90-minute multi-speaker podcast drafts without booking studios or hiring voice actors. Perfect for experimenting with episode formats, dialogue pacing, and guest interactions before final production.

Key Benefits

Rapid prototypingCost-effective testingFormat experimentation

Draft episodes • Test dialogue flow • Experiment with speaker dynamics

📚

Audiobook Narration

Publishing

Authors and publishers can generate multi-character audiobook recordings with up to four distinct voices per story. Unlike traditional single-speaker narration, each character gets their own consistent voice throughout the entire book.

Key Benefits

Multi-character voicesConsistent narrationCost reduction

Character dialogue • Story narration • Voice consistency across chapters

🎓

Educational Content & Training

Education

Teachers and course designers transform text lessons into engaging spoken dialogues between professors and students. Makes e-learning materials more dynamic and accessible, especially beneficial for auditory learners.

Key Benefits

Interactive learningAuditory accessibilityEngaging content

Lecture dialogues • Q&A sessions • Interactive learning scenarios

🌏

Language Learning & Bilingual Content

Education

With native English and Chinese support, create bilingual dialogues for language practice, listening comprehension, and immersive learning. Generate roleplay conversations between teachers and students directly from text scripts.

Key Benefits

Bilingual supportImmersive practiceNatural pronunciation

Language practice • Listening comprehension • Cultural dialogue exchange

🎮

Game Development & Interactive Stories

Entertainment

Game designers prototype in-game dialogue between multiple characters during early narrative design. Test pacing, tone, and emotional delivery without professional voice actors, accelerating the creative process.

Key Benefits

Rapid prototypingNarrative testingCharacter development

Character dialogue • Narrative pacing • Emotional tone testing

Accessibility & Assistive Technology

Accessibility

Convert long documents, articles, or reports into natural, conversational audio. Makes content easier to consume for visually impaired users or anyone who prefers listening over reading lengthy materials.

Key Benefits

Visual accessibilityAudio preferenceContent consumption

Document conversion • Article narration • Report accessibility

⚠️

VibeVoice Research-First, Responsible Usage

VibeVoice is designed primarily for research and creative experimentation. While technically MIT licensed, we recommend limiting use to non-commercial prototyping, academic research, and clearly disclosed AI-generated content.

Research & prototyping
Creative experimentation
Always disclose AI generation

VibeVoice AI Limitations: Know the Boundaries

Understanding VibeVoice AI's current limitations helps set appropriate expectations

🌍

Language Scope

medium impact

VibeVoice is currently optimized for English and Chinese. While it may produce outputs in other languages, results can be unstable or unintelligible. Cross-lingual generalization shows promise but remains experimental.

Impact

Non-English/Chinese applications should be treated as experimental

EN/ZH training focus • Limited multilingual data • Experimental cross-lingual transfer

🗣️

Overlapping Speech

high impact

The system does not support simultaneous speakers. All conversations assume turn-taking, which limits realism in debates, panel discussions, or natural group conversations where interruptions are common.

Impact

Cannot model realistic multi-party conversations with interruptions

Sequential generation only • No simultaneous speech modeling • Turn-based architecture

🎵

Background Sound & Non-Speech Audio

medium impact

Designed strictly for speech synthesis without background noise, music, or sound effects. Occasional artifacts like faint 'background music' may appear due to training data but are uncontrollable noise, not features.

Impact

Cannot add atmosphere or soundscape to generated audio

Speech-only training • Uncontrolled background artifacts • No environmental audio modeling

Computational Cost

high impact

Despite ultra-low-rate tokenizer efficiency, long-form synthesis remains computationally heavy. Generating tens of minutes requires powerful GPUs (7–24GB VRAM) and significant runtime, making real-time deployment impractical.

Impact

High-end hardware required for practical use

7-24GB VRAM requirement • Non-real-time generation • GPU-intensive processing

⚠️

Risk of Misuse

critical impact

Like all high-quality voice synthesis, VibeVoice could enable deepfakes, impersonation, fraud, or disinformation. The research team explicitly warns against deceptive or harmful applications and requires disclosure of AI-generated content.

Impact

Potential for harmful misuse requires responsible deployment

High-fidelity synthesis • Identity spoofing risk • Requires ethical safeguards

🔬

Research-Only Recommendation

critical impact

Official guidance emphasizes research and development use only, not direct commercial deployment. Further testing, safeguards, and development are needed before production applications like voice assistants or commercial audiobooks.

Impact

Not ready for commercial production without additional safeguards

Research-stage technology • Requires safety testing • Needs production hardening

🔧VibeVoice Additional Technical Considerations

Male voices may sound more robotic than female voices due to training data distribution
Singing and musicality are weak as they weren't explicitly trained
Long outputs require extended generation time; streaming capabilities in development
Prosody and emotional control are limited compared to specialized single-speaker models
🚨

VibeVoice Responsible Use Disclaimer

This technology is intended for research purposes only. Users must disclose AI-generated content, avoid deceptive applications, and implement appropriate safeguards. Commercial deployment requires additional testing, safety measures, and ethical review. The creators are not responsible for misuse.

Run VibeVoice AI Yourself

VibeVoice AI hardware guidance and quick commands

VibeVoice Hardware
  • 1.5B → ~7–10GB VRAM (≈90 min)
  • Large → ~18–24GB VRAM (≈45 min, more natural voices)
VibeVoice Setup
docker run --gpus all --rm -it nvcr.io/nvidia/pytorch:24.07-py3
git clone https://github.com/microsoft/VibeVoice.git
cd VibeVoice && pip install -e .
python demo/gradio_demo.py --model_path microsoft/VibeVoice-1.5B --share

Models: VibeVoice-1.5B, VibeVoice-Large (Hugging Face)

VibeVoice Roadmap
  • Streaming 0.5B low-latency variant
  • Improved multilingual stability
  • Emotion & prosody control
  • VibePod end‑to‑end podcast tool
  • Safety, watermarking, responsible usage guides

VibeVoice AI FAQ: Frequently Asked Questions

Everything you need to know about using VibeVoice AI effectively

🎯
Capabilities

How long can VibeVoice generate speech?

The 1.5B model supports up to 90 minutes of continuous audio, while the 7B model supports about 45 minutes with higher naturalness and richer prosody. Both maintain coherent dialogue throughout the entire generation.
👥
Capabilities

How many speakers can I include in one audio?

VibeVoice natively supports up to four distinct speakers. Each speaker can be assigned a text script and optional voice prompt to maintain consistent timbre and role identity throughout the conversation.
🌍
Languages

Which languages does VibeVoice support?

VibeVoice is primarily trained for English and Chinese, delivering the best quality in these languages. Other languages may produce unstable or unintelligible outputs as cross-lingual capabilities remain experimental.
🎵
Technical

Does VibeVoice generate background music or sound effects?

No. VibeVoice is strictly a speech synthesis system. Occasionally, faint music-like artifacts may appear due to training data, but these are not controllable features and should be treated as unintended noise.
💻
Hardware

Can VibeVoice run on consumer hardware?

Yes, but requirements depend on model size: • 1.5B model: ~7–10GB VRAM (RTX 3060/3070) • 7B model: ~18–24GB VRAM (RTX 3090/4090) Generation speed is slower than commercial services, especially for long audio.
💼
Usage

Can I use VibeVoice for commercial projects?

VibeVoice uses the MIT License, which is technically permissive. However, the research team explicitly recommends research and development use only, due to risks of misuse. Commercial deployment should include strong safeguards and disclosure practices.
🗣️
Technical

Does VibeVoice support overlapping speech?

Not currently. All generated conversations assume speakers take turns in sequence. Overlapping dialogue like interruptions or simultaneous group conversation is not modeled in the current architecture.
⚖️
Comparison

How does VibeVoice compare to ElevenLabs or Google TTS?

VibeVoice is open source, runs locally, and specializes in long-form multi-speaker content. Commercial services offer more consistent quality, real-time speed, and broader language support, but are closed-source and subscription-based.
🎶
Troubleshooting

Why does background music sometimes appear?

The model can reproduce training artifacts from its dataset. Background music is a known but uncontrolled artifact that appears occasionally. This is not a feature and cannot be controlled or eliminated.
🏃
Troubleshooting

Why does Chinese sometimes sound rushed or unstable?

Chinese generation can be improved by using English punctuation for better pacing, or switching to the larger 7B model which has better prosody control and cross-lingual stability.
⏱️
Performance

How long does audio generation take?

On a 12GB consumer GPU, generating a few minutes of audio may take several minutes to process. Very long outputs (30+ minutes) require proportionally longer processing time due to the computational complexity.
🎧
Quality

How can I improve the quality of generated audio?

For best results: use clear, well-punctuated scripts; specify speaker roles consistently; use the 7B model for higher quality; and ensure adequate VRAM. English content generally produces more stable results than other languages.

VibeVoice Quick Reference

Model Comparison

1.5B: 90 min, 7-10GB
7B: 45 min, 18-24GB

Supported Languages

✅ English (native)
✅ Chinese (native)
⚠️ Others (experimental)

Key Features

Up to 4 speakers
Long-form context
Open source (MIT)

Best For

Research & prototyping
Multi-speaker content
Local deployment
Built by Microsoft Research • Open Source (MIT). Generated voices may not always be realistic or safe for impersonation. Always disclose AI‑generated content.