Production-ready speech recognition for your product, within minutes.

Customized speech engines that outperform off-the-shelf models on real-life data, deployed where you need them, in just a few clicks.

01Why Blynt

General-purpose speech recognition breaks in your domain.

1.1The reality check

It's the real world out there. Not your lab.

You picked the best off-the-shelf speech model for your product.
On your real-world data, your acoustics, your users, your jargon, it breaks.

Medical Scribe

on-the-go clinical dictation

Challenging AcousticsFar-field, noisy settings, mic quality
Real-life AudiencesChildren, elderly, code-switching
Custom JargonMedical entities, 2K+ values
>25%WER
Interactive Story Box

educative voice agent for children

Challenging AcousticsFar-field, noisy settings
Real-life AudiencesChildren
Custom JargonThe adventure universe
>35%WER
Smart Glasses Voice Assistant

voice control for music playback

Challenging AcousticsOn-the-go, wind noise, competing speech
Real-life AudiencesNon-native speakers, code-switching
Custom JargonMusic catalog, millions of entities
>20%WER

Word error rate with frontier, enterprise models

1.2Build or buy?

Considering building a custom speech engine in house ?

The hard truth: building a speech team in house is a heavy long term-commitment and reaching production readiness takes at least $3M and 12 months.

[ 1. Build the team ]

Speech expertise is scarce and costly.

lead time·3 months
cost·team of 3 · $1M / year
[ 2. Fine-tune a base model ]

Navigate the open-weight model landscape, collect & annotate data, build pipelines, fix failure modes.

lead time·6 months
cost·$1M in data & training infra
[ 3. Deploy & scale ]

Scale real-time GPU workloads, and build customized models for other geos and products.

lead time·3 months
cost·$1M in devops & serving infra
02How it works

Your custom speech engine in minutes, not months.

Years of research and engineering, turned into a 4 step process.

STEP 1 · BASE MODEL

Not all speech models were born equal.

Pick the base model that best matches your use case.

Use case
Streaming
Languages0 of 34
en-USfr-FRde-DEes-ESit-IT+ 29
STEP 2 · ACOUSTIC ADAPTATION

An engine tuned to your acoustic domain in just a few clicks.

Compose your model from our library of pre-trained adapters.

Product category
Distance
Environment
Audience
custom-speech-engine-pediatrician
step 2 of 4
Base model
Qwen3-ASRscribeliveen-USfr-FR
STEP 3 · DOMAIN JARGON

Use our built-in entities, or import your own jargon.

Drop in a CSV of your domain terms and toggle on the built-in entities you need.

Built-in entities
weightheighttemperatureemailphone numberPIN codeamount of money
Custom entities
drop .csv files to register custom entities
parsed entities
medications1,250
conditions250
procedures127
custom-speech-engine-pediatrician
step 3 of 4
Base model
Qwen3-ASRscribeliveen-USfr-FR
Acoustic adapters
mobile appfar-fieldofficeadultschildren
STEP 4 · RUN-TIME CONTEXT

Steer transcription with text, in real time.

Inject text at any point - session start or mid-conversation - to lock in the words that matter most.

POST/v1/transcribe
{  "type": "update_session_context",  "session_context": [    { "name": "pediatrician", "value": "Dr. Amanda Klein" },    { "name": "patient", "value": "Léa Martin, 6 y.o." },    { "name": "conditions", "value": ["asthma", "otitis media"] }  ]}
custom-speech-engine-pediatrician
step 4 of 4
Base model
Qwen3-ASRscribeliveen-USfr-FR
Acoustic adapters
mobile appfar-fieldofficeadultschildren
Domain entities
builtin · 3
weightheighttemperature
custom · 3
medications 1,250conditions 250procedures 127
03Evaluate & Deploy

Evaluate on data that matters,
then ship it.

Run your engine against built-in sets or your own audio, and see WER and latency side by side with closed-API alternatives. When you're happy, deploy on your terms.

EVAL SET
10%20%30%40%120ms180ms240ms300mslatency p95 (ms) →↑ WER (worse)123Qwen3-ASR · + run-time context175ms · 5.0% WER · your pick
Latest configurationClosed-API baselinesCustomization layers: base → + adapters → + entities → + runtime
custom-speech-engine
Base model
Qwen3-ASRscribeen-US · fr-FR
Acoustic adapters
3adapters
Domain entities
6entities
5.0%WER
175msp95 latency
Deployment options
fastest path
Hosted API
Per-minute pricing.
your VPC
Dedicated deployment
AWS, GCP, or Azure. Your data stays in your tenant.
on-device
Embedded runtime
Efficient models, optimized for on-device inference.
Latency & scale
<300ms
end-to-end p95
1,000+
concurrent sessions
Plug in via
livekitLiveKit Agents
pipecatPipecat
blyntBlynt SDK
REST · gRPC · WebSocket
Ready to see what Blynt can do for your product?Let's talk
04Mission

Voice that works for everyone, everywhere.

Every product and every user is different.
Building Sonos Voice Control, our team delivered speech recognition customized down to the individual user and the individual turn. Our mission is to bring that level of control to your product, across all three levels of speech understanding.

SemanticConversationalEnvironmental
Semantic UnderstandingL1

Real-time transcription that gets it right when it matters

ASRContextual BiasingDomain Adaptation
Conversational IntelligenceL2

Real-time conversational dynamics: who said what, when.

Turn-takingDiarization
Environmental AwarenessL3

What happens beyond what is being said by whom.

Acoustic scene classificationEmotion classification
▲ WE'RE HIRING · ANNECY, PARIS, REMOTE

Join us across ML Research, Infrastructure, ML Ops & Dev Rel.

Build your custom speech engine in hours, not months.

20%absolute WER gains from customization
<300msp95 latency in production, end-to-end.
$3Msavings w.r.t to building in house