VoiceRun STT Router
VoiceRun STT Router is one realtime WebSocket interface for VoiceRun STT Model and external speech-to-text models. Select a primary model and optional ordered fallback models, switch models mid-session, and choose how turns end; audio input and transcription events stay the same.
wss://api.voicerun.com/v1/stt
How It Works#
- Connect — open the WebSocket with your VoiceRun API key.
- Configure — send
session.updatewith a primarymodeland optionalfallback_models. The Router answerssession.updated, thensession.readyonce the provider is connected. - Stream — send base64 audio in
input_audio_buffer.appendmessages. You can start right aftersession.update. - Receive — replace the current hypothesis on
transcription.delta; consume the completed turn fromturn.ended.turn_detectiondecides what ends a turn; see Turn detection. - Close — either close the WebSocket directly, or send
session.closefor a graceful shutdown and wait forsession.closed. The protocol message is optional.
The gateway normalizes provider protocols. Provider-specific messages never escape the socket.
session.close does not commit or flush unfinished audio. Wait for the final turn.ended before
closing a live session. When streaming a finite recording, append trailing silence so provider
endpointing can finish the last turn; see Messages and
Examples.
Authentication and billing#
One VoiceRun API key works for every enabled model:
- Sign in and open your VoiceRun profile.
- Under Personal Access Tokens, click Add PAT and copy the generated token.
- Use that PAT as
VOICERUN_API_KEY; keep it server-side and out of source control.
ws = await websockets.connect( "wss://api.voicerun.com/v1/stt", additional_headers={"Authorization": "Bearer YOUR_VOICERUN_API_KEY"}, )
Browser clients that cannot set an Authorization header may use ?token=YOUR_VOICERUN_API_KEY. Avoid putting keys in URLs outside browser-only integrations.
{ "type": "session.update", "model": "voicerun-asr-realtime-v1", "fallback_models": ["nova-3", "gpt-4o-transcribe"], "language": "en" }
Usage is charged to your VoiceRun credit balance at the published rate for the selected model. VoiceRun STT Model is free during its preview; external models continue to use their published rates. VoiceRun manages provider credentials, and provider credential fields are rejected by the public API.
Automatic model fallback#
fallback_models is an ordered list of up to five model names. If the active provider fails at runtime, the Router settles that model's accepted audio, starts the next available model at its published rate, and emits session.model_changed.
The Router retains up to eight seconds of input audio and replays the audio since the last turn boundary into the fallback model, so an in-progress turn can continue without repeating finished ones. Replayed audio is charged to the fallback model because that model processes it. Each configured model is attempted at most once.
Fallback models inherit the session's shared configuration. If the new model does not support a setting used by the primary model, its adapter ignores that setting; the incompatibility does not prevent failover. Choose models with compatible behavior when a particular setting is essential to the conversation.
Fallback is not attempted for invalid configuration, an invalid VoiceRun API key, insufficient credits, or a concurrency limit. The fallback chain cannot be changed after the session becomes active.
Switching models mid-session#
Send session.update with a different model on an active session. By default the switch waits for the current turn to end; apply: "now" switches at once. The Router starts the new model beside the current one and moves over only once it has connected, so a model that cannot connect never interrupts the session. The new model is billed at its own published rate from the moment of the switch. If it later fails, the Router returns to the previous model before trying fallback_models. See Switching models mid-session.
Turn detection#
turn_detection decides when a turn ends, with every model:
turn_detection | The turn ends when |
|---|---|
provider (default) | the model's own endpointing produces a final |
client | your client sends input_audio_buffer.commit |
silero | the Router hears a short silence after speech |
smart_turn | the Smart Turn model judges the pause complete, waiting longer when it sounds unfinished |
semantic | a semantic model judges the transcript a complete reply to the assistant's last utterance |
silero, smart_turn, and semantic run in the Router on its own copy of the audio, so they work with any model, sample rate, and audio format, and add speech.started / speech.stopped events. See Turn detection for the options and events.
Supported Models#
| Provider | Models |
|---|---|
| VoiceRun | voicerun-asr-realtime-v1 |
| Deepgram | nova-3, flux-general-en, flux-general-multi |
| OpenAI | gpt-4o-transcribe, gpt-4o-mini-transcribe |
| ElevenLabs | scribe_v2_realtime |
| Cartesia | ink-whisper |
| Soniox | stt-rt-v4 |
chirp_3 | |
| Qwen | qwen3-asr-flash, qwen3-asr-flash-realtime |
| xAI | grok-stt |
| Inworld | inworld/inworld-stt-1 |
| Gradium | gradium-default-stt |
| Tencent | tencent-16k |
Undated aliases may resolve to a tested provider snapshot while retaining the same public model selector.
This list covers the standalone STT Router only. VoiceRun voice agents select STT models from a separate catalog — see Speech to Text for the models available under Deployment.spec.stt.
Audio Format#
Input defaults to raw PCM16 little-endian, mono, at 16 kHz. Set input_audio_format to pcm16 or mulaw and sample_rate to a value from 8,000 through 48,000 Hz; VoiceRun converts it to the rate required by the selected engine.
Chunks around 20–100 ms work well for live audio. For a finite recording, append trailing silence so the final turn can end, or send input_audio_buffer.commit under turn_detection: "client".
Unified Events#
All models use the VoiceRun STT v1 vocabulary:
| Direction | Event | Purpose |
|---|---|---|
| Client → server | session.update | Configure the session, change settings, or switch models |
| Client → server | input_audio_buffer.append | Append base64 audio |
| Client → server | input_audio_buffer.commit | Mark the end of a turn; ends it under turn_detection: "client" |
| Client → server | turn.context | Set the assistant's latest utterance for semantic turn detection |
| Client → server | session.close | Finish gracefully |
| Server → client | session.created | Socket is ready for configuration |
| Server → client | session.updated | Configuration was accepted |
| Server → client | session.ready | The provider is connected |
| Server → client | session.update_pending | A model switch is starting or waiting for the turn to end |
| Server → client | session.update_failed | A model switch could not be applied; the current model continues |
| Server → client | session.model_changed | The live session moved to another model |
| Server → client | speech.started / speech.stopped | Voice activity, under silero, smart_turn, and semantic |
| Server → client | transcription.delta | Current full hypothesis; replace, do not concatenate |
| Server → client | turn.ended | Final transcript for a completed turn |
| Server → client | turn.context.updated | turn.context was applied |
| Server → client | warning | A setting the model cannot honor, or a turn-detection fallback |
| Server → client | session.closed | Answer to session.close |
| Server → client | error | Authentication, validation, or provider failure |
audio.append, audio.commit, and endpoint.context remain accepted as input aliases. Output always uses the v1 event names above.
VoiceRun STT Model additionally supports dynamic per-turn context, language constraints, and semantic server-side turn-taking. See VoiceRun STT Model.
