VoiceRun STT Router

VoiceRun STT Router is one realtime WebSocket interface for VoiceRun STT Model and external speech-to-text models. Select a primary model and optional ordered fallback models, switch models mid-session, and choose how turns end; audio input and transcription events stay the same.

wss://api.voicerun.com/v1/stt

How It Works#

  1. Connect — open the WebSocket with your VoiceRun API key.
  2. Configure — send session.update with a primary model and optional fallback_models. The Router answers session.updated, then session.ready once the provider is connected.
  3. Stream — send base64 audio in input_audio_buffer.append messages. You can start right after session.update.
  4. Receive — replace the current hypothesis on transcription.delta; consume the completed turn from turn.ended. turn_detection decides what ends a turn; see Turn detection.
  5. Close — either close the WebSocket directly, or send session.close for a graceful shutdown and wait for session.closed. The protocol message is optional.

The gateway normalizes provider protocols. Provider-specific messages never escape the socket.

session.close does not commit or flush unfinished audio. Wait for the final turn.ended before closing a live session. When streaming a finite recording, append trailing silence so provider endpointing can finish the last turn; see Messages and Examples.

Authentication and billing#

One VoiceRun API key works for every enabled model:

  1. Sign in and open your VoiceRun profile.
  2. Under Personal Access Tokens, click Add PAT and copy the generated token.
  3. Use that PAT as VOICERUN_API_KEY; keep it server-side and out of source control.
ws = await websockets.connect( "wss://api.voicerun.com/v1/stt", additional_headers={"Authorization": "Bearer YOUR_VOICERUN_API_KEY"}, )

Browser clients that cannot set an Authorization header may use ?token=YOUR_VOICERUN_API_KEY. Avoid putting keys in URLs outside browser-only integrations.

{ "type": "session.update", "model": "voicerun-asr-realtime-v1", "fallback_models": ["nova-3", "gpt-4o-transcribe"], "language": "en" }

Usage is charged to your VoiceRun credit balance at the published rate for the selected model. VoiceRun STT Model is free during its preview; external models continue to use their published rates. VoiceRun manages provider credentials, and provider credential fields are rejected by the public API.

Automatic model fallback#

fallback_models is an ordered list of up to five model names. If the active provider fails at runtime, the Router settles that model's accepted audio, starts the next available model at its published rate, and emits session.model_changed.

The Router retains up to eight seconds of input audio and replays the audio since the last turn boundary into the fallback model, so an in-progress turn can continue without repeating finished ones. Replayed audio is charged to the fallback model because that model processes it. Each configured model is attempted at most once.

Fallback models inherit the session's shared configuration. If the new model does not support a setting used by the primary model, its adapter ignores that setting; the incompatibility does not prevent failover. Choose models with compatible behavior when a particular setting is essential to the conversation.

Fallback is not attempted for invalid configuration, an invalid VoiceRun API key, insufficient credits, or a concurrency limit. The fallback chain cannot be changed after the session becomes active.

Switching models mid-session#

Send session.update with a different model on an active session. By default the switch waits for the current turn to end; apply: "now" switches at once. The Router starts the new model beside the current one and moves over only once it has connected, so a model that cannot connect never interrupts the session. The new model is billed at its own published rate from the moment of the switch. If it later fails, the Router returns to the previous model before trying fallback_models. See Switching models mid-session.

Turn detection#

turn_detection decides when a turn ends, with every model:

turn_detectionThe turn ends when
provider (default)the model's own endpointing produces a final
clientyour client sends input_audio_buffer.commit
silerothe Router hears a short silence after speech
smart_turnthe Smart Turn model judges the pause complete, waiting longer when it sounds unfinished
semantica semantic model judges the transcript a complete reply to the assistant's last utterance

silero, smart_turn, and semantic run in the Router on its own copy of the audio, so they work with any model, sample rate, and audio format, and add speech.started / speech.stopped events. See Turn detection for the options and events.

Supported Models#

ProviderModels
VoiceRunvoicerun-asr-realtime-v1
Deepgramnova-3, flux-general-en, flux-general-multi
OpenAIgpt-4o-transcribe, gpt-4o-mini-transcribe
ElevenLabsscribe_v2_realtime
Cartesiaink-whisper
Sonioxstt-rt-v4
Googlechirp_3
Qwenqwen3-asr-flash, qwen3-asr-flash-realtime
xAIgrok-stt
Inworldinworld/inworld-stt-1
Gradiumgradium-default-stt
Tencenttencent-16k

Undated aliases may resolve to a tested provider snapshot while retaining the same public model selector.

This list covers the standalone STT Router only. VoiceRun voice agents select STT models from a separate catalog — see Speech to Text for the models available under Deployment.spec.stt.

Audio Format#

Input defaults to raw PCM16 little-endian, mono, at 16 kHz. Set input_audio_format to pcm16 or mulaw and sample_rate to a value from 8,000 through 48,000 Hz; VoiceRun converts it to the rate required by the selected engine.

Chunks around 20–100 ms work well for live audio. For a finite recording, append trailing silence so the final turn can end, or send input_audio_buffer.commit under turn_detection: "client".

Unified Events#

All models use the VoiceRun STT v1 vocabulary:

DirectionEventPurpose
Client → serversession.updateConfigure the session, change settings, or switch models
Client → serverinput_audio_buffer.appendAppend base64 audio
Client → serverinput_audio_buffer.commitMark the end of a turn; ends it under turn_detection: "client"
Client → serverturn.contextSet the assistant's latest utterance for semantic turn detection
Client → serversession.closeFinish gracefully
Server → clientsession.createdSocket is ready for configuration
Server → clientsession.updatedConfiguration was accepted
Server → clientsession.readyThe provider is connected
Server → clientsession.update_pendingA model switch is starting or waiting for the turn to end
Server → clientsession.update_failedA model switch could not be applied; the current model continues
Server → clientsession.model_changedThe live session moved to another model
Server → clientspeech.started / speech.stoppedVoice activity, under silero, smart_turn, and semantic
Server → clienttranscription.deltaCurrent full hypothesis; replace, do not concatenate
Server → clientturn.endedFinal transcript for a completed turn
Server → clientturn.context.updatedturn.context was applied
Server → clientwarningA setting the model cannot honor, or a turn-detection fallback
Server → clientsession.closedAnswer to session.close
Server → clienterrorAuthentication, validation, or provider failure

audio.append, audio.commit, and endpoint.context remain accepted as input aliases. Output always uses the v1 event names above.

VoiceRun STT Model additionally supports dynamic per-turn context, language constraints, and semantic server-side turn-taking. See VoiceRun STT Model.

transcriptionwebsocketsttspeech-to-text