POST
/realtime/transcription_sessionsCreate transcription session
Create an ephemeral API token for use in client-side applications with the
Realtime API specifically for realtime transcriptions.
Can be configured with the same session parameters as the transcription_session.update client event.
It responds with a session object, plus a client_secret key which contains
a usable ephemeral API token that can be used to authenticate browser clients
for the Realtime API.
Returns the created Realtime transcription session object, plus an ephemeral key.
- RetriesRetries up to 2×, 500ms backoff, 30s timeout.
Create an ephemeral API key with the given session configuration.
turn_detectionobjectoptional
Configuration for turn detection. Can be set to `null` to turn off. Server VAD means that the model will detect the start and end of speech based on audio volume and respond at the end of user speech.
input_audio_noise_reductionobjectoptional
Configuration for input audio noise reduction. This can be set to `null` to turn off.
Noise reduction filters audio added to the input audio buffer before it is sent to VAD and the model.
Filtering the audio can improve VAD and turn detection accuracy (reducing false positives) and model performance by improving perception of the input audio.
input_audio_formatstringoptional
The format of input audio. Options are `pcm16`, `g711_ulaw`, or `g711_alaw`.
For `pcm16`, input audio must be 16-bit PCM at a 24kHz sample rate,
single channel (mono), and little-endian byte order.
input_audio_transcriptionobjectoptional
Configuration for input audio transcription. The client can optionally set the language and prompt for transcription, these offer additional guidance to the transcription service.
includearray<string>optional
The set of items to include in the transcription. Current available items are:
`item.input_audio_transcription.logprobs`
200Session created successfully.
client_secretobjectrequired
Ephemeral key returned by the API. Only present when the session is
created on the server via REST API.
modalitiesarrayoptional
The set of modalities the model can respond with. To disable audio,
set this to ["text"].
input_audio_formatstringoptional
The format of input audio. Options are `pcm16`, `g711_ulaw`, or `g711_alaw`.
input_audio_transcriptionobjectoptional
Configuration of the transcription model.
turn_detectionobjectoptional
Configuration for turn detection. Can be set to `null` to turn off. Server
VAD means that the model will detect the start and end of speech based on
audio volume and respond at the end of user speech.