Generate high-accuracy captions for a video and attach them as a text track with the Mux Robots API. Use premium captions when accuracy matters, such as for accessibility or compliance.
Generate captions for a Mux asset and automatically attach them as a text track. Premium captions use a higher-accuracy speech model than Mux's standard auto-generated captions, and add optional speaker labels, word-level timestamps, and custom phrase hints for proper nouns and jargon. Captions are generated directly from the asset's audio, so no existing track is required. See the Generate Premium Captions API referenceAPI for the full endpoint specification. See Mux Robots pricing for unit costs.
If the asset already has a text track in the same language as the new captions, or with the same name, the request is rejected by default. Set replace_existing_tracks to delete conflicting tracks first. See Managing existing tracks.
To produce captions, this workflow needs an audio-only static rendition of the asset. If the asset already has one, it's reused. If not, the workflow creates one to process the audio and deletes it once the job completes, so you aren't charged for extra storage.
generate-premium-captions jobcurl https://api.mux.com/robots/v0/jobs/generate-premium-captions \
-H "Content-Type: application/json" \
-X POST \
-d '{
"parameters": {
"asset_id": "YOUR_ASSET_ID",
"language_code": "en",
"phrases": ["Mux", "API"]
}
}' \
-u ${MUX_TOKEN_ID}:${MUX_TOKEN_SECRET}This request is asynchronous. The POST returns immediately with the job in pending status and does not include results. We strongly recommend listening for the robots.job.generate_premium_captions.completed webhook: the payload contains the full completed job, so no follow-up API call is needed. If webhooks aren't an option, you can poll GET /robots/v0/jobs/generate-premium-captions/{JOB_ID} with the id from the response until the status is completed.
| Parameter | Type | Description |
|---|---|---|
asset_id | string | Required. The Mux asset ID of the video to caption. |
language_code | string | BCP 47 language code of the audio (e.g. en, es). Auto-detected when omitted. See language support for Mux Robots. |
replace_existing_tracks | string | What to do when the asset already has a conflicting text track: fail (the default), replace_all, or replace_generated. See Managing existing tracks. |
replace_existing | boolean | Deprecated. Use replace_existing_tracks instead. true behaves as replace_all and false as fail. |
track_name | string | Custom name for the uploaded Mux text track. Defaults to "{Language} (Generated)" using the resolved language code. |
include_speakers | boolean | When true, speaker labels are identified and added to each caption cue. Useful for interviews, podcasts, and multi-speaker content. Defaults to false. |
include_words | boolean | When true, word-level timestamps are exported as a JSON file accessible via temporary_words_url in the output. Billed at a higher unit rate. Defaults to false. |
upload_to_mux | boolean | Whether to upload the generated captions as a new text track on the asset. Defaults to true. When false, no track is created (and replace_existing_tracks must be fail); the captions remain available via temporary_srt_url. |
phrases | array of strings | Best-effort list of words or short phrases (proper nouns, product names, jargon) likely to appear in the audio, used to bias recognition toward correct spellings. Up to 100 phrases, each up to 50 characters. Does not guarantee exact output. |
The outputs object is included in the job once its status is completed. You'll receive it on the robots.job.generate_premium_captions.completed webhook (recommended), or you can fetch it with GET /robots/v0/jobs/generate-premium-captions/{JOB_ID}. It contains:
| Field | Type | Description |
|---|---|---|
track_id | string | Mux text track ID of the newly uploaded caption track. Omitted when upload_to_mux is false. |
language_code | string | Resolved language code of the generated captions (may differ from the requested code when auto-detected). |
temporary_srt_url | string | Temporary pre-signed URL to download the generated SRT file. Expires 7 days after the job completes. |
temporary_words_url | string | Temporary pre-signed URL to download the word-level timestamp JSON. Present when include_words is true. Expires 7 days after the job completes, so download and store it for long-term access. |
replaced_tracks | array of objects | Every track deleted before the new track was created. Absent when nothing was deleted. See Managing existing tracks. |
replaced_track_id | string | Deprecated. Use replaced_tracks instead. The ID of the first entry in replaced_tracks, kept for backwards compatibility. |
This example uses a 6-minute asset with include_words: false: at 500 units per minute, the job consumes 3,000 units. This is the payload delivered to the robots.job.generate_premium_captions.completed webhook, and the same shape you get from GET /robots/v0/jobs/generate-premium-captions/{JOB_ID}:
{
"data": {
"id": "rjob_yza567",
"workflow": "generate-premium-captions",
"status": "completed",
"units_consumed": 3000,
"parameters": {
"asset_id": "YOUR_ASSET_ID",
"language_code": "en",
"replace_existing_tracks": "fail",
"include_speakers": false,
"include_words": false,
"upload_to_mux": true,
"phrases": ["Mux", "API"]
},
"outputs": {
"track_id": "track_en_abc123",
"language_code": "en",
"temporary_srt_url": "https://storage.googleapis.com/..."
}
}
}When upload_to_mux is true (the default), the caption track is automatically attached to your asset, and viewers will see the new language option in the player's caption menu.
Use replace_existing_tracks to control what happens when the asset already has a text track in the same language (ignoring region, so en matches en-US) or with the same name as the new track.
| Value | Behavior |
|---|---|
fail (default) | Nothing is deleted, and the new track isn't added. |
replace_all | Conflicting tracks are deleted, then the new track is added. |
replace_generated | Conflicting tracks auto-generated by Mux Video are deleted. If any other track conflicts, nothing is deleted and the new track isn't added. |
Use replace_generated to upgrade Mux's auto-generated captions without touching tracks you uploaded. Tracks created by an earlier Robots job count as uploaded, so replacing one needs replace_all.
outputs.replaced_tracks.422. If it only shows up later, the job errors instead. That happens when a conflicting track is added while the job runs, or when you omit language_code, since conflicts can't be checked until the language is detected. Neither case is billed.language_code, tracks are only deleted when the language is detected with high confidence. Otherwise the job errors and suggests setting language_code.replace_existing is deprecated. true behaves as replace_all and false as fail.When include_words is true, download the file at temporary_words_url to get word-level timestamps. It's a JSON array of token objects in playback order. The array interleaves spacing tokens between words so you can reconstruct the exact text, and audio_event tokens capture non-speech sounds.
| Field | Type | Description |
|---|---|---|
text | string | The token's text. For word tokens this includes any trailing punctuation (e.g. "Matt,", "is."). For spacing tokens it's a single space (" "). For audio_event tokens it's a bracketed non-speech cue (e.g. "[laughs]"). |
start | number | Start time of the token, in seconds from the start of the media (fractional, e.g. 26.38). |
end | number | End time of the token, in seconds. |
type | string | One of word, spacing, or audio_event. |
speaker_id | string | The speaker the token is attributed to, formatted speaker_N (zero-indexed). Set only when include_speakers is true. |
[
{ "text": "Hey,", "start": 0.2, "end": 0.42, "type": "word", "speaker_id": "speaker_0" },
{ "text": " ", "start": 0.42, "end": 0.45, "type": "spacing", "speaker_id": "speaker_0" },
{ "text": "I'm", "start": 0.45, "end": 0.6, "type": "word", "speaker_id": "speaker_0" },
{ "text": " ", "start": 0.6, "end": 0.63, "type": "spacing", "speaker_id": "speaker_0" },
{ "text": "a", "start": 0.63, "end": 0.7, "type": "word", "speaker_id": "speaker_0" },
{ "text": " ", "start": 0.7, "end": 0.73, "type": "spacing", "speaker_id": "speaker_0" },
{ "text": "Mux", "start": 0.73, "end": 0.98, "type": "word", "speaker_id": "speaker_0" },
{ "text": " ", "start": 0.98, "end": 1.01, "type": "spacing", "speaker_id": "speaker_0" },
{ "text": "Robot,", "start": 1.01, "end": 1.4, "type": "word", "speaker_id": "speaker_0" },
{ "text": " ", "start": 1.4, "end": 1.43, "type": "spacing", "speaker_id": "speaker_0" },
{ "text": "beep", "start": 1.43, "end": 1.7, "type": "word", "speaker_id": "speaker_0" },
{ "text": " ", "start": 1.7, "end": 1.73, "type": "spacing", "speaker_id": "speaker_0" },
{ "text": "boop.", "start": 1.73, "end": 2.05, "type": "word", "speaker_id": "speaker_0" },
{ "text": "[beeps]", "start": 2.1, "end": 2.4, "type": "audio_event", "speaker_id": "speaker_1" },
{ "text": "Beep", "start": 2.5, "end": 2.72, "type": "word", "speaker_id": "speaker_1" },
{ "text": " ", "start": 2.72, "end": 2.75, "type": "spacing", "speaker_id": "speaker_1" },
{ "text": "boop", "start": 2.75, "end": 2.98, "type": "word", "speaker_id": "speaker_1" },
{ "text": " ", "start": 2.98, "end": 3.01, "type": "spacing", "speaker_id": "speaker_1" },
{ "text": "yourself!", "start": 3.01, "end": 3.5, "type": "word", "speaker_id": "speaker_1" }
]