The transcription pipeline
Each pre-recorded request follows the same sequence.Queue (job scheduling)
Before processing starts, the job enters a queue. Wait time depends on:
- How many transcriptions are already running on Gladia’s clusters
- Overall system load
- Your plan’s concurrency limits
Pre-processing
Once the job is dispatched, the audio is prepared for transcription:
- Audio normalization (format, codec, bitrate, sample rate). Some formats take longer to convert — see conversion time by format
- Voice Activity Detection (VAD) to isolate speech from silence
Language discovery
If you do not specify a language, Gladia detects which language(s) are spoken. That improves accuracy but can add latency.If you already know the language, set it in
language_config.languages to skip this phase. See recommended parameters and language detection.Inference
Speech is converted to text. Inference time scales with audio length and content complexity, and is usually the most predictable part of the pipeline.
Post-processing and formatting
After inference, the result is refined. Depending on your configuration, this may include:
- Diarization (speaker separation)
- Sentence structuring, custom vocabulary / custom spelling
- Other enabled Audio Intelligence add-ons
Concurrency and plan limits
Processing speed and queue time also depend on your plan:
When concurrency limits are reached, additional requests wait until a slot is free. Full details: Concurrency and rate limits.
Retry policy: when to resubmit (and when not to)
Pre-recorded transcription is asynchronous. The job is accepted when either of these happens:- Your
POST /v2/pre-recordedreturns HTTP 200 (with a jobid) - You receive the
transcription.createdwebhook
queued or processing is not a failed submission.
Errors and when to retry
Atranscription.error event — via webhook or callback — or an error status on GET means processing failed for that job ID. At that point a new POST is allowed — but not every error is worth retrying:
- Transient / infrastructure-style failures (timeouts, temporary 5xx, short-lived capacity issues): retry with backoff can help
- Request or media issues (invalid / unreachable audio URL, unsupported or corrupt file, bad parameters): fix the input first — resubmitting the same payload will fail again and still create a new billable job if accepted
id) before deciding.
Recommended client pattern
Treat acceptance as final for that audio
On 200 or
transcription.created, store the job id and stop submitting that audio.No deduplication today
Gladia does not currently deduplicate jobs. Each successfulPOST creates a new job and is billed separately, even if:
- the audio file is the same
custom_metadatais the same (for example your ownvideoId/botId)- a previous job for that content is still queued, processing, or already done
Until a server-side deduplication option exists, your integration should treat “one accepted job per piece of content” as a client-side responsibility.
Related
Pre-recorded quickstart
Upload, create a job, poll or use webhooks/callbacks
Concurrency & rate limits
Plan limits, 429s, and queue capacity
Supported files & duration
Formats, conversion time, and duration limits
POST /v2/pre-recorded
Initiate an async transcription job
GET /v2/pre-recorded/:id
Poll job status and results
transcription.created
Webhook when a job is accepted
transcription.error (webhook)
Webhook when processing fails
transcription.error (callback)
Callback when processing fails