Skip to main content
When you call Gladia’s pre-recorded API, each request passes through several internal stages before the final transcript is ready. Knowing these steps helps you interpret response times, tune your configuration, and avoid retry patterns that create duplicate jobs and extra billing.

The transcription pipeline

Each pre-recorded request follows the same sequence.

Queue (job scheduling)

Before processing starts, the job enters a queue. Wait time depends on:
  • How many transcriptions are already running on Gladia’s clusters
  • Overall system load
  • Your plan’s concurrency limits
Queuing keeps capacity fair across users. During high demand, or when your account hits its concurrency limit, short delays are normal — especially on short files, where queue time can be a large share of the total duration.

Pre-processing

Once the job is dispatched, the audio is prepared for transcription:
  • Audio normalization (format, codec, bitrate, sample rate). Some formats take longer to convert — see conversion time by format
  • Voice Activity Detection (VAD) to isolate speech from silence

Language discovery

If you do not specify a language, Gladia detects which language(s) are spoken. That improves accuracy but can add latency.If you already know the language, set it in language_config.languages to skip this phase. See recommended parameters and language detection.

Inference

Speech is converted to text. Inference time scales with audio length and content complexity, and is usually the most predictable part of the pipeline.

Post-processing and formatting

After inference, the result is refined. Depending on your configuration, this may include:
  • Diarization (speaker separation)
  • Sentence structuring, custom vocabulary / custom spelling
  • Other enabled Audio Intelligence add-ons

Concurrency and plan limits

Processing speed and queue time also depend on your plan: When concurrency limits are reached, additional requests wait until a slot is free. Full details: Concurrency and rate limits.

Retry policy: when to resubmit (and when not to)

Pre-recorded transcription is asynchronous. The job is accepted when either of these happens:
  • Your POST /v2/pre-recorded returns HTTP 200 (with a job id)
  • You receive the transcription.created webhook
Both mean Gladia has queued the job. Processing is not finished yet. Once a job is accepted, Gladia continues trying to process it on our side for up to 2 hours after we receive the request. During incidents or high load, queue and processing times can be longer than usual — a job that is still queued or processing is not a failed submission.
Do not resubmit the same audio if you already got a 200 or a transcription.created webhook. Keep that job ID and wait for completion via polling, success/error webhooks, or callback. Re-POSTing does not make the transcript arrive faster: it creates a new job that competes for the same capacity and can worsen the backlog.
If the HTTP client times out but Gladia still accepted the job, you may still receive transcription.created. Prefer that webhook (or a follow-up GET by id if you already stored one) before re-POSTing the same audio.

Errors and when to retry

A transcription.error event — via webhook or callback — or an error status on GET means processing failed for that job ID. At that point a new POST is allowed — but not every error is worth retrying:
  • Transient / infrastructure-style failures (timeouts, temporary 5xx, short-lived capacity issues): retry with backoff can help
  • Request or media issues (invalid / unreachable audio URL, unsupported or corrupt file, bad parameters): fix the input first — resubmitting the same payload will fail again and still create a new billable job if accepted
Inspect the error details (webhook or callback payload, and/or GET on the job id) before deciding.
1

Submit the job

Call POST /v2/pre-recorded.
2

Treat acceptance as final for that audio

On 200 or transcription.created, store the job id and stop submitting that audio.
3

Wait for completion

Complete via poll, success/error webhook, or callback.
4

Retry only when needed

Only re-POST if acceptance failed (no 200 and no transcription.created), or after transcription.error (webhook or callback) / status error when the error is retryable.
Prefer exponential backoff with jitter on transport / 5xx / 429 failures. Avoid tight loops (for example, retrying every few minutes while the original job is still in flight).

No deduplication today

Gladia does not currently deduplicate jobs. Each successful POST creates a new job and is billed separately, even if:
  • the audio file is the same
  • custom_metadata is the same (for example your own videoId / botId)
  • a previous job for that content is still queued, processing, or already done
Until a server-side deduplication option exists, your integration should treat “one accepted job per piece of content” as a client-side responsibility.

Pre-recorded quickstart

Upload, create a job, poll or use webhooks/callbacks

Concurrency & rate limits

Plan limits, 429s, and queue capacity

Supported files & duration

Formats, conversion time, and duration limits

POST /v2/pre-recorded

Initiate an async transcription job

GET /v2/pre-recorded/:id

Poll job status and results

transcription.created

Webhook when a job is accepted

transcription.error (webhook)

Webhook when processing fails

transcription.error (callback)

Callback when processing fails