> ## Documentation Index
> Fetch the complete documentation index at: https://docs.gladia.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Transcription process & retry policy

> How pre-recorded jobs move through the Gladia pipeline, concurrency impact, and when to retry safely

When you call Gladia’s pre-recorded API, each request passes through several internal stages before the final transcript is ready. Knowing these steps helps you interpret response times, tune your configuration, and avoid retry patterns that create duplicate jobs and extra billing.

## The transcription pipeline

Each pre-recorded request follows the same sequence.

<Steps titleSize="h3">
  <Step title="Queue (job scheduling)" icon="list">
    Before processing starts, the job enters a queue. Wait time depends on:

    * How many transcriptions are already running on Gladia’s clusters
    * Overall system load
    * Your plan’s [concurrency limits](/chapters/limits-and-specifications/concurrency)

    Queuing keeps capacity fair across users. During high demand, or when your account hits its concurrency limit, short delays are normal — especially on short files, where queue time can be a large share of the total duration.
  </Step>

  <Step title="Pre-processing" icon="sliders">
    Once the job is dispatched, the audio is prepared for transcription:

    * Audio normalization (format, codec, bitrate, sample rate). Some formats take longer to convert — see [conversion time by format](/chapters/limits-and-specifications/supported-formats#conversion-time)
    * Voice Activity Detection (VAD) to isolate speech from silence
  </Step>

  <Step title="Language discovery" icon="language">
    If you do not specify a language, Gladia detects which language(s) are spoken. That improves accuracy but can add latency.

    If you already know the language, set it in `language_config.languages` to skip this phase. See [recommended parameters](/chapters/pre-recorded-stt/recommended-parameters) and [language detection](/chapters/language/language-detection).
  </Step>

  <Step title="Inference" icon="microchip">
    Speech is converted to text. Inference time scales with audio length and content complexity, and is usually the most predictable part of the pipeline.
  </Step>

  <Step title="Post-processing and formatting" icon="wand-magic-sparkles">
    After inference, the result is refined. Depending on your configuration, this may include:

    * Diarization (speaker separation)
    * Sentence structuring, custom vocabulary / custom spelling
    * Other enabled [Audio Intelligence](/chapters/pre-recorded-stt/audio-intelligence) add-ons
  </Step>
</Steps>

## Concurrency and plan limits

Processing speed and queue time also depend on your plan:

| Plan type | Usage | Max Transcriptions in concurrency <br />(pre-recorded) | Max Transcriptions in concurrency (Live) |
| - | - | - | - |
| **Enterprise** | Unlimited | On demand | On demand |
| **Paid** | Unlimited | 25 | 30 |
| **Free** | €50 credits (one-time) | 3 | 1 |

When concurrency limits are reached, additional requests wait until a slot is free. Full details: [Concurrency and rate limits](/chapters/limits-and-specifications/concurrency).

## Retry policy: when to resubmit (and when not to)

Pre-recorded transcription is **asynchronous**. The job is **accepted** when either of these happens:

* Your `POST /v2/pre-recorded` returns **HTTP 200** (with a job `id`)
* You receive the [`transcription.created`](/api-reference/v2/pre-recorded/webhook/created) webhook

Both mean Gladia has queued the job. Processing is not finished yet.

Once a job is accepted, Gladia continues trying to process it on our side for **up to 2 hours** after we receive the request. During incidents or high load, queue and processing times can be longer than usual — a job that is still `queued` or `processing` is not a failed submission.

<Warning>
  **Do not resubmit** the same audio if you already got a **200** or a `transcription.created` webhook. Keep that job ID and wait for completion via [polling](/api-reference/v2/pre-recorded/get), [success/error webhooks](/chapters/pre-recorded-stt/quickstart), or [callback](/api-reference/v2/pre-recorded/init). Re-POSTing does not make the transcript arrive faster: it creates a **new** job that competes for the same capacity and can worsen the backlog.
</Warning>

| Situation | What to do |
| - | - |
| `POST` returned **200**, or you received **`transcription.created`** | Do **not** resubmit. Poll / wait for success or error on that job `id`. |
| `POST` failed (timeout, connection error, **5xx**, no response) **and** no `transcription.created` | Treat as **not accepted**. Safe to retry the `POST` with backoff. |
| [`transcription.error`](/api-reference/v2/pre-recorded/webhook/error) webhook, [`transcription.error`](/api-reference/v2/pre-recorded/callback/error) callback, or job status **`error`** | The accepted job failed. You **may** submit again, but check the error first — retry only when it makes sense for that failure. |
| Job status is **`queued`** or **`processing`** | Do **not** resubmit. Keep waiting on the same `id`. |
| HTTP **429** (concurrency / rate limit) | Back off and retry later. See [Concurrency and rate limits](/chapters/limits-and-specifications/concurrency). |

<Tip>
  If the HTTP client times out but Gladia still accepted the job, you may still receive `transcription.created`. Prefer that webhook (or a follow-up GET by `id` if you already stored one) before re-POSTing the same audio.
</Tip>

## Errors and when to retry

A `transcription.error` event — via [webhook](/api-reference/v2/pre-recorded/webhook/error) or [callback](/api-reference/v2/pre-recorded/callback/error) — or an `error` status on [GET](/api-reference/v2/pre-recorded/get) means processing failed for that job ID. At that point a new `POST` is allowed — but **not every error is worth retrying**:

* **Transient / infrastructure-style failures** (timeouts, temporary 5xx, short-lived capacity issues): retry with backoff can help
* **Request or media issues** (invalid / unreachable audio URL, unsupported or corrupt file, bad parameters): fix the input first — resubmitting the same payload will fail again and still create a new billable job if accepted

Inspect the error details (webhook or callback payload, and/or GET on the job `id`) before deciding.

## Recommended client pattern

<Steps titleSize="h3">
  <Step title="Submit the job">
    Call `POST /v2/pre-recorded`.
  </Step>

  <Step title="Treat acceptance as final for that audio">
    On **200** *or* `transcription.created`, store the job `id` and stop submitting that audio.
  </Step>

  <Step title="Wait for completion">
    Complete via poll, success/error webhook, or callback.
  </Step>

  <Step title="Retry only when needed">
    Only re-POST if acceptance failed (no 200 and no `transcription.created`), or after **`transcription.error`** (webhook or callback) / status `error` **when the error is retryable**.
  </Step>
</Steps>

Prefer exponential backoff with jitter on transport / 5xx / 429 failures. Avoid tight loops (for example, retrying every few minutes while the original job is still in flight).

## No deduplication today

Gladia does **not** currently deduplicate jobs.

Each successful `POST` creates a **new** job and is **billed separately**, even if:

* the audio file is the same
* `custom_metadata` is the same (for example your own `videoId` / `botId`)
* a previous job for that content is still queued, processing, or already done

<Note>
  Until a server-side deduplication option exists, your integration should treat “one accepted job per piece of content” as a client-side responsibility.
</Note>

## Related

<CardGroup cols={2}>
  <Card title="Pre-recorded quickstart" icon="play" href="/chapters/pre-recorded-stt/quickstart">
    Upload, create a job, poll or use webhooks/callbacks
  </Card>

  <Card title="Concurrency & rate limits" icon="gauge-high" href="/chapters/limits-and-specifications/concurrency">
    Plan limits, 429s, and queue capacity
  </Card>

  <Card title="Supported files & duration" icon="file-audio" href="/chapters/limits-and-specifications/supported-formats">
    Formats, conversion time, and duration limits
  </Card>

  <Card title="POST /v2/pre-recorded" icon="paper-plane" href="/api-reference/v2/pre-recorded/init">
    Initiate an async transcription job
  </Card>

  <Card title="GET /v2/pre-recorded/:id" icon="download" href="/api-reference/v2/pre-recorded/get">
    Poll job status and results
  </Card>

  <Card title="transcription.created" icon="bell" href="/api-reference/v2/pre-recorded/webhook/created">
    Webhook when a job is accepted
  </Card>

  <Card title="transcription.error (webhook)" icon="triangle-exclamation" href="/api-reference/v2/pre-recorded/webhook/error">
    Webhook when processing fails
  </Card>

  <Card title="transcription.error (callback)" icon="reply" href="/api-reference/v2/pre-recorded/callback/error">
    Callback when processing fails
  </Card>
</CardGroup>
