---
title: Reliability
description: How a job is watched, retried, parked, timed out and refunded, and what we do not do.
---

# Reliability

Once a generation has a job id, the server owns its progress and credit settlement. Your client can disconnect, reconnect and read that same job; the way you wait does not change how the model runs.

## The poller

The poller reads jobs in batches of up to 20 and processes up to 10 at a time per poller. Jobs without a provider callback use a 5-second polling cadence. For callback-enabled video jobs, a verified provider callback wakes an immediate status read and a 20-second polling backstop reconciles missed callbacks; the next eligible poll time is persisted with the job.

![A provider callback wakes an immediate read; the poller remains a backstop](../assets/diagrams/callback-wakeup.svg)

| Observation | Result | Scope |
| --- | --- | --- |
| Working provider callback | About 3 status reads per render, previously about 15 | Measured on production. |
| Broken callback host | Job still settled within about 26 seconds in the exercised case | Measured on production; not a latency guarantee. |
| First provider with callbacks | Kling | Other jobs keep their existing polling path. |

Your client does not need to duplicate this provider polling. Follow the Nolgia job with a status read, [long-poll](./jobs.html#long-poll-with-wait) or [SSE](./streaming.html).

## Automatic retries

For queued image and audio execution, the server allows up to two transient re-attempts. The same credit hold stays attached across attempts; only the terminal outcome settles it. This applies to transient execution failures, not a guarantee that an invalid request or a content-policy refusal will be retried successfully.

For video, the poller follows the provider's accepted request, checks the deadline and completes or fails the Nolgia job. A provider's own internal retries are not a Nolgia retry setting. If an initial video submission meets a provider capacity wall, the job can remain queued while the poller re-attempts that submission within its deadline.

| Situation | Server behavior | Client behavior |
| --- | --- | --- |
| Transient image/audio execution failure | Up to two re-attempts under the same hold. | Keep following the same job. |
| Accepted video still rendering | Continue status reconciliation until completion or deadline. | Keep following the same job. |
| Provider capacity before video acceptance | Queue and retry the upstream submit within the model deadline. | Display `status_message`; do not create a replacement. |
| Client request refused with `429 rate_limit` | No generation accepted by that request. | Wait, then resubmit with the same key. |
| Duplicate body and key within five minutes | `409` points to the accepted job; no second bill. | Follow the returned `job_id`. |

## Provider outages

A queued job can expose `status_detail` so you can distinguish an unavailable provider from a healthy provider with no capacity. Neither value is a terminal error; the existing credit hold stays held while the job waits. Show the accompanying `status_message` instead of diagnosing the provider yourself.

| `status_detail` | Why the job waits | Bound and terminal outcome |
| --- | --- | --- |
| `provider_down` | The provider is unavailable; parked work resumes when it recovers. | The default outage parking bound is 6 hours from `created_at`; expiry fails the job and fully refunds the hold. |
| `upstream_at_capacity` | The provider is healthy but all its simultaneous-generation slots are occupied. | Submission is re-attempted within the ordinary model poll deadline; expiry fails with a timeout and refund. |

A capacity wait can become outage parking if the provider becomes unavailable. The six-hour parking bound applies to the outage state, not to every queued job.

> [!NOTE]
> You are never charged for an operational failure: provider errors end as `job_failed`, and an exhausted render deadline or outage parking deadline ends as `timeout`, with a full refund. A provider-billed content-policy refusal (`prompt_nsfw` or `ip_detected`) can be charged. The recorded `failure.credits_refunded` is authoritative: `true` means refunded, `false` means charged, and absent or null means the ledger outcome is not known. See [Pricing and credits](./billing.html).

<!-- gen:fields schema=JobFailure -->
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `code` | `GenerationErrorCode` | No | Stable refinement of `kind`, present on every job that failed after this field shipped and derived from the recorded reason for older ones.… |
| `kind` | string | Yes | What stopped the job.… |
| `message` | string | Yes | Human-readable reason, safe to show the customer as-is. The same text as `error.detail`. |
| `credits_refunded` | boolean, nullable | No | What the credit ledger actually did with this job's credit hold, recorded when the hold settled.… |
<!-- /gen -->

## Deadlines

The default video deadline is 15 minutes from `created_at`. Certain models have a 25-minute override; video restore starts with a 25-minute base and scales with source duration. These are server render budgets, distinct from how long one client HTTP request stays open.

| Work or request | Deadline or window | Outcome when it expires |
| --- | --- | --- |
| Default asynchronous video job | 15 minutes from `created_at` | Job fails with `timeout`; credits are refunded. |
| `minimax-h3` | 25 minutes | Job fails with `timeout`; credits are refunded. |
| `wan-3.0` | 25 minutes | Job fails with `timeout`; credits are refunded. |
| `wan-3.0-prime` | 25 minutes | Job fails with `timeout`; credits are refunded. |
| `seedance-2.5` | 25 minutes | Job fails with `timeout`; credits are refunded. |
| Video restore | 25-minute base plus 4 minutes per whole source second, rounded up, capped at 6 hours; unknown source duration uses the base | Job fails with `timeout`; credits are refunded. |
| Provider-outage parking | Default 6 hours from `created_at` | Explicit failure and a full refund. |
| One `GET /jobs/{id}/wait` request | `timeout_seconds`, at most 900 seconds | HTTP `408 timeout`; the job can still be running. |

<!-- gen:params op=waitForJob -->
| Parameter | In | Required | Description |
| --- | --- | --- | --- |
| `id` | path | Yes | Job UUID. |
| `timeout_seconds` | query | No |  |
<!-- /gen -->

> [!TIP]
> A `408` is a closed wait window, not a failed generation. Wait again on the same id; starting a new generation would create separate work.

## Model fallbacks

There is no general model fallback; the narrow exception is a content-policy refusal on OpenRouter-routed Seedance rows, which is re-routed to an alternative route.

## Backup domains

There is no backup domain; the production API host is `https://api.nolgia.ai/v1`.

The spec also publishes staging as a separate environment, not as a failover destination:

<!-- gen:spec-servers -->
| Environment | Base URL |
| --- | --- |
| Production | `https://api.nolgia.ai/v1` |
| Staging | `https://api.stg.nolgia.ai/v1` |
| Local development | `http://localhost:8080/v1` |
<!-- /gen -->

## The ten-minute window after a deploy

> [!WARNING]
> For about ten minutes after an API deploy, a brand-new model id can answer `400 validation` with “unknown … model” from an instance on the previous revision while another instance already accepts it. This was observed during rollout. Verify the id against the catalog, wait about a minute and retry; do not “fix” a newly published id that is already correct. Ordinary validation errors still require correcting the request.

## Timeouts, side by side

| Timeout | Where it runs | What stops | What continues |
| --- | --- | --- | --- |
| `timeout_seconds` on `/jobs/{id}/wait` | Server, for one HTTP request | The wait returns `408 timeout`. | The generation and credit hold. |
| The helper's `maxPollTime` (TypeScript) or `max_poll_time` (Python) | Client; default 1,800,000 milliseconds / 1,800 seconds | Client polling stops with a timeout error. | Server generation; successful work still spends credits. |
| Poller render deadline | Server | The generation job is failed and the hold refunded. | The terminal job remains readable. |
| Provider-outage parking deadline | Server | Waiting for provider recovery ends with a failure and refund. | The terminal job remains readable. |

The `subscribe`/`submit` helpers (TypeScript and Python 0.1.2; the Rust crate keeps its own) enforce only their client wait budget. Their timeout does not cancel a submitted job; see [Synchronous: subscribe](./subscribe.html).

## Error responses

| Code | Where to read it | What to do |
| --- | --- | --- |
| `rate_limit` | HTTP `429` problem at submit | Wait for capacity or the quota reset, then retry. |
| `timeout` | HTTP `408` problem from wait | Wait again; the generation has not necessarily failed. |
| `timeout` | `failure.code` on a failed job | The render budget is exhausted; inspect the refund and decide whether to submit again. |
| `job_failed` | `failure.code` on a failed job | Read `failure.message` and the recorded refund outcome. |
| `prompt_nsfw`, `ip_detected` | `failure.code` on a failed job, or `422` on an immediate refusal | Change the content or references; refund behavior comes from the recorded settlement. |
| `validation` | HTTP `400` problem at submit | Correct the request, except for the verified new-model rollout case above. |

## Next steps

:::cards
- [Concurrency limits](./concurrency-limits.html): Read your plan's live generation ceiling.
- [Errors](./errors.html): Branch on precise problem and failure codes.
- [Callbacks and webhooks](./callbacks.html): Understand callback wake-ups and how to observe completion.
:::
