Model APIs

Concurrency limits

On this page
  1. How it works
  2. Limits by plan
  3. Viewing your limit
    1. Parameters
  4. When you hit it
    1. Error responses
  5. Increasing your limit
  6. Next steps

Your plan sets how many generations can run at once. Read the live count before scheduling a batch, and treat a concurrency refusal as a signal to wait for a slot.

How it works #

One shared limit counts active generations across image, audio and video, including queued and running jobs whose credit hold owns a slot. The API checks the limit when you submit; already queued jobs are not dropped when you reach it. A refused request has no new job to poll, so wait for an existing job to finish before submitting that request again.

Limits by plan #

Plan Concurrent generations
Free 1
Starter 2
Pro 4
Studio 8
Team 8
Enterprise 8

These limits apply across the three modalities together, rather than separately to each model. For example, an image and a video running at the same time use both of a Starter plan's slots.

Viewing your limit #

Call GET /me and read generation_limits.concurrent_max and generation_limits.concurrent_active. The examples print only that part of the account response.

shell
curl --fail-with-body -sS https://api.nolgia.ai/v1/me \
  -H "Authorization: Bearer $NOLGIA_TOKEN"
shell
nolgia account me
TypeScript
import { createNolgiaClient } from "@nolgia/sdk";

const nolgia = createNolgiaClient(process.env.NOLGIA_TOKEN!);
const { data: me, error } = await nolgia.GET("/me");
if (error) throw new Error(`${error.title}: ${error.detail ?? ""}`);
console.log({ generation_limits: me.generation_limits });
Python
import os
import json
from nolgia import AuthenticatedClient
from nolgia.api.auth import get_current_user
from nolgia.models import GenerationLimits, User

client = AuthenticatedClient(base_url="https://api.nolgia.ai/v1", token=os.environ["NOLGIA_TOKEN"])
me = get_current_user.sync(client=client)
if not isinstance(me, User):
    raise SystemExit(f"refused: {me}")
if not isinstance(me.generation_limits, GenerationLimits):
    raise SystemExit("generation limits were not returned")
print(json.dumps({"generation_limits": me.generation_limits.to_dict()}, indent=2))
Rustsrc/main.rs
use nolgia_client::ClientBuilder;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let client = ClientBuilder::new("https://api.nolgia.ai/v1")
        .bearer_token(std::env::var("NOLGIA_TOKEN")?)
        .build()?;
    let me = client.get_current_user().send().await?.into_inner();
    println!("{:#?}", me.generation_limits);
    Ok(())
}

This is the generation_limits portion of the production account response captured on 2026-09-21.

JSON200 OK — generation_limits
{
  "generation_limits": {
    "concurrent_active": 0,
    "concurrent_max": 8
  }
}
Field Type Required Description
concurrent_max integer Yes Maximum generations this account may run at once on its effective plan.
concurrent_active integer Yes Generations currently running across image, audio, and video.

Parameters #

GET /me has no path or query parameters. Send your bearer token; the response describes the authenticated account in its current organization context.

When you hit it #

The generation submit returns 429 Too Many Requests with problem code rate_limit. This example is built from the Error schema and the handler's exact detail text for two active generations on a two-slot plan; it is not a captured response.

JSON429 — example from the Error schema
{
  "type": "about:blank",
  "title": "Generation Concurrency Limit Reached",
  "status": 429,
  "detail": "You have 2 generations running, which is the concurrent maximum of 2 for your plan. Wait for one to finish or upgrade your plan to run more at once.",
  "code": "rate_limit"
}
Field Type Required Description
code string No Machine-readable error code.…
type string Yes A URI reference identifying the problem type.
title string Yes
status integer Yes
detail string No
request_id string No

Wait for a running job to finish, then resubmit. rate_limit means refused for now, not refused outright: retry later and the same request will be accepted when capacity is available. Keep the same Idempotency-Key for retries so a previously accepted request is not turned into an intentional second generation.

There is a separate per-account quota refusal with the same 429 status and rate_limit code. When that response includes Retry-After-Reset, its value is an RFC 3339 reset timestamp; wait until that time before retrying. A concurrency refusal does not promise that header or a fixed delay.

Error responses #

Operation HTTP status Code Action
Read your account 401 No generation code is required Replace the invalid or expired token; do not retry it unchanged.
Submit at the concurrency ceiling 429 rate_limit Wait for a running job to finish, then resubmit.
Submit above an account quota 429 rate_limit Honor Retry-After-Reset when supplied.

Increasing your limit #

Choose a plan with more concurrent generations on Pricing. In an organization, the organization's plan determines the effective limit; a member's personal subscription does not raise it. See Organizations for organization context and shared credits.

Next steps #