Model APIs
Concurrency limits
On this page
Your plan sets how many generations can run at once. Read the live count before scheduling a batch, and treat a concurrency refusal as a signal to wait for a slot.
How it works #
One shared limit counts active generations across image, audio and video, including queued and running jobs whose credit hold owns a slot. The API checks the limit when you submit; already queued jobs are not dropped when you reach it. A refused request has no new job to poll, so wait for an existing job to finish before submitting that request again.
Limits by plan #
| Plan | Concurrent generations |
|---|---|
| Free | 1 |
| Starter | 2 |
| Pro | 4 |
| Studio | 8 |
| Team | 8 |
| Enterprise | 8 |
These limits apply across the three modalities together, rather than separately to each model. For example, an image and a video running at the same time use both of a Starter plan's slots.
Viewing your limit #
Call GET /me and read generation_limits.concurrent_max and generation_limits.concurrent_active. The examples print only that part of the account response.
curl --fail-with-body -sS https://api.nolgia.ai/v1/me \
-H "Authorization: Bearer $NOLGIA_TOKEN"
nolgia account me
import { createNolgiaClient } from "@nolgia/sdk";
const nolgia = createNolgiaClient(process.env.NOLGIA_TOKEN!);
const { data: me, error } = await nolgia.GET("/me");
if (error) throw new Error(`${error.title}: ${error.detail ?? ""}`);
console.log({ generation_limits: me.generation_limits });
import os
import json
from nolgia import AuthenticatedClient
from nolgia.api.auth import get_current_user
from nolgia.models import GenerationLimits, User
client = AuthenticatedClient(base_url="https://api.nolgia.ai/v1", token=os.environ["NOLGIA_TOKEN"])
me = get_current_user.sync(client=client)
if not isinstance(me, User):
raise SystemExit(f"refused: {me}")
if not isinstance(me.generation_limits, GenerationLimits):
raise SystemExit("generation limits were not returned")
print(json.dumps({"generation_limits": me.generation_limits.to_dict()}, indent=2))
use nolgia_client::ClientBuilder;
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let client = ClientBuilder::new("https://api.nolgia.ai/v1")
.bearer_token(std::env::var("NOLGIA_TOKEN")?)
.build()?;
let me = client.get_current_user().send().await?.into_inner();
println!("{:#?}", me.generation_limits);
Ok(())
}
This is the generation_limits portion of the production account response captured on 2026-09-21.
{
"generation_limits": {
"concurrent_active": 0,
"concurrent_max": 8
}
}
| Field | Type | Required | Description |
|---|---|---|---|
concurrent_max |
integer | Yes | Maximum generations this account may run at once on its effective plan. |
concurrent_active |
integer | Yes | Generations currently running across image, audio, and video. |
Parameters #
GET /me has no path or query parameters. Send your bearer token; the response describes the authenticated account in its current organization context.
When you hit it #
The generation submit returns 429 Too Many Requests with problem code rate_limit. This example is built from the Error schema and the handler's exact detail text for two active generations on a two-slot plan; it is not a captured response.
{
"type": "about:blank",
"title": "Generation Concurrency Limit Reached",
"status": 429,
"detail": "You have 2 generations running, which is the concurrent maximum of 2 for your plan. Wait for one to finish or upgrade your plan to run more at once.",
"code": "rate_limit"
}
| Field | Type | Required | Description |
|---|---|---|---|
code |
string | No | Machine-readable error code.… |
type |
string | Yes | A URI reference identifying the problem type. |
title |
string | Yes | |
status |
integer | Yes | |
detail |
string | No | |
request_id |
string | No |
Wait for a running job to finish, then resubmit. rate_limit means refused for now, not refused outright: retry later and the same request will be accepted when capacity is available. Keep the same Idempotency-Key for retries so a previously accepted request is not turned into an intentional second generation.
There is a separate per-account quota refusal with the same 429 status and rate_limit code. When that response includes Retry-After-Reset, its value is an RFC 3339 reset timestamp; wait until that time before retrying. A concurrency refusal does not promise that header or a fixed delay.
Error responses #
| Operation | HTTP status | Code | Action |
|---|---|---|---|
| Read your account | 401 |
No generation code is required | Replace the invalid or expired token; do not retry it unchanged. |
| Submit at the concurrency ceiling | 429 |
rate_limit |
Wait for a running job to finish, then resubmit. |
| Submit above an account quota | 429 |
rate_limit |
Honor Retry-After-Reset when supplied. |
Increasing your limit #
Choose a plan with more concurrent generations on Pricing. In an organization, the organization's plan determines the effective limit; a member's personal subscription does not raise it. See Organizations for organization context and shared credits.

