Keep a subject's identity and a place's appearance in reusable Library records. A character carries its name, canonical description and reference images; a location does the same for a room, storefront or set. Generation attaches their references and folds their canonical descriptions into the prompt verbatim.
Read readiness before selecting a reference: ready means one front-facing face. needs_front_view, no_face, multiple_faces and unverified are advisory states, not a promise that the identity gate will pass. GET /characters/{id} also includes recent fidelity_history; list and write responses omit that history.
Field
Type
Required
Description
id
string
Yes
name
string
Yes
canonical_description
string
No
The character's canonical written description - the asset bible entry (face, hair, build, wardrobe, anchor) reused VERBATIM by every downstream generation.…
primary_reference_asset_id
string, nullable
No
The reference asset that anchors the character's identity: sheet actions attach it as the character reference and identity scoring runs against it. Defaults to the first reference asset when unset.
reference_assets
array of Asset
Yes
Reference images (image assets) in display order, each with a fresh signed URL.
voice
CharacterVoice
No
Used by POST /generate/audio (character_id) and by video generation with use_character_voice true; absent when the character has no voice.
readiness
string
No
Whether the character's PRIMARY reference can anchor its identity, derived from the face verdict stored when the reference joined the character.…
fidelity_history
array of CharacterFidelitySample
No
The character's most recent identity-gated renders, newest first: the identity_score the Aura gate measured against the primary reference on each image or video generated with this character_id.…
A location needs no training. Its canonical description stays unchanged across scenes, and its primary reference uses an image-reference slot on an image model or an element-reference slot on a video model. A location without an image contributes its description only.
Primary image becomes the face reference; requires an Aura-compatible model with room for the reference and num_images: 1
Primary image becomes an element reference
character_ids
Ordered cast of up to four; lead supplies the face reference, then the other cast references follow
Ordered cast of up to four; references keep cast order after the caller's own references; @Name is rewritten to the member's @ImageN slot
use_character_voice
Not an image field
Opt in to the lead's clip voice as the next audio reference; requires a model with an available audio slot and an acceptable clip length
location_id
Adds the place's description and reference
Adds the place's description and an element reference
The image request fields are generated from the contract:
Field
Type
Required
Description
character_id
string, nullable
No
One of your characters (GET /characters, created on the Create Characters page).…
character_ids
array of string, nullable
No
An ORDERED cast of up to 4 of your characters rendered together (two-person dialogue scenes, duets, family commercials).…
location_id
string, nullable
No
One of your locations (GET /locations): the place this render is set in.…
The video fields include the explicit voice opt-in:
Field
Type
Required
Description
character_id
string, nullable
No
One of your characters (GET /characters).…
use_character_voice
boolean, nullable
No
Attaches the lead character's (character_id, or the first of character_ids) voice clip as a reference audio track (audio_asset_ids, emitted after your own audio tracks and audio_urls, taking the next @Audio slot) and adds a voice line to the prompt.…
character_ids
array of string, nullable
No
An ORDERED cast of up to 4 of your characters rendered together in one clip.…
location_id
string, nullable
No
One of your locations (GET /locations): the place this clip is set in.…
When both character_id and character_ids are supplied, the single id must be a member of the cast. The first member is the lead, including for use_character_voice. Voice opt-in may change the cost on models billed by the reference audio's length; quote the full request first.
202 with a portrait job, canonical_description and eight rolled axes
Model
grok-imagine-image
Billing
The equivalent image generation price
Next step
Wait for the portrait, then create the character with POST /characters
Autopilot rolls heritage, age band, hair color, build, hair style, eye color, wardrobe and an anchor detail. Consecutive rolls differ on at least two of the four core axes. It starts a portrait job; it does not create the character record for you. Store its returned canonical_description unchanged when creating that record.
This illustrative response is built from AutopilotCharacterResponse, not a production capture.
JSON202 — example from the response schema
{"job":{"id":"b9f4319a-239e-4907-a9bb-406be1f0c2b8","user_id":"4961eaef-70a6-4e9b-b746-c86ad3e62077","modality":"image","model":"grok-imagine-image","status":"queued","created_at":"2026-09-21T04:00:00Z","updated_at":"2026-09-21T04:00:00Z"},"canonical_description":"A middle-aged woman with short dark curls, brown eyes, a sturdy build, a navy work jacket and a small silver brooch.","axes":{"heritage":"mixed heritage","age_band":"middle-aged","hair_color":"dark","build":"sturdy","hair_style":"short curls","eye_color":"brown","wardrobe":"navy work jacket","anchor":"small silver brooch"}}
Field
Type
Required
Description
job
Job
Yes
canonical_description
string
Yes
The rolled canonical written description. Store it on the character record unchanged when creating the character from this roll - it is the continuity contract for every downstream generation.
axes
object
Yes
The variety-engine roll: all 8 axis values (heritage, age_band, hair_color, build, hair_style, eye_color, wardrobe, anchor).…
Field
Type
Required
Description
id
string
Yes
user_id
string
Yes
modality
Modality
Yes
model
string
Yes
status
JobStatus
Yes
asset
Asset
No
failure
JobFailure
No
created_at
string
Yes
updated_at
string
Yes
For an existing character, POST /characters/{id}/sheets creates a split, turnaround or expression sheet on gpt-image-2, billed like that model's image generation. It needs at least one reference image. Wearable references and a named wardrobe outfit share a budget of three extra references; overflow is refused instead of trimming the outfit.
Field
Type
Required
Description
kind
CharacterSheetKind
Yes
wearable_reference_asset_ids
array of string
No
Image assets (owned by the caller) showing exact wearable items - garments, jewelry, eyewear, footwear - that must render EXACTLY as shown on the character in every view: never redesigned, restyled, recolored, or reinterpreted.…
outfit
string, nullable
No
The name of one of the character's wardrobe outfits (Character.wardrobe).…
project_id
string, nullable
No
Files the generated sheet into this caller-owned project.
At most one re-roll; the better-scoring attempt is delivered
Multi-character result
identity_score is the weakest cast member's score
Video sampling
Frames at 2 fps; each member uses its best matching face/frame
Unavailable scoring
Identity fields are absent or null; the reference still conditioned the generation
Field
Type
Required
Description
identity_score
number, nullable
No
ArcFace-family cosine similarity against the generation's identity reference (Aura identity gate).…
identity_rerolls
integer, nullable
No
Number of completed automatic identity re-rolls behind this render (0 or 1 — the gate re-rolls at most once, then delivers the better-scoring attempt). Present only alongside identity_score.
identity_gate_passed
boolean, nullable
No
Whether identity_score clears the 0.60 identity gate.…
character_scores
array of AssetCharacterScore, nullable
No
Per-character identity outcome of a multi-character cast render (character_ids), in cast order: each member's own ArcFace-family cosine (its best-matching face in the render, or the best frame of a clip) and whether it clears the 0.60 gate.…
Field
Type
Required
Description
character_id
string
Yes
The cast member (GET /characters/{id}).
identity_score
number
Yes
This member's cosine against the delivered render.
identity_gate_passed
boolean
Yes
Whether this member clears the 0.60 identity gate.
Characters without a reference image contribute only their description and remain unscored. Locations are not scored by the face gate.