AI enrichment
Every file admitted to the library is analyzed. That analysis is what separates the library from a folder with a search box: it is why you can ask for photos that look like this one, photos good enough to print, or photos of a person whose name nobody typed in.
This page covers what analysis produces, how to re-run and customize it, and the three searches that depend on it.
What analysis writes
Analysis runs after finalize and lands in the file's
file_metadata.analysis. Reading a file once it is ready gets you all of it:
Code
The fields that matter most:
title, description | Short and human-facing. Written for a caption rather than for search. |
retrieval_description | The long form, written to be searched. This is what the embedding is built from. |
environment | Where the shot was taken (setting, venue type, background quality), each with a confidence percentage. |
people.individuals | One entry per clearly visible person, with estimated age, gender, role and mood. Blurry and background people are deliberately omitted. |
marketing_suitability | Which channels the shot suits, the audience it reads to, and its overall appeal. |
photo_quality | Fourteen scores from 1 to 5: exposure, sharpness, framing, storytelling and the rest, plus an overall_score that may carry one decimal. |
tags, flags, mediaTags | Controlled vocabulary. See below. |
nudges | Suggested edits, each with a ready-to-send prompt. |
Almost everything here is filterable. See the analysis filters.
Your controlled vocabulary
tags, flags and mediaTags are not free text. The model may use only values from your
knowledge bases: the Media Tags and Flags
sets, plus the ones your product uses. Invented values are rejected rather than
stored.
The consequence is worth knowing: editing a knowledge base changes how every future upload is described. If the analysis is not using your terminology, correct the vocabulary rather than the prompt.
Re-run analysis
Re-run analysis when the vocabulary changed, a schema changed, or the file arrived before either existed:
Code
Re-analysis runs a model on every file you name. It consumes credits and counts against
the expensive rate-limit tier. Re-analyzing a whole library is
expensive, so filter to the files that need it and check GET /me for your remaining
credits first.
Analysis is asynchronous, exactly like the first pass. The file returns to pending and
reaches ready again when it finishes.
To correct analysis by hand rather than re-running it, PUT /files/{id}/analysis merges
a partial analysis object into what is already there. Use it to fix one incorrect field
without paying to redo the rest.
Custom analysis schemas
The standard analysis answers a general set of questions. A schema adds your own: fields
the platform has no way to infer, such as a product SKU or a shot type. Schemas live at
/media-analysis-schemas and require the module-custom-analysis entitlement.
Test one before committing to it:
Code
This is a dry run: it shows what the schema would produce and writes nothing. Iterate on the schema against a handful of representative files, then re-analyze in bulk once the schema is right.
Schema output is filterable through custom:
Code
Core fields (photo_quality, tags, flags and the rest listed above) are never
redefined by a schema, so a schema can only add fields.
Documents
Documents get text extraction rather than a visual analysis. The platform extracts a markdown rendition of the file, which is what search indexes and what agents retrieve against when the file backs a knowledge base. The platform maintains the rendition, and re-running analysis re-extracts it.
Search by meaning
Analysis produces embeddings from retrieval_description, and three searches run on
them.
Text to image
Describe what you want rather than what it is called:
Matching rows carry a similarity score and come back ordered by it.
similarityThreshold moves the cutoff: raise it for precision, lower it for recall. It
combines with every other filter, so "photos similar to this, from this album, rated 4 or
better" is a single request.
Similar images
Code
Compares stored embeddings against one seed file. It returns up to 24 results, 6 by default, and only files above a fixed similarity floor.
An empty array is a normal answer. It means one of three things: the seed's analysis has not finished, the seed is a generated image (which never gets an embedding), or nothing in the library is close enough. Do not treat it as an error.
Find this face
This takes two calls. First, post an image to get the face's embedding. The request is
multipart/form-data rather than JSON, and the field is image:
Then pass it to the library as face_embedding:
Code
Matching rows carry face_similarity. A lower face_threshold is stricter, because the
value is a distance rather than a score.
Two constraints apply: the image must contain a detectable face or the first call returns 400, and the upload is capped at 10 MB and images only. Blurry unassigned faces are skipped during matching, so a face search will not surface photos where the person is an unrecognizable smudge in the background.
To find someone you have already named, face_person is cheaper and exact. See people
and faces.
One-shot AI helpers
Four endpoints run a model and answer immediately, with no thread and no agent. All of
them consume credits and count against the expensive tier.
POST /assist/photo-selector | Picks the best photos for a stated purpose. Takes a query, optional criteria and a minQuality floor. It reads photo_quality, so it will not hand back a technically poor shot. |
POST /assist/captions-generator | Social captions for a photo, by platform. Takes an image_file_id or a description. |
POST /assist/magic-write | Writes or rewrites text from a brief, with optional tone and platform. |
POST /assist/canvas-assistant | Design-surface assistance. |
Use photo-selector when your integration needs "a good photo of X" rather than a list
to page through: it does the query, the quality filter and the ranking in one call.
Next steps
- People and faces turns detected faces into named people.
- Organizing and finding assets lists the full filter vocabulary.
- Agents and conversations covers knowledge bases, and asking questions in a thread.
