Router9
Documentation
API Reference

Image Tool

Image understanding, generation, and editing over the API — the Skill tool for harnesses, distinct from Studio's Images workbench

Image Description

Analyze images using AI vision models.

Endpoint

POST /v1/image/describe

Request

curl https://api.router9.com/v1/image/describe \
  -H "Authorization: Bearer sk-r9k-your-key" \
  -H "Content-Type: application/json" \
  -d '{
    "mediaUrl": "https://example.com/photo.jpg",
    "prompt": "What is in this image?"
  }'
ParameterTypeRequiredDescription
mediaUrlstringOne of mediaUrl/mediaBase64URL to image
mediaBase64stringOne of mediaUrl/mediaBase64Base64-encoded image
promptstringNoSpecific question about the image. Defaults to a general description
modelstringNoOverride the default vision model (must be allowlisted)

Response

{
  "model": "google/gemini-3.5-flash",
  "text": "The image shows a modern office space with developers working at their desks, multiple monitors displaying code editors."
}

This response used to be wrapped in { "success": true, "tool": ..., "result": {...} }. The wrapper is gone — read text and model from the top level.

Image Generation

Generate images from a text prompt.

Endpoint

POST /v1/image/generations

Request

curl https://api.router9.com/v1/image/generations \
  -H "Authorization: Bearer sk-r9k-your-key" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "A futuristic city skyline at sunset, digital art",
    "aspect_ratio": "16:9",
    "output_format": "webp"
  }'
ParameterTypeRequiredDescription
promptstringYesDescription of the image to generate
modelstringNoOverride the default image-generation model (must be allowlisted)
nintegerNoNumber of images. Capped per plan: 1 on Free, 4 on Pro, 10 on Max
aspect_ratiostringNo1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9
output_formatstringNopng, jpeg, webp, svg. Defaults to webp
seedintegerNoFixes the seed so the same inputs reproduce the same image
qualitystringNoauto, low, medium, high
backgroundstringNoauto, transparent, opaque. Transparency requires png or webp
resolutionstringNo512, 1K, 2K, 4K. Capped per plan: 1K on Free, 2K on Pro, 4K on Max
sizestringNoExplicit pixels, e.g. 2048x2048. Overrides resolution where supported
output_compressionintegerNo0–100, webp and jpeg only
response_formatstringNob64_json (the only value served today)
negative_promptstringNoWhat the image must not contain. Diffusion models only
cfg_scalenumberNoHow strictly the model follows the prompt. Range varies by model
style_presetstringNoA house look, e.g. photographic, anime. Vocabulary varies by model — list it with GET /v1/image/styles
enhance_promptbooleanNoHave a language model expand the prompt before drawing. A second billed call; see below
project_idstringNoProject to file the result into — an id from GET /v1/projects. Omit it, or send one that is not yours, and it lands in your default project
parent_asset_idstringNoA picture in your library, by id. The result joins its lineage, and with no image the picture is read from your own storage

The last three are accepted only by models whose provider takes them, and the set differs even between two model families from the same vendor. Sending one to a model that does not take it returns 400 naming the provider, rather than quietly ignoring it — a dropped cfg_scale returns an image that is wrong in a way you cannot attribute to anything. Call GET /v1/public/image-models to see what the model you are using accepts.

/v1/image/generations is text-to-image only. Sending images returns 400 pointing at /v1/image/edits.

Prompt enhancement

enhance_prompt: true runs one extra chat completion that expands a short prompt into a detailed description, then draws from that. Reach for it when you have an instruction rather than a description; skip it when you already said exactly what you want, because the rewrite adds detail you did not ask for.

The text it produced comes back as enhanced_prompt in the response body, and also in the X-Router9-Enhanced-Prompt header — the body is authoritative, since the header is sanitized and capped at 512 characters.

It fails open. If the rewrite cannot be produced for any reason, your own prompt is drawn and nothing extra is billed. There is no separate flag to check: enhanced_prompt is absent exactly when the original was used.

It is ignored on /v1/image/edits. An edit's prompt describes a change, and whatever you do not mention is expected to survive unchanged; an expanded prompt reintroduces the whole scene as instructions.

It is a second billed call, on its own line in your usage — one row for the rewrite and one for the picture.

The enhancement is deliberately one explicit flag rather than something the gateway does silently: it is a second billed call, and a caller composing several calls can write the rewrite as its own step — one visible price per call, and the expanded prompt a value the later calls can reuse.

Parameter support varies by model. A value this endpoint accepts can still be refused upstream; that error is returned verbatim so you can see which model rejected what.

Response

{
  "id": "imggen_01J8Z3C4V9",
  "created": 1714000000,
  "model": "google/gemini-3.1-flash-image-preview",
  "data": [
    {
      "b64_json": "UklGRiQAAABXRUJQ...",
      "media_type": "image/webp",
      "revised_prompt": "A futuristic city skyline at sunset, digital art"
    }
  ],
  "enhanced_prompt": "A futuristic city skyline at sunset, digital art — glass towers catching low amber light, elevated walkways, a haze of traffic below",
  "timing": { "total_ms": 8412, "upstream_ms": 8207 }
}

Images are returned as base64 bytes, never URLs. Use media_type to decode — do not assume PNG, since the default output format is WebP.

enhanced_prompt appears only when you asked for a rewrite and one was produced.

id names this generation in a support thread, and matches the request id in your own logs. timing splits the total against the upstream call, so a slow generation can be attributed without guessing. The vendor's own request id, when they report one, is in the X-Router9-Upstream-Trace-Id response header.

If the model declines to draw, the call still returns 200 with an empty data and a refusal string — and is still billed, because the upstream charged for the attempt either way.

Image Editing

Edit existing images, or guide a generation with references.

Endpoint

POST /v1/image/edits

Editing used to be POST /v1/image/generations with a non-empty images array. That shape is gone: /generations now refuses references, and edits have their own route.

Request (JSON)

Up to four references, as HTTP(S) URLs or data: URLs:

curl https://api.router9.com/v1/image/edits \
  -H "Authorization: Bearer sk-r9k-your-key" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "make this a watercolor painting",
    "images": ["https://example.com/photo.jpg"]
  }'

What each image is for

The same photo can be the picture to change, a person to keep, or a look to borrow. Say which with {"url", "role"} in place of a bare URL:

RoleMeaning
editThe picture to change. Whatever the prompt does not mention stays as it is. First position only.
subjectA person or object to keep, placed in a new picture
referenceBorrow only what the prompt asks for: its style, its layout, a detail
curl https://api.router9.com/v1/image/edits \
  -H "Authorization: Bearer sk-r9k-your-key" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "this living room, in this palette",
    "images": [
      {"url": "https://example.com/room.jpg", "role": "edit"},
      {"url": "https://example.com/palette.jpg", "role": "reference"}
    ]
  }'

A picture already in your library can be named by id instead of URL — {"asset_id": "...", "role": "reference"}. It is read from your own storage, and it is not saved a second time.

A bare URL, and every multipart part, takes the model's default: reference, or the model's only role where it serves just one (so an unlabelled request behaves exactly as it did before roles existed). Not every model serves every role. The catalog's referenceRoles lists them per model, and a role the model cannot serve is refused with 400 image_reference_unsupported. It is not drawn as something else.

Request (multipart/form-data)

The shape an OpenAI SDK sends. image may be repeated, or spelled image[]:

curl https://api.router9.com/v1/image/edits \
  -H "Authorization: Bearer sk-r9k-your-key" \
  -F prompt="make this a watercolor painting" \
  -F [email protected]

Every generation parameter above applies here too. The response is the same shape as /generations.

The gateway downloads referenced URLs itself rather than passing them upstream, so private and internal addresses are refused.

mask is accepted by the parser and refused: no image provider configured on this deployment can inpaint. It is refused rather than ignored, because a dropped mask returns a whole-image edit that looks like a successful masked one.

Over MCP this is a separate image_edit tool, taking the same fields and resolving them the same way. images there is a list of strings — a URL, a data: URL, or a library asset_id — with the roles in image_roles, a second list in the same order. Each picture in the tool's structured result carries its asset_id, as each data item does here.

Upscaling

Enlarge an existing image without redrawing it. No prompt: the subject, composition and detail stay as they are, at a higher resolution.

Endpoint

POST /v1/image/upscale
curl https://api.router9.com/v1/image/upscale \
  -H "Authorization: Bearer sk-r9k-your-key" \
  -H "Content-Type: application/json" \
  -d '{
    "image": "https://example.com/photo.jpg",
    "scale": 4
  }'
ParameterTypeRequiredDescription
imagestringYes, unless parent_asset_idHTTP(S) URL or data: URL of the picture to enlarge
scaleintegerNoEnlargement factor. Providers offer specific factors rather than any number, and one fixed factor is common — a factor the configured provider does not offer is refused rather than rounded. Omit for the provider's own
output_formatstringNopng or webp
modelstringNoAllowlisted image model
project_idstringNoProject to file the result into — an id from GET /v1/projects. Omit it, or send one that is not yours, and it lands in your default project
parent_asset_idstringNoA picture in your library, by id. The result joins its lineage, and with no image the picture is read from your own storage

Background removal

Cut the subject out, leaving the background transparent.

Endpoint

POST /v1/image/background-removal
curl https://api.router9.com/v1/image/background-removal \
  -H "Authorization: Bearer sk-r9k-your-key" \
  -H "Content-Type: application/json" \
  -d '{ "image": "https://example.com/product.jpg" }'
ParameterTypeRequiredDescription
imagestringYes, unless parent_asset_idHTTP(S) URL or data: URL of the picture
output_formatstringNopng or webp. Both carry transparency; jpeg is not offered, because a cut-out subject saved without an alpha channel comes back on a black rectangle
modelstringNoAllowlisted image model
project_idstringNoProject to file the result into — an id from GET /v1/projects. Omit it, or send one that is not yours, and it lands in your default project
parent_asset_idstringNoA picture in your library, by id. The result joins its lineage, and with no image the picture is read from your own storage

Both answer in the same { id, created, model, data, timing } shape the generation endpoints do (each data item carrying its library asset_id), and both count against the Image Generation quota — the allowance is about how many pictures an account produces, not which operation produced them.

Neither ships with a default model. No model reachable through the gateway's bootstrap credentials can upscale or cut out a background, so both answer 503 until an operator configures a provider that can. That is deliberate: a default naming a model which would fail upstream is a promise the API would be making on that model's behalf.

Over MCP these are the image_upscale and image_background_remove tools, with the same image / parent_asset_id fields.

Catalog

GET /v1/public/image-models   → { "models": [...], "planLimits": {...} }

Public. Each model carries a constraints block — the aspect ratios, formats, resolutions and features the provider actually accepts, the roles an input image may play on /edits (referenceRoles) and how many it takes (maxReferences), plus the diffusion knobs (negativePrompt, cfgScale, stylePresets, seed) and what the transform endpoints take (upscaleFactors, transformFormats) — so a client can discover a limit without having to be refused by it. A knob the model does not take is reported as {"accepted": false} with no range, because a range of 0–0 reads as an offer. planLimits is separate and stated per tier, because those are the caps your plan adds on top of the vendor's.

Style vocabulary

GET /v1/image/styles   → { "styles": [...], "models": [...] }

The style_preset values each configured model accepts. styles is the default model's set — what you are validated against when you name no model — and models carries the same per model, because two families from one vendor routinely disagree: one takes seventeen presets and the other takes none.

A model that does not accept style_preset at all publishes an empty list rather than the provider's, so nothing here is a value that would be refused.

Errors

Every image endpoint answers failures in the shared media error envelope. See Errors.

Credits

Image Skill calls consume credits from your plan's monthly credit allowance — the same allowance used for chat completions. There are no separate per-Skill quotas. Image models are priced per image or per megapixel, so n and resolution multiply what a single call costs. See Credits & Usage and Tool Usage & Credits.

On this page