Skip to content

textGen — Third-Party Text Generation API

Synchronous LLM text generation for trusted internal Scenarix backends. One endpoint, two providers: BEDROCK (Claude and anything else Converse serves) and BEDROCK_MANTLE (OpenAI models such as GPT-5.4). Both are the same AWS account and credential.

Endpoint

POST /api/apiViaKey/textGen/generate

API-key auth only — there is no JaduAuth mount. Studio code should call BedrockService.generateText or BedrockMantleService.generateText directly rather than going through HTTP.

Header Required Notes
x-api-key yes Must appear in the comma-separated API_REQUEST_KEYS env var
x-user-id yes Recorded as context.userId in analytics. Caller-asserted — trusted callers only
Content-Type yes application/json

Request

The payload is deliberately shaped like an assetGen model config, so a consumer that already builds assetGen payloads builds these the same way. The types are textGen's own (textGen.types.ts) and are independent of the assetGen job system. Prompts ride in inputs[]; jsonSchema lives in modelMetaData because it is output-format config, not a user input.

Every field below is required except systemPrompt. Output is always structured JSON, so there is no plain-text mode to reason about on either side. id, name, modelTitle, outputType, and pollingIntervalMs are accepted but unused. Unrecognised inputs[] entries (e.g. temperature, images) are ignored, not rejected.

curl -X POST https://<host>/api/apiViaKey/textGen/generate \
  -H "x-api-key: $API_KEY" \
  -H "x-user-id: 8f2c1b0e-1234-4a56-9876-abcdef012345" \
  -H "Content-Type: application/json" \
  -d '{
    "modelConfig": {
      "id": "text-claude-haiku-4-5",
      "name": "claude-haiku-4-5",
      "modelTitle": "Claude Haiku 4.5",
      "provider": "BEDROCK",
      "modelType": "TEXT_GEN",
      "outputType": "TEXT",
      "pollingIntervalMs": 0,
      "modelMetaData": {
        "bedrockModelId": "arn:aws:bedrock:us-east-1:408921634255:inference-profile/global.anthropic.claude-haiku-4-5-20251001-v1:0",
        "jsonSchema": {
          "schema": { "type": "object", "properties": { "logline": { "type": "string" } }, "required": ["logline"] },
          "name": "LoglineResult"
        }
      },
      "inputs": [
        { "id": "prompt", "label": "Prompt", "type": "TEXT", "isRequired": true, "value": "Write a logline for a heist film." },
        { "id": "systemPrompt", "label": "System Prompt", "type": "TEXT", "isRequired": false, "value": "You are a screenwriter." }
      ]
    },
    "event": "PARTNER_LOGLINE_GENERATION",
    "source": "STORY_DESK",
    "context": { "projectId": "proj-1", "storyId": "story-1" }
  }'

jsonSchema is mandatory — omitting it is a 400. Every response is structured JSON; there is no plain-text output mode. systemPrompt is optional, but when sent it must be a non-empty string; omitting it means no system role is sent upstream at all.

event is a free-form string, recorded as sent — use whatever names your own use cases need. source is different: it must be one of the LogSource members in src/analyticsLogger.ts, so the set of calling systems stays a discoverable registry and a typo cannot silently create a second analytics bucket. A new calling system needs a one-line addition to that enum.

The model id field is not validated against an allowlist — see "Choosing a model" below.

context is free-form JSON, also recorded as sent — send whatever identifiers your system uses. A userId key inside context is ignored: the authenticated x-user-id header always wins.

If inputs[] contains more than one entry with id: "prompt" (likewise systemPrompt), the first match wins; the duplicates are ignored rather than rejected.

Choosing a model

The model is a caller parameter, not a Studio deploy. Any non-empty model id is accepted and passed upstream as sent, so switching models — including across vendors — is a change to your own payload.

First pick the provider, because it determines which field carries the model id:

Provider Model id field Serves Transport
BEDROCK bedrockModelId Claude, Llama, gpt-oss, … Converse on bedrock-runtime
BEDROCK_MANTLE mantleModelId OpenAI GPT-5.4 Responses API on bedrock-mantle

The two catalogues barely overlap, and a model on one is generally not reachable on the other — see "OpenAI models" below. Sending bedrockModelId with provider: "BEDROCK_MANTLE" (or the reverse) is a 400 naming the field, not a silent upstream failure.

For BEDROCK, both id forms work:

Form Example
Inference-profile ARN arn:aws:bedrock:us-east-1:408921634255:inference-profile/global.anthropic.claude-haiku-4-5-20251001-v1:0
Bare vendor.model id openai.gpt-oss-120b-1:0

The id must be enabled for the AWS account. An id that is wrong, or real but not provisioned, fails upstream as a 400 carrying the provider's own message about it — so a typo surfaces immediately in testing rather than being pre-screened here.

BedrockModelId and BedrockMantleModelId in src/shared/sharedTypes.ts remain as convenience registries for Studio's own call sites; neither is the set of permitted values.

OpenAI models (GPT-5.4)

GPT-5.4 runs through provider: "BEDROCK_MANTLE", not BEDROCK. Per its Bedrock model card it supports neither the bedrock-runtime endpoint nor the Converse API — only the Responses API on bedrock-mantle, at openai/v1/responses. A BEDROCK request naming it therefore cannot succeed, whatever the id's form; in particular there is no inference-profile ARN for it (In-Region only, no Geo or Global routing).

"GPT-5.4 Low" is the model id plus a reasoningEffort:

"modelConfig": {
  "provider": "BEDROCK_MANTLE",
  "modelType": "TEXT_GEN",
  "modelMetaData": {
    "mantleModelId": "openai.gpt-5.4",
    "reasoningEffort": "low",
    "jsonSchema": { "schema": { "…": "…" }, "name": "LoglineResult" }
  },
  "inputs": [
    { "id": "prompt", "value": "Write a logline for a heist film." },
    { "id": "systemPrompt", "value": "You are a screenwriter." }
  ]
}

Everything else is identical to a Claude call: same auth, same required jsonSchema, same optional systemPrompt, same response envelope, same error mapping. Three differences worth knowing:

  • jsonSchema.name is required upstream by the Responses API, though optional in this contract. A request without one is sent as structured_output rather than rejected.
  • strict schema mode is not enabled. It would require additionalProperties: false plus every property listed in required, which existing caller schemas do not guarantee.
  • additionalProperties: false is not required here, but is on BEDROCK. Converse rejects an object schema without it (For 'object' type, 'additionalProperties' must be explicitly set to false); mantle accepts either. A schema that works here may need that key added before the same payload works on BEDROCK.

Credentials and region are shared with the Converse path — AWS_BEARER_TOKEN_BEDROCK and BEDROCK_REGION (us-east-1, which mantle supports). No separate key to provision.

reasoningEffort

Optional, one of none / minimal / low / medium / high, sent inside modelMetaData. Works on both providers, carried in whichever form the transport expects — Converse's additionalModelRequestFields.reasoning_effort, or Responses' reasoning.effort.

Only reasoning models read it — Claude ignores it — so it is safe to omit, which is the common case for BEDROCK and unusual for BEDROCK_MANTLE. none means "send no reasoning field at all" rather than a value to forward, because OpenAI's models reject it; that is true on both providers, so none means one thing everywhere.

Cost tracking

USD cost is resolved by looking the model id up in SharedConfigs.LLMCostConfig. A model with no entry there still runs, and its token usage is still logged, but with no cost field — the server logs No LLMCostConfig entry for model naming the id. If you adopt a model for sustained use, add its pricing so spend keeps rolling up.

Note that Bedrock's rates for a model are its own, not the vendor's direct rates: openai.gpt-5.4 is priced $2.75 / $0.275 / $16.50 per 1M (in / cached read / out) against OpenAI direct's $2.50 / $0.25 / $15.00, so it carries a separate entry from the gpt-5.4 used by Studio's own OpenAI-direct call sites. GovCloud is billed ~20% higher again and is not modelled.

Response

{
  "isSuccess": true,
  "message": "Text generated successfully",
  "data": {
    "outputText": "…",
    "usage": { "input_tokens": 1200, "output_tokens": 340, "total_tokens": 1540 },
    "logId": "b3f1…"
  }
}

outputText is always a JSON string matching the supplied jsonSchema — parse it client-side.

USD cost is computed from SharedConfigs.LLMCostConfig and stored on the analytics log, not returned. Use logId to correlate against the analyticsLogs collection.

Errors

All errors use the standard envelope: { "isSuccess": false, "message": "…", "data": {} }.

Status Cause
400 Validation failure — missing/empty prompt, an empty systemPrompt when one is sent, unknown provider or wrong modelType, an empty model id or one sent under the wrong provider's field name, reasoningEffort outside the supported set, missing jsonSchema, missing event, or a source outside the LogSource enum. Also an upstream rejection of the request itself — an unknown or not-enabled model id, malformed jsonSchema, or an over-long prompt. Not retryable
401 Missing or invalid x-api-key, or missing x-user-id
502 Upstream failure, throttling (429) or timeout (408). This is the retry signal — retry with backoff
500 Unexpected server error

Error messages are intentionally generic and never echo the upstream provider's message, which can contain server-internal detail. The full upstream status and message are recorded in the server logs under errorType: TextGenUpstreamError, along with userId, provider and modelId; correlate by logId (on success) or by request timestamp and route.

Design notes

  • Synchronous. No job document, no collection, no polling, no webhooks, no realtime. Server-side timeout is SharedConstants.TEXT_GEN_TIMEOUT_MS (120s); executeWithTimeoutRetry retries once, so worst-case wall time is roughly double that.
  • No credits. Token usage and USD cost are recorded in analyticsLogs only.
  • Not exposed: images[], previousMessages[], temperature, maxTokens. The underlying services support them; this API does not surface them in v1.
  • The model is caller-chosen, by design. The model id was originally validated against the BedrockModelId enum, which meant a calling backend could not so much as trial a different model without a Studio code change, review and deploy. Bedrock's catalogue changes on AWS's schedule and now spans multiple vendors, so the enum was a standing source of friction with little safety return: an unknown or not-enabled id fails upstream regardless, with the provider's own message naming it. Only an empty id is rejected here, because that alone would build a request URL with no model segment and fail for a reason unrelated to the cause. The trade-off accepted is that an unpriced model logs usage without cost — surfaced as a warning rather than silently.
  • Provider is a closed enum, unlike the model id. Each provider is a distinct transport with its own client and request shape, so an unknown value has no code path to serve it. This is also why the two model id fields are named differently: the commonest caller mistake is reusing a Converse payload and only swapping provider, and a shared field name would forward that to an upstream 400 instead of failing here.
  • BEDROCK_MANTLE is separate from BEDROCK rather than a model id within it. Same AWS account and credential, but a different endpoint, wire protocol and capability set — GPT-5.4 supports neither bedrock-runtime nor Converse. It is equally not part of OpenAIService, whose client is a static singleton: pointing that at mantle would redirect every OpenAI call in the studio.
  • Do not set OPENAI_BASE_URL to the mantle host. Bedrock's GPT-5.4 setup instructions suggest exporting it, but OpenAIService takes its baseURL from the SDK default, which reads that variable — so setting it redirects every direct OpenAI call in the studio to Bedrock, where the OpenAI key is not a valid credential (401). This route does not need it: BedrockMantleService sets its own baseURL explicitly.
  • No rate limiting and no per-key identity. x-api-key is a flat env allowlist and x-user-id is caller-asserted, so a valid key can act as any user. Acceptable for trusted internal infrastructure; must be revisited before any external partner is given a key.
  • source is recorded but not yet consumed. It is written to every analyticsLogs document, but no dashboard or aggregation reads it today — the existing Text-LLM cost rollups group by event. Third-party spend therefore appears blended into the studio's own totals until something groups by source. It is validated against the LogSource enum, so it cannot be a typo, but it is still caller-chosen (a caller may legitimately send SCENARIX_STUDIO), so it must never be relied on as a trust boundary or as proof of origin. Deriving it from the API key would fix that.

Full design: 2026-08-11-textgen-third-party-api-design.md, which lives outside this repository — in the parent studio/ workspace at studio/docs/superpowers/specs/, one level above this git root (studio-backend/). A clone of studio-backend alone will not contain it.