textGen — Third-Party Text Generation API¶
Synchronous LLM text generation for trusted internal Scenarix backends. One endpoint, two providers:
BEDROCK (Claude and anything else Converse serves) and BEDROCK_MANTLE (OpenAI models such as
GPT-5.4). Both are the same AWS account and credential.
Endpoint¶
POST /api/apiViaKey/textGen/generate
API-key auth only — there is no JaduAuth mount. Studio code should call BedrockService.generateText
or BedrockMantleService.generateText directly rather than going through HTTP.
| Header | Required | Notes |
|---|---|---|
x-api-key |
yes | Must appear in the comma-separated API_REQUEST_KEYS env var |
x-user-id |
yes | Recorded as context.userId in analytics. Caller-asserted — trusted callers only |
Content-Type |
yes | application/json |
Request¶
The payload is deliberately shaped like an assetGen model config, so a consumer that already builds
assetGen payloads builds these the same way. The types are textGen's own (textGen.types.ts) and are
independent of the assetGen job system. Prompts ride in inputs[]; jsonSchema lives in
modelMetaData because it is output-format config, not a user input.
Every field below is required except systemPrompt. Output is always structured JSON, so there is no
plain-text mode to reason about on either side. id, name, modelTitle, outputType, and
pollingIntervalMs are accepted but unused.
Unrecognised inputs[] entries (e.g. temperature, images) are ignored, not rejected.
curl -X POST https://<host>/api/apiViaKey/textGen/generate \
-H "x-api-key: $API_KEY" \
-H "x-user-id: 8f2c1b0e-1234-4a56-9876-abcdef012345" \
-H "Content-Type: application/json" \
-d '{
"modelConfig": {
"id": "text-claude-haiku-4-5",
"name": "claude-haiku-4-5",
"modelTitle": "Claude Haiku 4.5",
"provider": "BEDROCK",
"modelType": "TEXT_GEN",
"outputType": "TEXT",
"pollingIntervalMs": 0,
"modelMetaData": {
"bedrockModelId": "arn:aws:bedrock:us-east-1:408921634255:inference-profile/global.anthropic.claude-haiku-4-5-20251001-v1:0",
"jsonSchema": {
"schema": { "type": "object", "properties": { "logline": { "type": "string" } }, "required": ["logline"] },
"name": "LoglineResult"
}
},
"inputs": [
{ "id": "prompt", "label": "Prompt", "type": "TEXT", "isRequired": true, "value": "Write a logline for a heist film." },
{ "id": "systemPrompt", "label": "System Prompt", "type": "TEXT", "isRequired": false, "value": "You are a screenwriter." }
]
},
"event": "PARTNER_LOGLINE_GENERATION",
"source": "STORY_DESK",
"context": { "projectId": "proj-1", "storyId": "story-1" }
}'
jsonSchema is mandatory — omitting it is a 400. Every response is structured JSON; there is no
plain-text output mode. systemPrompt is optional, but when sent it must be a non-empty string;
omitting it means no system role is sent upstream at all.
event is a free-form string, recorded as sent — use whatever names your own use cases need. source
is different: it must be one of the LogSource members in src/analyticsLogger.ts, so the set of
calling systems stays a discoverable registry and a typo cannot silently create a second analytics
bucket. A new calling system needs a one-line addition to that enum.
The model id field is not validated against an allowlist — see "Choosing a model" below.
context is free-form JSON, also recorded as sent — send whatever identifiers your system uses. A
userId key inside context is ignored: the authenticated x-user-id header always wins.
If inputs[] contains more than one entry with id: "prompt" (likewise systemPrompt), the first
match wins; the duplicates are ignored rather than rejected.
Choosing a model¶
The model is a caller parameter, not a Studio deploy. Any non-empty model id is accepted and passed upstream as sent, so switching models — including across vendors — is a change to your own payload.
First pick the provider, because it determines which field carries the model id:
| Provider | Model id field | Serves | Transport |
|---|---|---|---|
BEDROCK |
bedrockModelId |
Claude, Llama, gpt-oss, … |
Converse on bedrock-runtime |
BEDROCK_MANTLE |
mantleModelId |
OpenAI GPT-5.4 | Responses API on bedrock-mantle |
The two catalogues barely overlap, and a model on one is generally not reachable on the other —
see "OpenAI models" below. Sending bedrockModelId with provider: "BEDROCK_MANTLE" (or the
reverse) is a 400 naming the field, not a silent upstream failure.
For BEDROCK, both id forms work:
| Form | Example |
|---|---|
| Inference-profile ARN | arn:aws:bedrock:us-east-1:408921634255:inference-profile/global.anthropic.claude-haiku-4-5-20251001-v1:0 |
Bare vendor.model id |
openai.gpt-oss-120b-1:0 |
The id must be enabled for the AWS account. An id that is wrong, or real but not provisioned, fails upstream as a 400 carrying the provider's own message about it — so a typo surfaces immediately in testing rather than being pre-screened here.
BedrockModelId and BedrockMantleModelId in src/shared/sharedTypes.ts remain as convenience
registries for Studio's own call sites; neither is the set of permitted values.
OpenAI models (GPT-5.4)¶
GPT-5.4 runs through provider: "BEDROCK_MANTLE", not BEDROCK. Per its
Bedrock model card
it supports neither the bedrock-runtime endpoint nor the Converse API — only the Responses API on
bedrock-mantle, at openai/v1/responses. A BEDROCK request naming it therefore cannot succeed,
whatever the id's form; in particular there is no inference-profile ARN for it (In-Region only, no
Geo or Global routing).
"GPT-5.4 Low" is the model id plus a reasoningEffort:
"modelConfig": {
"provider": "BEDROCK_MANTLE",
"modelType": "TEXT_GEN",
"modelMetaData": {
"mantleModelId": "openai.gpt-5.4",
"reasoningEffort": "low",
"jsonSchema": { "schema": { "…": "…" }, "name": "LoglineResult" }
},
"inputs": [
{ "id": "prompt", "value": "Write a logline for a heist film." },
{ "id": "systemPrompt", "value": "You are a screenwriter." }
]
}
Everything else is identical to a Claude call: same auth, same required jsonSchema, same optional
systemPrompt, same response envelope, same error mapping. Three differences worth knowing:
jsonSchema.nameis required upstream by the Responses API, though optional in this contract. A request without one is sent asstructured_outputrather than rejected.strictschema mode is not enabled. It would requireadditionalProperties: falseplus every property listed inrequired, which existing caller schemas do not guarantee.additionalProperties: falseis not required here, but is onBEDROCK. Converse rejects an object schema without it (For 'object' type, 'additionalProperties' must be explicitly set to false); mantle accepts either. A schema that works here may need that key added before the same payload works onBEDROCK.
Credentials and region are shared with the Converse path — AWS_BEARER_TOKEN_BEDROCK and
BEDROCK_REGION (us-east-1, which mantle supports). No separate key to provision.
reasoningEffort¶
Optional, one of none / minimal / low / medium / high, sent inside modelMetaData. Works on
both providers, carried in whichever form the transport expects — Converse's
additionalModelRequestFields.reasoning_effort, or Responses' reasoning.effort.
Only reasoning models read it — Claude ignores it — so it is safe to omit, which is the common case
for BEDROCK and unusual for BEDROCK_MANTLE. none means "send no reasoning field at all" rather
than a value to forward, because OpenAI's models reject it; that is true on both providers, so none
means one thing everywhere.
Cost tracking¶
USD cost is resolved by looking the model id up in SharedConfigs.LLMCostConfig. A model with no
entry there still runs, and its token usage is still logged, but with no cost field — the server
logs No LLMCostConfig entry for model naming the id. If you adopt a model for sustained use, add its
pricing so spend keeps rolling up.
Note that Bedrock's rates for a model are its own, not the vendor's direct rates: openai.gpt-5.4 is
priced $2.75 / $0.275 / $16.50 per 1M (in / cached read / out) against OpenAI direct's
$2.50 / $0.25 / $15.00, so it carries a separate entry from the gpt-5.4 used by Studio's own
OpenAI-direct call sites. GovCloud is billed ~20% higher again and is not modelled.
Response¶
{
"isSuccess": true,
"message": "Text generated successfully",
"data": {
"outputText": "…",
"usage": { "input_tokens": 1200, "output_tokens": 340, "total_tokens": 1540 },
"logId": "b3f1…"
}
}
outputText is always a JSON string matching the supplied jsonSchema — parse it client-side.
USD cost is computed from SharedConfigs.LLMCostConfig and stored on the analytics log, not returned.
Use logId to correlate against the analyticsLogs collection.
Errors¶
All errors use the standard envelope: { "isSuccess": false, "message": "…", "data": {} }.
| Status | Cause |
|---|---|
| 400 | Validation failure — missing/empty prompt, an empty systemPrompt when one is sent, unknown provider or wrong modelType, an empty model id or one sent under the wrong provider's field name, reasoningEffort outside the supported set, missing jsonSchema, missing event, or a source outside the LogSource enum. Also an upstream rejection of the request itself — an unknown or not-enabled model id, malformed jsonSchema, or an over-long prompt. Not retryable |
| 401 | Missing or invalid x-api-key, or missing x-user-id |
| 502 | Upstream failure, throttling (429) or timeout (408). This is the retry signal — retry with backoff |
| 500 | Unexpected server error |
Error messages are intentionally generic and never echo the upstream provider's message, which can
contain server-internal detail. The full upstream status and message are recorded in the server logs
under errorType: TextGenUpstreamError, along with userId, provider and modelId; correlate by
logId (on success) or by request timestamp and route.
Design notes¶
- Synchronous. No job document, no collection, no polling, no webhooks, no realtime. Server-side
timeout is
SharedConstants.TEXT_GEN_TIMEOUT_MS(120s);executeWithTimeoutRetryretries once, so worst-case wall time is roughly double that. - No credits. Token usage and USD cost are recorded in
analyticsLogsonly. - Not exposed:
images[],previousMessages[],temperature,maxTokens. The underlying services support them; this API does not surface them in v1. - The model is caller-chosen, by design. The model id was originally validated against the
BedrockModelIdenum, which meant a calling backend could not so much as trial a different model without a Studio code change, review and deploy. Bedrock's catalogue changes on AWS's schedule and now spans multiple vendors, so the enum was a standing source of friction with little safety return: an unknown or not-enabled id fails upstream regardless, with the provider's own message naming it. Only an empty id is rejected here, because that alone would build a request URL with no model segment and fail for a reason unrelated to the cause. The trade-off accepted is that an unpriced model logs usage without cost — surfaced as a warning rather than silently. - Provider is a closed enum, unlike the model id. Each provider is a distinct transport with its
own client and request shape, so an unknown value has no code path to serve it. This is also why the
two model id fields are named differently: the commonest caller mistake is reusing a Converse
payload and only swapping
provider, and a shared field name would forward that to an upstream 400 instead of failing here. BEDROCK_MANTLEis separate fromBEDROCKrather than a model id within it. Same AWS account and credential, but a different endpoint, wire protocol and capability set — GPT-5.4 supports neitherbedrock-runtimenor Converse. It is equally not part ofOpenAIService, whose client is a static singleton: pointing that at mantle would redirect every OpenAI call in the studio.- Do not set
OPENAI_BASE_URLto the mantle host. Bedrock's GPT-5.4 setup instructions suggest exporting it, butOpenAIServicetakes its baseURL from the SDK default, which reads that variable — so setting it redirects every direct OpenAI call in the studio to Bedrock, where the OpenAI key is not a valid credential (401). This route does not need it:BedrockMantleServicesets its own baseURL explicitly. - No rate limiting and no per-key identity.
x-api-keyis a flat env allowlist andx-user-idis caller-asserted, so a valid key can act as any user. Acceptable for trusted internal infrastructure; must be revisited before any external partner is given a key. sourceis recorded but not yet consumed. It is written to everyanalyticsLogsdocument, but no dashboard or aggregation reads it today — the existing Text-LLM cost rollups group byevent. Third-party spend therefore appears blended into the studio's own totals until something groups bysource. It is validated against theLogSourceenum, so it cannot be a typo, but it is still caller-chosen (a caller may legitimately sendSCENARIX_STUDIO), so it must never be relied on as a trust boundary or as proof of origin. Deriving it from the API key would fix that.
Full design: 2026-08-11-textgen-third-party-api-design.md, which lives outside this repository
— in the parent studio/ workspace at studio/docs/superpowers/specs/, one level above this git
root (studio-backend/). A clone of studio-backend alone will not contain it.