Asset Gen — In-Place Auto-Retry of FAILED Provider Results¶
When a provider call inside AssetGenJobService.runJob comes back FAILED with a transient error, the job is re-run in place with the same jobId before anything is persisted. The user sees one job that either completes or fails once; no intermediate FAILED write, no credit refund, and no notification fires between attempts.
Code: src/assetGen/assetGenFailedRetry.ts. Config: SharedConfigs.ASSET_GEN_FAILED_RETRY_* in src/shared/sharedConfigs.ts.
Where it runs¶
runJob splits into two pieces:
runJob
└─ runWithFailedRetry(job, attemptJob => dispatchToProvider(attemptJob), executionOptions?.failedRetryPolicy)
└─ dispatchToProvider ← the provider switch (Replicate, FalAI, Gemini, OpenAI, …), one call per attempt
└─ B2 upload → post-processing → thumbnail → upsertJob → quality-analysis harness ← runs once, on the final result
dispatchToProvider returns the job with COMPLETED / FAILED / PROCESSING set by the provider service. Only a FAILED return is considered for retry. Errors thrown by a provider service propagate untouched and are not retried.
When a retry happens¶
All of the following must be true (see shouldRetry):
- The provider returned
status === FAILED. - Retry budget remains:
retriesDone < maxRetries. - The job is eligible —
isAssetGenFailedRetryEnabled(job): job.projectLinks.projectIdis set (project-linked jobs only, no ad-hoc / playground jobs), andjob.modelConfig.outputType === 'IMAGE'.- The error is retryable —
isRetryableAssetGenError(job.error): error.messagematches none ofASSET_GEN_FAILED_NON_RETRYABLE_ERROR_PATTERNS(checked first), and- matches at least one of
ASSET_GEN_FAILED_RETRYABLE_ERROR_PATTERNS.
Eligibility conditions are AND-ed inside isAssetGenFailedRetryEnabled; add new gates there, not at the call site.
Retryable patterns (case-insensitive)¶
Taken from provider errors observed on failed workbench image jobs that succeeded on a plain re-run:
| Pattern | Typical source |
|---|---|
^fetch failed$ |
Network drop to Gemini |
^Connection error\.?$ |
OpenAI SDK network drop (surfaced as OPENAI_ERROR by OpenAIService.createOpenAIPrediction) |
"status": "UNAVAILABLE" |
Gemini 503 |
RESOURCE_EXHAUSTED |
Gemini 429 |
Image download failed \(5\d\d |
5xx fetching our own B2 input image |
Read timed out |
Provider-side timeout fetching an input |
\b5\d\d Server Error |
Generic upstream 5xx |
Unable to process input image |
Transient input-handling failure |
No image generated from Google Gemini API |
Empty Gemini response (unless a safety reason is attached, see below) |
Non-retryable patterns (checked first, case-insensitive)¶
| Pattern | Why |
|---|---|
finishReason: \w*SAFETY |
Safety block is deterministic; same prompt will block again |
PROHIBITED_CONTENT |
Same |
blockReason: |
Prompt-level block from promptFeedback |
GoogleGeminiService now appends finishReason / blockReason to its "No image generated" error, so a safety-blocked empty response is told apart from a genuinely empty one and is not retried.
Policy and limits¶
// Default
SharedConfigs.ASSET_GEN_FAILED_RETRY_POLICY = { maxRetries: 1, delaysMs: [0] };
// Hard ceiling applied to every policy, including per-caller overrides
SharedConfigs.ASSET_GEN_FAILED_RETRY_MAX_RETRIES_CAP = 3;
maxRetries— number of re-runs after the first attempt. A job can run at mostmaxRetries + 1times.delaysMs— retry N waitsdelaysMs[N-1]ms before re-dispatching. If the array is shorter thanmaxRetries, the last entry is reused. An empty array means no delay.- Cap — a policy asking for more than the cap is clamped and logs
AssetGenFailedRetryPolicyClamped. The cap exists because retries run inside the originalrunJobpromise while the stored status is stillPROCESSING;AssetGenJobTimeoutServicemarks jobsFAILEDafterSTUCK_JOB_TIMEOUT_MS(30 min), so the retry loop must always finish well before that sweep could race the final upsert.
Per-use-case override¶
Callers of runJob / createJob can pass their own policy through execution options:
await AssetGenJobService.runJob(userId, job, characterModelConfigList, {
failedRetryPolicy: { maxRetries: 2, delaysMs: [500, 2000] },
});
// Disable in-place retry for one call
await AssetGenJobService.runJob(userId, job, undefined, { failedRetryPolicy: { maxRetries: 0, delaysMs: [] } });
The override tunes the budget only. The job still has to pass isAssetGenFailedRetryEnabled and the error still has to be retryable.
What is recorded on the job¶
Retries write jobMetaData.autoRetry (type AssetGenFailedRetryState). It is absent when no retry happened.
jobMetaData.autoRetry = {
attempts: 1, // number of re-runs performed
errors: [ // the error of each attempt that was retried, oldest first
{ code: 'GOOGLE_GEMINI_ERROR', message: 'fetch failed', at: '2026-09-22T10:15:03.412Z' },
],
};
Before each re-run the failed attempt's error field is removed and status is reset to PROCESSING, so the provider sees the same job shape it saw the first time. History is kept in place because every retry re-runs the identical job (same model, same params). See the "Future work" section for when that should change.
If the final attempt also fails, the job is persisted as FAILED with the last error in job.error and the earlier ones in autoRetry.errors.
Logging¶
Both events use logger.warn with an errorType from ErrorType (src/shared/errorTypes.ts):
| Event | errorType |
Fields |
|---|---|---|
| A retry is about to run | AssetGenFailedRetry |
jobId, retryNumber, maxRetries, provider, modelConfigId, errorMessage |
A caller's maxRetries exceeded the cap |
AssetGenFailedRetryPolicyClamped |
jobId, requestedMaxRetries, cap |
To measure how often retries rescue a job, correlate AssetGenFailedRetry logs against the final status of the same jobId, or query jobs where jobMetaData.autoRetry exists.
What is not covered¶
- Async providers (webhook / poll paths).
runWithFailedRetryonly sees whatdispatchToProviderreturns at submit time. If a provider returnsPROCESSINGand later reportsFAILEDvia webhook (WebhooksService.processWebhookAssetGenJobResponse) or viapollJobStatus, that failure is persisted as-is. Both sites carry an@todo (assetGenFailedRetry)comment describing how a re-dispatch would work. Not needed today because every project-linkedIMAGEjob runs synchronously (Gemini / OpenAI, and FalAI / Replicate withwaitForCompletion). - Non-image outputs (video, audio, 3D) and jobs without a
projectId. - Thrown errors. Only a returned
FAILEDjob is retried. - Mock mode.
completeJobWithMockResultreturns before the retry wrapper.
Future work¶
- Archive-per-attempt history. In-place
autoRetryis correct only while a retry re-runs the identical job. If a retry ever changes model or params (e.g. fall back to a different provider), archive the failed attempt as its own job document with a freshjobIdand aretryOfpointer, kept out oflistJobsand analytics, instead of growingautoRetry. Decided Sept 2026; see the@todoinrunWithFailedRetry. - Webhook / poll retry. See the
@todoonprocessWebhookAssetGenJobResponseand inpollJobStatus. - Pattern maintenance. New retryable messages should be added to
ASSET_GEN_FAILED_RETRYABLE_ERROR_PATTERNSonly after being observed to succeed on a plain re-run; deterministic failures go in the non-retryable list.
Tests¶
tests/assetGen/assetGenFailedRetry.test.ts— case-insensitive pattern matching (retryable vs. safety-blocked), and cap clamping of an oversized policy (attemptcalled cap + 1 times,autoRetry.attempts === cap).tests/assetGen/assetGen.service.test.ts—runJobend-to-end: afetch failedGemini result is re-dispatched once, the completed job is persisted once withautoRetry.attempts === 1.