Skip to content

Asset Gen — In-Place Auto-Retry of FAILED Provider Results

When a provider call inside AssetGenJobService.runJob comes back FAILED with a transient error, the job is re-run in place with the same jobId before anything is persisted. The user sees one job that either completes or fails once; no intermediate FAILED write, no credit refund, and no notification fires between attempts.

Code: src/assetGen/assetGenFailedRetry.ts. Config: SharedConfigs.ASSET_GEN_FAILED_RETRY_* in src/shared/sharedConfigs.ts.


Where it runs

runJob splits into two pieces:

runJob
  └─ runWithFailedRetry(job, attemptJob => dispatchToProvider(attemptJob), executionOptions?.failedRetryPolicy)
       └─ dispatchToProvider   ← the provider switch (Replicate, FalAI, Gemini, OpenAI, …), one call per attempt
  └─ B2 upload → post-processing → thumbnail → upsertJob → quality-analysis harness   ← runs once, on the final result

dispatchToProvider returns the job with COMPLETED / FAILED / PROCESSING set by the provider service. Only a FAILED return is considered for retry. Errors thrown by a provider service propagate untouched and are not retried.

When a retry happens

All of the following must be true (see shouldRetry):

  1. The provider returned status === FAILED.
  2. Retry budget remains: retriesDone < maxRetries.
  3. The job is eligible — isAssetGenFailedRetryEnabled(job):
  4. job.projectLinks.projectId is set (project-linked jobs only, no ad-hoc / playground jobs), and
  5. job.modelConfig.outputType === 'IMAGE'.
  6. The error is retryable — isRetryableAssetGenError(job.error):
  7. error.message matches none of ASSET_GEN_FAILED_NON_RETRYABLE_ERROR_PATTERNS (checked first), and
  8. matches at least one of ASSET_GEN_FAILED_RETRYABLE_ERROR_PATTERNS.

Eligibility conditions are AND-ed inside isAssetGenFailedRetryEnabled; add new gates there, not at the call site.

Retryable patterns (case-insensitive)

Taken from provider errors observed on failed workbench image jobs that succeeded on a plain re-run:

Pattern Typical source
^fetch failed$ Network drop to Gemini
^Connection error\.?$ OpenAI SDK network drop (surfaced as OPENAI_ERROR by OpenAIService.createOpenAIPrediction)
"status": "UNAVAILABLE" Gemini 503
RESOURCE_EXHAUSTED Gemini 429
Image download failed \(5\d\d 5xx fetching our own B2 input image
Read timed out Provider-side timeout fetching an input
\b5\d\d Server Error Generic upstream 5xx
Unable to process input image Transient input-handling failure
No image generated from Google Gemini API Empty Gemini response (unless a safety reason is attached, see below)

Non-retryable patterns (checked first, case-insensitive)

Pattern Why
finishReason: \w*SAFETY Safety block is deterministic; same prompt will block again
PROHIBITED_CONTENT Same
blockReason: Prompt-level block from promptFeedback

GoogleGeminiService now appends finishReason / blockReason to its "No image generated" error, so a safety-blocked empty response is told apart from a genuinely empty one and is not retried.

Policy and limits

// Default
SharedConfigs.ASSET_GEN_FAILED_RETRY_POLICY = { maxRetries: 1, delaysMs: [0] };

// Hard ceiling applied to every policy, including per-caller overrides
SharedConfigs.ASSET_GEN_FAILED_RETRY_MAX_RETRIES_CAP = 3;
  • maxRetries — number of re-runs after the first attempt. A job can run at most maxRetries + 1 times.
  • delaysMs — retry N waits delaysMs[N-1] ms before re-dispatching. If the array is shorter than maxRetries, the last entry is reused. An empty array means no delay.
  • Cap — a policy asking for more than the cap is clamped and logs AssetGenFailedRetryPolicyClamped. The cap exists because retries run inside the original runJob promise while the stored status is still PROCESSING; AssetGenJobTimeoutService marks jobs FAILED after STUCK_JOB_TIMEOUT_MS (30 min), so the retry loop must always finish well before that sweep could race the final upsert.

Per-use-case override

Callers of runJob / createJob can pass their own policy through execution options:

await AssetGenJobService.runJob(userId, job, characterModelConfigList, {
  failedRetryPolicy: { maxRetries: 2, delaysMs: [500, 2000] },
});

// Disable in-place retry for one call
await AssetGenJobService.runJob(userId, job, undefined, { failedRetryPolicy: { maxRetries: 0, delaysMs: [] } });

The override tunes the budget only. The job still has to pass isAssetGenFailedRetryEnabled and the error still has to be retryable.

What is recorded on the job

Retries write jobMetaData.autoRetry (type AssetGenFailedRetryState). It is absent when no retry happened.

jobMetaData.autoRetry = {
  attempts: 1,                       // number of re-runs performed
  errors: [                          // the error of each attempt that was retried, oldest first
    { code: 'GOOGLE_GEMINI_ERROR', message: 'fetch failed', at: '2026-09-22T10:15:03.412Z' },
  ],
};

Before each re-run the failed attempt's error field is removed and status is reset to PROCESSING, so the provider sees the same job shape it saw the first time. History is kept in place because every retry re-runs the identical job (same model, same params). See the "Future work" section for when that should change.

If the final attempt also fails, the job is persisted as FAILED with the last error in job.error and the earlier ones in autoRetry.errors.

Logging

Both events use logger.warn with an errorType from ErrorType (src/shared/errorTypes.ts):

Event errorType Fields
A retry is about to run AssetGenFailedRetry jobId, retryNumber, maxRetries, provider, modelConfigId, errorMessage
A caller's maxRetries exceeded the cap AssetGenFailedRetryPolicyClamped jobId, requestedMaxRetries, cap

To measure how often retries rescue a job, correlate AssetGenFailedRetry logs against the final status of the same jobId, or query jobs where jobMetaData.autoRetry exists.

What is not covered

  • Async providers (webhook / poll paths). runWithFailedRetry only sees what dispatchToProvider returns at submit time. If a provider returns PROCESSING and later reports FAILED via webhook (WebhooksService.processWebhookAssetGenJobResponse) or via pollJobStatus, that failure is persisted as-is. Both sites carry an @todo (assetGenFailedRetry) comment describing how a re-dispatch would work. Not needed today because every project-linked IMAGE job runs synchronously (Gemini / OpenAI, and FalAI / Replicate with waitForCompletion).
  • Non-image outputs (video, audio, 3D) and jobs without a projectId.
  • Thrown errors. Only a returned FAILED job is retried.
  • Mock mode. completeJobWithMockResult returns before the retry wrapper.

Future work

  • Archive-per-attempt history. In-place autoRetry is correct only while a retry re-runs the identical job. If a retry ever changes model or params (e.g. fall back to a different provider), archive the failed attempt as its own job document with a fresh jobId and a retryOf pointer, kept out of listJobs and analytics, instead of growing autoRetry. Decided Sept 2026; see the @todo in runWithFailedRetry.
  • Webhook / poll retry. See the @todo on processWebhookAssetGenJobResponse and in pollJobStatus.
  • Pattern maintenance. New retryable messages should be added to ASSET_GEN_FAILED_RETRYABLE_ERROR_PATTERNS only after being observed to succeed on a plain re-run; deterministic failures go in the non-retryable list.

Tests

  • tests/assetGen/assetGenFailedRetry.test.ts — case-insensitive pattern matching (retryable vs. safety-blocked), and cap clamping of an oversized policy (attempt called cap + 1 times, autoRetry.attempts === cap).
  • tests/assetGen/assetGen.service.test.ts — runJob end-to-end: a fetch failed Gemini result is re-dispatched once, the completed job is persisted once with autoRetry.attempts === 1.