Skip to content

Commit 0b35cd7

Browse files
committed
fix(agent): retry the selected model before falling back
Retry on fail used to wrap the whole fallback chain, so tries 3 with fallbacks B and C ran A, B, C three times over. A builder who lists fallbacks wants the selected model retried and the fallbacks tried once each after its last try fails, which is also how LiteLLM orders retries and fallbacks and how OpenRouter treats each model in its list. The executor keeps the retry policy. Each try is now told where it sits in it through the node metadata (`BlockNodeMetadata.retry`, with the executor's own `isFinalTry` judgment), and the Agent handler keeps the fallbacks out of the candidate list until the final try. Every earlier try runs the primary alone and lets the failure escape for the policy to replay. Blocks without fallbacks, and blocks with retry off, behave as before; other handlers ignore the field.
1 parent c77f940 commit 0b35cd7

7 files changed

Lines changed: 169 additions & 23 deletions

File tree

‎apps/docs/content/docs/workflows/blocks/agent.mdx‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -87,8 +87,8 @@ Some settings live under advanced, or appear only for models that support them:
8787
- **Reasoning effort / Thinking level.** For models with extended reasoning, how much the model thinks before answering. Higher is more thorough but slower and costs more tokens.
8888
- **Prompt caching.** For Anthropic Claude models, reuses the system prompt and tool definitions between runs instead of re-reading them every time. Cached input costs a tenth of the normal rate, but writing the cache costs 1.25x, so leave it off for one-off runs and turn it on when the same agent runs repeatedly. The cache covers a prefix only if it reaches 1,024 tokens (2,048 on Haiku) — below that Anthropic ignores it and nothing changes. Entries expire after five minutes of no use.
8989
- **API key.** Your key for the chosen provider. Hidden on hosted Sim, which supplies one.
90-
- **Fallback models.** An ordered list of models to try when the request to the selected model fails, whether the provider is overloaded, rate-limited, or down. Sim tries the 2nd choice, then the 3rd, and so on, and `<agent.model>` reports the model that answered. Hosted models need no setup. A model that needs its own key takes it from a workspace environment variable you pick on the row; a model on the same provider as the selected model reuses the block's key. Azure, Bedrock, and Vertex models can only be fallbacks for a selected model of the same family, since they use that model's credentials. The Auto model cannot be a fallback. A fallback runs with the selected model's settings where its provider accepts them: temperature and max output tokens are clamped to the fallback's limits, and when the fallback has a reasoning effort, thinking level, or verbosity setting that the selected model's value does not fit, the row shows that field so you can pick a value for it, otherwise the provider's default applies.
91-
- **Retry on fail.** Runs the block again after a failure, up to a maximum number of tries with a wait between them. Fallback models work inside each try: one try walks the selected model and then every fallback, and only when all of them fail does the next try begin. A failure that happens after the model already called a tool runs that conversation again on the next model, so keep fallbacks and retry off for agents whose tools must not repeat.
90+
- **Fallback models.** An ordered list of models to try when the request to the selected model fails, whether the provider is overloaded, rate-limited, or down. Sim tries the 2nd choice, then the 3rd, and so on, once each, and `<agent.model>` reports the model that answered. Hosted models need no setup. A model that needs its own key takes it from a workspace environment variable you pick on the row; a model on the same provider as the selected model reuses the block's key. Azure, Bedrock, and Vertex models can only be fallbacks for a selected model of the same family, since they use that model's credentials. The Auto model cannot be a fallback. A fallback runs with the selected model's settings where its provider accepts them: temperature and max output tokens are clamped to the fallback's limits, and when the fallback has a reasoning effort, thinking level, or verbosity setting that the selected model's value does not fit, the row shows that field so you can pick a value for it, otherwise the provider's default applies.
91+
- **Retry on fail.** Retries the selected model after a failure, up to a maximum number of tries with a wait between them. When its tries run out, the fallback models are tried in order, once each, with no wait before the first of them. A fallback is never retried. A failure that happens after the model already called a tool runs that conversation again on the next try or the next model, so keep fallbacks and retry off for agents whose tools must not repeat.
9292

9393
OpenAI and Gemini cache automatically at no extra cost and need no setting; their discount is already reflected in what you are charged.
9494

@@ -150,5 +150,5 @@ The Agent reads the message from Start with `<start.input>` and returns a result
150150
<FAQ items={[
151151
{ question: "How does max output tokens work with Anthropic models?", answer: "The Agent block uses each Anthropic model's full max output token limit by default (for example, 64,000 tokens). You can override this with the Max Output Tokens setting. For non-streaming requests that exceed the SDK's internal threshold, the provider automatically uses internal streaming to avoid timeouts." },
152152
{ question: "Can I use the Agent block with a custom or self-hosted model?", answer: "Yes. Use any Ollama or VLLM-compatible model by typing the model name directly into the model combobox, as long as it exposes a compatible API endpoint." },
153-
{ question: "What happens when my model's provider is down?", answer: "Add fallback models under Additional fields. When the request to the selected model fails, Sim tries the 2nd choice, then the 3rd, in order, and the block succeeds if any of them answers. The log detail shows the model that answered and the models it fell back from. Turn on Retry on fail as well to repeat the whole sequence after a wait." },
153+
{ question: "What happens when my model's provider is down?", answer: "Add fallback models under Additional fields. When the request to the selected model fails, Sim tries the 2nd choice, then the 3rd, in order, and the block succeeds if any of them answers. The log detail shows the model that answered and the models it fell back from. Turn on Retry on fail as well to give the selected model a few tries first; the fallbacks are tried once each after its last try fails." },
154154
]} />

‎apps/sim/blocks/blocks/agent.ts‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -37,7 +37,7 @@ const logger = createLogger('AgentBlock')
3737
/** Model the agent block falls back to when `model` is unset or the auto pseudo-model. */
3838
const AGENT_FALLBACK_MODEL = 'claude-sonnet-5'
3939

40-
const FALLBACK_MODELS_DESCRIPTION = `Ordered models tried in sequence when the request to the selected model fails. Each row is { model, apiKey?, reasoningEffort?, thinkingLevel?, verbosity? }; apiKey, when present, must be a whole {{ENV_VAR}} reference, and a tuning value must be one the row model declares. sim-auto is not allowed. Max ${MAX_FALLBACK_MODELS}.`
40+
const FALLBACK_MODELS_DESCRIPTION = `Ordered models tried in sequence, once each, when the request to the selected model fails; with Retry on fail, after the selected model's tries run out. Each row is { model, apiKey?, reasoningEffort?, thinkingLevel?, verbosity? }; apiKey, when present, must be a whole {{ENV_VAR}} reference, and a tuning value must be one the row model declares. sim-auto is not allowed. Max ${MAX_FALLBACK_MODELS}.`
4141
const MODELS_WITH_REASONING_EFFORT = getModelsWithReasoningEffort()
4242
const MODELS_WITH_VERBOSITY = getModelsWithVerbosity()
4343
const MODELS_WITH_THINKING = getModelsWithThinking()

‎apps/sim/executor/execution/block-executor.retry.test.ts‎

Lines changed: 30 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -121,6 +121,36 @@ describe('BlockExecutor retry', () => {
121121
expect(execute).toHaveBeenCalledTimes(1)
122122
})
123123

124+
it('tells each try where it sits in the policy, and a block without one nothing', async () => {
125+
const block = createBlock({ enabled: true, maxTries: 3, waitBetweenTriesMs: 0 })
126+
const execute = vi
127+
.fn()
128+
.mockRejectedValueOnce(new Error('one'))
129+
.mockRejectedValueOnce(new Error('two'))
130+
.mockResolvedValueOnce({ ok: true })
131+
const state = new ExecutionState()
132+
const executor = buildExecutor(block, { canHandle: () => true, execute }, state)
133+
134+
await executor.execute(createContext(state), createNode(block), block)
135+
136+
expect(execute.mock.calls.map(([, , , metadata]) => metadata.retry)).toEqual([
137+
{ attempt: 1, maxTries: 3, isFinalTry: false },
138+
{ attempt: 2, maxTries: 3, isFinalTry: false },
139+
{ attempt: 3, maxTries: 3, isFinalTry: true },
140+
])
141+
expect(execute.mock.calls[0][3].nodeId).toBe(block.id)
142+
143+
const plain = createBlock()
144+
const executePlain = vi.fn().mockResolvedValue({ ok: true })
145+
const plainState = new ExecutionState()
146+
await buildExecutor(
147+
plain,
148+
{ canHandle: () => true, execute: executePlain },
149+
plainState
150+
).execute(createContext(plainState), createNode(plain), plain)
151+
expect(executePlain.mock.calls[0][3]).not.toHaveProperty('retry')
152+
})
153+
124154
it('replays any failure and succeeds on a later try', async () => {
125155
const block = createBlock(enabled)
126156
const execute = vi

‎apps/sim/executor/execution/block-executor.ts‎

Lines changed: 19 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -44,6 +44,7 @@ import {
4444
import {
4545
type BlockHandler,
4646
type BlockLog,
47+
type BlockRetryAttempt,
4748
type BlockState,
4849
type ExecutionContext,
4950
getNextExecutionOrder,
@@ -277,11 +278,12 @@ export class BlockExecutor {
277278
* token is drained, so a replay cannot duplicate output the client has
278279
* already seen.
279280
*/
280-
const output = await this.runHandlerWithRetry(blockCtx, block, blockLog, () =>
281-
handler.executeWithNode
282-
? handler.executeWithNode(blockCtx, block, resolvedInputs, nodeMetadata)
283-
: handler.execute(blockCtx, block, resolvedInputs, nodeMetadata)
284-
)
281+
const output = await this.runHandlerWithRetry(blockCtx, block, blockLog, (retry) => {
282+
const invocationMetadata = retry ? { ...nodeMetadata, retry } : nodeMetadata
283+
return handler.executeWithNode
284+
? handler.executeWithNode(blockCtx, block, resolvedInputs, invocationMetadata)
285+
: handler.execute(blockCtx, block, resolvedInputs, invocationMetadata)
286+
})
285287

286288
completedHandlerCost = readTrustedExecutionCost(output)
287289

@@ -546,15 +548,20 @@ export class BlockExecutor {
546548
* Rethrows the final try's error so the caller's catch — and with it the error
547549
* port — behaves exactly as it does for a block that never retried. Retrying
548550
* only ever delays the existing outcome; it never changes it.
551+
*
552+
* Each try is told where it sits in the policy (`BlockRetryAttempt`). The
553+
* policy stays here: a handler cannot ask for another try or skip the wait,
554+
* it can only hold work for the try after which no other follows, the way the
555+
* Agent block keeps its fallback models for the final try.
549556
*/
550557
private async runHandlerWithRetry<T>(
551558
ctx: ExecutionContext,
552559
block: SerializedBlock,
553560
blockLog: BlockLog | undefined,
554-
invoke: () => Promise<T>
561+
invoke: (retry: BlockRetryAttempt | undefined) => Promise<T>
555562
): Promise<T> {
556563
const policy = resolveBlockRetryPolicy(block)
557-
if (!policy) return invoke()
564+
if (!policy) return invoke(undefined)
558565

559566
const shouldAccumulateFunctionCost = block.metadata?.id === BlockType.FUNCTION
560567
let accumulatedFunctionCost: TrustedExecutionCost | undefined
@@ -563,7 +570,11 @@ export class BlockExecutor {
563570
for (;;) {
564571
tries++
565572
try {
566-
const output = await invoke()
573+
const output = await invoke({
574+
attempt: tries,
575+
maxTries: policy.maxTries,
576+
isFinalTry: tries >= policy.maxTries,
577+
})
567578
if (!shouldAccumulateFunctionCost || !accumulatedFunctionCost || !isRecordLike(output)) {
568579
return output
569580
}

‎apps/sim/executor/handlers/agent/agent-handler.test.ts‎

Lines changed: 67 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -616,6 +616,73 @@ describe('AgentBlockHandler', () => {
616616
expect(blockLog).toMatchObject({ modelFallbacks: ['gpt-4o'] })
617617
})
618618

619+
it('holds the fallbacks on a try the executor will replay', async () => {
620+
mockExecuteProviderRequest.mockRejectedValueOnce(new Error('overloaded'))
621+
const blockLog = openLog()
622+
blockLog.modelFallbacks = ['stale-from-earlier-run']
623+
624+
await expect(
625+
handler.execute(
626+
{ ...mockContext, blockLogs: [blockLog] },
627+
mockBlock,
628+
{ ...baseInputs, fallbackModels: [{ model: 'claude-sonnet-5' }] },
629+
{ nodeId: mockBlock.id, retry: { attempt: 1, maxTries: 3, isFinalTry: false } }
630+
)
631+
).rejects.toThrow('overloaded')
632+
633+
expect(mockExecuteProviderRequest).toHaveBeenCalledTimes(1)
634+
expect(mockAgentLogger.info).toHaveBeenCalledWith('Fallback models held for the final try', {
635+
blockId: mockBlock.id,
636+
attempt: 1,
637+
maxTries: 3,
638+
})
639+
expect(mockAgentLogger.warn).not.toHaveBeenCalledWith(
640+
'Agent model failed; trying fallback',
641+
expect.anything()
642+
)
643+
expect(blockLog.modelFallbacks).toBeUndefined()
644+
})
645+
646+
it('walks the chain on the final try, and on a block that never retries', async () => {
647+
mockExecuteProviderRequest
648+
.mockRejectedValueOnce(new Error('overloaded'))
649+
.mockResolvedValueOnce(providerResponse('claude-sonnet-5'))
650+
.mockRejectedValueOnce(new Error('overloaded'))
651+
.mockResolvedValueOnce(providerResponse('claude-sonnet-5'))
652+
const inputs = { ...baseInputs, fallbackModels: [{ model: 'claude-sonnet-5' }] }
653+
654+
const onFinalTry = await handler.execute(mockContext, mockBlock, inputs, {
655+
nodeId: mockBlock.id,
656+
retry: { attempt: 3, maxTries: 3, isFinalTry: true },
657+
})
658+
const withoutPolicy = await handler.execute(mockContext, mockBlock, inputs, {
659+
nodeId: mockBlock.id,
660+
})
661+
662+
expect(mockExecuteProviderRequest).toHaveBeenCalledTimes(4)
663+
expect((onFinalTry as { model: string }).model).toBe('claude-sonnet-5')
664+
expect((withoutPolicy as { model: string }).model).toBe('claude-sonnet-5')
665+
expect(mockAgentLogger.info).not.toHaveBeenCalledWith(
666+
'Fallback models held for the final try',
667+
expect.anything()
668+
)
669+
})
670+
671+
it('says nothing about held fallbacks when none are configured', async () => {
672+
mockExecuteProviderRequest.mockRejectedValueOnce(new Error('overloaded'))
673+
674+
await expect(
675+
handler.execute(mockContext, mockBlock, baseInputs, {
676+
nodeId: mockBlock.id,
677+
retry: { attempt: 1, maxTries: 2, isFinalTry: false },
678+
})
679+
).rejects.toThrow('overloaded')
680+
expect(mockAgentLogger.info).not.toHaveBeenCalledWith(
681+
'Fallback models held for the final try',
682+
expect.anything()
683+
)
684+
})
685+
619686
it('never falls back on a deep-research follow-up turn', async () => {
620687
mockExecuteProviderRequest.mockRejectedValueOnce(new Error('overloaded'))
621688

‎apps/sim/executor/handlers/agent/agent-handler.ts‎

Lines changed: 33 additions & 11 deletions
Original file line numberDiff line numberDiff line change
@@ -75,7 +75,13 @@ import type {
7575
ToolInput,
7676
} from '@/executor/handlers/agent/types'
7777
import { parseResponseFormat } from '@/executor/handlers/shared/response-format'
78-
import type { BlockHandler, ExecutionContext, StreamingExecution, UserFile } from '@/executor/types'
78+
import type {
79+
BlockHandler,
80+
BlockNodeMetadata,
81+
ExecutionContext,
82+
StreamingExecution,
83+
UserFile,
84+
} from '@/executor/types'
7985
import { collectBlockData } from '@/executor/utils/block-data'
8086
import { stringifyJSON } from '@/executor/utils/json'
8187
import { projectResolvedSecretDiagnosticContent } from '@/executor/utils/resolved-secret-content-projection'
@@ -283,7 +289,8 @@ export class AgentBlockHandler implements BlockHandler {
283289
async execute(
284290
ctx: ExecutionContext,
285291
block: SerializedBlock,
286-
inputs: AgentInputs
292+
inputs: AgentInputs,
293+
nodeMetadata?: BlockNodeMetadata
287294
): Promise<BlockOutput | StreamingExecution> {
288295
ctx.mcpBlockId = block.id
289296
const providerErrorRegistry = ctx.resolvedSecretTraceRegistry?.forkForInputPaths(
@@ -483,19 +490,33 @@ export class AgentBlockHandler implements BlockHandler {
483490
}
484491

485492
/**
493+
* Retry on fail retries the selected model; the fallbacks join only on the
494+
* try after which the executor promises no other. Until then a failure of
495+
* the primary is left to escape, so the executor's policy can replay it.
496+
*
486497
* A follow-up turn of a deep-research interaction lives on the primary's
487498
* provider; another model has none of that conversation, so a green answer
488499
* from it would be built on a fresh context. Such a request never falls back.
489500
*/
490-
const fallbackCandidates = modelInputs.previousInteractionId
491-
? []
492-
: normalizeFallbackModels(filteredInputs.fallbackModels).filter(
493-
(candidate) => candidate.model.toLowerCase() !== model.toLowerCase()
494-
)
495-
if (modelInputs.previousInteractionId && filteredInputs.fallbackModels?.length) {
501+
const configuredFallbacks = normalizeFallbackModels(filteredInputs.fallbackModels)
502+
const retry = nodeMetadata?.retry
503+
const fallbacksHeld = retry !== undefined && !retry.isFinalTry
504+
const fallbackCandidates =
505+
modelInputs.previousInteractionId || fallbacksHeld
506+
? []
507+
: configuredFallbacks.filter(
508+
(candidate) => candidate.model.toLowerCase() !== model.toLowerCase()
509+
)
510+
if (configuredFallbacks.length > 0 && modelInputs.previousInteractionId) {
496511
logger.info('Fallback models skipped for a deep-research follow-up turn', {
497512
blockId: block.id,
498513
})
514+
} else if (configuredFallbacks.length > 0 && fallbacksHeld) {
515+
logger.info('Fallback models held for the final try', {
516+
blockId: block.id,
517+
attempt: retry.attempt,
518+
maxTries: retry.maxTries,
519+
})
499520
}
500521
const candidates: ModelCandidate[] = [
501522
{ model, apiKey: modelInputs.apiKey, isPrimary: true },
@@ -2397,8 +2418,9 @@ export class AgentBlockHandler implements BlockHandler {
23972418
*
23982419
* Which model serves the request is decided here, inside one handler
23992420
* invocation. How many invocations the block gets is the executor's retry
2400-
* policy, which wraps this whole chain: with retry on, every try walks the
2401-
* chain again from the primary.
2421+
* policy, and the caller keeps the fallbacks out of the candidate list until
2422+
* the final try, so with retry on the block runs the primary alone on every
2423+
* earlier try and walks the whole chain once.
24022424
*
24032425
* Falling through is deliberately as indiscriminate as block retry
24042426
* (`isRetryableBlockError`): a provider error carries no status, so an
@@ -2678,7 +2700,7 @@ export class AgentBlockHandler implements BlockHandler {
26782700
* the executor pushes the entry before running the handler with `endedAt`
26792701
* still empty, which is what tells it apart from earlier runs of the same
26802702
* block in a loop or an earlier retry. An empty list clears the field, since
2681-
* a retry that succeeds on the primary reuses the entry a failed try wrote.
2703+
* every try reuses one entry and only the final try can write failed models.
26822704
*
26832705
* A model id can itself come from a resolved reference, so the names are
26842706
* projected through the same secret registry as every other diagnostic and

‎apps/sim/executor/types.ts‎

Lines changed: 16 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -756,6 +756,22 @@ export interface BlockNodeMetadata {
756756
originalBlockId?: string
757757
isLoopNode?: boolean
758758
executionOrder?: number
759+
/** Where this invocation sits in the block's retry policy; absent when the block has none. */
760+
retry?: BlockRetryAttempt
761+
}
762+
763+
/**
764+
* One try of a block under its retry policy, told to the handler so it can hold
765+
* work for the last try. `isFinalTry` is the executor's own judgment, not
766+
* `attempt >= maxTries` recomputed by the handler: what makes a try final is the
767+
* policy's business, and a handler that fails on a non-final try is promised
768+
* another invocation for any retryable error.
769+
*/
770+
export interface BlockRetryAttempt {
771+
/** 1-based. */
772+
attempt: number
773+
maxTries: number
774+
isFinalTry: boolean
759775
}
760776

761777
export interface BlockHandler {

0 commit comments

Comments
 (0)