You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
fix(agent): retry the selected model before falling back
Retry on fail used to wrap the whole fallback chain, so tries 3 with
fallbacks B and C ran A, B, C three times over. A builder who lists
fallbacks wants the selected model retried and the fallbacks tried once
each after its last try fails, which is also how LiteLLM orders retries
and fallbacks and how OpenRouter treats each model in its list.
The executor keeps the retry policy. Each try is now told where it sits
in it through the node metadata (`BlockNodeMetadata.retry`, with the
executor's own `isFinalTry` judgment), and the Agent handler keeps the
fallbacks out of the candidate list until the final try. Every earlier
try runs the primary alone and lets the failure escape for the policy to
replay. Blocks without fallbacks, and blocks with retry off, behave as
before; other handlers ignore the field.
Copy file name to clipboardExpand all lines: apps/docs/content/docs/workflows/blocks/agent.mdx
+3-3Lines changed: 3 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -87,8 +87,8 @@ Some settings live under advanced, or appear only for models that support them:
87
87
-**Reasoning effort / Thinking level.** For models with extended reasoning, how much the model thinks before answering. Higher is more thorough but slower and costs more tokens.
88
88
-**Prompt caching.** For Anthropic Claude models, reuses the system prompt and tool definitions between runs instead of re-reading them every time. Cached input costs a tenth of the normal rate, but writing the cache costs 1.25x, so leave it off for one-off runs and turn it on when the same agent runs repeatedly. The cache covers a prefix only if it reaches 1,024 tokens (2,048 on Haiku) — below that Anthropic ignores it and nothing changes. Entries expire after five minutes of no use.
89
89
-**API key.** Your key for the chosen provider. Hidden on hosted Sim, which supplies one.
90
-
- **Fallback models.** An ordered list of models to try when the request to the selected model fails, whether the provider is overloaded, rate-limited, or down. Sim tries the 2nd choice, then the 3rd, and so on, and `<agent.model>` reports the model that answered. Hosted models need no setup. A model that needs its own key takes it from a workspace environment variable you pick on the row; a model on the same provider as the selected model reuses the block's key. Azure, Bedrock, and Vertex models can only be fallbacks for a selected model of the same family, since they use that model's credentials. The Auto model cannot be a fallback. A fallback runs with the selected model's settings where its provider accepts them: temperature and max output tokens are clamped to the fallback's limits, and when the fallback has a reasoning effort, thinking level, or verbosity setting that the selected model's value does not fit, the row shows that field so you can pick a value for it, otherwise the provider's default applies.
91
-
-**Retry on fail.**Runs the block again after a failure, up to a maximum number of tries with a wait between them. Fallback models work inside each try: one try walks the selected model and then every fallback, and only when all of them fail does the next try begin. A failure that happens after the model already called a tool runs that conversation again on the next model, so keep fallbacks and retry off for agents whose tools must not repeat.
90
+
- **Fallback models.** An ordered list of models to try when the request to the selected model fails, whether the provider is overloaded, rate-limited, or down. Sim tries the 2nd choice, then the 3rd, and so on, once each, and `<agent.model>` reports the model that answered. Hosted models need no setup. A model that needs its own key takes it from a workspace environment variable you pick on the row; a model on the same provider as the selected model reuses the block's key. Azure, Bedrock, and Vertex models can only be fallbacks for a selected model of the same family, since they use that model's credentials. The Auto model cannot be a fallback. A fallback runs with the selected model's settings where its provider accepts them: temperature and max output tokens are clamped to the fallback's limits, and when the fallback has a reasoning effort, thinking level, or verbosity setting that the selected model's value does not fit, the row shows that field so you can pick a value for it, otherwise the provider's default applies.
91
+
-**Retry on fail.**Retries the selected model after a failure, up to a maximum number of tries with a wait between them. When its tries run out, the fallback models are tried in order, once each, with no wait before the first of them. A fallback is never retried. A failure that happens after the model already called a tool runs that conversation again on the next try or the next model, so keep fallbacks and retry off for agents whose tools must not repeat.
92
92
93
93
OpenAI and Gemini cache automatically at no extra cost and need no setting; their discount is already reflected in what you are charged.
94
94
@@ -150,5 +150,5 @@ The Agent reads the message from Start with `<start.input>` and returns a result
150
150
<FAQitems={[
151
151
{ question: "How does max output tokens work with Anthropic models?", answer: "The Agent block uses each Anthropic model's full max output token limit by default (for example, 64,000 tokens). You can override this with the Max Output Tokens setting. For non-streaming requests that exceed the SDK's internal threshold, the provider automatically uses internal streaming to avoid timeouts." },
152
152
{ question: "Can I use the Agent block with a custom or self-hosted model?", answer: "Yes. Use any Ollama or VLLM-compatible model by typing the model name directly into the model combobox, as long as it exposes a compatible API endpoint." },
153
-
{ question: "What happens when my model's provider is down?", answer: "Add fallback models under Additional fields. When the request to the selected model fails, Sim tries the 2nd choice, then the 3rd, in order, and the block succeeds if any of them answers. The log detail shows the model that answered and the models it fell back from. Turn on Retry on fail as well to repeat the whole sequence after a wait." },
153
+
{ question: "What happens when my model's provider is down?", answer: "Add fallback models under Additional fields. When the request to the selected model fails, Sim tries the 2nd choice, then the 3rd, in order, and the block succeeds if any of them answers. The log detail shows the model that answered and the models it fell back from. Turn on Retry on fail as well to give the selected model a few tries first; the fallbacks are tried once each after its last try fails." },
/** Model the agent block falls back to when `model` is unset or the auto pseudo-model. */
38
38
constAGENT_FALLBACK_MODEL='claude-sonnet-5'
39
39
40
-
constFALLBACK_MODELS_DESCRIPTION=`Ordered models tried in sequencewhen the request to the selected model fails. Each row is { model, apiKey?, reasoningEffort?, thinkingLevel?, verbosity? }; apiKey, when present, must be a whole {{ENV_VAR}} reference, and a tuning value must be one the row model declares. sim-auto is not allowed. Max ${MAX_FALLBACK_MODELS}.`
40
+
constFALLBACK_MODELS_DESCRIPTION=`Ordered models tried in sequence, once each, when the request to the selected model fails; with Retry on fail, after the selected model's tries run out. Each row is { model, apiKey?, reasoningEffort?, thinkingLevel?, verbosity? }; apiKey, when present, must be a whole {{ENV_VAR}} reference, and a tuning value must be one the row model declares. sim-auto is not allowed. Max ${MAX_FALLBACK_MODELS}.`
0 commit comments