You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: apps/docs/content/docs/workflows/blocks/agent.mdx
+4-1Lines changed: 4 additions & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -25,7 +25,7 @@ Answer in two sentences, cite the doc you used, and never guess a price.
25
25
26
26
### Model
27
27
28
-
The model that runs the step. Defaults to `claude-sonnet-4-6`. Type or pick any model from OpenAI, Anthropic, Google, xAI, Groq, Cerebras, DeepSeek, Azure, AWS Bedrock, Google Vertex, or OpenRouter, or a local model through Ollama or VLLM.
28
+
The model that runs the step. Defaults to `claude-sonnet-5`. Type or pick any model from OpenAI, Anthropic, Google, xAI, Groq, Cerebras, DeepSeek, Azure, AWS Bedrock, Google Vertex, or OpenRouter, or a local model through Ollama or VLLM.
29
29
30
30
For a custom cloud deployment, enter its provider prefix and model ID: `azure/my-deployment`, `azure-anthropic/my-deployment`, `bedrock/my-inference-profile`, or `vertex/my-gemini-model`. The prefix selects the provider and shows its credential fields even when the ID is absent from the catalog. Bedrock accepts full inference profile ARNs after `bedrock/`; Vertex uses the Gemini API and accepts Google model resource names. The deployment must support the selected provider's API. Custom IDs have no catalog pricing or token limits.
31
31
@@ -87,6 +87,8 @@ Some settings live under advanced, or appear only for models that support them:
87
87
-**Reasoning effort / Thinking level.** For models with extended reasoning, how much the model thinks before answering. Higher is more thorough but slower and costs more tokens.
88
88
-**Prompt caching.** For Anthropic Claude models, reuses the system prompt and tool definitions between runs instead of re-reading them every time. Cached input costs a tenth of the normal rate, but writing the cache costs 1.25x, so leave it off for one-off runs and turn it on when the same agent runs repeatedly. The cache covers a prefix only if it reaches 1,024 tokens (2,048 on Haiku) — below that Anthropic ignores it and nothing changes. Entries expire after five minutes of no use.
89
89
-**API key.** Your key for the chosen provider. Hidden on hosted Sim, which supplies one.
90
+
- **Fallback models.** An ordered list of models to try when the request to the selected model fails, whether the provider is overloaded, rate-limited, or down. Sim tries the 2nd choice, then the 3rd, and so on, and `<agent.model>` reports the model that answered. Hosted models need no setup. A model that needs its own key takes it from a workspace environment variable you pick on the row; a model on the same provider as the selected model reuses the block's key. Azure, Bedrock, and Vertex models can only be fallbacks for a selected model of the same family, since they use that model's credentials. The Auto model cannot be a fallback. A fallback runs with the selected model's settings where its provider accepts them: temperature and max output tokens are clamped to the fallback's limits, and when the fallback has a reasoning effort, thinking level, or verbosity setting that the selected model's value does not fit, the row shows that field so you can pick a value for it, otherwise the provider's default applies.
91
+
-**Retry on fail.** Runs the block again after a failure, up to a maximum number of tries with a wait between them. Fallback models work inside each try: one try walks the selected model and then every fallback, and only when all of them fail does the next try begin. A failure that happens after the model already called a tool runs that conversation again on the next model, so keep fallbacks and retry off for agents whose tools must not repeat.
90
92
91
93
OpenAI and Gemini cache automatically at no extra cost and need no setting; their discount is already reflected in what you are charged.
92
94
@@ -148,4 +150,5 @@ The Agent reads the message from Start with `<start.input>` and returns a result
148
150
<FAQitems={[
149
151
{ question: "How does max output tokens work with Anthropic models?", answer: "The Agent block uses each Anthropic model's full max output token limit by default (for example, 64,000 tokens). You can override this with the Max Output Tokens setting. For non-streaming requests that exceed the SDK's internal threshold, the provider automatically uses internal streaming to avoid timeouts." },
150
152
{ question: "Can I use the Agent block with a custom or self-hosted model?", answer: "Yes. Use any Ollama or VLLM-compatible model by typing the model name directly into the model combobox, as long as it exposes a compatible API endpoint." },
153
+
{ question: "What happens when my model's provider is down?", answer: "Add fallback models under Additional fields. When the request to the selected model fails, Sim tries the 2nd choice, then the 3rd, in order, and the block succeeds if any of them answers. The log detail shows the model that answered and the models it fell back from. Turn on Retry on fail as well to repeat the whole sequence after a wait." },
0 commit comments