OpenClaw Model Failover Guide: Configure Reliable Fallbacks
A model outage should not turn every OpenClaw workflow into a manual rescue. Providers rate-limit accounts, credentials expire, and a model can be temporarily overloaded even when the Gateway itself is healthy. OpenClaw's model failover system can keep a turn moving by rotating credentials within a provider and then trying configured fallback models.
Failover is not a license to put random models in a list and hope for the best. The model that answers after a fallback may have different context limits, tool behavior, latency, or safety characteristics. The useful configuration is an intentional chain: a primary for normal work, one or two compatible backups, and clear rules for when an explicit model choice must fail instead of silently changing.
Fallback is most useful when the backup preserves the job's contract. Before you add a candidate, test a representative prompt with the same tools and output format, then note any differences in latency or context capacity. Availability is valuable, but a backup that produces unusable output only moves the incident downstream.
The rest of this guide explains the current failover path, shows a safe configuration pattern, and covers agent, session, and cron behavior.
How OpenClaw model failover works
OpenClaw handles failure in two stages:
- Auth-profile rotation within the current provider. If the provider has multiple eligible profiles, OpenClaw can try another profile using its cooldown and rotation rules.
- Model fallback across candidates. If the provider is exhausted and the error is eligible for failover, OpenClaw advances to the next model in the configured fallback chain.
The runner also performs bounded same-model recovery for eligible transient failures before it rotates profiles or advances to another model. Not every error belongs in the failover bucket. An invalid model ID, a policy restriction, or an unsupported request option may need configuration changes rather than another attempt.
Fallback execution is turn-local. If openai/gpt-5.5 answers with anthropic/claude-sonnet-4-6 because the first candidate failed, the next turn still starts from the selected primary. A backup answering once does not silently rewrite the user's model preference.
For pure provider overload, OpenClaw can retry the turn-local candidate chain with exponential backoff while no tool execution or assistant output has started. If the complete chain is exhausted, the run returns a structured failure rather than pretending that the request succeeded.
Configure a default fallback chain
Set model references explicitly as provider/model. A minimal JSON5 configuration looks like this:
// ~/.openclaw/openclaw.json
{
agents: {
defaults: {
model: {
primary: "openai/gpt-5.5",
fallbacks: [
"anthropic/claude-sonnet-4-6",
"ollama/llama3.3"
]
}
}
}
}
Use model IDs that are actually configured and available to your providers. The names above are examples, not a guarantee that your account has access to them. If you use a local provider, confirm the local model ID and endpoint before adding it to a production chain.
Prefer a short list of compatible candidates. A fallback with a much smaller context window may fail on the same prompt. A model without the required vision or tool capabilities may produce a misleading second failure. For a tool-using agent, check the whole chain against the tools, structured output, media, and context requirements of your workflow.
After editing, validate the configuration through the supported path rather than replacing the entire file with a copied example. OpenClaw's configuration is schema-validated, and the project recommends inspecting exact fields before making targeted changes. The existing OpenClaw configuration guide covers safe edits and reload behavior.
Know when fallback applies
The source of the model selection matters. These cases do not all behave alike.
Configured default
agents.defaults.model.primary is the normal starting point for an agent. Its fallbacks list is used when the primary provider fails with an eligible error.
Per-agent model
An entry under agents.entries can define its own model. That agent's model is strict unless its model object includes its own fallback list. Make the policy obvious instead of relying on an assumption:
{
agents: {
entries: {
support: {
model: {
primary: "anthropic/claude-sonnet-4-6",
fallbacks: ["openai/gpt-5.5"]
}
},
billing: {
model: {
primary: "openai/gpt-5.5",
fallbacks: []
}
}
}
}
}
Here, support can use its configured backup. billing is explicitly strict: if its primary fails, the failure is surfaced instead of routing the request to another model.
User session override
An explicit user selection is different from an automatic default. /model, the model picker, session_status(model=...), and sessions.patch record a user session selection. If that exact model fails before producing a reply, OpenClaw reports the failure rather than silently answering from an unrelated fallback.
That behavior is useful when you are testing a model, comparing outputs, or need a specific provider for a regulated workflow. Do not mistake a user override for a configured reliability policy.
Cron jobs and fallback policy
A cron payload's model is a job primary, not a permanent user-session override. By default, the job can use configured fallbacks. A job can also define its own fallback list when the automation needs a different reliability or cost policy.
Use an empty list to make the cron run strict:
{
payload: {
model: "openai/gpt-5.5",
fallbacks: []
}
}
A strict job is appropriate when a report must be generated by one approved model, or when changing models would make the result invalid. A non-strict job is better for routine notifications and maintenance work where completing the task matters more than using one exact model. Whichever policy you choose, document it with the job so a later operator does not interpret a fallback as an unexpected model change.
For scheduled work with tools, also consider whether the backup model understands the same tool contracts. A model that can answer text but cannot reliably call the required tool is not a useful fallback.
Cooldowns, overloads, and diagnosis
OpenClaw applies cooldowns to failing auth profiles rather than repeatedly hammering a credential that is rate-limited or temporarily unhealthy. If every profile for a provider fails with a failover-worthy error, the runner advances to the next model candidate.
When a run fails, start with the model and provider status rather than adding more fallbacks. Check:
- the exact
provider/modelreference and whether the provider is configured; - the auth profile status and cooldown reason;
- whether the error is a transient rate limit, overload, network failure, or a non-retryable configuration error;
- whether the fallback has the context, tools, and media features the turn needs;
- whether a user or agent-level strict override is preventing fallback.
OpenClaw's model FAQ recommends explicit provider/model references and targeted model changes. If you need to repair the surrounding Gateway configuration, use the OpenClaw Gateway health-check workflow and run the documented status and diagnostic commands before changing unrelated settings.
Do not use a fallback list to hide a broken primary indefinitely. A chain can preserve availability while masking expired credentials or a removed model. Review failure notices and provider usage regularly, and remove candidates that are no longer supported.
Security rules for fallback models
A fallback model is part of the agent's behavior, not a security boundary. OpenClaw's security guidance treats a tool-enabled Gateway as one trusted boundary and warns that smaller or heavily quantized models are more vulnerable to prompt injection.
Use these safeguards:
- Prefer a comparably capable fallback for agents that can execute commands, browse, send messages, or modify files.
- Keep sandboxing and strict tool allowlists enabled; do not broaden permissions just because the primary is unavailable.
- Do not put credentials, private prompts, or sensitive work into an untrusted provider merely to keep a task running.
- Use strict model selection for workflows whose output must come from an approved provider.
- Record which model answered when auditability matters; fallback notices distinguish the selected model from the model that answered.
The OpenClaw security guide is a useful companion for reviewing channel access, trust boundaries, and tool exposure. If different users do not trust one another, separate Gateways and credentials are safer than trying to solve the problem with model selection alone.
FAQ
Does failover permanently change the selected model?
No. A fallback is normally used for the current turn. The next turn starts from the selected primary unless an operator explicitly changes the selection or an automatic recovery mechanism clears a temporary override.
Why did my explicit /model selection not fall back?
An explicit user session selection is strict by design. This prevents OpenClaw from silently answering with a different model when you asked for one particular provider/model.
Should I add many fallback models?
Usually no. Keep the chain short and compatible. More candidates increase the chance of a mismatch in tools, context, output format, or cost, and can make diagnosis harder.
Can cron jobs use fallback models?
Yes. A cron payload model is a job primary and uses configured fallbacks unless the payload supplies its own fallbacks. Set fallbacks: [] when the job must fail rather than switch models.
Is a fallback safe for a tool-enabled agent?
Only if you evaluate it like the primary. Confirm its tool behavior and context limits, keep sandbox and allowlist controls in place, and avoid weaker or untrusted models for sensitive actions.




Comments
Loading comments…