
Consider a procurement assistant that extracts invoice lines and prepares an ERP update. Its primary model times out. The gateway switches to another model, receives valid JSON, and reports success. Yet the replacement may have processed the invoice in a different geography, omitted a credit line, or generated a second update for a transaction already submitted.
A successful HTTP response says very little about whether that recovery was acceptable.
Enterprise model failover needs two decisions: which deployment is eligible to receive this request, and which workflow step can safely resume? Approve fallback routes against the data and task requirements before an outage. During recovery, restart only from an explicit checkpoint, keep action authorization outside the model, and stop or queue work when no eligible route remains.
This is an architecture for bounded recovery in applications you control. If a hosted agent executes tools inside a provider-managed session, first establish what execution state and cancellation guarantees that service exposes. A generic HTTP gateway cannot reconstruct business state it never observed.
A fallback route includes more than a model name
Treat the routing unit as a versioned deployment bundle: provider, endpoint or inference profile, model/version where exposed, processing boundary, retention configuration, adapter, prompt version, and output-validation path. Approval applies to that bundle for a particular workflow, not to a model family in the abstract.
The distinction has concrete consequences. Amazon Bedrock’s cross-Region inference documentation separates geography-scoped profiles from global profiles. An API request’s source region does not, by itself, establish where inference will run. Record the profile’s allowed destinations, not just the hostname your application calls.
Microsoft’s Foundry deployment documentation likewise distinguishes global, data-zone, and geography-based processing. Data stored at rest and data processed during inference have different location semantics. Translate the actual deployment terms into your allowed processing boundary; a friendly label such as “EU” or “regional” is insufficient evidence for a narrower country-level requirement.
Apply the same scrutiny to auxiliary paths. A routing classifier, request logger, trace exporter, or fallback cache can receive confidential context even if the selected inference endpoint is approved. Prefer routing from trusted workflow metadata when it is enough. If a classifier must inspect content, it needs its own approved data boundary.
Keep the bundle and its owner in the enterprise AI system registry. A changed region profile, retention setting, adapter, or model version should invalidate affected route approvals until the relevant checks pass again.
Filter eligible routes before optimizing latency or price
A routing score must not trade a hard restriction against a performance benefit. Build an eligible set first:
1 | eligible routes = approved data destinations |
Health, quota, latency, and price help select within that set. They cannot add a disallowed route to it. Missing or expired approval evidence should remove a candidate; an outage should not relax the rule.
For each route, the admission check needs specific evidence:
| Requirement | What the runtime must establish |
|---|---|
| Data handling | The assembled request’s classification is permitted by the provider, processing scope, retention configuration, and logging path. |
| Input compatibility | All required modalities and context fit. Count tokens with the destination’s tokenizer or a validated conservative estimator, reserving output capacity. |
| Semantic compatibility | The adapter preserves instructions, document boundaries, units, null values, and required evidence. |
| Output compatibility | The route can produce an output accepted by the same application validator and business invariants. |
| Task evidence | The tested bundle passed this workflow’s acceptance criteria at the proposed authority level. |
| Execution context | Tenant, user delegation, tool allowlist, and approval requirements remain valid at dispatch. |
For mixed context, derive handling requirements from the actual assembled payload, including attachments, retrieved text, memory, and tool results. The data-classification architecture for AI assistants supplies that upstream decision. Do not let a caller’s unverified public label override confidential retrieved content.
Here is an illustrative admission matrix for the invoice workflow. These routes are hypothetical; they are not vendor capability claims.
| Candidate after the primary fails | Evidence available | Decision |
|---|---|---|
| Standby deployment of the same pinned model | Approved processing scope; tested adapter; current invoice eval; sufficient capacity | Eligible for the existing extraction step |
| Different model at another approved provider | Same task gate passed; required document inputs supported; matching handling commitments | Eligible through its tested adapter |
| Fast global endpoint | Task tests pass, but its possible processing locations exceed the invoice policy | Exclude before transmitting the request |
| Small model approved only for text summaries | No evidence for image extraction or credit-line accuracy | Offer a separately defined summary-only workflow, if useful and permitted |
| No admissible route | Every candidate is unavailable or fails a hard condition | Queue within an explicit expiry or hand off to a person |
The smaller model is not a transparent replacement. Changing from invoice extraction to a human-reviewed summary changes the product’s behavior. Make that transition visible, disable the ERP-write path, and evaluate the narrower workflow separately. Human review does not make an otherwise prohibited data transfer acceptable.
Separate transport compatibility from task equivalence
Two deployments can accept the same API request and behave differently on the exact cases that matter: negative amounts, handwritten corrections, conflicting supplier identifiers, missing pages, or instructions embedded inside an invoice.
For this example, schema validation might establish that total is a number. Business checks still need to establish currency consistency, line-total reconciliation within a declared rounding tolerance, and whether the extracted supplier matches the purchase order. A plausible explanation from a second model is not a substitute for those checks.
Evaluate each approved bundle end to end: preprocessing, prompt, model, adapter, and validator. Use the same task acceptance criteria across alternatives, while allowing implementation differences such as provider-specific prompts. The existing enterprise AI eval-gate pattern covers the release decision; the additional requirement here is to force traffic through every fallback path during evaluation.
Do not silently truncate input to fit a smaller context window. Do not remove an attachment because the standby cannot process images. A reviewed preprocessing path may make a route eligible, but that creates a different bundle with its own evidence.
Runtime invalid output also needs a bounded policy. You might permit one evaluated repair attempt for malformed JSON. A failed business invariant should follow a defined exception path. Cycling through models until one returns a convenient answer turns validation into selection bias.
Resume from a checkpoint, not from an interrupted transcript
Keep durable business state in the orchestrator. Model conversation state is useful context, but it should not be the only record of which actions occurred.
For the invoice example, separate these checkpoints:
1 | input snapshot recorded |
Record the input references, document versions, accepted outputs, and operation identifiers necessary to resume. Sensitive snapshots still need access controls and retention limits; “durable” does not mean keeping everything indefinitely.
| Failure point | Recovery boundary |
|---|---|
| Inference failed before any output was accepted or tool dispatched | Another eligible route may retry that inference step within budget. |
| Text was streamed to the user, then the connection broke | Mark the response incomplete. If restarting, present a replacement as a new attempt rather than appending a different model’s continuation. |
| A tool-call payload arrived only partially | Discard the incomplete proposal. Never execute partial tool arguments. |
| A complete proposal was accepted but not dispatched | Recover the stored proposal and recheck current authorization; regeneration creates a new proposal. |
| ERP execution was submitted but its result is unknown | Reconcile using the operation identifier before permitting another write. |
| ERP commit is confirmed, but the final explanation failed | Generate the explanation from the recorded receipt; do not rerun the write. |
Switch models at a complete, recorded step boundary. If a new proposal changes approved arguments, obtain a new approval. If a route is approved only for a lower authority level, the orchestrator must enforce that reduction regardless of what the model requests.
Give each inference attempt a distinct identifier and let the orchestrator atomically accept only one result for a checkpoint. A slow primary can return after the standby has succeeded. Its late response must not overwrite the accepted result or dispatch another tool call. Cancellation is useful, but the single-winner rule is what protects workflow state when cancellation is delayed or ineffective.
These action semantics belong in enterprise AI tool contracts. The routing layer’s responsibility is to preserve them across inference attempts. An idempotency key must identify the business operation, not each newly generated attempt to perform it.
The OWASP AI Agent Security Cheat Sheet supports the underlying separation: scoped tool access, explicit authorization for sensitive operations, and controls around irreversible actions. The checkpoint arrangement above is an implementation proposal, not a mechanism a model API automatically supplies.
When migrating conversation context, rebuild it from approved messages and recorded tool results through the destination adapter. Do not assume a provider’s opaque session identifier, hidden state, or reasoning artifact is portable. Recheck source access and freshness before reuse if recovery has been delayed.
Classify the failure before choosing a recovery
A catch-all except: try_next_model() handler collapses several different decisions into one.
For transport timeouts and transient service errors, a bounded retry or eligible standby can be reasonable. For a 429, use the provider’s error details to distinguish transient throttling from an exhausted allocation, and respect the applicable retry guidance. A request that will not fit the remaining deadline belongs in a queue or an explicit failure response.
Authentication failures, authorization denials, and malformed requests should raise an operational fault, not trigger an uncontrolled tour of every endpoint. A provider refusal or content-policy block needs a separate application disposition. Automatically seeking a more permissive model would change the policy outcome under the cover of availability recovery.
Circuit breakers help stop traffic to unhealthy destinations. They do not certify a destination’s suitability. Azure API Management’s backend documentation describes priority pools and circuit breakers, including handling Retry-After. It also notes that breaker decisions are local to gateway instances rather than synchronized globally. Account for the implementation’s actual behavior when sizing retry traffic.
An application eligibility filter can feed several homogeneous backend pools, each containing routes approved for the same workflow and data constraints. That is easier to audit than one universal pool whose members have different processing boundaries and capabilities.
Spend one deadline and one attempt budget
Retries multiply when the client, gateway, SDK, and orchestrator each apply their own defaults. If three layers each permit three total attempts, a single logical request can produce up to 27 downstream attempts. Assign one component ownership of the overall budget and configure the other layers to respect it.
AWS’s retry-with-backoff guidance explains why retries need backoff, suitable timeouts, and idempotent operations. For model calls, also bound output generation, concurrent work, and aggregate spending across attempts. A timed-out request may still consume remote resources; cancelling a local wait is not proof that remote inference stopped or that it will not be billed.
An illustrative 12-second request budget might allocate:
| Stage | Maximum allocation |
|---|---|
| Admission and input preparation | 1.0 s |
| Primary inference attempt | 4.0 s |
| Backoff and route decision | 0.5 s |
| Standby inference attempt | 4.0 s |
| Output validation and response assembly | 1.5 s |
| Scheduling and network margin | 1.0 s |
These are design allocations, not measured performance or universal timeout recommendations. Implement an absolute deadline using a monotonic clock inside the request handler. At each attempt, reserve time for validation and delivery. If the primary consumes five seconds instead of four, reduce the remaining inference allowance or stop; do not restart the twelve-second clock.
Track timeout phases separately. Time to first token, idle time between chunks, and total generation time answer different questions. A stream that emits one token every few seconds can avoid an idle timeout while missing the business deadline.
Budget capacity as carefully as time. A standby that accepts a health probe may lack enough quota for diverted production traffic. Set per-tenant concurrency limits, bounded queues, and load-shedding rules before the incident. Two routes sharing the same account quota, regional service, network egress, or credentials may fail together. Record those shared dependencies instead of counting endpoints as independent redundancy.
Prove recovery with forced failures
Run drills against the actual adapters and routing policy. Use synthetic or appropriately protected test data; shadowing traffic to a second provider is itself a data transfer that needs approval.
| Injected condition | Evidence required before enabling automatic recovery |
|---|---|
| Primary times out; fastest standby is outside the allowed processing scope | No request body leaves for that standby; the exclusion reason is recorded. |
| Standby lacks a required modality or sufficient context capacity | Admission rejects the route without silently dropping input. |
Every primary request receives 429 | Attempts, queue depth, and concurrent standby requests remain bounded under load. |
| Connection closes during a streamed tool-call proposal | No tool executes; the attempt ends as incomplete. |
| ERP commits but its response is lost | Reconciliation finds the existing operation; no second business write occurs. |
| Primary recovers while fallback traffic is active | A staged return respects capacity limits and does not move running steps between models. |
Also test the empty eligible set. It is a normal state of this architecture. The application should report that the task is queued, needs manual handling, or cannot complete within its constraints. It must not quietly widen the destination allowlist.
Measure recovery at the workflow level. Record each attempt’s route-bundle version, policy decision, failure class, checkpoint, timing, and known usage; mark unknown remote usage for later reconciliation. Keep confidential payloads out of routine routing logs. These fields extend the agent audit-log architecture without requiring a second transcript store.
Separate three outcomes on the operations dashboard: successful completion under the original contract, completion through an explicitly reduced workflow, and work deferred or stopped. Report quality and latency by route. A blended success rate can hide a standby that remains reachable but produces unusable results.
The platform team should own route configuration, capacity, and retry behavior. The workflow owner should approve task equivalence and degraded behavior; the data or security control owner should approve handling boundaries. Put the approved combinations in versioned configuration and exercise them before an outage. At incident time, operators should be selecting a tested recovery path, not deciding which promises the application can abandon.