Purchasing Guides / Model Routing & Fallback

Model Routing & Fallback

A purchasing guide to buying a workflow that chooses an approved AI model for each task and switches when a model is unavailable.

Yearly running cost

Assuming 120,000 requests a year

Fixed model
$2.8K / year
Model usage only
Hosted PAYG
$3K+ / year
This workflow
$2.8K / year
Machine usage only; setup, hosting and review are extra.
Where do you want to use model routing?

Use models that pass your task tests, keep calls within your rules and return a clear error when no allowed model can answer.

One request includes up to 4,000 input tokens and 500 output tokens.

Your team approves model choices and rollout. No automatic tool execution, secret sharing, permission changes or production rollout. Safety refusals and disallowed data must stop, not trigger a bypass.

What is included

Run your samples through approved routes and return the response, model, attempts and usage. Test task quality and failures before live use.

Start with bounded text requests and server-approved task types. No image generation, web browsing, tool execution or arbitrary agent loops in the starting scope.

Models in the starting pool: up to 2 models.

Task types in the starting scope: up to 3 task types.

Billed model attempts per request: up to 2 attempts.

Apps in the connected scope: up to 1 app.

How reliable should the responses and fallbacks be?

Good routing completes your tasks, handles permitted failures and never escapes your data or spending rules.

How these standards are measured

Test your real tasks and controlled outages separately. Count failed requests, check answer quality and verify that every attempt respects the same permissions and budget.

Targets for your selected standard
What is checkedTarget
Tasks completed correctlyAgree reference answers, extraction fields or an independent grading rubric for each task. Count all in-scope requests, including errors and failed output validation; a successful HTTP response is not a correct answer.≥90%
Recoverable failures handledUse a controlled outage with an approved healthy spare and sufficient pre-agreed budget. Keep missed and failed recovery cases in the denominator. Safety refusals, policy blocks and exhausted budgets must stop; they are not outages to bypass.≥95%
Model and data rules respectedCheck every attempt, including retries and errors. Untrusted prompt text cannot modify allowlists, tenant access, data policy or fallback destinations.≥100%
Budget limits respectedReserve the maximum possible billed cost before dispatch using verified tariffs and token limits. Account for concurrent requests, timed-out calls and uncertain billing; do not release an uncertain reservation as though no cost occurred.≥100%
Complete routing recordsRecord all requests and attempts without storing sensitive prompt contents by default. Unknown provider usage stays pending reconciliation, not zero.≥100%

These are pilot targets, not a promise of universal accuracy or uptime. Freeze the task rubric and approved routes before testing. A cost band is not a hard spending ceiling. Fail closed when no eligible route remains.

How quickly do you need each response?

Choose how quickly the complete response should come back.

Timing details

Time authentication, policy checks, routing, model generation, queues, retries and the stored result. Routing overhead and time to first token are separate; neither proves the complete response target.

The target applies to at least 95% of agreed test runs, with 10 in progress at a time.

How much do you want to spend per request?

Choose the machine budget for the response, including fallback attempts.

Cost details

Includes approved model calls, retries and shared hosting at the reference workload. Larger prompts, premium models or a hosted gateway change the cost. A tighter cap can leave too little budget for a fallback.

Reference machine cost: $0.026 – 0.029 per request at 10000 requests a month. The selected cap is a target to test, not a replacement for this estimate.

Where do you want it to run?

Run the routing service in your cloud or on your machine. Choose whether requests may go to approved external models.

Data and access details

Runs in an account you control, with server-side keys, access rules and request logs.

Route only to the models and endpoints your team permits. Enforce your data and spending rules for every attempt.

Keep keys and routing permissions on the server. Private mode uses models you supply on local hardware; model licences, capacity and running cost need testing. It cannot connect to the remote app in this starting scope.

How do you want to use it?

Try a test page, call from your own agent or use a dedicated test agent. The connected scope adds an API for your app.

Anything else your provider should know?

Optional. Your choices are included automatically.

Common questions

When is a standard gateway enough?

Use OpenRouter or a standard self-hosted gateway if its routing and controls already fit. Buy this workflow for your own task tests, private processing or app-specific rules and integration.

Will routing always make AI cheaper?

No. Hard tasks, retries and premium models may cost more. Compare actual task pass rates, latency and billed usage with your fixed-model baseline before choosing a route.

Can a fallback ignore a safety refusal?

No. A refusal or blocked data policy must stop the request. An approved spare is for permitted recoverable failures, not weaker restrictions.

Can all models run on my own machine?

No. Use only models you can legally host with adequate hardware. Private inference speed, quality and cost need a pilot; hosted proprietary models are not local downloads.

Download RenX

Get the app.

Install, sign up, and start on your free plan with welcome credit included. No credit card required.

On a platform not listed? Leave your email and we'll notify you when a build is available.