Yearly running cost Assuming 120,000 requests a year
Fixed model $2.8K / year Model usage only
Hosted PAYG $3K+ / year
This workflow $2.8K / year Machine usage only; setup, hosting and review are extra. Where do you want to use model routing? Use models that pass your task tests, keep calls within your rules and return a clear error when no allowed model can answer.
One request includes up to 4,000 input tokens and 500 output tokens.
Try my sample requests Test routing before connecting an app $864 – 2,640 setup Connect my app Use tested routing in my existing app $1,296 – 3,840 setup
Which app should it connect to?My app API Your team approves model choices and rollout. No automatic tool execution, secret sharing, permission changes or production rollout. Safety refusals and disallowed data must stop, not trigger a bypass.
What is included Run your samples through approved routes and return the response, model, attempts and usage. Test task quality and failures before live use.
Start with bounded text requests and server-approved task types. No image generation, web browsing, tool execution or arbitrary agent loops in the starting scope.
Models in the starting pool: up to 2 models.
Task types in the starting scope: up to 3 task types.
Billed model attempts per request: up to 2 attempts.
Apps in the connected scope: up to 1 app.
How reliable should the responses and fallbacks be? Good routing completes your tasks, handles permitted failures and never escapes your data or spending rules.
Standard Test your common tasks, fallback behaviour and model and spending rules. Tasks completed correctly: ≥90% Recoverable failures handled: ≥95% Included Strict Add harder task samples, failure combinations and independent policy and billing tests. Tasks completed correctly: ≥95% Recoverable failures handled: ≥98% Setup +$216 – 600 My own targets Agree more models, modalities, enterprise controls or your own task-quality requirements. Quote separately
How these standards are measured Test your real tasks and controlled outages separately. Count failed requests, check answer quality and verify that every attempt respects the same permissions and budget.
Targets for your selected standard What is checked Target Tasks completed correctlyAgree reference answers, extraction fields or an independent grading rubric for each task. Count all in-scope requests, including errors and failed output validation; a successful HTTP response is not a correct answer. ≥90% Recoverable failures handledUse a controlled outage with an approved healthy spare and sufficient pre-agreed budget. Keep missed and failed recovery cases in the denominator. Safety refusals, policy blocks and exhausted budgets must stop; they are not outages to bypass. ≥95% Model and data rules respectedCheck every attempt, including retries and errors. Untrusted prompt text cannot modify allowlists, tenant access, data policy or fallback destinations. ≥100% Budget limits respectedReserve the maximum possible billed cost before dispatch using verified tariffs and token limits. Account for concurrent requests, timed-out calls and uncertain billing; do not release an uncertain reservation as though no cost occurred. ≥100% Complete routing recordsRecord all requests and attempts without storing sensitive prompt contents by default. Unknown provider usage stays pending reconciliation, not zero. ≥100%
These are pilot targets, not a promise of universal accuracy or uptime. Freeze the task rubric and approved routes before testing. A cost band is not a hard spending ceiling. Fail closed when no eligible route remains.
How quickly do you need each response? Choose how quickly the complete response should come back.
Within 60 seconds Included Within 30 seconds Setup +$144 – 480 Within 15 seconds Setup +$288 – 840
Timing details Time authentication, policy checks, routing, model generation, queues, retries and the stored result. Routing overhead and time to first token are separate; neither proves the complete response target.
The target applies to at least 95% of agreed test runs, with 10 in progress at a time.
How much do you want to spend per request? Choose the machine budget for the response, including fallback attempts.
Up to $0.05 Included Up to $0.03 Setup +$72 – 360 Up to $0.02 Setup +$216 – 600
Cost details Includes approved model calls, retries and shared hosting at the reference workload. Larger prompts, premium models or a hosted gateway change the cost. A tighter cap can leave too little budget for a fallback.
Reference machine cost: $0.026 – 0.029 per request at 10000 requests a month. The selected cap is a target to test, not a replacement for this estimate.
Where do you want it to run? Run the routing service in your cloud or on your machine. Choose whether requests may go to approved external models.
Run it onYour cloud Local environment AI model accessApproved model API No external calls
Data and access details Runs in an account you control, with server-side keys, access rules and request logs.
Route only to the models and endpoints your team permits. Enforce your data and spending rules for every attempt.
Keep keys and routing permissions on the server. Private mode uses models you supply on local hardware; model licences, capacity and running cost need testing. It cannot connect to the remote app in this starting scope.
How do you want to use it? Try a test page, call from your own agent or use a dedicated test agent. The connected scope adds an API for your app.
A routing test page Try approved requests and inspect the model, response, attempts and cost. Included My existing AI agent Let your agent request a response through the approved routing rules. Setup +$72 – 240 A dedicated test agent Use a separate agent to try requests and inspect routing results. Setup +$144 – 480