Yearly running cost Assuming 12,000 test cases a year
By hand $54K / year 1,200 hours of work
Hosted service $2.4K+ / year
This workflow $396 / year Machine usage only; setup, hosting and review are extra. How should your AI outputs be tested? Start with saved answers or connect the results from your test pipeline. Compare versions using the same agreed rules.
One test case includes up to 1 saved answer and 5 scoring rules.
Evaluate saved answers Import examples and compare the results $864 – 2,640 setup Connect my test pipeline Report regressions after a test run $1,296 – 3,840 setup
Which test pipeline should supply outputs?GitHub Actions GitLab CI Existing API Other Score and report only. Never deploy a model, change production prompts or treat a judge score as a safety guarantee.
What is included Read saved prompts, answers and reference evidence. Apply your calibrated rules and export a case-level report with version comparisons.
The base scores supplied text outputs and reference evidence. Generating new answers, testing live tools, full red teaming and regulatory certification are separate scopes.
Prompt, answer and reference evidence: up to 2,000 combined words per test case.
How well should the evaluator spot problems? Choose targets for agreement with experts, missed failures and false alarms. These measure the evaluator, not the AI product's overall accuracy.
Standard Calibrated scoring with visible errors and version comparisons. Agreement with expert labels: ≥95% Known failures detected: ≥90% Included Strict Broader independent calibration and a second check on disputed decisions. Agreement with expert labels: ≥98% Known failures detected: ≥95% Setup +$288 – 840 My own targets Agree your own acceptance requirements with the provider. Quote separately
How these standards are measured Calibrate against independent expert labels, then test on held-out outputs. Include known failures, acceptable answers and attempts to manipulate the judge. Missing evidence and tool errors must never become passes.
Targets for your selected standard What is checked Target Agreement with expert labelsCompare decisions with a separately labelled holdout, not with the judge's own explanations. Balance rule and outcome groups so an easy majority cannot hide errors. ≥95% Known failures detectedCount detected known failures divided by all independently labelled failures. Report per failure type; judge errors count as missed decisions, not successful detection. ≥90% Acceptable answers passedCount correctly passed acceptable cases divided by all acceptable cases. Flagging everything cannot satisfy this measure. ≥95% Complete result recordsRequire case ID, source answer, dataset and rubric version, judge version, outcome and reason or explicit error. Never silently omit failed cases. ≥100%
Your team approves the test rules and release decision. A passing score does not establish safety or compliance outside the tested cases.
How quickly should a case be scored? Choose the time to check a saved answer and return its scores and reasons.
Within 60 seconds Included Within 30 seconds Setup +$144 – 480 Within 15 seconds Setup +$288 – 840
Timing details Measure all selected scoring rules, queueing, retries and report storage per case. Generating the answer under test and human adjudication are separate.
The target applies to at least 95% of agreed test runs, with 2 in progress at a time.
How much do you want to spend per test case? Choose the AI budget for scoring a saved answer against your rules.
Up to $0.10 Included Up to $0.07 Setup +$72 – 360 Up to $0.05 Setup +$216 – 600
Cost details Includes judge-model calls, retries and allocated hosting. Calls that generate candidate answers are separate. Comparing model versions scores each saved answer as its own case.
Reference machine cost: $0.05 – 0.08 per test case at 1000 test cases a month. The selected cap is a target to test, not a replacement for this estimate.
Where do you want it to run? Choose where the evaluation records live and whether approved test text may go to an external judge model.
Run it onYour cloud Local environment AI model accessApproved model API No external calls
Data and access details Runs in a cloud account you control, with access controls and logs.
Only approved source fields go to the selected external model. Agree access and retention first.
Remove secrets and unnecessary personal data before evaluation. Use local files and a private judge for no external calls; connecting a hosted test pipeline needs network access.
How do you want to use it? Read a results page, ask your existing agent to evaluate outputs or use a dedicated evaluation agent.
A results web page Inspect case-level scores, reasons and differences between versions. Included My existing AI agent Ask your existing agent to score saved outputs and retrieve regression reports. Setup +$72 – 240 A dedicated agent Use a dedicated evaluation agent to explain failures without changing production. Setup +$144 – 480