Yearly running cost Assuming 12,000 test runs a year
By hand $60K / year 1,000 hours of work
Subscription $1.5K+ / year
This workflow $230 / year Machine usage only; setup, hosting and review are extra. What should the test workflow cover? Test your main user flows, or also check preview builds and file bug reports.
Write and maintain tests Run tests in your CI and propose repairs for review. $1,296 – 3,720 setup Add preview checks and bug reports Also test pull request previews and file bugs with evidence. $1,944 – 5,760 setup
Which CI runs your tests?GitHub Actions GitLab CI CircleCI Jenkins Buildkite Other Your engineers merge test changes and decide what counts as a bug. It never weakens a test to make a real bug pass and never touches production data.
What is included Write browser tests for the user flows you describe, run them in your CI, sort each failure into an app bug, a test that needs updating or an unstable environment, and open a repair pull request when an intended change broke a test.
Start with your most important user flows written in plain language, a test environment with test accounts and the CI you already use. The wider option also runs tests on preview builds of each pull request, files bug reports with steps, screenshots and a recording, and keeps a map of which flows are covered.
How reliable should the tests and failure reports be? Set targets for test coverage, bugs caught and correct failure explanations.
Standard Working tests for the agreed flows, with unclear failures handed to an engineer rather than guessed. Flows with a passing test: ≥90% Planted bugs caught: ≥95% Included Strict A second check on each failure explanation and repair, and more planted-bug tests. Flows with a passing test: ≥95% Planted bugs caught: ≥98% Setup +$288 – 840 My own targets Agree your own acceptance requirements with the provider. Quote separately
How these standards are measured Test with the agreed user flows, with bugs planted in a test build of your app and with past failed runs whose cause is known. A repair that changes what a test checks so that a real bug passes counts as a critical failure. Measure planted bugs caught separately from failures sorted correctly.
Targets for your selected standard What is checked Target Flows with a passing testAgreed user flows that end in a passing test which checks the agreed outcome, divided by all agreed flows. ≥90% Planted bugs caughtBugs planted in a test build of your app that make a covered test fail, divided by all planted bugs in covered flows. ≥95% Failures sorted correctlyFailed runs sorted the same way as your engineer would sort them, as an app bug, a test that needs updating or an unstable environment, divided by all failed runs in the test set. ≥85% Test fixes acceptedTest repair pull requests your engineers merge without rework, divided by all repair pull requests. A repair that lets a real bug pass counts as a critical failure. ≥70%
Test results and failure explanations help your engineers decide; they are not a release approval. Your team still decides what ships.
How quickly should a failed run be explained? Choose how soon after a test fails the explanation and, where needed, a repair pull request appear. The test run itself and human review time are separate.
Within 10 minutes Included Within 5 minutes Setup +$144 – 480 Within 2 minutes Setup +$288 – 840
Timing details Time from a failed run to a posted explanation or repair pull request, including reading the trace, screenshots and recent changes, model calls, a confirming rerun, queueing and retries. Confirm suite size and runner capacity with your provider.
The target applies to at least 95% of agreed test runs, with 8 in progress at a time.
How much do you want to spend per test run? Set the average AI processing budget per test run. Passing tests do not need model calls.
Up to $0.10 Included Up to $0.05 Setup +$72 – 360 Up to $0.03 Setup +$216 – 600
Cost details Includes model calls for writing new tests and explaining failures, retries and a shared hosting allocation. Your CI runner minutes, test environment and the calling agent are separate.
Reference machine cost: $0.04 – 0.08 per test run at 1000 test runs a month. The selected cap is a target to test, not a replacement for this estimate.
Where do you want it to run? Run it in your cloud or on your own server. Choose whether test traces, screenshots and code may go to an approved AI service.
Run it onYour cloud Local environment AI model accessApproved model API Private model only
Data and access details Runs in a cloud account you control, next to your test environment, with access controls and logs.
Only flow descriptions, test code, traces and screenshots from the test environment go to the selected external model. Agree access and retention first.
Use test accounts and test data only; never point it at production customer accounts. Give it permission to open pull requests on test files only. Private model only keeps code and screenshots on your hardware; the workflow still connects to your own code host and test environment.
How do you want to use it? Choose where you want to use it. You can select more than one.
Results in my CI and pull requests Tests live in your repository and run in your CI. Failures get a short explanation and test repairs arrive as pull requests. Included My existing AI agent Let your coding agent ask for a test of a new flow or an explanation of a failed run. Setup +$72 – 240 A test dashboard See which user flows are covered, recent failures, unstable tests and which repairs were accepted. Setup +$144 – 480