Yearly running cost Assuming 12,000 responses a year
By hand $18K / year 600 hours of work
Subscription Quote needed
This workflow $552 / year Machine usage only; setup, hosting and review are extra. How will you supply the training examples? Teach a small model your repeatable task, using examples your team has approved.
One response includes up to 2,048 input tokens and 512 output tokens.
Train from approved example files Your task, model adapter and running endpoint $1,728 – 4,800 setup Read an approved dataset folder Versioned imports and reviewed dataset updates $2,304 – 6,480 setup
Which approved dataset source should it read?Amazon S3 Other Use licensed models and authorised examples. Human reviewers remain responsible for consequential outputs. No autonomous decisions, external actions or silent model updates.
What is included Prepare the supplied examples, establish a prompted baseline, run bounded fine-tuning trials and deploy the accepted adapter. Include a reproducible recipe, held-out results and rollback.
Start with an approved dataset for a bounded text task. Deliver the adapter, training recipe, tests and running endpoint. Dataset labelling, broad research and foundation-model training are separate.
Approved training examples: up to 1,000 examples.
Training trials: up to 3 trials.
Included trial compute allowance: up to 20 USD.
How reliably should it perform your task? Test real task completion, not whether the model sounds more specialised.
Standard A domain adapter with independent task, format and unsupported-answer tests. Tasks completed correctly: ≥90% Valid output format: ≥99% Included Strict More challenge tests and a verification pass before returning each response. Tasks completed correctly: ≥95% Valid output format: ≥99.5% Setup +$432 – 1,440 My own targets Agree task-specific release and regression thresholds. Quote separately
How these standards are measured Compare the fine-tuned model with a well-prompted base model on unseen tasks. Score useful completion and factual support separately. Keep test data out of training and do not deploy a worse model just because it was fine-tuned.
Targets for your selected standard What is checked Target Tasks completed correctlyHave independent reviewers score unseen tasks against all required rubric items. Empty or incomplete responses fail answerable tasks; refusals cannot inflate success. ≥90% Valid output formatValidate every response against the agreed schema before repair. Count malformed, missing and truncated outputs as failures. ≥99% Supported factual claimsCheck factual output claims against the supplied context. Score required-fact coverage through task success so an empty answer cannot pass by making no claims. ≥98% Missing information handledUse genuinely unanswerable tasks and adversarial requests. Require an explicit request for missing evidence instead of fabricated facts, without refusing answerable tasks. ≥95%
A fine-tune is not a live knowledge store. Provide current facts in the input; test abstention, privacy leakage and forgetting. Reject regressions on protected tasks and keep rollback available.
How quickly should each response be ready? Choose the complete response time, including loading and retries.
Within 5 minutes Included Within 3 minutes Setup +$144 – 480 Within 2 minutes Setup +$288 – 840
Timing details Measure request to validated output at the agreed workload. Include cold starts and queues. Faster settings may need warm capacity and a higher measured running budget.
The target applies to at least 95% of agreed test runs, with 2 in progress at a time.
How much do you want to spend per response? Choose the machine budget for inference and output checks.
Up to $0.20 Included Up to $0.15 Setup +$72 – 360 Up to $0.10 Setup +$216 – 600
Cost details Reference cost includes allocated GPU, CPU, memory, load and idle time, retries and shared hosting. Training is a separate setup activity; new hardware, human review and later retraining are separate.
Reference machine cost: $0.07 – 0.11 per response at 1000 responses a month. The selected cap is a target to test, not a replacement for this estimate.
Where do you want it to run? Deploy in your cloud account or on your hardware. Decide whether training data can use an approved service.
Run it onYour cloud Local environment AI model accessApproved model API No external calls
Data and access details Deploy in a compute account you control, with authentication, limits and a rollback version.
Use an approved training service for authorised examples; export the accepted adapter to your controlled inference deployment.
Approved training may use Tinker; the delivered model runs in your own controlled deployment. No external calls means private training and inference after provisioning, not merely downloading an adapter after sending your data out.
How do you want to use it? Use a model test page, call from your AI agent or add a dedicated agent.
A model test page Send a task, inspect its output and compare model versions. Included My existing AI agent Call your specialist model through a controlled tool. Setup +$72 – 240 A dedicated agent Use a dedicated agent around the deployed model endpoint. Setup +$144 – 480