Purchasing Guides / Specialist Language Model Fine-tuning

Specialist Language Model Fine-tuning

A purchasing guide to buying a model adapted to your examples, tested on your tasks and deployed behind an interface you control.

Yearly running cost

Assuming 12,000 responses a year

By hand
$18K / year
600 hours of work
Subscription
Quote needed
This workflow
$552 / year
Machine usage only; setup, hosting and review are extra.
How will you supply the training examples?

Teach a small model your repeatable task, using examples your team has approved.

One response includes up to 2,048 input tokens and 512 output tokens.

Use licensed models and authorised examples. Human reviewers remain responsible for consequential outputs. No autonomous decisions, external actions or silent model updates.

What is included

Prepare the supplied examples, establish a prompted baseline, run bounded fine-tuning trials and deploy the accepted adapter. Include a reproducible recipe, held-out results and rollback.

Start with an approved dataset for a bounded text task. Deliver the adapter, training recipe, tests and running endpoint. Dataset labelling, broad research and foundation-model training are separate.

Approved training examples: up to 1,000 examples.

Training trials: up to 3 trials.

Included trial compute allowance: up to 20 USD.

How reliably should it perform your task?

Test real task completion, not whether the model sounds more specialised.

How these standards are measured

Compare the fine-tuned model with a well-prompted base model on unseen tasks. Score useful completion and factual support separately. Keep test data out of training and do not deploy a worse model just because it was fine-tuned.

Targets for your selected standard
What is checkedTarget
Tasks completed correctlyHave independent reviewers score unseen tasks against all required rubric items. Empty or incomplete responses fail answerable tasks; refusals cannot inflate success.≥90%
Valid output formatValidate every response against the agreed schema before repair. Count malformed, missing and truncated outputs as failures.≥99%
Supported factual claimsCheck factual output claims against the supplied context. Score required-fact coverage through task success so an empty answer cannot pass by making no claims.≥98%
Missing information handledUse genuinely unanswerable tasks and adversarial requests. Require an explicit request for missing evidence instead of fabricated facts, without refusing answerable tasks.≥95%

A fine-tune is not a live knowledge store. Provide current facts in the input; test abstention, privacy leakage and forgetting. Reject regressions on protected tasks and keep rollback available.

How quickly should each response be ready?

Choose the complete response time, including loading and retries.

Timing details

Measure request to validated output at the agreed workload. Include cold starts and queues. Faster settings may need warm capacity and a higher measured running budget.

The target applies to at least 95% of agreed test runs, with 2 in progress at a time.

How much do you want to spend per response?

Choose the machine budget for inference and output checks.

Cost details

Reference cost includes allocated GPU, CPU, memory, load and idle time, retries and shared hosting. Training is a separate setup activity; new hardware, human review and later retraining are separate.

Reference machine cost: $0.07 – 0.11 per response at 1000 responses a month. The selected cap is a target to test, not a replacement for this estimate.

Where do you want it to run?

Deploy in your cloud account or on your hardware. Decide whether training data can use an approved service.

Data and access details

Deploy in a compute account you control, with authentication, limits and a rollback version.

Use an approved training service for authorised examples; export the accepted adapter to your controlled inference deployment.

Approved training may use Tinker; the delivered model runs in your own controlled deployment. No external calls means private training and inference after provisioning, not merely downloading an adapter after sending your data out.

How do you want to use it?

Use a model test page, call from your AI agent or add a dedicated agent.

Anything else your provider should know?

Optional. Your choices are included automatically.

Common questions

Should I fine-tune or improve my prompts?

Start with a strong prompt and a baseline test. Fine-tune when repeated examples show a stable behaviour or format problem; use retrieval when the issue is access to current facts.

Do I receive the model?

You receive the trained adapter, recipe, base-model reference and deployed endpoint, subject to the agreed licences. You do not receive ownership of the base model or the training platform.

Will it definitely outperform the base model?

No. Agree a held-out release gate before training. If the adapter does not meet it, retain the better baseline and review the data or scope rather than deploying a regression.

Can all training stay private?

Choose private training and inference on suitable infrastructure. Sending examples to a hosted training API and later exporting the model is not private training.

Can I update it later?

Keep new examples versioned and rerun the same evaluation before release. Later training, compute and provider support need a separate agreed budget.

Download RenX

Get the app.

Install, sign up, and start on your free plan with welcome credit included. No credit card required.

On a platform not listed? Leave your email and we'll notify you when a build is available.