Purchasing Guides / Domain Speech Transcription

Domain Speech Transcription

A purchasing guide to buying a workflow that turns your industry recordings into accurate transcripts, with your terminology, speaker labels and audio-linked review.

Yearly running cost

Assuming 2,400 recordings a year

By hand
$72K / year
2,400 hours of work
Hosted PAYG
$403.2+ / year
This workflow
$484 / year
Machine usage only; setup, hosting and review are extra.
Where are your recordings?

Turn recorded speech into a transcript using your industry terminology.

One recording includes up to 30 audio minutes and 2 speakers and 50 glossary terms.

Obtain recording permissions and review important wording against the audio. No speaker identity verification, clinical decisions, legal certification or automatic publication.

What is included

Transcribe authorised recordings, highlight uncertain wording and link each segment to its audio. Export corrected text, subtitles and structured timestamps.

Start with clear English recordings and your glossary. Add a read-only recording folder if your team needs repeat imports. Live captions, translation and meeting summaries are separate.

Glossary token limit: up to 500 tokens.

Timestamp tolerance: up to 2 seconds.

How accurate should the transcript be?

Measure ordinary wording and important terms separately.

How these standards are measured

Test unseen recordings against independently corrected transcripts. Measure word errors, missed and invented technical terms, and speaker attribution separately. Silence must not produce invented speech.

Targets for your selected standard
What is checkedTarget
Word accuracyCompute word error rate from substitutions, deletions and insertions against the independently corrected transcript; word accuracy is the remaining proportion. Keep the same normalisation rules and report negative scores rather than hiding insertion errors.≥95%
Technical terms capturedCount correctly transcribed spoken glossary terms against all reference occurrences. Missing, uncertain or silently dropped terms count as misses.≥95%
Technical terms correctCheck output glossary terms against the audio, including recordings where those terms are absent. A glossary must not make the model invent a product name.≥98%
Correct speaker attributionCompare reference speech time after matching anonymous speaker labels. Missed speech and speech assigned to the wrong speaker count as errors. Test silence and overlapping speech separately.≥90%

Score raw machine output before human corrections. Speaker labels are anonymous, not identity verification. Check timestamps against the agreed tolerance and reject incomplete audio.

How quickly should the transcript be ready?

Choose the time from a complete recording to a reviewable transcript.

Timing details

Include decoding, transcription, retries, speaker labels and exports. Upload time and human correction are separate; private hardware must meet the same measured target.

The target applies to at least 95% of agreed test runs, with 2 in progress at a time.

How much do you want to spend per recording?

Choose the machine budget for transcription and automated checks.

Cost details

Includes speech API calls, retries and shared hosting at the reference volume. Human correction, storage beyond the agreed retention period and new hardware are separate.

Reference machine cost: $0.30 – 0.50 per recording at 200 recordings a month. The selected cap is a target to test, not a replacement for this estimate.

Where do you want it to run?

Run the workflow in your cloud or on your server. Choose whether recordings can reach an approved speech service.

Data and access details

Runs in a cloud account you control, with access controls and logs.

Send authorised recordings to your approved speech provider. Agree retention, region and account settings first.

Local deployment does not automatically keep audio offline. No external calls requires private speech and speaker models, downloaded before operation, with separate licence and hardware checks.

How do you want to use it?

Review transcripts on a web page, call the workflow from your agent or use a dedicated agent.

Anything else your provider should know?

Optional. Your choices are included automatically.

Common questions

Why not use Deepgram directly?

Use the API or a transcription app if it already meets your needs. Buy this workflow when you need your own glossary, recording source, review interface and measured domain accuracy.

Does this train a new speech model?

No. The baseline configures an existing speech model and your glossary. Weight training needs a separate dataset, rights check, evaluation and agreed scope.

Does it know who is speaking?

It assigns anonymous speaker labels. It does not verify identity from a voice, and labels still need checking.

Can recordings stay offline?

Yes, with private speech and speaker models on suitable hardware. That is a different deployment from using a hosted API and must be measured; proprietary model licences are not included.

Will it work for every accent or language?

Agree the language, audio conditions and representative test recordings first. The reference scope is clear English business audio, not every recording.

Download RenX

Get the app.

Install, sign up, and start on your free plan with welcome credit included. No credit card required.

On a platform not listed? Leave your email and we'll notify you when a build is available.