Which recordings do you want to turn into clips? Pick complete moments from interviews, podcasts and explainers, then add captions and your brand style.
One video batch includes up to 30 source-video minutes and 3 clip drafts.
Clip recordings I upload Find moments, add captions and prepare exports $1,800 – 5,040 setup Pick up recordings from my folder Add a source-library connection and batch queue $2,232 – 6,240 setup
Where are your source recordings?Google Drive Other Review every clip before sharing. No automatic publishing, voice cloning or invented speech. Views, engagement and virality are not guaranteed.
What is included Upload an authorised recording and receive editable clip drafts with source timestamps, caption templates and portrait framing. Keep the source and edit settings for later changes.
Includes topic-led moment selection, editable cut points, caption templates, portrait framing and downloadable video with subtitle files. Use a speaker crop, split-screen or full-frame layout; uncertain crops stay flagged for review. Sports highlights and visual-only storytelling need different models and scope.
Maximum clip duration: up to 60 seconds.
Minimum clip duration: at least 15 seconds.
Maximum source height: up to 1,080 pixels.
Maximum frame rate: up to 30 frames per second.
Maximum source file: up to 1 GiB.
Caption templates: 3 styles.
Brand setup: 1 brand.
Caption timing tolerance: up to 200 milliseconds.
Minimum useful clips when suitable moments exist: at least 2 clips.
Maximum audio-video drift: up to 100 milliseconds.
Reference render capacity per worker: 4 physical CPU cores.
Reference render memory per worker: 8 GiB.
What should make a clip ready for review? A useful short clip makes sense on its own, keeps the speaker's meaning and needs little caption or framing correction.
Standard Use reusable caption styles, editable moments and checked portrait framing. Clips accepted for light editing: ≥80% Caption word accuracy: ≥97% Included Strict Add more checks for difficult captions, cut points and framing. Clips accepted for light editing: ≥90% Caption word accuracy: ≥99% Setup +$360 – 1,080 My own targets Agree your own clip requirements with the provider. Quote separately
How these standards are measured Test on unseen recordings with editor-labelled moments. Score complete meaning, caption accuracy, safe framing and timing. Count rejected drafts and recordings with no usable output, not just successful exports. Never stitch speech into a claim the speaker did not make.
Targets for your selected standard What is checked Target Clips accepted for light editingAn editor accepts the topic, complete meaning, context and cut points using a rubric agreed before testing. Count every requested draft, including duplicates and failures. Report no-output recordings and useful-clip yield separately. ≥80% Caption word accuracyCompare captions with a human transcript using word substitutions, deletions and insertions. Preserve names, numbers and negation; meaning-changing errors block sharing even if the overall score passes. ≥97% Important content kept in frameInspect frames sampled across every output clip plus all cut and layout changes. Check the speaker, slides and caption safe areas; unresolved crop failures count as failures, not excluded samples. ≥98% Caption timing within toleranceCompare caption boundaries with the reference audio timing. Check audio-video drift independently and reject exports with missing or clipped speech. ≥95% Exports pass technical checksValidate duration, dimensions, frame rate, readable captions, decodable video and intact audio for every requested export. Deliver subtitle files and cut settings; a successful render alone does not prove editorial quality. ≥100%
Meaning-changing edits, fabricated speech or unauthorised footage block sharing. Technical success and an AI score do not replace editorial approval.
How quickly do you need the clips? Choose how soon the previews and exports should be ready. Your review time is separate.
Within 20 minutes Included Within 10 minutes Setup +$144 – 480 Within 5 minutes Setup +$288 – 840
Timing details Start after the complete source file is available to the workflow. Include audio extraction, transcription, selection, captions, framing, queueing, retries and rendering. Upload time is shown separately; a requested new edit or batch is another run.
The target applies to at least 95% of agreed test runs, with 2 in progress at a time.
How much do you want to spend per video batch? Choose the machine budget for finding moments, transcribing speech and rendering the clip drafts.
Up to $7.00 Included Up to $4.00 Setup +$72 – 360 Up to $2.50 Setup +$216 – 600
Cost details Includes transcription, model calls, rendering, retries and shared hosting at the stated volume. A video batch is not a published clip: count failed runs and re-renders. Human editing, original recording and the calling agent are separate.
Reference machine cost: $2.34 – 6.34 per video batch at 10 video batches a month. The selected cap is a target to test, not a replacement for this estimate.
Where do you want it to run? Run in your cloud or on a machine you control. Choose whether audio and transcripts can go to approved AI APIs.
Run it onYour cloud Local environment AI model accessApproved model API No external calls
Data and access details Runs in a cloud account you control, with access controls and logs.
Audio and selected transcript context go to approved speech and text APIs. Agree retention before processing.
Only process recordings and brand assets you are allowed to use. Offline mode uses local files and private speech and text models, without cloud-folder pickup.
How do you want to use it? Use a preview web page, connect your existing agent or run a dedicated clipping agent.
A review web page Preview clips, adjust cut points, correct captions and export. Included My existing AI agent Ask your agent to prepare a clip batch and retrieve the previews. Setup +$72 – 240 A dedicated agent Use a dedicated agent to collect recordings and prepare clip drafts. Setup +$144 – 480