Verified against ChatGPT · 2026-08-12
Write a 60-second voiceover script that's timed to the second, not just word-counted
Produces a voiceover script broken into timed segments matched to a realistic spoken pace, with pacing notes and a natural breath point marked, so it actually fits the video length instead of running long once a narrator reads it aloud.
The prompt
Ready to copy — highlighted parts are example details you can swap.
Write a voiceover script for a 60 seconds explainer video. This needs to actually fit the runtime when read aloud at a natural pace, not just look like the right length on the page — most voiceover scripts run long because they're word-counted instead of time-tested. PRODUCT OR SUBJECT A budgeting app feature that automatically splits a paycheck into savings, bills, and spending buckets. VIDEO STRUCTURE (what's on screen) 0-8s: problem (a chaotic pile of bills), 8-25s: app opens and auto-splits the paycheck, 25-45s: buckets update in real time as spending happens, 45-60s: app icon and download CTA. TARGET AUDIENCE People in their 20s who've never used a budgeting app and are mildly skeptical that one would actually help. READ PACE 150 words per minute, a relaxed conversational read rather than a rushed ad-read pace. RULES Budget the script against a spoken pace of 150 words per minute, a relaxed conversational read rather than a rushed ad-read pace. words per minute, not the faster pace of silent reading — calculate the actual word count ceiling for 60 seconds at that pace before writing a single line, and state that ceiling at the top of your output so the budget is checked, not assumed. Break the script into segments that match the on-screen structure given, and mark an approximate timestamp at the start of each segment (0:00, 0:08, 0:15, and so on) so it's clear which line should be playing under which visual. Write for the ear, not the eye: short sentences, no subordinate clauses stacked three deep, and no sentence a narrator would have to reread to figure out where the emphasis goes. Mark one natural breath point per sentence longer than about twelve words with a forward slash, so a voice talent has an explicit cue rather than guessing where to pause. End on a single clear call to action that fits in one breath, not a sentence that tries to both summarize the video and ask for the action. WHAT NOT TO DO Do not write a script that reads well silently but would run over 60 seconds at a spoken pace — if the draft comes in over budget, cut content rather than asking the narrator to speed up, since a sped-up read undercuts clarity and trust. Do not stack more than one idea per on-screen segment; if a segment needs two ideas, that's a sign the on-screen structure needs an extra beat, and you should flag that rather than cramming both ideas into the given timestamp. OUTPUT FORMAT 1. The word count ceiling you calculated, and the actual word count of your draft, so I can see it's within budget. 2. The full script broken into timestamped segments with breath marks. 3. One line flagging any place the on-screen structure was too tight for the content and what you cut to make it fit.
Customize
Optional — swap in your own details for the highlighted parts above.
Why this works
Voiceover scripts that run long almost never fail because the writer misjudged the topic — they fail because word count and spoken duration are different units, and a model asked to write "a 60-second script" defaults to producing a script that looks the right length on the page, since it has no built-in mechanism forcing it to check word count against a spoken-pace ceiling unless explicitly told to calculate one before drafting. Requiring the ceiling to be computed and stated up front, then checked against the actual draft's word count, converts an implicit assumption into a verifiable number the writer can catch before ever handing the script to a narrator, which matters because the cost of discovering a script runs long is highest at the recording session, not at the drafting stage. Timestamping segments against the given on-screen structure keeps the audio and visual tracks synchronized at the level of the actual deliverable, rather than producing a generic paragraph of narration that someone downstream has to manually chop and match to the storyboard — a task that introduces its own errors. The instruction to write for the ear rather than the eye targets a specific and common defect in model-written voiceover: sentences with stacked subordinate clauses read fine silently but force a narrator to make an unplanned interpretive choice about emphasis or pause mid-recording, and marking explicit breath points removes that ambiguity by giving the voice talent a concrete cue instead of a guess. Instructing the model to cut content rather than ask the narrator to speed up when a draft runs over budget matters because a rushed read measurably reduces comprehension and perceived trustworthiness in explainer-video research, so the fix belongs in the script, not in the performance.
What you get back
Word budget: 60 seconds at 150 wpm = approx. 150 words. Draft: 147 words. [0:00] Bills pile up. / Paychecks disappear before you notice. [0:08] What if your money sorted itself the moment it landed? [0:25] Meet AutoSplit — it divides every paycheck into savings, / bills, and spending, automatically.
Verified against
ChatGPT GPT-5.1 · 2026-08-12
Changelog
- 2026-08-12 — Initial publish, verified against ChatGPT GPT-5.1.
Need this built into your business?
If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.
EXPLORE WHAT SCULT BUILDS
