Verified against Gemini · 2026-07-25
Turn a long lecture or podcast recording into structured, checkable study notes
An audio-only multimodal prompt that turns a raw lecture, podcast, or interview recording into notes organized by concept and speaker, with timestamps and confidence flags on anything the audio makes genuinely hard to make out, instead of a single wall-of-text transcript.
The prompt
Ready to copy — highlighted parts are example details you can swap.
Listen to the audio I've uploaded above. This is a 55-minute university lecture on macroeconomic inflation targeting, roughly 55 minutes. WHO'S SPEAKING one primary lecturer, plus occasional student questions from the floor — if a speaker's identity isn't stated explicitly, use a consistent label like Speaker 1 rather than guessing a name from tone or accent. STRUCTURE organize by the concepts the lecturer actually covers, in the order covered, not a generic textbook-chapter outline INSTRUCTIONS 1. Organize notes by concept, in the order the recording actually covers them — don't impose a generic outline template that doesn't match how this specific recording is structured. 2. Note vocal emphasis where it's a real signal, not decoration: if the speaker raises their voice, repeats a point, or says something like "this matters" or "remember this," flag that line in the notes as emphasized — that's exactly the kind of moment a flat transcript throws away. 3. Attribute every claim and every question to the correct speaker. In a lecture with audience questions, do not let a tentative student question get folded into the notes as if it were an authoritative statement from the lecturer — keep questions and answers clearly separated and attributed. 4. Include a timestamp on every major note so I can jump back to the original audio to verify anything that matters. 5. Where the audio is mumbled, overlapping, or a term is unfamiliar enough that you're genuinely guessing at the word, mark it [unclear ~mm:ss] rather than transcribing a confident-sounding guess. A wrong term stated with total confidence is worse for someone studying from these notes than an honest gap. 6. enough detail that someone who missed the lecture could answer a quiz on it afterward 7. Where the recording references something visual that isn't itself in the audio — "as shown on this slide," "look at the diagram I'm pointing to" — note that a visual reference was made at that timestamp, so I know a slide deck or handout exists that these notes alone can't fully capture, rather than silently dropping the reference because there's no way to describe an image you can't see. 8. If the same concept is explained twice using two different explanations or analogies (common when a lecturer notices confusion and re-explains), keep both versions in the notes rather than merging them into one — the second explanation is often there specifically because the first one didn't land for part of the audience, and losing that alternate framing removes exactly the version that might work better for someone reviewing the notes later. OUTPUT FORMAT Notes grouped by concept/topic with timestamps, a separate short list of anything flagged as emphasized by the speaker, a list of every timestamp where a visual reference was made without accompanying audio description, and a final section naming every [unclear] span so I know exactly what to go back and re-listen to myself.
Customize
Optional — swap in your own details for the highlighted parts above.
Why this works
Gemini processes audio natively as a single input rather than running a separate speech-to-text pass and handing the LLM a flattened transcript, which is what lets it carry tone and emphasis cues — a raised voice, a repeated phrase, a paused "this will be on the exam" — into the notes as signals of importance instead of losing that information the moment audio becomes plain text. Requiring consistent speaker attribution, rather than a best-guess name inferred from voice, matters specifically for lecture recordings with audience Q&A: misattributing a hesitant student's half-formed question to the lecturer as if it were an authoritative claim quietly corrupts study notes in a way that's hard to catch later, since the notes read equally confidently either way. Organizing by the concepts actually covered, in the order the recording covers them, rather than a generic template, preserves the lecturer's own logical scaffold — which is very often the actual structure being tested, and a reorganized outline can accidentally erase the reasoning chain that connected one concept to the next. The [unclear] flagging rule targets a real and common failure specific to lecture and podcast audio: a mumbled technical term, an unfamiliar name, or a moment of cross-talk gets transcribed as the closest plausible-sounding word rather than flagged as uncertain, and a student who doesn't know the material yet has no way to tell a hallucinated term from a real one unless the model is explicit about where its own confidence actually drops. Flagging visual references the audio alone can't resolve — "as shown on this slide" — matters because an audio-only prompt has no access to what was actually on screen, and silently absorbing that reference into the notes as if nothing was missing would let a student believe the notes are complete when a meaningful piece of the explanation was never available to capture in the first place; naming the gap is what tells them a slide deck or handout still needs to be found and reviewed separately. Preserving a second explanation of the same concept, rather than merging it into the first, respects a specific pedagogical signal: a lecturer typically only re-explains something because the first framing visibly didn't land, and collapsing both into one merged version optimizes for compactness at the exact point where redundancy was actually doing useful work for the audience the explanation didn't reach the first time.
Verified against
Gemini Gemini 3 Pro · 2026-07-25
Changelog
- 2026-07-25 — Initial publish, verified against Gemini 3 Pro on a 55-minute recorded university lecture with audience Q&A.
Need this built into your business?
If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.
EXPLORE WHAT SCULT BUILDS
