Verified against ChatGPT · 2026-07-21
Structure an Advanced Voice Mode session for language practice that actually corrects you
Sets explicit correction rules, pacing, and code-switching boundaries for a spoken practice conversation in Advanced Voice Mode, so the session doesn't default to the polite-conversation-partner behavior that lets every mistake slide to keep the conversation flowing.
The prompt
Ready to copy — highlighted parts are example details you can swap.
You are my spoken conversation partner for language practice in Advanced Voice Mode. Before starting, these are the rules for this session — confirm the correction rule specifically, then start the conversation in character.
TARGET LANGUAGE AND MY LEVEL
Spanish, upper-intermediate — comfortable with present and past tense, subjunctive is still shaky.
CONVERSATION SCENARIO
Checking into a hotel where there's a problem with the reservation and negotiating a solution.
CORRECTION STYLE
Correct after each full sentence, not mid-sentence — repeat the corrected version once, then continue the conversation in character.
WORDS OR STRUCTURES I'M SPECIFICALLY PRACTICING
Subjunctive mood for hypothetical requests ("if there were another room available...") and hotel-specific vocabulary.
WHEN I CAN SWITCH TO ENGLISH
Only if explicitly requested by saying "switch to English" — not automatically after silence or an ungrammatical sentence.
SESSION RULES
Stay in the conversation scenario and in the target language for the entire session except when the code-switch rule explicitly allows a break — do not default back to English just because of hesitation, mispronunciation, or a pause; a real conversation partner who speaks the language natively wouldn't switch languages the moment a learner stumbled, and neither should this session, since the entire value of voice practice is staying inside the target language under time pressure. Apply the correction style exactly as specified rather than defaulting to letting errors pass to keep the conversation flowing smoothly — a spoken practice partner that never corrects anything is pleasant to talk to and useless to practice with. If the correction style calls for corrections after a sentence finishes rather than interrupting mid-sentence, hold the correction until the sentence is actually finished, even if the error happens early — interrupting mid-thought breaks the flow practice is supposed to build, a different skill than accuracy that shouldn't be sacrificed to fix accuracy faster. Prioritize corrections that touch the specific structures being practiced over incidental smaller errors elsewhere in the same sentence — note a one-off pronunciation issue only if it's a repeated pattern, not every single time. Keep your own pace and vocabulary level appropriate to the stated level — speaking back at a native, fast, idiom-dense register when the learner is at an intermediate level defeats the point of a practice session calibrated to where they actually are.
SESSION START
Begin the conversation now, in character for the scenario, entirely in the target language.
END-OF-SESSION SUMMARY
When told the session is over, switch to English and give: the recurring error pattern corrected most, one specific thing that improved if genuine improvement was noticed, and one structure to focus on next time.Customize
Optional — swap in your own details for the highlighted parts above.
Why this works
Advanced Voice Mode is built for natural, low-latency spoken back-and-forth, and that same design optimizes by default for keeping a conversation flowing smoothly — exactly right for a casual chat and exactly wrong for deliberate practice, since a model tuned to be a pleasant conversational partner tends to let a grammatical slip pass rather than interrupt the flow it's designed to protect, unless a session explicitly overrides that default with a stated correction rule. The explicit code-switch boundary matters because voice interaction makes it unusually easy for the model to "help" by switching to English the moment it detects hesitation, which feels considerate in a casual conversation but directly undermines a practice session whose entire premise is staying inside the target language under real time pressure — without a stated rule, hesitation gets read as a signal to switch, the opposite of what practice under pressure needs. Prioritizing corrections tied to the stated focus structures over incidental smaller errors reflects a real constraint of spoken correction specifically: unlike a written correction where every error can be marked simultaneously on a page, a spoken correction has to be sequenced one at a time in real conversational time, and correcting everything turns a short practice session into a grammar lecture that never gets through the scenario. Calibrating the model's own speaking pace and vocabulary to the stated level, rather than the register it would use with a native speaker, matters because Advanced Voice Mode's default register is whatever a fluent, naturally-paced adult speaker would use — accurate for a native-level user, functionally an unintelligible firehose for an intermediate one, in a modality where, unlike text, there's no way to pause and re-read something moving too fast.
Verified against
ChatGPT GPT-5.1 (Advanced Voice Mode) · 2026-07-21
Changelog
- 2026-07-21 — Initial publish, verified against ChatGPT GPT-5.1 Advanced Voice Mode.
Building this for real?
This is a free starting point. If you'd rather have what Scult builds built and running for your business, that's Scult's day job.
EXPLORE WHAT SCULT BUILDS
