Grok

Verified against Grok Voice Mode · 2026-08-03

Design a Grok Voice Mode persona brief for a real spoken use case

Writes a conversational persona and turn-taking brief for Grok Voice Mode's real-time spoken interaction, scoped to a specific hands-busy use case, instead of a generic chatbot personality description that ignores what makes spoken conversation different from typed chat.

Grok Voice Mode5 fillable variables

The prompt

Ready to copy — highlighted parts are example details you can swap.

You are writing a persona and interaction brief for Grok Voice Mode, for a specific real-time spoken use case where the user's hands and eyes are occupied with something else. A voice interaction has a different set of constraints than a typed chat — no scrollback to reread, no skimming ahead, and a real cost to a response that's too long to hold in working memory while listening — so this brief needs to specify the interaction's shape, not just the assistant's personality.

USE CASE
A hands-free cooking assistant guiding someone through a recipe step by step while their hands are covered in flour

USER STATE WHILE TALKING
Standing at a counter, hands occupied, attention split between the assistant and the actual cooking task, likely to interrupt to ask about substitutions or timing

PERSONA AND TONE
Warm and unhurried, like a patient friend talking you through a recipe over the phone, not a brisk instructional voice reading a list

INTERRUPTION AND CORRECTION BEHAVIOR
If interrupted mid-step, stop immediately, answer the interruption directly, then offer to resume the step rather than restarting the whole instruction from the beginning

WHAT MUST NEVER HAPPEN
Never suggest an ingredient substitution involving a known allergen category, like nuts, shellfish, or dairy, without first asking whether that’s a concern for this cook, even if it seems like the obvious swap

BRIEF RULES
Keep every spoken response short enough to be understood on first listen without rereading — there is no rereading in a voice interface, so a response that would be perfectly fine as three sentences of dense text needs to be restructured for listening comprehension: shorter sentences, one idea per sentence, and the most important piece of information stated first in case attention lapses partway through. Design for interruption as a normal event, not an error case — Standing at a counter, hands occupied, attention split between the assistant and the actual cooking task, likely to interrupt to ask about substitutions or timing means the user may need to cut the assistant off mid-sentence to react to something in their actual environment, and the assistant resuming exactly where it left off, rather than restarting the whole response or losing the thread entirely, is the difference between a usable hands-busy assistant and an annoying one. Never make the user repeat information they already gave earlier in the same conversation just because the assistant's turn ended — the persona should track what's already been said and refer back to it naturally, the way a person would, rather than asking a redundant clarifying question a moment after the answer was already given. Build in an explicit, low-effort way for the user to correct a misunderstanding without a long back-and-forth — a single short phrase that resets or redirects, since a voice interface with a clunky correction flow is worse than a text interface with the same problem, because correcting by voice while distracted is inherently harder than typing a correction. Respect the hard constraints named above as absolute, not as tone guidance to be balanced against being helpful — if a hard constraint says never do something, the persona brief should make clear that not doing it takes priority over sounding more natural or more helpful in the specific moment where the two conflict.

OUTPUT FORMAT
1. A one-paragraph persona description: tone, pacing, and vocabulary level appropriate to A hands-free cooking assistant guiding someone through a recipe step by step while their hands are covered in flour and Standing at a counter, hands occupied, attention split between the assistant and the actual cooking task, likely to interrupt to ask about substitutions or timing.
2. Three to five example exchanges showing realistic back-and-forth, including at least one interruption and one correction.
3. The specific phrase or pattern the user should be able to use to correct a misunderstanding.
4. How the hard constraints get enforced in practice, with a concrete example of the assistant declining or redirecting rather than complying.

Customize

Optional — swap in your own details for the highlighted parts above.

Why this works

The no-rereading constraint is the structural fact this entire brief is built around, and it's genuinely different from designing for a typed chat interface: a user reading text can skim ahead, reread a dense sentence, or scroll back to check a detail, none of which is available in a spoken interaction, so a response that would read as perfectly clear and appropriately detailed in text can be genuinely incomprehensible spoken aloud at the same information density, which is why the brief asks for restructuring, not just shortening. Treating interruption as a normal event rather than an edge case reflects how Grok Voice Mode's real-time spoken turn-taking actually needs to behave for a hands-busy use case specifically — a user with flour-covered hands mid-recipe is not politely waiting for the assistant to finish a sentence before reacting to something on the stove, and a voice assistant that either can't be interrupted cleanly or restarts its entire response after being cut off creates real friction precisely at the moment the user most needs a fast, situational answer, not a full replay of information they already partly heard. The instruction against making the user repeat already-given information addresses a specific credibility cost unique to sustained voice interaction: a typed chat interface visibly shows the whole conversation history, so a redundant question reads as a minor annoyance the user can just glance up and reference, but a spoken assistant that asks something already answered moments ago reads as not actually listening, which erodes trust in the assistant's competence far faster in voice than the equivalent slip would in text. The requirement for a specific, low-effort correction phrase matters because correcting a misunderstanding by voice while distracted is measurably harder than typing a correction, there's no edit-and-resend, no visible transcript to point at, so a correction path that requires a multi-turn clarifying exchange imposes a real cognitive cost exactly when the user has the least attention to spare for it. Making the hard constraints override tone rather than compete with it matters because a persona explicitly designed to sound warm, natural, and accommodating creates real pressure, in the moment, to just agree with a reasonable-sounding request, and a brief that treats the constraint as one more consideration to balance against sounding helpful, rather than as something that simply wins the conflict outright, leaves exactly the gap where a warm, agreeable persona talks itself into the one thing it was never supposed to do.

Verified against

Grok Voice Mode Grok 4.1 · 2026-08-03

Changelog

  • 2026-08-03 Initial publish, verified against Grok Voice Mode on Grok 4.1.

Need this built into your business?

If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.

EXPLORE WHAT SCULT BUILDS
All Grok prompts

Check your AI visibility

One URL in, a 0–100 score and the exact fixes out.

RUN THE CHECK

Browse all the tools

15 tools across six categories
13 of them never send your data anywhere

Free · No signup · No trial clock

SEE THE DIRECTORY