Gemini

Verified against Gemini · 2026-06-28

Translate and culturally localize the text inside a photo, not just the words

A multimodal prompt for translating a photographed menu, sign, or label that asks for cultural and practical localization alongside the literal translation, instead of a flat word-for-word swap.

Gemini (Gemini 3 Pro)Gemini app

The prompt

Ready to copy — highlighted parts are example details you can swap.

Here's a photo of a restaurant menu board, originally in Japanese. Translate it into English for a first-time visitor from the US with no dietary restrictions knowledge of the local cuisine.

For each item/line in the photo:
1. Give the literal translation.
2. Then a localized version if the literal translation would be confusing, misleading, or just odd to a English reader — explain in one line why you changed it (an idiom, a measurement unit, a culturally specific dish or reference).
3. Flag anything you can't read clearly in the photo as [unclear] rather than guessing text that isn't really there.
4. Convert any units, currency, or sizes to what's standard for a first-time visitor from the US with no dietary restrictions knowledge of the local cuisine, noting the original alongside the conversion.

If something in the source has no real equivalent in English/a first-time visitor from the US with no dietary restrictions knowledge of the local cuisine (a dish, an idiom, a cultural reference), say so explicitly instead of forcing an approximate translation that changes the meaning.
Customize the highlighted detailsoptional — the prompt above already works

Why this works

Reading the photo directly rather than routing through a separate OCR step lets Gemini use surrounding visual context (menu section headers, price formatting, layout) to disambiguate text that a plain OCR-then-translate pipeline would get wrong in isolation. The request for a localized version alongside the literal one directly targets the most common failure of flat translation: an idiom, a dish name, or a unit gets translated word-for-word into something technically accurate but meaningless or misleading to the reader, and asking the model to explain why it changed something keeps that judgment call visible instead of hidden inside a single 'translation.' The [unclear] flag matters specifically for real photos, which are often at an angle, partially glared-over, or slightly blurry — a model translating confidently from a misread character produces a wrong dish description with no signal that anything was uncertain.

What you get back

Line 3: 焼き餃子 (yakigyoza) Literal: "Fried gyoza" Localized: "Pan-fried pork dumplings" — kept as "gyoza" is understood in English now, but added "pan-fried" since "fried" alone suggests deep-fried to most US readers, which this isn't. Line 7: [unclear] — price partially obscured by a sticker, showing only "¥8_0".

Verified against

Gemini Gemini 3 Pro · 2026-06-28

Changelog

  • 2026-06-28 Initial publish, verified against Gemini 3 Pro on a photographed Japanese restaurant menu.

Need this built into your business?

If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.

EXPLORE WHAT SCULT BUILDS
All Gemini prompts

Check your AI visibility

One URL in, a 0–100 score and the exact fixes out.

RUN THE CHECK

Browse all the tools

15 tools across six categories
13 of them never send your data anywhere

Free · No signup · No trial clock

SEE THE DIRECTORY