marketing-video-automation
Automate marketing/demo videos for mobile, web, Electron, and macOS/Tauri apps: deterministic storyboard, screen capture, zoom/GIF/WebP polish. Use for 'marketing video', 'App-Demo aufnehmen', 'Screen Studio style', 'App Store preview video'.
Works with
--- name: marketing-video-automation description: Automate marketing/demo videos for mobile, web, Electron, and macOS/Tauri apps: deterministic storyboard, screen capture, zoom/GIF/WebP polish. Use for 'marketing video', 'App-Demo aufnehmen', 'Screen Studio style', 'App Store preview video'. license: MIT --- # Marketing Video Automation Produces a demo/marketing video by automating the app itself rather than free-hand clicking during the recording. A freely improvising agent driving the app live, on camera, is the unreliable path — jitter, mis-clicks, and dead air all end up baked into the footage with no way to cut them out. The reliable path is a **storyboard**: a script the agent writes and rehearses first, then replays verbatim while a separate recorder captures the screen. ## 1. Pick the platform branch | App runs as | Automate with | Reference | |---|---|---| | iOS/Android app (simulator or device) | Maestro | [`mobile-maestro.md`](mobile-maestro.md) | | Web app in a browser | Playwright CLI | [`web-playwright.md`](web-playwright.md) | | Electron app | Playwright (experimental Electron support) | [`web-playwright.md`](web-playwright.md) | | Native macOS app, or Tauri (WKWebView) desktop app | AppleScript/JXA or Hammerspoon + native screen capture | [`desktop-macos.md`](desktop-macos.md) | A Tauri app is Rust with a WKWebView front end, not Chromium — Playwright/browser automation does not see the window as it actually renders. Treat it as a native macOS app: the desktop branch, not the web branch, even though the UI is built with web tech. Load the matched reference file before continuing; it carries the install commands, script templates, and capture commands for that branch. **Done when:** the branch is picked and its reference file is loaded. ## 2. Design the story, then write the storyboard **One clip, one benefit.** A clip that tours the whole app teaches nothing; a clip built as a three-beat story — starting problem/state, one legible action, a visibly better result — teaches one thing well. Prefer several short, single-benefit clips over one long walkthrough. The payoff has to land in the first few seconds: cut every intro, loading screen, and setup click that doesn't serve that one beat. **Every motion points at the beat.** Cursor moves, zooms, and highlights exist only to direct attention to what matters right now — never decoration. Zoom only where a detail would otherwise be illegible, move one thing at a time, and let each state change visibly connect its before → action → after, so the sequence reads as cause and effect rather than a slideshow of screens. Now write the exact, numbered sequence the video will show — every tap/click, every text entered, every wait, every expected resulting screen — as the script format native to the branch (Maestro YAML, a Playwright CLI script, an AppleScript/JXA file, or a Hammerspoon function). Use the branch's exploration tool (Maestro MCP, Playwright MCP, or `macos-automator-mcp`'s accessibility-tree inspection) to find stable selectors — accessibility labels, testIDs, menu items — never fixed screen coordinates, which break the moment a window resizes or a font scales. The opening steps have to reach a clean demo state: realistic, readable placeholder data, every window that belongs in the shot at a fixed size and position, no notifications, no error states, nothing private — and, for a desktop capture, nothing that *doesn't* belong in the shot: no stray windows, no Dock, no menu bar (see the desktop branch reference for how to stage that; it also covers the case where several windows are deliberately composed into one shot). If the app already has a way to load that state directly — a debug/demo menu, a seed script, a launch flag, a `?demo=1` URL — use it instead of reconstructing the state by re-typing fixture data through the UI on every take; it's faster, and it can't itself glitch on camera. If the app has no such hook yet, that's worth adding as a small dev-only feature before doing repeat marketing-video work on it, rather than fighting the same manual setup every time. Pace it like a demo, not a test: a brief settle beat before each key action, and enough hold time on the final screen for the result to actually register. A storyboard that fires every step back-to-back reads as an automated test, not something a person did. If the clip is meant to loop, make the first and last frame match so the loop doesn't visibly jump. Dry-run the storyboard without recording. Fix every flaky selector and every race (a tap that lands before the previous screen's animation finished) until it is boring: the same outcome, unattended, every time. **Done when:** the storyboard dry-runs to completion twice in a row with no manual intervention on the second run, its first frame is legible and its last frame reads as a clear result, and — for a looping deliverable — the first and last frames match. ## 3. Capture one raw take Start the recorder, replay the storyboard **unmodified**, stop the recorder. Never improvise a click or adapt the flow mid-recording — the storyboard from step 2 is the single source of truth for what's on screen; recording is a mechanical replay of it, not a fresh performance. If the app misbehaves during the take, stop, fix the storyboard or the app state, and take again from the top — don't patch around a bad take in post. **Done when:** one continuous raw video file exists spanning the full storyboard, and `ffprobe` confirms its duration is within a few seconds of the storyboard's expected run time. ## 4. Polish and export Apply zoom/cursor-highlight polish if the deliverable calls for that look, then export the deliverable that fits where it's going: silent autoplaying `<video>` for a website, GIF/WebP only where the platform specifically needs a static-embeddable format, and a spec-compliant MP4 for an App Store/Play Store preview. Any on-screen text — burned-in captions, callouts, an animated headline — is opt-in: default to none, and add it only when the user actually asks for it; that's separate from the clip's benefit also needing to exist as real page text nearby, which stands regardless. See [`post-production.md`](post-production.md) for exact commands, delivery markup, and format specs. **Done when:** every requested output file exists, and `ffprobe`/`gifsicle` confirm it opens cleanly at the target resolution — not just that the export command exited zero.
More SEO & Marketing skills
ai-video-generation
skills-101/superpowers
Generate AI videos with Google Veo, Seedance 2.0, HappyHorse, Wan, Grok and 40+ models via inference.sh CLI. Models: Veo 3.1, Veo 3, Seedance 2.0, HappyHorse 1.0, Wan 2.5, Grok Imagine Video, OmniHuman, Fabric, HunyuanVideo. Capabilities: text-to-video, image-to-video, reference-to-video, video editing, lipsync, avatar animation, video upscaling, foley sound. Use for: social media videos, marketing content, explainer videos, product demos, AI avatars. Triggers: video generation, ai video, text to video, image to video, veo, animate image, video from image, ai animation, video generator, generate video, t2v, i2v, ai video maker, create video with ai, runway alternative, pika alternative, sora alternative, kling alternative, seedance, happyhorse
ai-image-generation
skills-101/superpowers
Generate AI images with GPT-Image-2, FLUX, Gemini, Grok, Seedream, Reve and 50+ models via inference.sh CLI. Models: GPT-Image-2, FLUX Dev LoRA, FLUX.2 Klein LoRA, Gemini 3 Pro Image, Grok Imagine, Seedream 4.5, Reve, ImagineArt. Capabilities: text-to-image, image-to-image, inpainting, LoRA, image editing, upscaling, text rendering. Use for: AI art, product mockups, concept art, social media graphics, marketing visuals, illustrations. Triggers: flux, image generation, ai image, text to image, stable diffusion, generate image, ai art, midjourney alternative, dall-e alternative, text2img, t2i, image generator, ai picture, create image with ai, generative ai, ai illustration, grok image, gemini image, gpt image, openai image, chatgpt image
ai-avatar-video
skills-101/superpowers
Create AI avatar and talking head videos via inference.sh CLI. Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS). Also: OmniHuman, Fabric, PixVerse. Audio: Inworld TTS-2 (100+ languages, emotion steering for characters), ElevenLabs, Kokoro. Capabilities: audio-driven avatars, text-to-avatar, lipsync videos, talking head generation, virtual presenters, UGC content. Use for: AI presenters, explainer videos, virtual influencers, dubbing, marketing videos, UGC ads, gaming avatars, NPC dialogue. Triggers: ai avatar, talking head, lipsync, avatar video, virtual presenter, ai spokesperson, audio driven video, heygen alternative, synthesia alternative, talking avatar, lip sync, video avatar, ai presenter, digital human, ugc, ugc video, ugc ad, avatar ugc

