The Exact 5-Prompt Stack I Use to Generate High-Converting Ecom Video Ads Daily
The 5-Prompt Stack for High-Converting Ecommerce Video Ads
Building a video ad that actually converts is less about cinematic artistry and more about psychological sequencing. A high-converting e-commerce video ad isn't a single creative piece; it's a choreographed sequence of cognitive triggers designed to move a stranger from passive scrolling to an active purchase decision within thirty to sixty seconds. The challenge for most brands is not generating footage—it's structuring the narrative so that every second earns the next second of attention. This is where prompt engineering becomes a strategic lever, not just a technical skill.
The following five-prompt stack is designed as a pipeline: each output feeds into the next, and together they form a complete production workflow. You can adapt it to any product category—apparel, skincare, home goods, food, electronics—and any video length from 15-second social clips to 60-second hero spots. The logic beneath each prompt is what makes them reusable, so understanding the why matters as much as the wording.
Prompt 1: Product Deep-Dive and Audience Mapping
Before a single frame is generated, you need a structured brief that encodes the product's selling points into decision-relevant facts. Most brands write vague briefs—"premium," "elegant," "lifestyle-focused"—which gives a generative model little to work with. This prompt forces specificity by extracting the exact attributes a buyer would evaluate on a product page, then maps them onto two distinct audience segments: the primary buyer and the influencer/buyer pair (the person who recommends it).
Analyze [PRODUCT NAME] as follows:
1. Core Function: What problem does this solve?
2. 3 Differentiators: Why choose this over the top 2 competitors in [CATEGORY]?
3. Sensory Profile: Texture, sound, smell, weight — what would a user feel in the first 5 seconds of using it?
4. Price Positioning: Where does it sit in the category price band? What's the anchor comparison?
5. Primary Buyer Persona: Age range, income tier, lifestyle, top 3 purchase motivators (functional, emotional, social)
6. Recommendation Persona: Who would they tell about this product? What do they value that the buyer might not prioritize themselves?
Output as a structured JSON object with keys matching the above labels.This prompt produces a factual foundation that subsequent prompts can reference by key name, reducing ambiguity. The dual-persona structure is deliberate: e-commerce conversion often depends on two people's opinions—the one paying and the one who suggested it. Capturing both creates richer creative angles for the video script.
Prompt 2: Hook Engineering (First 3 Seconds)
The first three seconds determine whether someone keeps watching or scrolls past. On social platforms, this is not a soft metric—it is the conversion gate. A weak hook doesn't just lose engagement; it removes the ad from view before the selling points are even delivered. This prompt generates five distinct hook options across four proven structural archetypes:
Question Hook: Opens with a question that mirrors the viewer's latent desire ("What if your morning coffee routine took 30 seconds instead of 5?")
Contrast Hook: Juxtaposes the status quo against the post-product state
Sensory Hook: Leads with a specific tactile or auditory detail from the sensory profile in Prompt 1
Social Proof Hook: Opens with a concrete usage statistic or review-style line
Using the product brief above, write 5 hook scripts (max 20 words each) for the first 3 seconds of a [LENGTH]-second video ad. For each hook:
- Label its archetype (Question / Contrast / Sensory / Social Proof)
- Note which persona (Primary Buyer or Recommendation Persona) it primarily speaks to
- Identify the specific sensory detail or data point from the brief it references
Rank all 5 by predicted attention-retention strength and justify the top choice in one sentence.The ranking-and-justification requirement is important. It pushes the model past list-generation into evaluative reasoning, which produces more defensible creative choices than a flat menu of options. You're not just collecting hooks; you're getting an editorial recommendation with reasoning attached.
Prompt 3: Narrative Arc and Scene Scripting
With a chosen hook, this prompt structures the full video as a scene-by-scene script with explicit timing, camera direction, and voiceover/dialogue text. The structure follows a proven e-commerce arc: Problem → Product Reveal → Feature-Demo → Social Proof → CTA. Each scene is tied to a specific selling point from Prompt 1, so the script isn't generic—it's anchored in actual product attributes.
Using [CHOSEN HOOK] as Scene 1, write a full scene-by-scene script for a [LENGTH]-second video ad:
For each scene provide:
- Timecode range (e.g., 0:03–0:08)
- Visual direction (camera angle, movement, setting, on-screen subject)
- Voiceover or dialogue line (max 15 words per scene)
- On-screen text overlay (if applicable)
- The specific product attribute from the brief that this scene demonstrates
Structure: Problem (scenes 2–3) → Reveal (scene 4) → Feature-Demo (scenes 5–7) → Social Proof (scene 8) → CTA (final scene).
Total scenes should not exceed [MAX_SCENES] to fit the target duration.The word-count cap on voiceover lines is a practical constraint that prevents overwriting, which is the most common failure mode in generated ad scripts. It keeps each line tight and speakable within the timecode window. The attribute-mapping requirement ensures no scene is decorative; every visual moment is doing selling work.
Prompt 4: CTA Optimization and Platform Adaptation
A generic "Buy Now" CTA underperforms against a CTA that matches the platform's native interaction pattern. On Instagram Reels, the CTA might be "Save this for your next [USE CASE]"; on TikTok, it could be "Comment [KEYWORD] and I'll DM you the link." This prompt generates three CTAs per platform variant, each calibrated to the platform's primary user behavior (swipe-up, save, comment, share, tap-through).
Generate 3 CTA options for each of these platforms: Instagram Reels, TikTok, YouTube Shorts.
Constraints:
- Each CTA must reference a specific product attribute or use case from the brief
- Match the platform's dominant interaction (IG: save/share; TikTok: comment/duet; YT: tap/description)
- Max 12 words per CTA
- Include the psychological trigger each CTA leverages (urgency, exclusivity, convenience, social validation)
Output as a table: Platform | CTA Text | Trigger TypeThe trigger-type column is where this prompt adds analytical value. It's not just generating text; it's classifying the persuasion mechanism so you can A/B test with intention rather than guessing which variant will outperform.
Prompt 5: Asset Generation Spec (Visuals, Audio, Editing)
The final prompt translates the script into a production-ready asset specification that can be handed to a video generation tool (Runway, Pika, Sora, or an in-house pipeline) or a motion graphics team. It specifies resolution, aspect ratio, frame rate, audio layering, and color grading direction—all parameters that downstream tools require but creative writers rarely think about until after the script is done.
Convert the scene-by-scene script into a production spec:
1. Canvas: [9:16 / 16:9 / 1:1], resolution, fps (target 24 or 30)
2. Per-scene: image prompt for AI video generation (camera angle, lighting, color palette in hex), transition type to next scene
3. Audio: background music genre/BPM, voiceover tone and pacing note, SFX list per scene
4. Color Grade: dominant mood keyword + 3 reference brands whose aesthetic this should echo
5. Typography: font style for overlays, max characters per line
Output as a structured spec sheet. All image prompts must be under 80 words to fit standard generation tool input limits.The word-count constraint on image prompts reflects real-world tool limitations—most AI video generators have input token caps that make overly verbose scene descriptions either truncated or expensive. This prompt bakes in the practical constraints of the downstream pipeline, which is where most prompt stacks fail: they optimize for creative quality but not for tool compatibility.
How the Stack Works as a System
The five prompts are designed to be executed sequentially with each output feeding forward. Prompt 1 produces a JSON brief that Prompts 2–5 reference by key name. Prompt 2's ranked hooks feed directly into Prompt 3's scene script, which in turn feeds Prompt 4's CTA options and Prompt 5's production spec. The stack is modular—you can swap any single prompt without breaking the chain—as long as you preserve the JSON key names from Prompt 1.
One practical note: run each prompt with a temperature setting between 0.3 and 0.6 for e-commerce work. You want enough variation to generate multiple options per step, but not so much that outputs drift from the factual brief. A higher temperature (0.8+) works well for Prompt 2's hook generation, where you're explicitly seeking breadth; a lower one (0.3–0.4) is better for Prompts 3 and 5, where structural precision matters more than creative range.
The five-prompt stack turns video ad creation from an open-ended creative task into a structured production pipeline. Each prompt handles one decision point—product understanding, attention capture, narrative structure, conversion mechanics, and technical execution—and the JSON brief threads them together so that no step operates in isolation. That's what makes it repeatable: you're not generating five separate outputs; you're running a five-stage assembly line where each stage is constrained by the factual foundation laid down at the start.