Exact timing comes at the edit
A model’s supported clip lengths may not match the final ad duration. Trim or combine rendered clips, and measure the final output rather than promising exact timing from the generation prompt alone.
Product photo → narrated social ad
Use Creative Claw in ChatGPT to turn a product photo and approved copy into a vertical video ad. Build the visuals, generate the voiceover, combine them, and add captions to the final audio. A 15-second finished ad may require several generation and editing jobs; it is not necessarily one 15-second model call.
Bring
A product photo, approved benefits, audience, and call to action
Create
One reviewed vertical video with narration and captions
The workflow
Connect Creative Claw in ChatGPT and attach a clear product reference. Specify the audience, one real benefit, a call to action, and a target 9:16 format. Do not ask the model to invent customer experiences or product claims.
Plan a visual hook, a benefit demonstration, and a closing product shot. For a 15-second target, start with roughly 25–35 spoken words, then adjust to the actual voiceover length. Review the proposed models and estimated credits before generating.
Use the product reference for each applicable shot. Check packaging, proportions, logos, and continuity. Approve replacements and their cost before retrying shots that drift.
Generate the approved script as a separate voiceover, listen to it, and combine it with the approved clips. Trim and sequence the visuals to fit the target duration. Decide whether generated scene audio should remain under the narration.
Add synchronized captions after the narration and edit are settled. Check spelling, contrast, timing, and placement around the product. Review the final duration and portrait framing, then request the finished MP4 rather than separate source assets.
Before you render
A model’s supported clip lengths may not match the final ad duration. Trim or combine rendered clips, and measure the final output rather than promising exact timing from the generation prompt alone.
This workflow adds narration to product footage. A person speaking visibly on camera needs a compatible talking-presenter or lip-sync workflow and a separate quality check.
Include image preparation, each video shot, speech, editing, captioning, and possible retries. A per-clip generation estimate is not the total price of a finished ad.
Check current pricing, inspect the tool reference, or browse published media examples. Examples illustrate creative approaches; they are not proof that this exact starter prompt was run.
Practical answers
Yes. Generate the approved narration, combine it with the video, then caption the settled audio. These are separate operations; review the final words, timing, and placement before publishing.
Use 15 seconds as the edit target. Model clip durations vary, so the workflow may need trimming or assembly. Verify the final file duration instead of relying on the video prompt alone.
Not necessarily. Reference-guided video can distort text, logos, proportions, or small details. Review each shot against the original product photo and approve any regeneration before spending more credits.
This guide is for product footage with voiceover. A realistic person delivering the script on camera is a different workflow that requires suitable presenter generation or lip synchronization. Do not treat narration over footage as that result.
It depends on the models, clip count, duration, resolution, speech, editing, and retries. Ask Creative Claw for estimates before generation and track the operations used. There is no single fixed price for every 15-second ad.