Most AI generated commercials look like AI generated commercials: a pretty product spin, a synthetic voice, no story. The canvas in this case study is the opposite ambition; a complete cinematic brand film for a luxury watch, built like a real production: a hero product, cast, storyboards, acts with their own color science, lens choices per shot, and a narrative about a racing driver and his father's hands.
The project is the F1 Luxury Watch Film Template on Zoviz Canvas: 66 prompt nodes, around 30 generated images and over 30 video clips, all organized into a vertical 9:16 film titled "Every Second Has a Story." It is the largest canvas we've documented, and it reads less like a prompt collection and more like a film production pipeline; which is exactly why it's worth studying if you want to know how to make a commercial with AI that doesn't feel like one.
This article walks the pipeline in production order: the product, the cast, the stills, the storyboards, and the film itself, with the real prompts at every stage.
The story before the prompts
The film has a screenplay level structure baked into its prompts. A Formula One driver — whose face is never revealed — walks the corridor before the race. A flashback montage, "Hands That Held Him," tells his childhood entirely through his father's weathered hands. The race is the adrenaline peak. The podium resolves into quiet triumph, and the watch is the object that carries the memory. Racing acts run a carbon black and red color science; memory sequences run warm amber sepia with soft film grain.
None of that is decoration. Every one of those decisions — the hidden face, the hands only flashback, the per act color grades — turns out to be an AI production technique as much as a creative one, as you'll see below.
Stage 1: the hero product
Everything starts with the product. The very first node generates the watch itself:
Product photograph of an ultra-luxury men's automatic chronograph watch, round case 44mm, brushed and polished titanium alternating finish, deep midnight-blue sunburst dial, applied silver hour markers, integrated three-link bracelet, sapphire crystal with faint blue AR coating reflection, no visible text or logo, resting at a three-quarter angle on a dark reflective surface, single dramatic rim light, shallow depth of field, studio product photography, hyperreal detail, 8k
A second node then re lights it as a formal studio shot — "Transform into a professional luxury studio product shot: watch centered on a black glass surface with mirror reflection, single soft key light from upper left… color grade: cold steel blue with warm metallic highlight." If you're a real brand, this is where you'd drop in a photo of your actual product instead and let the transform prompt do the studio work — the same pattern as an AI product video pipeline, just aimed at film.
Want to turn one phone photo into six finished marketing assets?
See how a cluttered, badly-lit snapshot became a studio pack shot, three lifestyle scenes, a product video, and a full posting kit — all anchored to one approved image so nothing drifts between shots.
Read the full AI product photography pipeline →
Stage 2: a cast built for AI's weaknesses
The canvas defines two characters, and both character sheets are engineered around the same insight: AI drifts on faces, so don't depend on faces.
Character sheet, working-class father, age 47, weathered tan hands with visible calluses and fine scars, gentle worn face, slight grey-flecked beard, kind tired eyes, simple grey work shirt sleeves rolled up, neutral studio background — sheet emphasizes hand poses (tightening a strap, resting on a shoulder, gripping a tool) over facial coverage, since face stays largely unseen on screen
The driver gets the same treatment narratively: across the whole film his face stays hidden behind visors, helmets and framing. What reads as arthouse mystery is also the most robust character consistency strategy available in AI filmmaking today: a character defined by wardrobe, build, hands and objects survives regeneration far better than a face.
The realism layer is enforced in every still. The driver's keyframes all open with the same armor against the AI look:
Raw unretouched photograph, real photojournalism aesthetic, shot on ARRI Alexa with a vintage anamorphic lens — visible film grain, natural lens imperfection, no CGI sheen. Formula One driver, athletic 6'0" build, wearing a clean white racing suit with dark graphite piping down the arms and legs (blank sponsor space, modern F1-movie-era silhouette…)
Stage 3: storyboards as single images
Before any video is rendered, the film is previsualized the way real productions do it — except each storyboard sheet is itself one generated image containing a numbered grid of panels:
The storyboard prompts specify the grid ("2x2 grid of 4 numbered panels… 2x3 grid of 6 numbered panels… 5 numbered panels"), thin black borders, PANEL labels, a consistent grade across panels, and the same characters and locations panel to panel. Fifteen of the 66 prompts in this canvas are storyboard sheets. That ratio is the budget discipline of the whole pipeline: scenes get argued about, revised and locked as cheap multi panel images long before the expensive video nodes run — a working storyboard template for AI production.
Stage 4: the film, act by act
The video prompts are written like director's treatment pages. Here is Act III, the adrenaline peak, in full:
Ultra-premium luxury watch brand film, Act III of IV — "Every Second Has a Story." 9:16 vertical, four-shot sequence, the adrenaline peak of the film. High-contrast carbon-black-and-red color science, fast rhythmic editing energy even within a single generation — helmet visor closing under pit lights, the car launching off the grid in slow motion with sparks off the tires, a rapid fragmented montage of cornering/hands/eyes/pit-crew/flag/crowd, then a return to the podium where warm-gold light begins bleeding back in. Driver's face still not revealed. Lens language: 100mm macro on the visor, ultra-wide to 35mm on the launch, fragmented multi-angle coverage in the montage shot, 50mm on the podium return. Diegetic sound implied (engine, mechanical clicks, crowd roar building) — no music or VO. Tone: kinetic, high-stakes, building tension that resolves into quiet triumph.
Alongside each act prompt sits a structured shot list — a JSON array the video node consumes, one entry per shot with its own duration:
[ {"index": 1, "prompt": "9:16 vertical, Helmet visor closing, pit lights reflected across the tinted visor, no facial detail readable, garage background, carbon-black and red grade, macro 100mm, mechanical sound-driven motion", "duration": "4"}, {"index": 2, "prompt": "9:16 vertical, Formula One car launches from the grid, shot low and close front-on so the car fills a vertical frame, cinematic slow-motion, tires throwing sparks in lower frame, 35mm…"} … ]
That pairing — prose treatment plus machine readable shot list — is the canvas's answer to the biggest problem in AI video ads: single prompts produce single moods. Breaking an act into indexed shots with explicit durations, lenses and grades produces something that cuts like a commercial.
Watch three moments from the generated film: the corridor walk (a continuous handheld shot), the Act III race sequence, and the flashback montage "Hands That Held Him" — six shots spanning decades of a life, no faces, hands only.
And when a generated frame is almost right, the canvas doesn't regenerate from scratch — tiny edit prompts fix it in place. The shortest prompt in all 66 nodes: "Make his shirt white and short sleeved, wearing the watch."
The six techniques worth stealing
Generate the product once, reference it forever. One canonical hero product image feeds every wrist shot, macro and podium frame, the same way a character sheet anchors a cast.
Hide the faces. The unrevealed driver and the hands only father are consistency engineering disguised as art direction. If your AI commercial needs a recurring human, define them by everything except their face.
Fight the AI sheen explicitly. "Raw unretouched photograph… visible film grain, natural lens imperfection, no CGI sheen" appears across the stills. Luxury brand marketing lives on texture; these prompts demand it.
Storyboard in grids before you spend on video. Multi panel sheets are cheap, reviewable and force scene level thinking. Fifteen storyboard nodes preceded the thirty plus video generations.
Write acts, not clips. Each act prompt carries its own color science, lens language, sound intent and emotional arc — then a JSON shot list turns that treatment into indexed, timed shots.
Patch, don't reroll. A one line edit prompt on an existing image preserves everything you already approved.
No camera, no crew, no actor, one full short film in a single day.
See how a three minute episode called 06:00, complete with a consistent AI lead, fourteen scripted shots, and diegetic sound, was built entirely from prompts in Zoviz Studio.
Read the full prompt-by-prompt recipe →
Who this is for
Watch and jewelry brands are the obvious audience, but the template generalizes to any product that sells on emotion rather than specs: cars, fragrance, spirits, fashion, audio gear. The structure — hero product, anonymous cast, storyboards, acts with distinct color science, shot lists — is product agnostic. Swap the watch prompt, rewrite the story beats, keep the architecture. For a small brand, this is how to make an ad with genuine film craft and no film crew; for an agency, it's a previsualization and pitch machine that happens to output finished 9:16 film.
FAQ
Can you really make a commercial with AI?
Yes, and this canvas shows the realistic shape of it: not one magic prompt but a pipeline — product photography, character sheets, realism rules, storyboard sheets, act level film prompts and structured shot lists, with a human reviewing at every stage.
How do you keep characters consistent in an AI commercial?
This production's answer is radical and effective: design the film so faces barely matter. Character sheets emphasize hands, wardrobe and objects; the driver's face is never revealed. What must stay consistent is what AI keeps consistent best.
What makes AI video look cinematic instead of synthetic?
In this template, three things: realism language in every still prompt (film grain, lens imperfection, no CGI sheen), per act color science instead of one global look, and explicit lens choices per shot (100mm macro, 35mm launch, 50mm podium).
Why 9:16 vertical for a luxury film?
The film is built for where it will be watched: Reels, TikTok, Shorts and story placements. Every prompt in the canvas enforces vertical framing, down to shooting the F1 launch low and close so the car fills a vertical frame.
Do I need the racing story for my product?
No — the story is the variable, the architecture is the template. Keep the stages (product, cast, stills, storyboards, acts, shot lists) and replace the narrative with the one your brand owns.
Open the F1 Luxury Watch Film Template and trace it node by node: the watch, the character sheets, the storyboard grids, the act prompts and their shot lists. Then replace the watch with your product and the racetrack with your story — the pipeline already knows how to make the film.