即将出发
加载中...

No Camera. No Studio. No Singer. How to Make a Hollywood Style Music Video

No Camera. No Studio. No Singer. How to Make a Hollywood Style Music Video

We wrote a song, created a virtual star, and directed an eight-shot music video, all inside one browser tab. Using Zoviz’s AI video generator with Seedance 2.0, we produced “Come Alive,” an original electronic track featuring precise lip sync and a cinematic midnight-kitchen performance. This guide shares the complete workflow, including every prompt, setting, and screenshot.

See it live: Explore the song, reference pack, and all eight shots in the reusable Zoviz Studio project. To create another video in the same style, simply replace the shot prompts without rebuilding the entire project.

Create the first frame of your music video

The pipeline at a glance

  • Step 1. Write and generate the song (Voice Studio)
  • Step 2. Slice the track into shot sections
  • Step 3. Build the reference image pack (Quick Image Generator)
  • Step 4. Learn the Hollywood prompt format
  • Step 5. Generate the shots (Pro Studio + Seedance 2.0)
  • Step 6. Merge the shots into the final video
  • Then: costs, lessons, and pro tips

Watch the finished music video: See how “Come Alive” turned out before exploring the complete eight-shot workflow

Step 1. Write and generate the song

Go to Voice Studio, then AI Music. Pick a music style (we chose Electronic). Then do this:

  • Switch the lyrics mode to Write lyrics manually. This is the single most important decision in the whole project.
  • Manual lyrics mean you know every word in advance, which later lets you build lip sync maps and design shots around specific lines.
  • Write the lyrics with the video in mind. Every image in the song should be filmable.
When every lyric is written for the camera, the music stops being just a soundtrack and becomes the storyboard.

Our lyrics describe a midnight kitchen where the world wakes up:

[Verse]
Midnight in the kitchen, nothing moves at all
Just a sleeping city and shadows on the wall
A cup on the counter, steam begins to rise
Something in the silence is opening its eyes

[Pre-Chorus]
Hold your breath, the stillness breaks
Every little thing awakes

[Chorus]
We come alive, alive
Everything you thought was still is dancing tonight
We come alive, alive
Open up your eyes, the world is more than it seems
We come alive

[Bridge]
Milk becomes a dragon, honey starts to sing
Ordinary magic hiding in everything

[Chorus]
We come alive, alive
Everything you thought was still is dancing tonight
We come alive

Hit Generate Music (5 credits). We got a 2:01 electronic track with a female vocal in about two minutes. Download the MP3. You’ll need it for the next step.

  Need help writing lyrics? Follow the approach in our AI cartoon video screenplay case study

Step 2. Slice the track into shot sections

Seedance 2.0 generates clips of 4 to 15 seconds. A 2 minute song needs about eight shots.

  • Play the track once and note where each section starts.
  • Cut the MP3 into one audio file per shot at those timestamps.
  • Use any free audio editor (Audacity or a phone voice memo trimmer works fine).
  • If you prefer scripting it, ffmpeg does the same job in one pass. Example commands:
ffmpeg -i song.mp3 -ss 0  -t 10 -c copy shot01-intro.mp3
ffmpeg -i song.mp3 -ss 10 -t 12 -c copy shot02-verseA.mp3
ffmpeg -i song.mp3 -ss 22 -t 12 -c copy shot03-verseB.mp3
ffmpeg -i song.mp3 -ss 35 -t 10 -c copy shot04-prechorus.mp3
ffmpeg -i song.mp3 -ss 45 -t 15 -c copy shot05-chorus1.mp3
ffmpeg -i song.mp3 -ss 65 -t 15 -c copy shot06-bridge.mp3
ffmpeg -i song.mp3 -ss 80 -t 15 -c copy shot07-chorus2.mp3
ffmpeg -i song.mp3 -ss 105 -t 15 -c copy shot08-outro.mp3

Each slice becomes the reference audio for its shot. That’s what makes the artist’s lips move to your real track and the visual beats land on the real music.

Step 3. Build the reference image pack

This step is what separates a music video from a pile of disconnected clips. Build it before generating any video.

  • Use the Quick Image Generator.
Bring your lyrics to life with Zoviz AI
  • Use High quality and 9:16.
  • Generate five images total.
  • After the first artist image, attach it as a reference (Use Saved Images) for the other two artist angles. This locks her face from the start.

Image 1: artist, front portrait

Portrait of a young woman in her mid twenties, long honey blonde hair, warm brown eyes, wearing an oversized cream cable knit sweater, standing in a modern kitchen at night lit by warm under cabinet light and cool blue moonlight from a window, calm confident expression, photorealistic, cinematic

Image 2: artist, profile singing (attach Image 1 as reference)

Three quarter profile of the same young woman from the reference image, same long honey blonde hair and cream cable knit sweater, singing softly with eyes half closed, one hand near her chest, warm and cool split lighting in a dark modern kitchen at night, photorealistic, cinematic

Image 3: artist, full body (attach Image 1 as reference)

Full body shot of the same young woman from the reference image, long honey blonde hair, oversized cream cable knit sweater, dark jeans, barefoot, standing at a marble kitchen island at midnight with a glass of iced coffee on it, mix of cool blue moonlight and warm amber under cabinet glow, photorealistic, cinematic

Image 4: the location (no reference attached)

Wide shot of an empty modern home kitchen at midnight, marble island with a glass of iced coffee and a small glass milk bottle, big window with sleeping city lights outside, mix of cool blue moonlight and warm amber under cabinet glow, light steam rising from a mug, no people, photorealistic, cinematic, moody

Image 5: the color grade anchor

Cinematic still frame with no people: a dark midnight kitchen scene with a teal and amber color grade, deep shadows, warm highlights on marble and glass, gentle film grain, soft volumetric moonlight through a window, dreamy atmosphere, photorealistic
The location master: an empty midnight kitchen that every shot returns to

Download all five into a project folder. Cost: about 20 credits for the pack.

Perfect the frame without starting over

Step 4. Learn the Hollywood prompt format

This is the anatomy that turns AI clips into a directed music video. We learned it from Zoviz’s own template library and from how our AI cartoon video case study and cinematic AI video ad workflow structure their prompts. Every professional shot prompt has these blocks:

Block What it does
Identity lock "[Image 1] is the performer — preserve her exact identity in every shot:" followed by specific look details. Binds the video to your references.
Master track clause "[Audio 1] is the finished master track — the only audio; no invented music, no new vocals." Stops the model from composing its own soundtrack.
Lip sync priority In capital letters: "SHE SINGS THE VOCAL ON CAMERA — PRECISE LIPSYNC IS THE TOP PRIORITY." Plus "no cutaways mid word."
Lipsync map Timestamps mapped to the exact lyric lines in this clip's audio slice. This is why manual lyrics matter.
Director thesis One sentence stating the idea of the shot, like a director explaining it to a crew.
Visual world Set, light, palette — and explicit negatives: "No neon, no particles, no AI gloss."
Shot flow Second by second camera language: push in, orbit, rack focus, crash zoom on the beat.
Quality bar The closing standard: "expensive Hollywood pop video, filmic grain, no AI gloss."

Step 5. Generate the shots in Pro Studio

Open Pro Studio and select Seedance 2.0. For every shot, repeat this routine:

  • Set the duration to match the audio slice.
  • Set the aspect ratio to 9:16.
  • Pick a resolution. 480p (about 2 credits per second) is fine for drafts and keeps costs consistent. 720p doubles the cost for hero quality.
  • Upload the five reference images in the same order every time.
  • Upload the shot’s audio slice.
  • Keep Generate Audio on.
  • Paste the shot’s prompt and generate. Renders take 2 to 4 minutes.
Pro Studio: duration, aspect ratio, resolution, reference images, reference audio, prompt

Every singing shot opens with the same two blocks. Upload references in the order below so the numbers match.

Shared identity and audio block

[Image 1] [Image 2] [Image 3] are the performer, preserve her exact identity in every shot: woman in her mid twenties, long honey blonde hair, warm brown eyes, oversized cream cable knit sweater, natural glossy lips. [Image 4] is the location, the same modern midnight kitchen, marble island, big window, sleeping city outside. [Image 5] is the color grade, teal and amber midnight palette, filmic. [Audio 1] is the finished master track, the only audio. No invented music, no new vocals.

SHE SINGS THE VOCAL ON CAMERA. PRECISE LIPSYNC IS THE TOP PRIORITY. Her mouth articulates every syllable of [Audio 1] exactly on time. Face visible and sharp through every vocal line. No cutaways mid word.

Then append the shot specific section below. Here are all eight shots.

Shot 01 · Intro (10s) — the cold open

[Image 4] is the location — a modern midnight kitchen and apartment, marble island, big window, sleeping city skyline outside. [Image 5] is the color grade — teal and amber midnight palette, deep shadows, filmic. [Audio 1] is the finished master track — the only audio; no invented music, no added sounds. No one sings in this shot; no people in frame.

Music video route: cold open, establishing. Director thesis: the apartment holds its breath before the song wakes it.

Shot flow: 0-10s one continuous slow dolly through the dark apartment toward the kitchen island from [Image 4]; cool blue moonlight stripes the floor; light steam rises from a single mug on the marble; as the intro builds, reflections on the marble shimmer subtly with the synth; on the final beat the under cabinet light flickers once, like a pulse.

Visual world: exactly the kitchen from [Image 4] graded like [Image 5]. Gentle film grain, soft volumetric moonlight. No people, no creatures, no neon, no particles, no AI gloss.

Quality bar: A24 cold open, expensive cinematography, filmic.
Shot 01 as rendered: the sleeping kitchen, ending on the warm amber pulse

Shot 02 · Verse A (12s) — confession at the counter

LIPSYNC MAP:
0.0-6.0 "Midnight in the kitchen, nothing moves at all"
6.0-12.0 "Just a sleeping city and shadows on the wall"

Music video route: intimate performance. Director thesis: confession at the counter — she sings low and close, like telling the camera a secret at midnight.

Shot flow:
0-6s static medium shot across the marble island: she leans on her forearms singing softly to the lens, half her face in cool moonlight, half in warm under cabinet glow.
6-12s slow lateral track to the right keeping her eyes locked to the camera; the shadow of the window frame slides across the wall behind her; on the final word she lowers her gaze.

Visual world: the kitchen from [Image 4] graded like [Image 5]. Palette: teal, amber, cream. No creatures, no floating objects, no neon, no particles, no AI gloss.

Performance rules: hushed, intimate, minimal movement, immaculate diction; she is the only person on screen. Continuity: same woman, same sweater, same kitchen. Audio intent: [Audio 1] only, her mouth locked to it; faint room tone under the music. Quality bar: expensive Hollywood pop video, Billie Eilish bedroom intimacy, filmic grain, no AI gloss.
Shot 02: moonlight on one side, candlelight on the other, eyes to the lens

Shot 03 · Verse B (12s) — the room answers

LIPSYNC MAP:
0.0-6.0 "A cup on the counter, steam begins to rise"
6.0-12.0 "Something in the silence is opening its eyes"

Music video route: intimate performance. Director thesis: first signs of life — the room starts answering her.

Shot flow:
0-3s macro insert: the coffee surface trembles in delicate rings with the bass, steam curling up through moonlight.
3-6s rack focus from the cup to her face behind it; she sings the line watching the cup with quiet wonder.
6-12s slow handheld push toward her as the under cabinet lights breathe brighter on each beat; her expression turns from wonder to a slow smile on the final word, eyes lifting to the lens.

Visual world: the kitchen from [Image 4] graded like [Image 5]; realistic physics only. Palette: teal, amber, cream. No creatures, no floating objects, no neon, no particles, no AI gloss.

Performance rules: hushed, intimate, immaculate diction; she is the only person on screen. Continuity: same woman, same sweater, same kitchen. Audio intent: [Audio 1] only, her mouth locked to it; faint room tone under the music. Quality bar: expensive Hollywood pop video, cinematic macro photography, filmic grain, no AI gloss.
Shot 03: the mug in razor focus, then the push in landing on her smile

Shot 04 · Pre chorus (10s) — the inhale

LIPSYNC MAP:
0.0-5.0 "Hold your breath, the stillness breaks"
5.0-10.0 "Every little thing awakes"

Music video route: suspended moment. Director thesis: the video inhales before the drop — everything quiets and gathers around her.

Shot flow:
0-5s slow motion feel: she straightens up from the counter into the moonlight, hair drifting a half beat behind her movement, singing the line low; behind the window, the city lights begin dimming out one by one as if gathering power.
5-10s tight close up: she whispers the line almost to the lens; the kitchen practical lights fade down around her until only her face holds a soft glow; on the final syllable everything goes to black. Clean cut to black at the end.

Visual world: the kitchen from [Image 4] graded like [Image 5]. Palette: teal, amber, cream, deepening to black. No creatures, no floating objects, no neon, no particles, no AI gloss.

Performance rules: hushed, tense, controlled breathing, immaculate diction; she is the only person on screen. Continuity: same woman, same sweater, same kitchen. Audio intent: [Audio 1] only, her mouth locked to it; faint room tone under the music. Quality bar: expensive Hollywood pop video, trailer moment tension, filmic grain, no AI gloss.
Shot 04: the whisper before the drop — everything fades until only her face holds light

Shot 05 · Chorus 1 (15s) — the flagship

LIPSYNC MAP (verify against the slice and adjust ±0.5s):
0.0-1.0 rising synth, she inhales.
1.0-4.0 "We come alive, alive"
4.0-8.0 "Everything you thought was still is dancing tonight"
8.0-11.0 "We come alive, alive"
11.0-15.0 "Open up your eyes, the world is more than it seems"

Music video route: intimate midnight performance that erupts. Director thesis: the apartment wakes up with her — on every downbeat of the chorus, another light in the kitchen and across the skyline blooms awake, as if the city is breathing with the song.

Shot flow:
0-1.0s slow push in from wide: she stands at the marble island, eyes closed, the city dark behind the window.
1.0-4.0s medium push in: her eyes open on the first "alive"; the under cabinet lights bloom brighter in rhythm; she sings straight into the lens.
4.0-8.0s camera orbits 90 degrees around her as she turns through the moonlight; behind her, windows across the skyline flicker awake one by one, exactly on the beats.
8.0-11.0s chest up close frame: hair rim lit, a soft warm lens flare crosses; she smiles inside the line without breaking articulation.
11.0-15.0s she opens her arms on the final line; camera pulls back to a wide as every practical light in the apartment and half the skyline now glows; hold on her silhouette against the living city. Loopable.

Performance rules: intimate but euphoric, natural micro expressions, immaculate diction; she is the only person on screen. Continuity: same woman, same sweater, same kitchen. Audio intent: [Audio 1] only, her mouth locked to it; faint room tone under the music. Quality bar: expensive Hollywood pop video, anamorphic feel, filmic grain, no AI gloss, no glow.

Shot 06 · Bridge (15s) — the lyric stays a metaphor

LIPSYNC MAP:
0.0-7.0 "Milk becomes a dragon, honey starts to sing"
7.0-15.0 "Ordinary magic hiding in everything"

Music video route: playful performance. Director thesis: she plays with the real world and it plays back — the fantasy lyric stays a metaphor.

Shot flow:
0-7s she pours milk into her glass in golden practical light, grinning at the camera as the swirl blooms into a beautiful marbled cloud (real fluid dynamics only);
7-15s she lifts the glass in a toast to the lens, city fully lit behind her, camera slowly circling; honey jar and warm bokeh in the foreground.

Visual world: realistic physics only, palette from [Image 5]. No creatures, no floating objects, no neon, no particles, no AI gloss.

Performance rules: playful, flirtatious with the camera, immaculate diction. Continuity: same woman, same sweater, same kitchen. Quality bar: high end commercial tabletop photography meets pop video, no AI gloss.

Shot 07 · Chorus 2 (15s) — full release

LIPSYNC MAP: same chorus lines as shot 05, escalated delivery.

Music video route: performance, full release. Director thesis: she finally dances.

Shot flow:
0-4s wide: she spins barefoot across the living room floor in moonlight, singing full out;
4-8s handheld orbit counter to her spin, skyline blazing behind the window;
8-12s crash zoom punch in on "alive", then a second crash zoom on the next "alive" — the signature accents of the video;
12-15s she lands the last line breathless, laughing mid lyric without losing sync; every light in the apartment pulses once, together, on the final downbeat.

Visual world: same apartment continuing the kitchen from [Image 4], graded like [Image 5]. No creatures, no neon, no particles, no AI gloss.

Performance rules: euphoric, athletic, immaculate diction. Continuity: same woman, same sweater, same apartment. Quality bar: euphoric Hollywood chorus, Robyn Dancing On My Own energy, no AI gloss.

Shot 08 · Outro (15s) — the exhale

LIPSYNC MAP: final "We come alive" echoes; she mouths only the last "alive" at 10-12s, otherwise silent.

Music video route: closing shot. Director thesis: the world exhales — stillness returns, but warmer than before.

Shot flow:
0-6s slow pull back through the apartment as the practical lights dim gently one by one, in reverse order of how they woke;
6-12s she settles onto the window sill with her mug, silhouetted against the now softly glowing skyline, mouths the final "alive";
12-15s the steam from her mug curls once through the moonlight, she smiles at the lens, freeze on the last chord. Ends on a frame that mirrors shot 01. Loopable back to the intro.

Visual world: same apartment, graded like [Image 5]. No creatures, no neon, no particles, no AI gloss.

Quality bar: closing shot of a film, filmic grain, no AI gloss.

Step 6. Merge the shots

Download every finished clip in order. Each clip already carries its own slice of the song, so assembly is quick.

  • Import all eight clips into any editor (CapCut, iMovie, DaVinci Resolve).
  • Place them back to back in order.
  • Add a half second crossfade between each pair, on both picture and sound.
  • Export.

Prefer scripting the merge? One ffmpeg command handles the whole timeline. Each xfade offset is the running total of previous clip durations minus the transition length.

ffmpeg -i shot1.mp4 -i shot2.mp4 -i shot3.mp4 -i shot4.mp4 \
 -filter_complex "[0:v][1:v]xfade=transition=fade:duration=0.5:offset=9.54[v01];
 [v01][2:v]xfade=transition=fade:duration=0.5:offset=21.08[v012];
 [v012][3:v]xfade=transition=fade:duration=0.5:offset=32.62[vout];
 [0:a][1:a]acrossfade=d=0.5[a01];[a01][2:a]acrossfade=d=0.5[a012];
 [a012][3:a]acrossfade=d=0.5[aout]" \
 -map "[vout]" -map "[aout]" -c:v libx264 -crf 18 -c:a aac merged.mp4

For a guaranteed seamless soundtrack, replace all clip audio with the original master MP3 laid under the whole timeline instead. Since every clip was generated against its exact slice, lip sync still lands correctly.

Upscale your shots without losing detail

Costs, lessons, and pro tips

Rough credit cost for the full project:

  • Original song, 2:01, Voice Studio: 5 credits
  • Reference image pack, 5 images at High quality: 20 credits
  • Eight video shots at 480p, about 2 credits per second: about 220 credits
  • Total: about 245 credits

720p shots cost double per second. Keep resolution consistent across the final cut, or draft everything at 480p first and re render only the keepers.

Spend credits on consistency, not guesswork: test the timing at 480p, lock every recurring detail, and upscale only the shots that make the final cut.

Quick lessons from this production:

  • Verify your lip sync map timestamps by ear before generating. Sync is what sells the Hollywood feel.
  • Name everything that must stay identical: sweater, kitchen, camera move. The model transforms whatever you leave unspecified.
  • End dark shots on a clean cut to black. It gives you clean edit points, like our pre chorus, which lands the whole video on a held breath right before the drop.
  • Failed generations refund automatically, so a bad roll costs nothing but time.
  • Save your best prompt as a template (Save Template button) so the next video starts at step five.

Who this is for

This pipeline fits anything that needs a lip synced performance video without a shoot:

  • Independent musicians releasing a single with no budget for a video crew
  • Brand jingles and sonic logos that need a face
  • Social first cover videos and lyric visualizers
  • Virtual or AI artist projects testing a look before investing in a real production

The reference pack and prompt template are reusable. A second single from the same virtual artist only needs new shot prompts. The character sheet and the kitchen stay as they are.

Want the same pipeline for a spoken story instead of a song? See our cinematic AI video ad case study for the same shot by shot prompt structure applied to a dialogue driven ad.

FAQ

How do you get accurate lip sync in an AI generated music video?

Feed the video model the exact audio slice for that shot, not the full song. Write a lip sync map with timestamps tied to the literal lyric lines. State “PRECISE LIPSYNC IS THE TOP PRIORITY” in capitals in every prompt. Manual, hand written lyrics make this possible, since you know every word and its timing in advance.

How long can each shot be?

Seedance 2.0 generates clips from 4 to 15 seconds with audio. That’s why a 2 minute song is broken into eight shots of 10 to 15 seconds each, joined afterward in an editor.

How much does a full music video cost in Zoviz credits?

About 245 credits for a 2 minute video: 5 for the song, 20 for the five image reference pack, and roughly 220 for eight shots at 480p. Rendering at 720p roughly doubles the video cost.

Do I need video editing experience to finish this?

No. The final assembly step is placing eight clips back to back with a half second crossfade, which any basic editor handles with no technical skill. The ffmpeg command in this guide is an optional shortcut for anyone who prefers scripting it.

Can I reuse the same virtual artist for another song?

Yes. Keep the reference image pack and the shared identity and audio prompt block, then write new shot prompts for the new lyrics. The character sheet and the location are what stay constant between videos.

Open Zoviz Studio and roll camera on an artist the world has never seen.

获取独家资讯简报

立即订阅,最新资讯将直接发送到您的收件箱。千万不要错过!