Shot Director · Music video workflow

🎬 Music Video Director

A song runs 2–4 minutes. Neither Seedance 2.5 nor Grok Aurora will generate that in one go — and neither platform's docs explain how to assemble a full track from capped-length pieces while keeping the performer consistent and the cuts on the beat. That's what this page is for. Every figure here traces to official documentation, linked at the bottom.

Planning for
📓 Still shaping the idea? Start in the Builder

1 Write the track in Suno

The song comes first. It's the song that fixes your BPM, your total length and your section boundaries — and all three are inputs the segment planner below needs. Generate the video before the track and you're editing blind.

The two fields, and what belongs in each

Suno gives you a Style field and a Lyrics field, and the single most common mistake is putting everything in the first one. Style describes the sound; Lyrics carries the words plus the section markers that shape the arrangement.

FieldPut this in itKeep this out
Style Genre + era, mood, key instruments, vocal character, production feel Lyrics, section markers, long structural instructions
Lyrics Your actual words, plus [Verse] / [Chorus] style section markers Genre and instrument descriptions — they're weaker here than in Style

Style prompt builder

Suno's own guidance for v4.5 is that the Style field now takes conversational, detailed descriptions rather than bare comma-separated tags. The four components below are the reliable skeleton: genre and era, mood, instrumentation, vocals. Community testing puts the sweet spot at roughly 4–7 descriptors — fewer reads as generic, more starts pulling the model in competing directions.

Style prompt Copied!

Section markers in the Lyrics field

These are community convention, not documented syntax. Suno's own help pages describe what to write in the Lyrics box but never publish a list of bracket tags. The tags below are what the user community has converged on and they do influence output — but treat them as suggestions the model may ignore, not as an API. If a tag doesn't take, regenerate rather than assuming you typed it wrong.
TagWhat it marksNotes
[Intro]Instrumental openingAdd detail — [Intro: solo piano] lands better than the bare tag
[Verse]Main storytelling sectionNumber them ([Verse 1]) when you want distinct arrangements
[Pre-Chorus]Build into the hookUseful when the chorus keeps arriving flat
[Chorus]The hookThe most reliably obeyed tag
[Bridge]Contrasting mid-song sectionYour natural place for a visual change of scene
[Instrumental]Solo or break, no vocalsGives you a segment with no lip-sync to match
[Outro]EndingInclude one, or tracks tend to cut off abruptly

Vocal-delivery tags follow the same pattern and the same caveat — [Whispered], [Falsetto], [Belting], [Raspy], [Spoken word], [Harmonized] — and can be combined with a section marker, as in [Chorus: belting, powerful]. They are noticeably less consistent than structure tags.

Example style prompts

Pick a genre to filter. Each one is a complete Style field — copy it, then swap the mood or era to make it yours.

What goes wrong

MistakeInstead ofWrite
Too vague"sad song"Melancholic piano ballad, slow tempo, introspective female vocals, rainy-day mood
OverloadedA paragraph naming six instruments, a BPM, a key and a structureIndie pop, emotional female vocals, acoustic and electronic blend, bittersweet
Contradictory"calm aggressive metal"Heavy metal with melodic interludes and dynamic contrast
Music theory"120 BPM, C major, 4/4"Upbeat and energetic, driving rhythm, nostalgic 80s feel
Named artistsA living artist's name (Suno restricts these)The characteristics: "1980s production, dreamy female vocals, gated reverb"
Expect to regenerate. Community reports converge on three to six attempts before a track lands, so budget for it. Generate short previews while you're still testing a prompt and only commit to full-length renders once the vibe is right — and remember that re-rolling the song after you've planned your video means re-planning, because a new take almost always means a new tempo and new section boundaries.

2 Beats, bars & BPM

The planner below asks for a BPM, and every cut it suggests is measured in beats. If those words are fuzzy, this section is the one to read — and if you have the audio file but not the number, the analyser will work it out for you.

The four words that matter

TermWhat it isWhy your edit cares
Beat The steady pulse you'd tap your foot to. Usually the kick drum. The smallest unit a cut can land on without feeling early or late.
BPM Beats per minute — how fast that pulse runs. One beat lasts 60 ÷ BPM seconds. Converts musical time into the seconds a generation platform actually takes.
Bar A group of beats, almost always 4 in popular music. Beat 1 is the downbeat — the strong one. Cutting on a downbeat feels deliberate; cutting on beat 3 feels like a mistake.
Phrase A group of bars the music is built from — typically 2 bars (8 beats) or 4 bars (16). This is where the music itself changes. Cut here and the edit feels invisible.

Hear it

Two bars of four. The teal beat is the downbeat — the one that feels like a beginning. Play it against your track: if the clicks drift, your BPM is wrong.

Bar 1Bar 2
Tap tempo is the no-tools fallback: play your song and tap the button on every beat. Four taps gives a rough answer, sixteen gives a good one.

Analyse an audio file

Drop in your track and this will estimate its tempo. It runs entirely in your browser using the Web Audio API — the file is decoded in memory and never uploaded anywhere.

🎵 Drop an MP3 / WAV / M4A here or click to choose a file — nothing leaves your machine

3 Track analysis

Reads the same file you dropped in the Beats section above — one upload, one decode, no second drop zone. Waveform, loudness, true peak and a spectrogram, so you can tell whether a track is mastered sensibly before you spend a generation against it.

Drop a track in Beats, bars & BPM above to see its waveform, loudness and spectrum here.

4 Lyric sync — SRT export

Paste your lyrics — or have them guessed — then fix the timing by ear. Paste-your-own only ever aligns text you already know is correct to the audio, which is a much more reliable problem than guessing the words, and nothing is sent anywhere for it. Auto-detect is a separate, opt-in option below that actually transcribes — it downloads a speech-recognition model and runs it on your machine; still nothing is uploaded anywhere, but do read its note before using it, since guessing the words is a much harder problem than timing text you already have. Reads the same track loaded in Beats, bars & BPM above.

Drop a track above to sync lyrics against it.

5 Segment planner

Enter your song and pick a platform. This works out how many generations you need, and scaffolds a timestamp list cut on musical phrase boundaries (4/8/16 beats) instead of arbitrary durations — so each segment ends where the music actually breathes.

Segment list Copied!
How the cut length is chosen. Seconds-per-beat = 60 ÷ BPM. A phrase (4/8/16 beats) is that many beats long. The segment length is the largest whole number of phrases that still fits under the platform's cap — not the cap itself — so every cut (except the last, shorter tail segment) lands exactly on a phrase boundary instead of slicing through a beat.

6 Flip book preview

One slot per segment from the planner above. Load a start and end frame for each — the stills you're planning to generate from, or a quick reference/mood image if you haven't made them yet — and play them back in order to check continuity, color, and pacing before spending a single generation. Everything here runs in your browser; nothing is uploaded anywhere. Export the pass as a GIF or a video file to share for approval.

Load at least one frame below, then hit Play.
This previews stills, not motion. It's a fast way to sanity-check ordering, performer/ style consistency, and rough pacing across the whole song before you generate — it doesn't simulate what either platform's actual motion will look like between your start and end frames.

1 Seedance 2.5 vs Grok Aurora

Every figure here traces to each platform's official docs (linked in Sources) — nothing here is from a third-party write-up.

The decision that matters most for a music video: if a shot needs to visibly track real vocals or a specific piece of the actual track, Seedance is the tool that can reference it. Aurora's audio references are preset TTS voices for most accounts — it can still carry a music video's visual side perfectly well, but don't plan a lip-sync-to-the-real-track shot on Aurora unless you have trusted-partner access.

2 Stitching a full song together

Neither platform assembles a whole song for you. Here's what each piece actually does, and what's left to you.

  • New generation vs. extension. Use a fresh generation when the song moves to a new scene, location, or setup (verse → chorus with a location change). Use extension when you're continuing the same continuous shot or performance take past its cap — extension picks up from the last frame and both platforms return one combined file for that lineage.
  • Extension only stitches within its own lineage. If segment 3 was its own fresh generation (not an extension of segment 2), the platform will not join segments 2 and 3 into one file for you — that join happens in a normal video editor, cutting each generated/extended clip together in order. This page gets you to a correct, ordered clip list; final assembly is external.
  • Pin the aspect ratio before you start. Pick one aspect ratio for the whole video and use it on every single generation call. Neither platform enforces this across separate generations — if a segment drifts to a different ratio, it will not cut cleanly against its neighbors. (A ratio-locking task like editing/extension makes this automatic within a lineage; it does nothing to keep unrelated segments matching each other.)
  • Decide your reference set once, then reuse it exactly. Whatever image(s) define your performer, use the identical file(s) — same crop, same upload — across every segment of the whole song. See Performer consistency.
  • Chain extensions inside a scene, start fresh generations between scenes. A verse that stays on one continuous performance shot is a good candidate for one long extension chain (segment 1 → extend → extend …); a hard cut to a new location is a new generation with the same reference images.

3 Performer & style consistency

Builds directly on the reference-binding discipline already documented in the Seedance 2.5 guide — the same habits carry over to Aurora's reference tags.

  • One reference image, reused verbatim, for the whole song. Don't re-crop or re-generate your performer reference between segments — every regeneration risks a subtly different face, outfit, or proportions. Keep the exact same file.
  • Bind it explicitly in every single prompt. Seedance: @Image 1 defines the performer's face, hair and outfit only. Do not use its background. Aurora: the same idea via <IMAGE_1> — state what to take and what to leave out every time, not just in the first segment.
  • Write the outfit/style description into your additional-notes text every time, not just once — long projects drift if consistency notes are only stated at the start and then assumed.
  • Keep a running "bible" outside the tool — one plain-text block with the performer description, outfit, environment, and visual style, pasted into every segment's prompt. This page's segment list (above) is a good place to paste it once and copy per-segment.

4 Beat-sync reference

Seconds-per-beat = 60 ÷ BPM. A cut on every phrase boundary (not mid-beat) is what makes edits feel musical rather than arbitrary.

BPMTypical feelSec / beatSec / 4-beat phraseSec / 8-beat phrase
  • Express the cut in the platform's own timestamp syntax. Seedance honours integer-second timestamps directly in the prompt (0-4s: …) — see the Seedance guide's timestamp section. Aurora has no documented multi-shot or timestamp syntax within a single generation — each segment is one prompt describing that whole segment's action; don't invent a timestamp block for it.
  • Don't beat-sync high-frequency motion. A cut every 4 or 8 beats reads as intentional; trying to time a specific gesture to every individual beat inside one generation is asking more precision than either platform's docs claim to offer.

5 Worked example

A 3:30 (210s) track at 120 BPM — the planner's own defaults — worked all the way through.

🎵 3:30 @ 120 BPM, Seedance 2.5 (30s cap)

120 BPM → 0.5s/beat → an 8-beat phrase is 4s. The largest whole number of 4s phrases under 30s is 7 phrases = 28s — so every segment but the last is 28s, cut exactly on a phrase.

210s ÷ 28s → 7 segments of 0:28, plus a final 0:14 tail segment — 8 generations total. Verses and choruses that hold one continuous performance shot: chain those 28s segments as extensions. A hard scene change (new location, new outfit beat): start the next segment as a fresh generation instead, using the same performer reference image.

🎵 Same song, Grok Aurora (15s cap)

Same 4s phrase. The largest whole number of 4s phrases under 15s is 3 phrases = 12s.

210s ÷ 12s → 17 segments of 0:12, plus a final 0:06 tail segment — 18 generations total. More than twice the segment count of Seedance for the same song, at the same phrase length — worth factoring into which platform you pick for a longer track.

6 Sources

Grok Aurora facts on this page come directly from xAI's official developer docs:
Video overview
Video generation — model, duration, aspect ratio, resolution
Video extension
Reference-to-video — image/audio reference tagging

Seedance 2.5 facts come from the same BytePlus ModelArk docs used to build the Seedance 2.5 guide on this site — see that page's own Sources section for the exact links.

Suno guidance is split by how well it's sourced, and the split matters:
Suno Help Center — Detailed Style Instructions — official; the source for "the Style field takes conversational, detailed descriptions"
Suno Help Center — Better Prompts in Lyrics — official; note that it describes what to write in the Lyrics box and does not publish a list of bracket tags
Musci.io — Suno Prompts guide — community write-up; source of the 4-component structure, the 4–7 descriptor guideline, the section/voice tag list and the example prompts
The bracket tags are community convention, not documented syntax. They're widely used and they do influence output, but no Suno documentation defines them — so this page marks them as such rather than presenting them as an API.

Tempo detection here is a peak-interval histogram over a low-passed signal, computed locally with the Web Audio API. No audio is uploaded.