Shot Director · Music video workflow
A song runs 2–4 minutes. Neither Seedance 2.5 nor Grok Aurora will generate that in one go — and neither platform's docs explain how to assemble a full track from capped-length pieces while keeping the performer consistent and the cuts on the beat. That's what this page is for. Every figure here traces to official documentation, linked at the bottom.
The song comes first. It's the song that fixes your BPM, your total length and your section boundaries — and all three are inputs the segment planner below needs. Generate the video before the track and you're editing blind.
Suno gives you a Style field and a Lyrics field, and the single most common mistake is putting everything in the first one. Style describes the sound; Lyrics carries the words plus the section markers that shape the arrangement.
| Field | Put this in it | Keep this out |
|---|---|---|
| Style | Genre + era, mood, key instruments, vocal character, production feel | Lyrics, section markers, long structural instructions |
| Lyrics | Your actual words, plus [Verse] / [Chorus] style section markers |
Genre and instrument descriptions — they're weaker here than in Style |
Suno's own guidance for v4.5 is that the Style field now takes conversational, detailed descriptions rather than bare comma-separated tags. The four components below are the reliable skeleton: genre and era, mood, instrumentation, vocals. Community testing puts the sweet spot at roughly 4–7 descriptors — fewer reads as generic, more starts pulling the model in competing directions.
| Tag | What it marks | Notes |
|---|---|---|
[Intro] | Instrumental opening | Add detail — [Intro: solo piano] lands better than the bare tag |
[Verse] | Main storytelling section | Number them ([Verse 1]) when you want distinct arrangements |
[Pre-Chorus] | Build into the hook | Useful when the chorus keeps arriving flat |
[Chorus] | The hook | The most reliably obeyed tag |
[Bridge] | Contrasting mid-song section | Your natural place for a visual change of scene |
[Instrumental] | Solo or break, no vocals | Gives you a segment with no lip-sync to match |
[Outro] | Ending | Include one, or tracks tend to cut off abruptly |
Vocal-delivery tags follow the same pattern and the same
caveat — [Whispered], [Falsetto], [Belting],
[Raspy], [Spoken word], [Harmonized] — and can be
combined with a section marker, as in [Chorus: belting, powerful]. They are
noticeably less consistent than structure tags.
Pick a genre to filter. Each one is a complete Style field — copy it, then swap the mood or era to make it yours.
| Mistake | Instead of | Write |
|---|---|---|
| Too vague | "sad song" | Melancholic piano ballad, slow tempo, introspective female vocals, rainy-day mood |
| Overloaded | A paragraph naming six instruments, a BPM, a key and a structure | Indie pop, emotional female vocals, acoustic and electronic blend, bittersweet |
| Contradictory | "calm aggressive metal" | Heavy metal with melodic interludes and dynamic contrast |
| Music theory | "120 BPM, C major, 4/4" | Upbeat and energetic, driving rhythm, nostalgic 80s feel |
| Named artists | A living artist's name (Suno restricts these) | The characteristics: "1980s production, dreamy female vocals, gated reverb" |
The planner below asks for a BPM, and every cut it suggests is measured in beats. If those words are fuzzy, this section is the one to read — and if you have the audio file but not the number, the analyser will work it out for you.
| Term | What it is | Why your edit cares |
|---|---|---|
| Beat | The steady pulse you'd tap your foot to. Usually the kick drum. | The smallest unit a cut can land on without feeling early or late. |
| BPM | Beats per minute — how fast that pulse runs. One beat lasts 60 ÷ BPM seconds. | Converts musical time into the seconds a generation platform actually takes. |
| Bar | A group of beats, almost always 4 in popular music. Beat 1 is the downbeat — the strong one. | Cutting on a downbeat feels deliberate; cutting on beat 3 feels like a mistake. |
| Phrase | A group of bars the music is built from — typically 2 bars (8 beats) or 4 bars (16). | This is where the music itself changes. Cut here and the edit feels invisible. |
Two bars of four. The teal beat is the downbeat — the one that feels like a beginning. Play it against your track: if the clicks drift, your BPM is wrong.
Drop in your track and this will estimate its tempo. It runs entirely in your browser using the Web Audio API — the file is decoded in memory and never uploaded anywhere.
Reads the same file you dropped in the Beats section above — one upload, one decode, no second drop zone. Waveform, loudness, true peak and a spectrogram, so you can tell whether a track is mastered sensibly before you spend a generation against it.
Paste your lyrics — or have them guessed — then fix the timing by ear. Paste-your-own only ever aligns text you already know is correct to the audio, which is a much more reliable problem than guessing the words, and nothing is sent anywhere for it. Auto-detect is a separate, opt-in option below that actually transcribes — it downloads a speech-recognition model and runs it on your machine; still nothing is uploaded anywhere, but do read its note before using it, since guessing the words is a much harder problem than timing text you already have. Reads the same track loaded in Beats, bars & BPM above.
Enter your song and pick a platform. This works out how many generations you need, and scaffolds a timestamp list cut on musical phrase boundaries (4/8/16 beats) instead of arbitrary durations — so each segment ends where the music actually breathes.
One slot per segment from the planner above. Load a start and end frame for each — the stills you're planning to generate from, or a quick reference/mood image if you haven't made them yet — and play them back in order to check continuity, color, and pacing before spending a single generation. Everything here runs in your browser; nothing is uploaded anywhere. Export the pass as a GIF or a video file to share for approval.
Every figure here traces to each platform's official docs (linked in Sources) — nothing here is from a third-party write-up.
Neither platform assembles a whole song for you. Here's what each piece actually does, and what's left to you.
Builds directly on the reference-binding discipline already documented in the Seedance 2.5 guide — the same habits carry over to Aurora's reference tags.
@Image 1 defines the
performer's face, hair and outfit only. Do not use its background. Aurora: the same idea
via <IMAGE_1> — state what to take and what to leave out every time, not just
in the first segment.Seconds-per-beat = 60 ÷ BPM. A cut on every phrase boundary (not mid-beat) is what makes edits feel musical rather than arbitrary.
| BPM | Typical feel | Sec / beat | Sec / 4-beat phrase | Sec / 8-beat phrase |
|---|
0-4s: …) — see the Seedance
guide's timestamp section. Aurora has no documented multi-shot or timestamp syntax within
a single generation — each segment is one prompt describing that whole segment's action;
don't invent a timestamp block for it.A 3:30 (210s) track at 120 BPM — the planner's own defaults — worked all the way through.
120 BPM → 0.5s/beat → an 8-beat phrase is 4s. The largest whole number of 4s phrases under 30s is 7 phrases = 28s — so every segment but the last is 28s, cut exactly on a phrase.
210s ÷ 28s → 7 segments of 0:28, plus a final 0:14 tail segment — 8 generations total. Verses and choruses that hold one continuous performance shot: chain those 28s segments as extensions. A hard scene change (new location, new outfit beat): start the next segment as a fresh generation instead, using the same performer reference image.
Same 4s phrase. The largest whole number of 4s phrases under 15s is 3 phrases = 12s.
210s ÷ 12s → 17 segments of 0:12, plus a final 0:06 tail segment — 18 generations total. More than twice the segment count of Seedance for the same song, at the same phrase length — worth factoring into which platform you pick for a longer track.
Grok Aurora facts on this page come directly from xAI's official developer docs:
→ Video overview
→ Video generation — model, duration, aspect ratio, resolution
→ Video extension
→ Reference-to-video — image/audio reference tagging
Seedance 2.5 facts come from the same BytePlus ModelArk docs used to build the
Seedance 2.5 guide on this site — see that
page's own Sources section for the exact links.
Suno guidance is split by how well it's sourced, and the split matters:
→ Suno Help Center — Detailed Style Instructions — official; the source for "the Style field takes conversational, detailed descriptions"
→ Suno Help Center — Better Prompts in Lyrics — official; note that it describes what to write in the Lyrics box and does not publish a list of bracket tags
→ Musci.io — Suno Prompts guide — community write-up; source of the 4-component structure, the 4–7 descriptor guideline, the section/voice tag list and the example prompts
The bracket tags are community convention, not documented syntax. They're
widely used and they do influence output, but no Suno documentation defines them — so this
page marks them as such rather than presenting them as an API.
Tempo detection here is a peak-interval histogram over a low-passed signal, computed locally
with the Web Audio API.
No audio is uploaded.