Timed line data
Use optional trusted transcription or existing aligned lyrics.
Lyrics belong in the compositor
Direct cinematic, animated, or abstract footage, then add accurate timed words during final composition instead of asking a video model to spell them.
Timed transcription · lyric-safe shots · ASS typography · multilingual fonts
MiniMaxMusic Studio is an independent third-party service and is not the official MiniMax website.
Choose a generated song or upload an authorized MP3/WAV.
Select a 15- or 30-second highlight, output ratio, and visual direction.
Generate and edit a versioned storyboard before any H3 job is submitted.
Attach validated reference images and approve an immutable server quote.
Render each H3 shot independently, review it privately, and rerender only what needs work.
Compose a final H.264/AAC MP4 with the original song as the only master audio.
Generative video can suggest text-like shapes but should not be the authority for exact lyrics. The Studio explicitly tells H3 not to bake subtitles, logos, watermarks, UI, or lyrics into shots.
Timed lyric lines are shifted into the selected music range and rendered by the final FFmpeg compositor using ASS styles. This preserves spelling, timing, font choice, contrast, multilingual glyph support, and consistent placement.
Use optional trusted transcription or existing aligned lyrics.
Choose templates that reserve negative space for words.
The compositor image includes Noto CJK fonts for Chinese and mixed-language text.
Title, lyrics, and watermark remain configurable final-stage layers.
This separation produces more reliable words and more flexible creative revisions.
Highlight one memorable chorus line in a 15-second teaser.
Use 9:16 negative space for Reels, Shorts, or TikTok.
Render mixed-language timed lines with consistent typography.
Combine lower-complexity cover motion with accurate words.
No. H3 generates the visual shots. The final compositor renders exact timed lyrics from structured data.
You select a music range, generate a versioned storyboard, edit every shot, apply validated visual references, approve the storyboard, and request a server-side quote. No H3 job starts merely because a song was uploaded.
Successful private H3 shots are normalized and concatenated by a dedicated FFmpeg compositor. H3 audio is ignored; the selected portion of the original song becomes the only master soundtrack.
The first production formats are a 15-second teaser and a 30-second social clip in 16:9, 9:16, or 1:1. The project architecture is designed to support longer formats later without treating them as one opaque generation.