FLUX 3 Video Prompts: 40 Shot Briefs With Audio [2026]
August 7, 2026By Bilal Azhar
40 copy-paste shot briefs for the omni model that renders picture and synchronized sound in one pass. Audio cue technique, dialogue lines, credit math.
FLUX 3 is not the next Flux image model. It is Black Forest Labs' omni video model, a single flow model where video, audio, and image understanding are trained jointly, and the practical consequence is that one prompt produces picture and synchronized sound together. Dialogue with lip-sync, room tone, ambience, and effects all come out of the same pass. That changes how you write the prompt: a caption that describes a beautiful frame leaves the entire soundtrack to chance. The 40 shot briefs below all follow the same shape, and every one of them names its audio explicitly. FLUX 3 Video reached general availability on August 4, 2026, and runs on Morphed at 25 credits per second in 720p.
| If you want | Write this | Skip this |
|---|---|---|
| A creator piece to camera | one speaker, one short quoted line, named room tone | a script paragraph and three cutaways |
| A product ad read | product in hand, one claim spoken, one surface sound | adjective stacks like premium and stunning |
| An atmospheric scene | one motion source plus the ambience it makes | "cinematic, moody, 4k, highly detailed" |
| A vertical hook | 9:16, action centered, sound in the first second | wide framing with a small subject |
| A cheap iteration | 720p at 25 credits per second | 1080p drafts at 42 credits per second |
What Does an Omni Model Change About Prompting?
The short version: you now have a second channel to direct, and if you ignore it the model guesses. Older video models rendered silent picture, and audio arrived through a separate pass, a separate tool, and separate money. FLUX 3 generates both in one forward pass, which means the sound is committed at the same moment as the frame. You cannot fix it afterward without regenerating the whole clip.
That is why every prompt in this guide ends with an audio cue. It is the single highest-leverage habit for this model, and it is the thing most people copying old Flux image prompts get wrong.
Two more structural facts worth internalizing. First, FLUX weights the start of a prompt most heavily, so front-load the element you most want the model to get right. If the dialogue matters more than the kitchen, lead with the person speaking. Second, this is a video model, so describe change over time rather than a static composition. "A woman standing in a workshop" is a photograph. "A woman turns from the bench, wipes her hands, and looks up at the camera" is a shot.
Black Forest Labs has not published an official FLUX 3 prompt guide. The structures below come from the model's public behavior, provider example sets, and our own generation logs on FLUX 3 Video.
How Should You Structure a Shot Brief?
A shot brief is five slots in one to three sentences: subject, motion, setting and light, camera, audio cue. That is the whole structure. Anything past three sentences tends to produce a clip that tries to satisfy conflicting instructions and satisfies none of them.
| Slot | What goes there | Example fragment |
|---|---|---|
| Subject | who or what, with one identifying detail | a woodworker in a canvas apron |
| Motion | one clear action that unfolds | sets down a chisel and turns to camera |
| Setting and light | the physical space and its light source | a small workshop, low afternoon sun through dusty glass |
| Camera | one camera behavior, not three | static medium shot at chest height |
| Audio cue | dialogue, ambience, or effect, named | she says "This is the only one I keep." Sawdust-muted room tone |
Here is the same idea written badly and then written well.
Weak:
Beautiful cinematic video of a woodworker in a workshop, highly detailed, 4k, dramatic lighting, professional quality, trending on artstation
That prompt has no motion, no camera, and no sound. It also spends five of its words on legacy diffusion tags that carry no weight in FLUX 3. The result is usually a slow drift across a static-feeling scene with generic background hum.
Strong:
A woodworker in a canvas apron sets down a chisel and turns to camera in a small workshop, low afternoon sun through dusty glass. Static medium shot at chest height. She says "This is the only one I keep." Sawdust-muted room tone under her voice.
Writing dialogue that syncs. Three rules, in order of impact. Keep the line under about twelve words. Use one speaker per clip. Put the line in quotation marks so the model reads it as speech rather than description. If you need a second voice, make it off-camera and say so, because two on-camera speakers in one prompt is the most reliable way to get mouth movement on the wrong face.
The Audio Cue Cookbook: Dialogue, Ambience, or Effect?
Three cue types, three different phrasings. Mixing them up is the most common reason a clip sounds wrong even when it looks right.
| Cue type | Phrase it as | Weak version | Strong version |
|---|---|---|---|
| Dialogue | verb plus quoted line plus language | "she talks about the product" | she says in Spanish "Solo quedan tres." |
| Ambience | the room, then what makes the noise | "background sounds" | tiled bathroom reverb, a tap running in the next room |
| Effect | the object and the action making it | "cool sound design" | the metal latch clicks, then a low pneumatic hiss |
The pattern underneath all three: name the physical source, not the impression. Asking for "reverb" gives you a vague wash. Asking for "empty concrete stairwell" gives you reverb that matches the picture, because the model is inferring the acoustics from the space it is already rendering. That coupling between image and sound is the whole point of an omni model, and prompts that lean on it get better results than prompts that describe audio in mixing-desk vocabulary.
One more rule for ad work: put the sound event in the first second. Vertical viewers decide in under a second, and a clip whose only audio event lands at 0:04 has already lost them.
Prompts for Talking-Head and Creator Pieces
Play it with sound on. The line, the steam wand hiss behind it, and the room tone were all generated in the same pass as the picture; nothing was recorded or added afterward.
These are the clips FLUX 3 is genuinely good at, because the speech, the face, and the acoustics of the room are all decided together. Keep each to one speaker and one idea. 8 to 12 seconds is the sweet spot.
A woodworker in a canvas apron sets down a chisel and turns to camera in a small workshop, low afternoon sun through dusty glass. Static medium shot. She says "This is the only one I keep." Sawdust-muted room tone under her voice.
A barista leans on the counter of an empty cafe at opening, warm window light from the left, slow push-in to a medium close-up. He says "Most people grind way too fine." Espresso machine hiss and a distant fridge hum.
A software engineer sits in a home office at night, monitor glow on her face, static shot at eye level. She says "I deleted the whole file and started over." Keyboard clicks stop as she speaks, quiet room tone.
A gardener kneels beside a raised bed in early morning, dew on the leaves, handheld medium shot slightly below eye level. He says "Water the soil, never the leaves." Birdsong and a hose dripping onto gravel.
A tattoo artist looks up from a sketchbook in a studio lit by a single lamp, static close-up. She says "I turn down about half of what people ask for." Faint needle buzz from the next station, low room tone.
A chef in a white jacket stands at a pass and wipes a plate rim, overhead warm heat lamps, static medium shot. He says "The plate goes out hot or it does not go out." Kitchen clatter and ticket printer in the background.
A woman in a rain jacket speaks to camera on a wet balcony at dusk, city lights soft behind her, handheld medium close-up. She says "Nobody told me it would take this long." Rain on the railing and distant traffic.
A mechanic wipes his hands on a rag and steps back from an open hood, garage strip lighting, static waist-up shot. He says "That noise is not the belt." Compressor cycling and a radio playing faintly.
Prompts for Scripted UGC-Style Ad Reads
The structure that works: product visible in frame within the first second, one spoken claim, one sound the product itself makes. Do not stack three benefits into one line, and do not make the product fly. For a wider treatment of the format, see our guide to AI UGC video generation.
A woman holds a matte black water bottle to camera in a bright kitchen, morning light, static medium close-up on 9:16. She says "Ice at 6pm. Still ice at 6am." The cap clicks open, then a single ice cube shifts inside.
A man unboxes a pair of white running shoes on a wooden floor, overhead shot tilting down, soft daylight. He says "These are the ones I actually kept." Tissue paper crumples and the box lid slides off.
A woman presses a pump bottle of serum into her palm at a bathroom mirror, cool daylight, static close-up. She says "Two pumps, that is the whole routine." The pump clicks twice, tiled bathroom reverb.
A man sets a portable espresso maker on a camp table beside a van at sunrise, handheld medium shot. He says "Real espresso, no power." A pressurized hiss, then coffee dripping into a metal cup.
A woman zips a compact travel backpack closed on a hotel bed, warm lamp light, static three-quarter shot. She says "Ten days, no checked bag." A long zipper pull and the buckle snapping shut.
A man taps a wireless earbud case open at a desk and lifts one bud out, soft window light, tight macro shot. He says "It pairs before I get it in my ear." A magnetic snap and a soft connection chime.
A woman slides a cast iron pan onto an induction hob and drops in butter, kitchen downlights, static medium shot. She says "Preheat it dry. Always." Butter sizzles immediately, extractor fan hum behind.
Prompts for Ambience-Led Scenes
Here the sound leads and the picture supports it. These are useful as b-roll, as backgrounds under a voiceover you record yourself, and as atmosphere beds. No dialogue, so lip-sync variance is not a factor and takes are more consistent.
A kettle comes to a boil on a stove in an empty kitchen at dawn, steam rising into cold blue light, static locked-off shot. Rolling boil, then the click of the switch, then silence and a fridge hum.
Rain runs down a workshop window while a bench light burns behind it, no people, slow push-in. Steady rain on glass, a radiator ticking, distant thunder once.
A subway platform empties as the last passenger walks out of frame, fluorescent tubes flickering, static wide shot. Receding footsteps, tunnel wind, a rail creak.
Laundry moves in a dryer drum seen through the door, warm tumbling light, static tight shot. Rhythmic tumbling, a zipper hitting the drum every rotation, laundromat murmur.
A market stall vendor stacks oranges in the early sun, hands only in frame, slow tilt up. Crates knocking, a scale ratcheting, overlapping street chatter.
Wind moves through a wheat field at golden hour, wide static shot at knee height. Sustained wind through stalks, one distant dog bark, no music.
A record needle drops onto vinyl in a dim living room, macro shot on the tonearm. Surface crackle for two seconds, then a warm bass note, then room tone.
Prompts for Cinematic and Atmospheric Shots
Use 21:9 or 2:1 for these. One motion source, one camera move, one sound layer. Resist the urge to add a second action.
A lone figure crosses a flooded car park at night, headlights raking across standing water, low tracking shot from behind. Footsteps in shallow water and a distant siren.
A cable car rises through fog above a valley, slow drone pullback, cold flat light. Cable hum, wind buffeting the cabin, a metallic knock as it passes a pylon.
A woman in a red coat stops on an empty bridge and looks back, locked-off symmetrical frame, blue hour. Wind through cables, a single car passing behind her.
Snow falls onto a parked motorcycle outside a diner, neon sign reflected in the tank, static shot with slow snow drift. Neon buzz, muffled diner noise through glass, no wind.
A door opens onto a concrete stairwell and a figure descends, handheld follow from above, harsh overhead strip lighting. Long stairwell reverb on every footstep, a fire door slamming above.
Waves break against a harbour wall at dusk, static wide shot at 21:9, spray catching the last light. Heavy surf, a bell buoy, gulls in the distance.
Steam vents from a manhole as a taxi passes through it, low static shot at street level, sodium streetlight. Pressurized steam hiss, tires on wet asphalt, a horn two blocks away.
Prompts for Multilingual Dialogue
FLUX 3 handles multilingual dialogue with lip-sync upstream. Name the language explicitly, keep the line short, and generate a few takes before you commit. Support and quality vary by language, so treat these as starting points rather than guarantees.
A woman in a Madrid market holds up a bunch of herbs to camera, warm morning light, static medium shot. She says in Spanish "Solo quedan tres." Market chatter and crates being stacked.
A man in a Paris bakery slides a tray of baguettes onto a rack, static three-quarter shot, warm oven glow. He says in French "Encore cinq minutes." Oven fan hum and a metal tray scraping.
A woman in a Tokyo stationery shop lifts a fountain pen from a case, static close-up, cool shop lighting. She says in Japanese "This one writes finer." A pen cap clicking and quiet shop ambience.
A man on a Berlin balcony leans on the railing at dusk, handheld medium shot. He says in German "Das dauert noch." City traffic below and a tram bell.
A woman in a São Paulo kitchen stirs a pot and looks up, static medium shot, afternoon light. She says in Brazilian Portuguese "Prova isso." A spoon against the pot rim and an extractor fan.
Prompts for Vertical 9:16 Hooks
Center the action, keep the subject large in frame, and land a sound event inside the first second. These are written for 5 to 8 second generations.
A woman drops a handful of ice into a glass and pours cold brew over it, tight centered 9:16 shot, bright kitchen light. Ice cracking on contact, then liquid pouring.
A man rips open a foil pouch and tips powder into a shaker, centered 9:16 close-up, gym locker room light. Foil tearing loudly, then powder hitting plastic.
A hand slides a phone out of a hard case and the case snaps shut, tight centered 9:16 macro, white desk. A sharp plastic snap, then a soft desk thud.
A woman pulls a hoodie over her head and looks straight to camera, centered 9:16 medium shot, flat daylight. Fabric rustle, then she says "Okay, that is thicker than I expected."
A stack of pancakes gets hit with syrup from above, centered 9:16 overhead shot, warm kitchen light. Thick syrup pour, then a fork cutting through.
A man flicks open a folding knife and starts stripping wire, centered 9:16 close-up on the hands, workbench lamp. A metallic flick, then insulation peeling.
What We Noticed Testing Prompt Structures
These are qualitative observations from our own generation logs on Morphed, not benchmark scores. We ran matched prompt pairs, changing one variable at a time, across talking-head and product briefs at 720p.
Leading with the audio cue weakened the picture. Prompts that opened with the sound and mentioned the subject afterward produced clips where the visual felt incidental, with looser framing and vaguer motion. Front-loading the subject and keeping the audio cue last gave more consistent results. This matches the general behavior of FLUX models weighting prompt start most heavily.
Long dialogue lines drift. Lines running past roughly a dozen words held sync at the start and lost it toward the end, with the mouth still moving after the audio finished or vice versa. Splitting a long line into two clips and cutting between them was more reliable than fixing it with rewording.
Two on-camera speakers is a failure mode, not a limitation to work around. Every prompt naming two visible speakers produced at least one take where the wrong person's mouth moved. Making the second voice explicitly off-camera resolved it.
Naming the room beat naming the effect. "Tiled bathroom" produced more convincing acoustics than "reverb." "Sawdust-muted workshop" produced a drier, closer sound than "dry room tone." The model appears to derive acoustics from the space it is rendering, so describing the space does double duty.
Legacy diffusion tags did nothing useful. Adding "highly detailed, 4k, octane render, trending on artstation" to otherwise identical briefs changed nothing measurable and occasionally flattened motion, which we read as the tags pulling the model toward still-image behavior. Drop them.
Our own demo on FLUX 3 Video, "Gloves of the Day," is a 12-second piece to camera with spoken lines and workshop room tone. It follows exactly the structure above: subject first, one action, one short quoted line, room named rather than described.
How Should You Split Budget Between Drafts and Finals?
Draft at 720p, finish at 1080p, and let the credit gap do the work. On Morphed, FLUX 3 Video is 25 credits per second at 720p and 42 credits per second at 1080p, so 1080p costs 68% more per second for output you are going to throw away during iteration.
| Task | Resolution | Duration | Credits |
|---|---|---|---|
| Concept pass | 720p | 8s | 200 |
| Three concept passes | 720p | 8s each | 600 |
| Full-length draft | 720p | 20s | 500 |
| Final delivery | 1080p | 20s | 840 |
| Final delivery | 1080p | 8s | 336 |
The math that matters: three 8-second concept passes at 720p cost 600 credits, still less than a single 20-second 1080p final at 840. You can be wrong twice about your dialogue line and still spend less than one premature finish. Nail the brief at 8 seconds and 720p, then spend once at full length and full resolution.
Upstream, the same shape holds. Black Forest Labs prices the API around $0.17 per second for HD and $0.29 per second for Full HD, with a draft mode near $0.06 per second, and audio included at every tier rather than billed separately. Whatever surface you use, the discipline is identical: iterate cheap, finish once.
Two settings notes. Duration accepts whole seconds from 5 to 20, so there is no penalty for a 13-second clip if that is what the beat needs. Aspect ratio has seven options, 21:9 through 9:16, and choosing it before you write the brief will change how you frame the action.
What About FLUX 3 Image?
Not released as of August 7, 2026. Black Forest Labs announced the full FLUX 3 family on July 23, 2026, and shipped FLUX 3 Video to general availability on August 4. FLUX 3 Image was described at announcement as arriving in the following weeks, and no public API access existed at the time of writing.
Here is the honest state of the family:
| Component | Status as of Aug 7, 2026 | What is known |
|---|---|---|
| FLUX 3 Video | Generally available since Aug 4, 2026 | Up to 20s, native synchronized audio, text/image/video/keyframe inputs upstream |
| FLUX 3 Image | Announced, not released | "Coming weeks" per the July 23 announcement, no date given |
| FLUX 3 Dev (open weights) | Announced, not released | "Later this year," no license terms or parameter count disclosed |
| FLUX 3 Action | Early access | Robotics and action prediction, outside consumer creative tooling |
On quality, be skeptical of anyone claiming a decisive win. Early independent testing put FLUX 3 Video near a coin flip against leading competitors on side-by-side preference, around 52%. That is a real result and it is also not a landslide. Pick it for the native audio and the way it couples sound to the space it renders, not because a leaderboard told you it is the best video model available.
If you want a general survey of what else is worth using right now, our comparison of text-to-video generators covers the current field.
When Should You Pick a Different Model?
Four scenarios where FLUX 3 Video is the wrong tool, including one that costs us conversions to say out loud.
You need to animate a still image. FLUX 3 Video on Morphed is text-to-video only. The upstream model supports image-to-video, but our integration does not expose it. If your workflow starts from a product photo, a brand still, or a character sheet, use MiniMax H3, which takes image, video, and audio references and lets you assign each one a role.
You need identity to hold across many clips. Text-to-video means describing the same person in words every time, and words are a lossy way to specify a face. Two clips generated from identical descriptions will not produce the same person. Reference-driven models handle this properly.
You need typography that is actually readable. Video models render text poorly, and FLUX 3 is no exception. If your deliverable is a poster, a packaging mockup, or anything where a word has to be legible and correctly spelled, generate the still in Qwen Image 3 instead of hoping a video frame lands it.
You need one continuous take longer than 20 seconds. Twenty seconds is the ceiling for a single generation. Upstream, longer sequences are assembled through clip chaining that carries characters forward between generations. On Morphed you build the same thing manually: generate clips, cut them together, accept the small continuity drift at each seam. If your piece genuinely requires an unbroken 45-second take, no current text-to-video model will give it to you.
And the honest one: lip-sync quality varies between takes even with a well-formed brief. A line that lands perfectly on take one can slip on take two from the identical prompt. Budget for two or three takes on any clip where speech carries the message. That variance is why the draft-at-720p workflow above matters more here than on a silent model.
Where to Start
Pick one talking-head brief from the list above, swap in your own subject and your own quoted line, and generate it at 8 seconds and 720p for 200 credits. Look at whether the sound matches the room before you look at anything else, because that is the thing this model does that others do not, and it is the fastest signal that your brief is working.
Then change one slot at a time. Tighten the line, name the room more specifically, or move the camera from static to a slow push-in. Prompt engineering on an omni model is a variable-at-a-time exercise, not a matter of finding one magic sentence.
Start generating with FLUX 3 Video on Morphed, or create an account to get set up.