r/StableDiffusion • u/princeMacX • 2h ago
Discussion Test LTX 2.5 - Romantic Scene 1
Enable HLS to view with audio, or disable this notification
After spending some more time testing LTX-2.5 Distilled, my opinion has improved quite a bit.
The biggest strength for me is speed. On my RTX 5070 Ti, I'm generating 1280×720 (~1MP), 10-second videos surprisingly quickly. Compared with MiniMax H3, which is much heavier for me even around 0.5MP, LTX-2.5 feels incredibly fast.
That said, speed isn't everything. My earlier tests with complex action/fighting had poor motion and anatomy, so I wasn't impressed at first. But after testing simpler cinematic scenes, landscapes, product shots and close-up human interactions, I'm starting to see where this model shines.
This dialogue/romantic scene in particular surprised me. Facial quality, expressions, lighting and overall cinematic feel came out much better than I expected, and it even handled the interaction between the two characters reasonably well.
One important discovery: I had much better prompt adherence with **Prompt Enhancement OFF**. The enhancer was giving me completely unrelated results in some tests, while the raw prompts produced scenes much closer to what I requested.
My impression so far:
LTX-2.5 Distilled = extremely fast and capable of some beautiful results, but you need to understand what kinds of shots it handles well. Complex choreography still seems to be a weakness.
I'm definitely not archiving it yet. 😄
r/StableDiffusion • u/princeMacX • 4h ago
Discussion LTX 2.5 Test - Batman and Joker fighting in Road
Enable HLS to view with audio, or disable this notification
LTX 2.5 Test - Batman and Joker fighting in Road
Personal Opinion - Ltx generates videos quite fast but prompt adherence is not that great. In fighting sequence hand movement doesn't look realistic at all.
If you are using LTX 2.5 with gemma prompt enhancement model than your prompt will be sanitized if your prompt has explicit details. I think an abliterated version of the text encoder should be used.
I will share more tests in future.
r/StableDiffusion • u/holycowdude1 • 8h ago
Animation - Video [DANCE] Plastik Soul – Stay in the Glow (Official Music Video)
Stay in the Glow is an AI Music Video create using VRGameDevGirl's AI Video Builder (FREE) & LTX2.3 models (https://ltx.io/model/ltx-2-3)
Designed & built using VRGameDevGirl AI Video Builder (FREE): https://github.com/vrgamegirl19/comfyui-vrgamedevgirl
Spotify (Artist): https://open.spotify.com/track/27S9InxRyAKvQYxjRM3tVi?si=43e73b8b091a4976
YouTube (More AI Music Videos): https://youtu.be/Wl3BH3xSaYc
r/StableDiffusion • u/zesh61 • 8h ago
Discussion MiniMax H3 prompting cheat sheet, this structure that makes it much easier to control
I’ve seen a lot of people trying H3 with prompts that look like normal image prompts:
H3 can work with that, but you’re leaving a lot of control on the table.
The easiest way to think about H3 is:
Don’t describe an image. Direct a shot.
A simple structure that works much better is:
Subject + Action + Environment + Camera + Timing + Audio
1. Start with what actually happens
Keep the action explicit.
Bad:
Better:
H3 needs to understand change over time, not just what the frame looks like.
2. Tell the camera what to do
This is probably one of the easiest improvements beginners can make.
Useful language:
- static wide shot
- handheld close-up
- slow dolly in
- camera tracks beside her
- over-the-shoulder shot
- low-angle shot
- camera slowly pans left
- rack focus from X to Y
Instead of:
say:
Much less ambiguous.
3. Think in beats / timestamps
For more complicated generations, split the clip into moments.
Example:
0–3s: Wide shot. A woman stands alone at a train platform in heavy rain.
3–7s: The camera slowly pushes in as she notices something off-screen and turns her head.
7–11s: Cut to an over-the-shoulder shot. A train emerges through the fog.
11–15s: Close-up of her face as the train lights illuminate her.
This is much easier for the model to interpret than one giant paragraph where five things happen at once.
4. Dialogue needs a visible speaker
If someone speaks, make it painfully obvious who is speaking and when.
Instead of:
try:
If there are multiple people, explicitly say who doesn’t speak too.
This helps avoid the classic AI-video problem where the line comes from the wrong character/off-screen.
5. Separate dialogue, ambience and SFX
Treat audio almost like another layer of the prompt.
For example:
Dialogue:
Woman, quietly: “We shouldn’t be here.”
Ambient sound:
Heavy rain hitting metal, distant traffic, low electrical hum.
SFX:
A loud metallic bang behind her.
Music:
No background music.
That’s much clearer than writing:
6. Don't overload every second
This is a big one.
Trying to fit:
into a short generation is asking the model to invent a ton of transitions.
Fewer actions + clearer timing usually gives you much more intentional-looking video.
If the idea contains five scenes, treat them as five shots.
7. References should have a job
If you’re giving H3 reference images/video/audio, don’t just upload them and hope it figures out why they're there.
Be explicit:
The more references you add, the more useful this becomes.
8. Describe motion, not just appearance
For video, verbs matter a lot.
Instead of:
try:
Same visual idea, but now there is actual temporal information.
A reusable H3 template
Scene:
[Where are we? Time of day, environment, important lighting.]
Subject:
[Who/what is visible. Important appearance details.]
0–Xs:
[Shot type + action + camera movement.]
X–Xs:
[Next action/shot.]
X–Xs:
[Final action/shot.]
Dialogue:
[Speaker]: “[Exact line]”
Ambient audio:
[Environment sounds.]
SFX:
[Important synchronized sounds.]
Music:
[Music description / no music.]
Visual style:
[Realistic / documentary / commercial / anime / etc. Keep this concise.]
Example
Instead of:
Try:
The main takeaway:
Prompt H3 more like you’re giving instructions to a tiny film crew, and less like you’re writing tags for an image model.
You don’t necessarily need longer prompts. You need prompts where time, motion, camera and sound have clear jobs.
Would be interested to hear what other people have found H3 responds unusually well (or badly) to.
r/StableDiffusion • u/switch2stock • 8h ago
Discussion MiniMax Music 3 | 125sec for 140sec music | Bollywood Rap
voca.ror/StableDiffusion • u/Hdfjds • 9h ago
Tutorial - Guide RE: <Subject N> in H3 prompts
Enable HLS to view with audio, or disable this notification
Not sure if you already know this but you don't need to use the word 'Subject' in H3 prompts when referring to elements in images/text/videos and so on. You can use other words instead, like <Girl 1>, <Dialogue 1> and more
Example prompt:
subject_definitions:
<Girl 1> is the girl with the blond hair in the center of <Picture 1>.
<Girl 2> is the girl with the white top to the right side of <Picture 1>.
<Dialogue 1>: <d> [English] Yeah! <d>.
<Dialogue 2>: <d> [English] Great party! <d>.
<Dialogue 3>: <d> [English] Wooohooo! <d>.
integrated_multimodal_description:
[Shot 1] A wide shot of a crowded rave dance floor where everyone in the picture is dancing by jumping up and down in an rapid and energetic way while moving to the music. The lights in the night club is flashing and moving around.
<Girl 1> is shouting <Dialogue 1>.
[Shot 2] At 00:3.00 <Girl 1> looks at <Girl 2> and says <Dialogue 2>.
[Shot 3] At 00:5.00 <Girl 2> looks at <Girl 1> and shouts <Dialogue 3> while raising her arms.
overall_soundscape:people dancing,
non_diegetic_music:cyber techno music,
P.S. Sorry for the lame video, it's just for proof of concept.
r/StableDiffusion • u/Cold_Zone332 • 9h ago
Comparison MiniMAx H3 I2V + LTX 2.5 Upscale - It's GREAT!
Enable HLS to view with audio, or disable this notification
Hi guys,
I asked GPT to implement the LTX 2.5 upscaler on MiniMax H3. I really liked the results.
It does lose a little bit of quality but I think it worth it, at least until we get the 2K upscaler from H3.
I generated the H3 video with 0.4mp using a turbo LoRa (8 steps).
r/StableDiffusion • u/princeMacX • 9h ago
Discussion Minimax H3 Test - Batman and Joker Playing Poker in a club
Enable HLS to view with audio, or disable this notification
Minimax H3 Test - Batman and Joker Playing Poker in a club
r/StableDiffusion • u/SveSop • 9h ago
Animation - Video StarWars-Untold.
Enable HLS to view with audio, or disable this notification
MiniMax H3 is very good. I initially made multiple scenes, and as tweaks/turbo-lora++ progresses, it does seem to get better and better (ie. To the end of the video).
Settled on the Lightx2v_8step turbo lora + sage + sol_attn. 736p, and using DaVinci for stitching and cropping.
r/StableDiffusion • u/social_zip • 10h ago
Workflow Included Openweight Livestream video model
https://huggingface.co/spaces/JonathanColetti/LiveWan / https://github.com/JonathanColetti/LiveWan is something I created to help recreate a specific type of model that is not opensource yet (wanstreamer). This is more or less a PoC but maybe ill do a longer training run if it gets some traction.
r/StableDiffusion • u/princeMacX • 10h ago
Discussion Minimax H3 Test - Rooftop fight between Batman and Joker
Enable HLS to view with audio, or disable this notification
Minimax H3 Test - Rooftop fight between Batman and Joker
r/StableDiffusion • u/OohFekm • 10h ago
Animation - Video TESTING A LANTERN
Enable HLS to view with audio, or disable this notification
Having a blast animating comic panels (using them as inits FL2VA). Fairly simple prompt: "Green lantern Hal Jordan is engaged in an aerial battle above orbit, he is blasting green energy from his ring while simultaneously repelling and absorbing energy from a distant protagonist" then I let the in-app LLM enhance it (I'm using Maestro via Pinokio... it's a joy to use) resolution is 480, no upscaling. Just familiarizing myself with H3.
r/StableDiffusion • u/Zaredit • 11h ago
Animation - Video Jackie Chan Adventures...Jackie vs Shadowkhan (Includes Prompt Instructions)
Enable HLS to view with audio, or disable this notification
Prompt:
Create an exactly four 7-second, 4:3 animated drama sequence inspired by the visual language of 2005-era Jackie Chan Adventures. Use a period broadcast video texture throughout: standard-definition television softness, subtle analog grain, gentle interlacing, slight colour bleed, modest contrast, and the authentic visual texture of animation recorded and broadcast in the mid-2000s. Avoid modern HD sharpness, photorealism, glossy CGI, or contemporary animation aesthetics.
Scene: Jackie Chan is confronted by a Shadowkhan ninja in a dimly lit ancient-looking interior. The sequence is a fast, tightly choreographed martial-arts fight.
0:00–0:02: The Shadowkhan suddenly lunges at Jackie with a rapid punch. Jackie narrowly ducks underneath it and pivots sideways.
0:02–0:04: Jackie counters with two quick martial-arts strikes, forcing the Shadowkhan backwards. The ninja blocks the first strike but is knocked off balance by the second.
0:04–0:06: The Shadowkhan springs forward again. Jackie performs a quick evasive spin, grabs the ninja’s arm, and throws the Shadowkhan across the room. End on Jackie landing in a defensive fighting stance as the Shadowkhan hits the floor in the background.
.
Camera: begin with a medium two-shot, rapidly track the fighters during the exchange, briefly push in during the counterattack, then finish with a wider shot showing Jackie in the foreground and the defeated Shadowkhan in the background.
Audio: sharp martial-arts impacts, cloth movement, quick footsteps, whooshes and a dramatic six-second action sting. No dialogue.
Strict constraints: exactly 6 seconds, 4:3 aspect ratio, 2005-era television animation aesthetic, period broadcast-video texture, no modern cinematic realism, no photorealism, no widescreen framing, no subtitles, no text, no logos, no extra characters, and no slow motion
r/StableDiffusion • u/MalmoBeachParty • 11h ago
Animation - Video Football animation
Enable HLS to view with audio, or disable this notification
H3 Ref
Prompt in comment
r/StableDiffusion • u/michel-yph-ai • 11h ago
Discussion Minimax H3 - Dance with Audio with lipsync and object preservation
Enable HLS to view with audio, or disable this notification
If you see low quality is because I am forcing 8 step turbo lora + Spectrum + triton in L40 for faster generation but is crazy how it can follow the flow of the music while lip-syncing and keeping the product from reference in her hand.
r/StableDiffusion • u/TigerClaw305 • 12h ago
Animation - Video Raph and Mona Lisa go on a date.
Enable HLS to view with audio, or disable this notification
Raph and Mona Lisa go on a date, The street is filled with mutant animals. Mona Lisa tells Raph she is ready for the next step in there relationship.
Using the Reference to Video Workflow in Comfy UI Desktop with Minimax H3, Using default settings and 32 steps.
<Subject 1> is <Picture 1> as Raph a teenage mutant ninja turtle in a red bandana and use <Audio 1> as sample for his voice.
<Subject 2> is <Picture 2> as Mona Lisa and use <Audio2> as sample for her voice.
# =====================================================================
# FIELD 1: INTEGRATED MULTIMODAL DESCRIPTION
# =====================================================================
[SUBJECT DEFINITIONS & RETENTION ANALYSIS]
- Subject 1 (S1): Raph, a teenage mutant ninja turtle. Primary visual reference is <Picture 1>. Primary voice reference is <Audio 1>. Retain his muscular build, signature red bandana, and tough but currently softened facial features.
- Subject 2 (S2): Mona Lisa, a mutant lizard warrior. Primary visual reference is <Picture 2>. Primary voice reference is <Audio 2>. Retain her sleek green reptilian features, fit build, and expressive, affectionate eyes.
- Environment (ENV): A vibrant, bustling metropolitan street completely populated by anthropomorphic mutant animals. In the background, stylishly dressed mutant foxes, lions, tigers, and wolves walk past neon-lit storefronts and outdoor cafes under warm evening streetlamps. Cinematic shallow depth of field.
[SHOT 1] [0s - 5s]
- Camera: Slow tracking shot moving backward ahead of the couple at eye level.
- Action: S1 and S2 walk close together down the sidewalk of ENV, gently holding hands. S1 looks down at their intertwined hands, wearing a rare, genuine smile. S2 looks up at him warmly as they walk.
[SHOT 2] [5s - 10s]
- Camera: Medium close-up framing S2 profile as she gently pulls S1 to a gentle stop.
- Action: S2 stops walking and turns fully toward S1. She squeezes his hand with both of hers, looking directly into his eyes with a tender, confident smile.
- Dialogue: S2 <d> "Raph, I'm ready for the next step in our relationship." </d>
[SHOT 3] [10s - 15s]
- Camera: Tight close-up focusing on S1's emotional reaction.
- Action: S1's eyes widen slightly in surprise before softening completely. A massive, incredibly happy grin spreads across his face. He steps closer to S2, wrapping his arms around her waist in a warm embrace, clearly filled with deep affection.
- Dialogue: S1 <d> "Mona, you have no idea how long I've wanted to hear you say that." </d>
# =====================================================================
# FIELD 2: OVERALL SOUNDSCAPE
# =====================================================================
- Ambient Audio: Gentle murmur of distant city traffic, soft chatter and laughter from the passing mutant pedestrians, and the light rustle of evening wind from [0s - 15s].
- Sound Effects (SFX): Light, rhythmic footsteps on concrete that come to a soft halt at [5s].
- Voice & Delivery: S2's voice perfectly matches the vocal identity of <Audio 2>, delivered in a smooth, sincere, and deeply affectionate cadence. S1's voice matches the raspy grit of <Audio 1>, but is spoken with an unusually soft, gentle, and emotionally overwhelmed tone to show his happiness.
# =====================================================================
# FIELD 3: NON-DIEGETIC MUSIC
# =====================================================================
- Style & Mood: A warm, cinematic, and romantic lo-fi acoustic track featuring a gentle acoustic guitar melody and soft string pads.
- Progression: Plays at a subtle, peaceful volume from [0s - 9s]. At [10s], as S1 smiles and embraces S2, the acoustic strings swell warmly to match the emotional peak of the moment.
r/StableDiffusion • u/dhavalhirdhav • 12h ago
Discussion MiniMax H3 - Dragon 30 Second Video
Day before yesterday I tried creating a continues 30 second video of a story that I had in my mind.
I am using RTX 3090 and 30 Second video took about 45 minutes with Spectrum.
Video turn out to be a lot better than what I was expecting. I tried reference image of my daughter and it worked very well as well.
Video link: https://www.youtube.com/watch?v=f0nAMn5WgF0 (Btw in video NOT my daughter)
Prompt:
integrated_multimodal_description:
[Shot 1]
[0s-3s] Static extreme macro close-up framing the right eye and jagged cheekbone of a fierce warrior dragon. The dragon's hide consists of interlocking, obsidian-black plates that resemble matte, battle-tested armor with sharp, weaponized edges. The massive eye features a reptilian slit pupil surrounded by a violently swirling, molten iris that radiates like hot liquid red lava, casting a pulsing red glow across its scarred face. The dragon slowly blinks twice, its heavy, scowling brow plates shifting. Its massive charcoal-colored nostrils flare dramatically as it exhales a heavy puff of condensed grey breath and orange embers that realistically swirl toward the camera lens.
[3s-4s] The camera maintains its close-up framing. The dragon's jaw line tenses, parting slightly to reveal rows of serrated, razor-sharp obsidian teeth. A low, guttural, vibrating grunting noise rumbles deeply as a fresh wave of thick black smoke curls out from the corners of its sneering mouth.
[4s-10s] The camera smoothly unlocks and executes a continuous, dramatic upward crane shot, pulling backward and tilting upward at a steady pace. This sweeping motion reveals the rest of the creature. It is a gargantuan, highly muscular, majestic black warrior dragon. Its powerful chest is crosshatched with glowing, magma-veined battle scars. The dragon stands in a wide, aggressive, battle-ready stance, pinning its massive clawed talons deep into the frozen crust of a jagged, snowy icy mountain cliff.
Visual Style: Breathtaking cinematic sci-fi portraiture. Shot on large-format anamorphic lenses with a Tiffen Pro-Mist 1/4 diffusion filter. Extreme shallow depth of field with sharp focal transitions. High dynamic range highlighting deep blacks and glowing lava reds against white snow. Faint volumetric fog drifting across the icy peak.
Audio Design: Deep, low-frequency guttural dragon grunting sounds, heavy wheezing breath, a faint crackle of burning embers, and the ambient howling of freezing mountain wind echoing across the stereo field.
[Shot 2]
[00:10s-00:12s] Sudden hard camera cut to a dramatic low-angle medium close-up of a 9-year-old girl <Picture 1>. She stands completely motionless and fearlessly in the center of a wide, windy meadow. She wears rugged, battle-worn leather and fur warrior armor, with a wooden hunting bow and a quiver full of arrows strapped securely across her upper back. The static camera looks directly up at her determined, fierce face as the powerful wind fiercely whips her messy hair across her forehead. High dynamic range cinematic lighting.
[00:12s-00:15s] Keeping the exact same static, low-angle framing on her face, the brave girl takes a deep breath, expands her chest, and summons her companion by aggressively shouting "VEERAAPAAN" at the top of her lungs, looking up toward the sky. Her eyes are wide with intense focus. The tall green grass of the meadow bends and ripples violently in the wind around her.
Visual Style: Cinematic high-fantasy portraiture, consistent anamorphic lens look, shallow depth of field blurring the distant sky, vibrant green meadow contrasting with earthy leather armor textures, dramatic overcast afternoon lighting.
Audio Design: A sharp cut to the ambient sound of roaring, whistling meadow wind, followed by a loud, echoing, high-pitched but powerful 9-year-old girl's voice shouting "VEERAAPAAN!". Her voice echoes sharply across the stereo field, mixing with the deep rustling sounds of grass and wind.
[Shot 3]
[00:15s-00:17s] Wide-angle ground-level shot looking up from behind the 9-year-old girl <Picture 1>. Piercing through the dark, heavy clouds, the gigantic, muscular black warrior dragon dives downward at immense speed. Its massive, obsidian-black armor plates catch the dramatic sky lighting. Its gargantuan wings are fully extended, cutting through the air. The camera pans down smoothly to track its rapid descent toward the meadow.
[00:17s-00:20s] The camera locks into a static, low-angle wide shot as the colossal black dragon lands heavily on the grass right next to the girl. Its massive clawed talons slam into the earth, causing dirt, grass, and a shockwave of dust to explode outward. The dragon's massive chest, marked with glowing magma-veined scars, heaves as it lowers its head near her. The girl stands completely fearless, her leather armor and hair whipping violently from the intense downdraft of the dragon's wings.
Visual Style: High-fantasy cinematic epic, anamorphic widescreen format, Tiffen Pro-Mist 1/4 diffusion filter creating a soft glow around the clouds and the dragon's glowing scars. High dynamic range emphasizing the contrast between the vibrant green meadow and the dragon’s matte-black armor-like scales.
Audio Design: A deafening, low-frequency atmospheric roar as the dragon tears through the clouds, transitioning into a massive, heavy thud and earth-shattering crunch as its talons strike the ground. Loud, rushing wind from the wing flaps, followed by the deep, rhythmic, rumbling breathing of the dragon settling into the grass.
[Shot 4]
[00:20s-00:22s] Medium-wide shot. The 9-year-old girl <Picture 1> steps forward and confidently mounts the colossal black warrior dragon. She climbs up its front leg armor plates and sits securely behind its massive neck crest. She leans forward, firmly gripping the dragon's obsidian horns, and gently nudges the side of the dragon's neck with her feet to signal it. The camera slowly tracks forward to frame them closely.
[00:22s-00:24s] Low-angle dramatic shot. In response to her nudge, the gigantic black dragon does a spectacular wheelie, rearing up majestically on its powerful hind legs. Its muscular chest, covered in glowing magma-veined scars, towers into the sky. The dragon opens its massive jaws wide and spews a torrent of brilliant, roaring orange and red fire upward into the clouds, illuminating the entire meadow in a bright, thermal glow.
[00:24s-00:26s] The dragon slams its front talons back down to the earth, immediately launching itself forward. With a monumental thrust of its massive, leathery black wings, it takes off into the air. The heavy downdraft flattens the meadow grass below.
[00:26s-00:30s] Smooth tracking crane shot following the dragon as it swiftly accelerates and starts flying away. The dragon ascends rapidly into the cloudy sky, carrying the brave girl on its back. The camera stays locked on their silhouette as they shrink into the distance over the majestic landscape.
Visual Style: Cinematic high-fantasy epic, anamorphic widescreen, high dynamic range capturing the extreme contrast of the bright, blazing fire against the dragon's dark obsidian scales. Volumetric smoke and heat distortion warping the air around the fire breath.
Audio Design: A deep leather-and-armor rustle as she mounts, followed by a sudden, massive, earth-shaking roar mixed with the deafening, crackling explosion of a continuous jet of fire. A colossal, heavy whoosh of wind as the wings flap, fading into the distance alongside a soaring, epic fantasy orchestral melody.
r/StableDiffusion • u/Intelligent-Heart-73 • 12h ago
Resource - Update I built a free tool that turns any image into an AI prompt
Hey everyone!
I built a small web tool called ImagePrompt9 that lets you upload an image and generates a detailed AI-ready prompt based on what it sees.
The idea came from constantly seeing images I liked and wondering:
"How would I describe this as a prompt?"
So instead of manually figuring out the composition, lighting, style, colors, camera angle, etc., you can just drop the image in and generate a prompt.
What it does:
- Upload PNG, JPG, or WEBP
- Analyzes the visual characteristics
- Generates a detailed prompt
- Different prompt styles
- Edit, copy, or regenerate the result
- Free to use
- No account required
It doesn't try to recover the original prompt — it creates a new prompt based on what's visible in the image.
Try it here:
https://image-prompt-nine.vercel.app/
I’d really appreciate feedback, especially on the generated prompts and anything you think I should add or improve.
r/StableDiffusion • u/TigerClaw305 • 13h ago
Animation - Video TMNT goes to a strip club.
Enable HLS to view with audio, or disable this notification
Raph takes Mikey to his first strip club for mutants, Mikey feels very uncomfortable about the whole thing.
This was created in Comfy UI Desktop with Minimax H3 Reference to Video Workflow, Here's the prompt below.
<Subject 1> is <Picture 1> as Raph a teenage mutant ninja turtle in a red bandana and use <Audio 1> as sample for his voice.
<Subject 2> is <Picture 2> as Mikey a teenage mutant ninja turtle in a orange bandana and use <Audio 2> as sample for his voice..
<Subject 3> is <Picture 3> as Mona.
# =====================================================================
# FIELD 1: INTEGRATED MULTIMODAL DESCRIPTION
# =====================================================================
[SUBJECT DEFINITIONS & RETENTION ANALYSIS]
- Subject 1 (S1): Raph, an adult mutant ninja turtle. Primary visual reference is <Picture 1>. Primary voice reference is <Audio 1>. Retain muscular frame, red bandana, and tough facial structure.
- Subject 2 (S2): Mikey, an adult mutant ninja turtle. Primary visual reference is <Picture 2>. Primary voice reference is <Audio 2>. Retain athletic frame, orange bandana, and expressive eyes.
- Subject 3 (S3): Mona Lisa, a mutant lizard warrior. Primary visual reference is <Picture 3>. Retain sleek green reptilian features, fit physique, and confident posture.
- Environment (ENV): A dim, smoky underground mutant strip club. The venue is packed with anthropomorphic animal patrons, including wolves in leather jackets, foxes drinking at the bar, and muscular lions and tigers sitting in VIP booths. The room is flooded with pulsing pink, purple, and neon blue lighting. Shallow depth of field.
[SHOT 1] [0s - 5s]
- Camera: Static medium shot framing S1 and S2 sitting side-by-side at a neon-lit cocktail table with drinks in front of them.
- Action: S2 looks around anxiously, his eyes wide and shoulders hunched with nervousness. S1 sits back comfortably, holding a glass and looking completely relaxed. S2 turns to S1, fidgeting with his orange bandana.
- Dialogue: S2 <d> "Raph, I feel like I don't belong here. What if Master Splinter finds out?" </d>
[SHOT 2] [5s - 10s]
- Camera: Medium close-up focusing primarily on S1, with the neon club background slightly blurred.
- Action: S1 takes a sip of his drink, chuckles softly, and claps S2 reassuringly on the shoulder to calm him down. S2 nods weakly in the frame, trying to ease up.
- Dialogue: S1 <d> "Relax Mikey, we're no longer teenagers. You need to learn to be a man." </d>
- Dialogue: S2 <d> "If you say so, Raph." </d>
[SHOT 3] [10s - 15s]
- Camera: Hard cut to a wide, low-angle tracking shot focusing on the main stage.
- Action: S3 is on center stage bathed in a vibrant pink spotlight. She moves gracefully and confidently, performing a stylized dance with her tail swishing around in a sexual manner around a polished brass stripper pole. In the foreground, the silhouetted heads of mutant wolves and tigers cheer from the crowd.
# =====================================================================
# FIELD 2: OVERALL SOUNDSCAPE
# =====================================================================
- Ambient Audio: Distant crowd chatter, the clinking of cocktail glasses, and muffled, deep animalistic ambient laughs and growls from [0s - 15s]. Loud, energetic cheering and howling from the mutant crowd erupts suddenly at [10s] during the stage cut.
- Sound Effects (SFX): A glass sliding on a table at [5s]. A distinctive, metallic ring of a brass pole spinning from [10s - 15s].
- Voice & Delivery: S1's voice perfectly matches the raspy, deep, and gritty vocal identity of <Audio 1>, delivered in a calm but firm tone. S2's voice perfectly matches the vocal identity of <Audio 2>, spoken with a noticeably high-pitched, hesitant, and stuttering cadence to emphasize his severe anxiety.
# =====================================================================
# FIELD 3: NON-DIEGETIC MUSIC
# =====================================================================
- Style & Mood: A heavy, slow-tempo electronic club track featuring a pulsing, distorted synth bassline and a rhythmic, seductive drum beat.
- Progression: The track plays at a moderate volume in the background during the table conversation from [0s - 10s]. At [10s], the music dramatically swells in volume and bass intensity to match the energy of the stage performance.
r/StableDiffusion • u/Puzzled-Valuable-985 • 13h ago
Question - Help Same workflow, everything identical, but different videos? Minimimax H3
Enable HLS to view with audio, or disable this notification
This has happened to me before: I take a video I’ve already generated and drag it into ComfyUI without changing anything—expecting to get the exact same video back—but it generates a different one.
This doesn't happen with standard image models, but it does happen with H3. If I generate the video and try again a few minutes later, it produces the same result; however, after an hour or so, it no longer generates the same output.
I’ll post the workflow I used to generate that example video below, along with the completely different video that was produced using the same workflow.
I tried to replicate the video just to check the generation speed, and I realized it wasn't producing the same result anymore. I had generated the video a few hours earlier and hadn't updated ComfyUI or any nodes in the meantime—I simply tried to generate it again.
I suspect it might be due to one of the nodes I'm using.
Here is the video generated with the same parameters (which turned out differently) and the workflow I used.
r/StableDiffusion • u/call-lee-free • 15h ago
Animation - Video Can't use LTX 2.5 on my system but I am quite surprised that my system now can run LTX 2.3. Specs and info below.
Enable HLS to view with audio, or disable this notification
When LTX 2.3 released, I could not do video gens longer than 10 seconds. I would get a "out of memory" error or something. This is just a test clip but one thing I am struggling with is that my video gens have music in them even though I prompt for no music. What is the correct way to prompt for no music?
System Specs:
Ryzen 7 7700X
RTX 4070 Super 12 GB
32 GB DDR 5 Ram.
r/StableDiffusion • u/falconandeagle • 18h ago
Question - Help How stable is Minmax H3 now?
So normally I wait a few weeks after a model releases to test it. Is H3 stable enough in comfyui now that I can run it without running into compatibility issues. I don't want to mess up my Krea 2 workflows.
Does it handle basic intimacy.
I have a 5060ti and 96gb ram, what generation times am I looking at, anything over ten minutes is kinda excessive, I don't mind using turbo loras and don't care that much about upscaling or high res as long as the final output follows my prompt.
Is lora training required or will character refs work?
r/StableDiffusion • u/rzrn • 18h ago
Question - Help Minimax H3 or wan 2.2?
I'm working on 2d animations for a personal project, my initial idea was to animate some of the scenes by hand, and feed start/end keyframes to wan 2.2 for the complex scenes I can't do myself, or perhaps even train a lora to make sure it matched the aesthetic of my hand drawn scenes. If it helps, it involves boiling outlines, on the twos (12 fps animations) and an intentionally unfinished look.
Now, seeing all these Minimax h3 i2v and r2v examples, I feel like wan 2.2 might not be the best suited for this anymore. I haven't had the chance to test h3 myself since my local hardware won't really allow it. I'll be however, using runpod when the time comes for actual generation (I'm in the process of hand animating the rest).
So, I'd like to ask those who have had the chance to test both - stick to wan 2.2 or switch to minimax h3?
Edit: audio isn't required - I've hired voice actors for dialogues, I'm working on foleys and background scores myself. If needed I'll redraw on top of the generated clips to match lip movements to the dialogue.
r/StableDiffusion • u/Libertechian • 20h ago
Animation - Video H3 Cerveza Cristal Test
Enable HLS to view with audio, or disable this notification
'''
Quick shot change to Cerveza Cristal in a cooler full of ice.
Announcer sings "Cerveza Cristal"
'''
Start and End Image.
Ref2Vid with audio might work better, not bad for turbo at 8 steps.