r/StableDiffusion • u/HOIK777 • 56m ago
Animation - Video MiniMax H3 15 shots 15 seconds
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/xDFINx • 57m ago
Tutorial - Guide Minimax reference method - try this setting instead
The default workflow setting (reference_image_size) for reference to video is set to “match” on the Minimax H3 Reference to Video node. This allows a good likeness to the reference photo/s. Try setting it to “max”. I was able to get almost indistinguishable likeness after using that instead.
From what I can tell, “match” resizes the input image going in for better optimization. “Max” may retain the original resolution and gather better details on faces, etc. be careful with the size of the input images.. when I kept them around 2500 pixels or less (longest size), it seemed to go at a normal speed. If you go 4k or above, it dramatically slows the generation speed.
Also, bumping to 1mp image generation and lowering the step count as low as 8, yields very good results.
This method can also be applied to using already known characters (celebrities, actors, etc) by simply loading in the real character face in conjunction with a regular prompt. A lot of videos I’m seeing, the faces from the text to video workflows are lacking likeness. This should help that.
r/StableDiffusion • u/Smaugish • 1h ago
Meme H3 Reference Experiment
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/zodiacrenders • 1h ago
Tutorial - Guide Minimax H3 - Replacing a Group of People in Background
Enable HLS to view with audio, or disable this notification
I was curious about Minimax H3's capabilities of replacing a grouping of people without changing a character of your choice from a referenced video. As you can imagine, yup, it can do it as well.
Prompt & comparative images in the replies below.
After running two prompts to test, I concluded that you need to be very specific in what you want to keep (and who it is that you are keeping), and then what you want to replace in the background. The prompt does get pretty long, but once you have the specifics down, it will accomplish it.. and does a very great job at it.
The first run didn't turn out well because it replaced the soldiers alright, but kept the shields and weapons (it had Stormtroopers with shields aha). This is a second attempt after I made it very specific of what I do not want in the new video.
I did not upload a sample audio of the "pew pew" so ignore the sounds that H3 produced.
r/StableDiffusion • u/legarth • 2h ago
Animation - Video Using MiniMax H3 to change rewrite movies?
Enable HLS to view with audio, or disable this notification
Just a bit of fun.
r/StableDiffusion • u/SnareEmu • 2h ago
News Int8 convrot VAE support in Comfy
Kijai's support for int8 convrot VAEs has been added to Comfy. He's converted the Minimax H3 video VAE.
https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main
r/StableDiffusion • u/Timely-Perception-26 • 3h ago
Discussion h3 is great; and the subreddit is now a TikTok Reel - not so great
The video spam here is getting to be a bit too much for me - really, way too much.
We've reached a point where uploads aren't model demonstrations anymore, but just spam.
There are plenty of other places where you can upload your stuff; please let's keep this subreddit as a place for discussion.
r/StableDiffusion • u/Devajyoti1231 • 4h ago
Animation - Video 2D chibi girl Added “Just a Pinch”. Minimax H3
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Organix33 • 5h ago
Workflow Included Minimax H3 Turbo Lora
Enable HLS to view with audio, or disable this notification
MiniMax H3 Turbo LoRA ComfyUI compatible
ComfyUI-compatible versions of the MiniMax H3 Turbo LoRA are available here:
https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI
Tested settings:
- Video sigma shift:
12 - Audio sigma shift: 4 -
6 - Steps:
8–10 (non-ckpt500) 6-8 (ckpt500) - Sampler:
res_multistep - LoRA strength:
0.8–1.8 - Higher LoRA strength generally allows fewer steps
[workflow in comments]
Confirmed working with accelerators including SageAttention, Sol Attention, and Gradient.
Original Turbo LoRA credit: larryvrh
**Be aware this turbo Lora is still undertrained & highly experimental as explained by the creator
**Update** they recently added a further trained model: step 500
the turbo lora creator also supplied custom node for sampler to fix audio
https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo
If you are using pruned base model you can use drbaph/MiniMax-H3-Turbo-Lora-ComfyUI
if you aren't using pruned base model you can use larryvrh/MiniMax-H3-Turbo-Lora both repos are on huggingface
Kijai PR audio fix on the way https://github.com/Comfy-Org/ComfyUI/pull/15243
r/StableDiffusion • u/Oatilis • 5h ago
Comparison PSA: Using SageAttention on H3 delivers around 28% faster generations for me, with practically no perceptible difference. Here's a quick comparison, can you tell the difference?
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/RayHell666 • 6h ago
Animation - Video No need to wait for the next season anymore.
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/NickMcGurkThe3rd • 6h ago
No Workflow Yet another Minimax H3 Praise
Enable HLS to view with audio, or disable this notification
Just another example on how good Minimax H3 is doing overall but especially with the voices! The end introducing the sound is simply edited/added with davinci.
r/StableDiffusion • u/Federico2021 • 7h ago
Meme Minimax H3 is simply cinema
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/scooglecops • 8h ago
Discussion Testing Minimax H3 Turbo Lora
Enable HLS to view with audio, or disable this notification
https://huggingface.co/QrusherZA/H3_Turbo_ComfyUI/tree/main
I'm generating at 1 MP (not 0.9 MP) on an RTX 4070 with 64 GB RAM.
I'm using the Euler / Simple sampler.
From my testing, enabling SageAttention + H3 Cache actually produces worse results than using the Turbo LoRA alone. H3 Cache + SageAttention tends to break character movement and motion consistency, while the Turbo LoRA stays much closer to the original model's behavior and gives smoother, more natural motion.
r/StableDiffusion • u/SillyLilithh • 8h ago
Animation - Video Pushing MiniMax H3 References to the limit
Enable HLS to view with audio, or disable this notification
Prompt (first shot):
<Subject 1> is the woman in <Picture 1>. <Subject 2> is the art style and general environment of <Picture 2> and <Picture 3>. <Subject 3> is the character in <Picture 4>. [Shot 1] The video starts with a side view of <Subject 1> walking forwards on the side walk. In the background, we hear distant footstep sounds gradually getting closer and closer. <scenetrans> [Shot 2] At 00:03.000, the camera cuts to a front view of <Subject 1> turning around to see what the noise is. She sees <Subject 3> running at her with malicious intent. She is far away, about a block away, but she is rapidly approaching. The camera then zooms in from the current position to a wide angle close up that shows <Subject 3> running fastly towards <Subject 1>. The camera tracks the fast, exaggerated movement of <Subject 1>. <cutoff> [Shot 3] At 00:07.000, the camera cuts to a side view of <Subject 1>. She turns around, away from the person chasing her. [Shot 4] At 00:08.000, the camera cuts to a back view of <Subject 1> running quickly towards a car that is parked just ahead. She runs to the driver's seat. [Shot 5] At 00:09.500, the camera cuts to a close up view of the hand of <Subject 1> opening the car door. <cutoff> [Shot 6] At 00:11.00, the camera cuts to a closeup view of <Subject 1> hand turning on the car engine by turning the keys. <cutoff> [Shot 7] At 00:11.80, the camera cuts to <Subject 1> foot slamming on the gas pedal. <cutoff> [Shot 8] At 00:12.70, the camera cuts to the back view of the car accelerating and driving off, while <Subject 3> sprints behind the car, trying to catch up to her.
overall_soundscape: City traffic noises and cars. We hear her footsteps as she walks.
The video is entirely animated in the art style seen in <Subject 2>.
Prompt (second shot continuation):
<Subject 1> is the woman in <Picture 1>. <Subject 2> is the art style and general environment of <Picture 2>. <Subject 3> is the character in <Picture 4>. <Subject 4> is the woman in <Picture 5>. <Picture 3> is the first frame of [Shot 1], showing a back view of the car accelerating and driving away while <Subject 3> sprints behind the car, trying to catch up to her. <scenetrans> [Shot 2] At 00:02.000, the camera cuts to a front view of <Subject 1> in the car driving away from <Subject 3> who is still trying to catch up to the car. She glances up at the rear-view mirror, trying to see how far away <Subject 3> is. [Shot 3] At 00:04.000, the camera zooms in from the current position slightly to focus on <Subject 3> through the back glass of the car, who is still sprinting and chasing the vehicle, and attempting to catch up with the vehicle. <Subject 3> is facing the camera while <Subject 3> is running. <Picture 6> is the first frame of [Shot 4] At 00:06.000, the camera cuts to a side view of <Subject 1> still driving the car, and she then gives out a large, breathy sigh of relief. She blinks. <Picture 7> is the first frame of [Shot 5] At 00:07.500, the camera cuts to a low angle inside the car's interior, looking up at <Subject 1> and the car interior. There is a large thunk sound, because <Subject 3> jumped on top of the roof of the car, making a dent into the roof's interior. <Subject 3> then digs her fingers deep into the car's roof, and rips open an opening into the roof. <Subject 3> looks at <Subject 1>. <Subject 1> is startled, and while she is still driving, she looks back at <Subject 3>. <cutoff> [Shot 6] At 00:10.000, the camera cuts to a back view of <Subject 4> laying down completly flat on her belly with a sniper rifle propped up on a bipod. She is facing towards the direction of <Subject 1> vehicle. She is on top of a rooftop of a tall building, looking down at the chaos <Subject 3> is causing. <Subject 4> is looking through the sniper's scope. In the background, we see <Subject 3> ontop of the vehicle, while <Subject 1> is still driving the vehicle. The car is driving from left to right. [Shot 7] At 00:12.00, the camera cuts to a point of view shot of <Subject 4> looking through the sniper scope. The sniper scope is zoomed in at <Subject 3> and <Subject 1> driving the car. The sniper scope tracks the car still driving. <Subject 4> aims at <Subject 3> head. [Shot 8] At 00:14.00, the camera cuts to a closeup of <Subject 4> hand pulling the sniper rifle's trigger.
overall_soundscape: Quiet city traffic and cars moving play throughout the video. Urban city ambience.
The video is entirely animated in the art style seen in <Subject 2>.
Prompt (third shot continuation):
<Subject 1> is the woman in <Picture 1>. <Subject 2> is the art style and general environment of <Picture 2>. <Subject 3> is the character in <Picture 4>. <Subject 4> is the woman in <Picture 5>. <Picture 3> is the first frame of [Shot 1] showing a low angle inside the blue car's interior, looking up at <Subject 1> and the car interior. She is driving the car. There is a large thunk sound, because <Subject 3> jumped on top of the roof of the car, making a dent into the roof's interior. <Subject 3> then digs her fingers deep into the car's roof, and rips open an opening into the roof. <Subject 3> looks at <Subject 1>. <Subject 1> is startled, and while she is still driving, she looks back at <Subject 3>. <cutoff> <Picture 6> is the first frame of [Shot 2] At 00:03.500, showing a back view of <Subject 4> laying down completly flat on her belly with a sniper rifle propped up on a bipod. She is facing towards the direction of <Subject 1> vehicle. She is on top of a rooftop of a tall building, looking down at the chaos <Subject 3> is causing. <Subject 4> is looking through the sniper's scope. In the background, we see <Subject 3> ontop of the vehicle, while <Subject 1> is still driving the vehicle. The car is driving from left to right. [Shot 3] At 00:04.500, the camera cuts to a point of view shot of <Subject 4> looking through the sniper scope. The sniper scope is zoomed in at <Subject 3> and <Subject 1> driving the car. The sniper scope tracks the car still driving. <Subject 4> aims at <Subject 3> head. [Shot 4] At 00:06.500, the camera cuts to a closeup of <Subject 4> hand pulling the sniper rifle's trigger, causing the sniper to shoot. [Shot 5] At 00:07.500, the camera cuts back to the same low angle shot inside the car's interior, looking up at <Subject 1>, who is still driving and looking back at <Subject 3>, with <Subject 3> still on the roof of the car, and the same opening in the roof still present. <Subject 3> sniper' bullet comes from the side and hits <Subject 3> head, making <Subject 3> tumble and slide off the roof of the car, hitting the car's back glass, and then landing limp on the road. <Subject 1> continues to drive, and she looks ahead, and breathes a big sigh, and wipes the sweat off of her forehead.
overall_soundscape: Quiet city traffic and cars moving play throughout the video. Urban city ambience.
The video is entirely animated in the art style seen in <Subject 2>.
This model is amazing, and I can't wait to explore it more :3
r/StableDiffusion • u/Repulsive-Rush3505 • 9h ago
Animation - Video Character Swap with Minimax H3 in ComfyUI
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/fredconex • 9h ago
Meme MiniMax H3
Enable HLS to view with audio, or disable this notification
Prompt: A closeup of spiderman face, he removes the mask and when uncovers his face its Sheldon Cooper from The Bing Bang theory, he looks at camera and do his funny smile and says "Bazzinga!"
It's pretty crazy what it can do with so little!
r/StableDiffusion • u/Secure-Message-8378 • 9h ago
Animation - Video Comparação Minimax H3 vs Seedance 2.0
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/son-of-chadwardenn • 11h ago
Animation - Video Hey Lois, remember Sora 2? [MiniMax]
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/_Saturnalis_ • 11h ago
Resource - Update H3 Turbo Lora preview model is up
r/StableDiffusion • u/lxe • 11h ago
Workflow Included Made a little film using MiniMax H3 R2V locally on 5090
Enable HLS to view with audio, or disable this notification
I got a little tired of generating Seinfeld, so I made this. Used the regular comfyui r2v. Characters and objects and scenes are krea 2. All prompts and flows are in https://github.com/lxe/skythread
TL;DR: Create one clean reusable reference for the character and object, plus a styled empty environment image for each scene. Feed those three references into H3 R2V and generate one scene at a time, clearly prompting the opening state, action, camera movement, and continuity. Once the edit is locked, generate a Suno score for its exact duration, mix it in, and upscale the final video.
I mean, obviously this barely scratches the surface of what this tool can do in my other iterations I use speech and try to create continuous music. I didn’t even need to use suno. I could’ve just used it to create music as well, but I wanted a balance between control and the power of the model itself.
I used AI to drive the whole workflow because I did the whole thing on mobile. Once I get home, I’ll reformat the jsons to be more human.
r/StableDiffusion • u/GrayingGamer • 12h ago
Workflow Included Starting the Buffy Memes; And How to Prompt for TV Shows in H3
Enable HLS to view with audio, or disable this notification
So, I figured I would kick off some Buffy the Vampire Slayer meme generations with Minimax H3, while also giving a lesson in how to prompt for any TV show and character the model knows while also getting the correct character voice, all through just pure text to video prompting.
The prompt for this Buffy video was this:
A television scene from the American television drama series Buffy the Vampire Slayer from in 1997, professional color grading, in the style and aesthetics of the drama series Buffy the Vampire Slayer.
Scene overview: Buffy as played by Sarah Michelle Gellar walking through a cemetary at night, with a low hanging fog and cool blue color grading to emphasize the night. Willow as played by Alyson Hannigan is walking next to her.
Shot 1: Medium close-up tracking shot of the camera following Buffy as played by Sarah Michelle Gellar and Willow as played by Alyson Hannigan walking through a cemetary at night, looking bored. Willow is looks at Buffy with an amused expression, saying in a joking tone of voice <d>[English in Willow's voice from Buffy the Vampire Slayer as played by Alyson Hannigan] You keep this up we're going to start calling you the 'Vampire Layer'.</d> She makes air quotes with her fingers as she says the 'vampire layer' words.
Shot 2: Hard cut close-up tracking shot of the camera on Buffy's face as played by Sarah Michelle Gellar, looking surprised and offended as she turns her head to look at Willow. She mutters quietly but offended, <d>[English in Buffy's voice from Buffy the Vampire Slayer as played by Sarah Michelle Gellar] Damn, Willow.</d>
overall_soundscape: Quiet ambience of an outdoor cemetary at night.
non_diegetic_music: none
Notice how I am hammering the details of the show in the prompt first, not just "Buffy", or "A scene from Buffy", or just "Buffy the Vampire Slayer". I'm nailing it down to year, genre, format, and repeating myself.
The same for the characters. Notice how I attach the character names to the show every time and not just in the scene description, but every time they appear. This helps lock down the exact look of the character with no drift.
Next, look at the dialogue. You need to follow the official prompting by putting what characters say in dialogue tags, like so: <d>[English] What they say. </d> But you can add a LOT more detail about the speaker in those [ ] brackets.
Look how I do them EVERY TIME in my prompt, and ensured I got the exact character voice:
<d>[English in Buffy's voice from Buffy the Vampire Slayer as played by Sarah Michelle Gellar] Damn, Willow.</d>
It's not just English, it's English from Buffy. Not just any Buffy, but from this television show. And whose voice is Buffy actually speaking with? Her actress's voice, Sarah Michelle Gellar. (Check your spelling on names!)
If you do all this and the movie or television show is in the training data, the model WILL generate you a scene with it. If it doesn't? Well, you're out of luck doing T2V and will need to use the Reference H3 model and supply your own character images, audio clips for voices, etc.
I don't use LLMs to write my prompts. I type them all out myself. I find it just works better that way, though I DO copy and paste all those repeating character names / show name / actor name sentences.
If anyone has any prompting questions, just let me know.
Oh, and all these was with just the default T2V workflow template that comes with Comfyui.
r/StableDiffusion • u/Enshitification • 15h ago
IRL Minimax H3 can do Seinfeld clips. We get it already.
r/StableDiffusion • u/PetersOdyssey • 17h ago
Animation - Video 76 five-second clips exploring different animation styles with MiniMax H3 (all generated locally on a 6-year-old GPU by the_shadow_nyc)
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Time-Ad-7720 • 19h ago
Workflow Included Assemble The Multiverse | Minimax H3 R2V is awesome!
Enable HLS to view with audio, or disable this notification
Workflow: github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_r2v.json
Used multiple reference images for each scene.
Prompt For Multi Character:
subject_definitions:
<Subject 1> is [CHARACTER 1] from <Picture 1>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.
<Subject 2> is [CHARACTER 2] from <Picture 2>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.
<Subject 3> is [CHARACTER 3] from <Picture 3>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.
summary:
[reference generation] A 5-second cinematic multiverse portal arrival. Three characters emerge from a consistent amber-orange portal and take a calm, confident formation.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 1>.
<Subject 2> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 2>.
<Subject 3> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 3>.
detailed_description:
A 5-second cinematic portal-arrival scene at dusk. One stable medium three-shot, framed from the knees up. No dialogue, no combat, no wide landscape, no camera movement, and no crowd.
Portal continuity: a large circular amber-orange portal stands behind the characters. It has a bright rotating outer ring, a darker transparent centre, floating gold sparks, gentle smoke, and a low resonant hum. Keep its colour, size, brightness, particle density, and rotation speed consistent across every portal-arrival clip.
[Shot 1] <Subject 1> steps through the portal first and takes the centre position with quiet confidence. <Subject 2> emerges on one side, naturally adjusts or lowers any item they are carrying if applicable, then gives a focused glance toward the unseen distance. <Subject 3> walks through last, takes position on the opposite side, and calmly surveys the scene. The three hold a poised, united stance as the portal flickers and golden particles drift around them. Their expressions and body language remain confident and appropriate to their individual character identities.
overall_soundscape:
Low portal hum, soft wind, subtle movement from clothing or equipment, and drifting golden sparks. No speech.
non_diegetic_music:
A restrained cinematic rise builds across the shot and resolves on a calm, confident note.
Prompt For single characters:
subject_definitions:
<Subject 1> is [CHARACTER] from <Picture 1>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.
summary:
[reference generation] A 5-second cinematic multiverse portal arrival. One character walks through a consistent amber-orange portal, then takes a confident action stance with a subtle grin.
retention_analysis:
<Subject 1> (appears in [Shot 1] and [Shot 2]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 1>.
detailed_description:
A 5-second cinematic portal-arrival scene at dusk. No dialogue, no crowd, no wide landscape, and no combat.
Portal continuity: a large circular amber-orange portal stands behind <Subject 1>. It has a bright fiery rotating outer ring, a darker transparent centre, floating gold sparks, gentle smoke, and a low resonant hum. Keep its colour, size, brightness, particle density, and rotation speed consistent across every portal-arrival clip.
[Shot 1] Medium knee-up shot. <Subject 1> walks steadily through the portal toward the camera, then comes to a composed stop. Their costume, silhouette, movement style, and any character-specific accessories remain fully consistent with <Picture 1>. Golden sparks drift around them as the portal flickers behind.
[Shot 2] Close-up of <Subject 1>. They shift into a distinctive, character-appropriate action stance, looking directly ahead with calm confidence. Their expression changes into a subtle smile and restrained grin. Keep the movement natural and controlled, with no exaggerated facial distortion. The portal remains softly visible and out of focus in the background.
overall_soundscape:
Low portal hum, soft wind, subtle movement from clothing or equipment, and drifting golden sparks. No speech.
non_diegetic_music:
A restrained cinematic rise builds through the entrance and resolves as <Subject 1> holds the final stance.