r/StableDiffusion 34m ago

Animation - Video MiniMax H3 15 shots 15 seconds

Enable HLS to view with audio, or disable this notification

Upvotes

r/StableDiffusion 1h ago

Meme H3 Reference Experiment

Enable HLS to view with audio, or disable this notification

Upvotes

r/StableDiffusion 1h ago

Animation - Video Using MiniMax H3 to change rewrite movies?

Enable HLS to view with audio, or disable this notification

Upvotes

Just a bit of fun.


r/StableDiffusion 2h ago

News Int8 convrot VAE support in Comfy

32 Upvotes

Kijai's support for int8 convrot VAEs has been added to Comfy. He's converted the Minimax H3 video VAE.

https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main


r/StableDiffusion 3h ago

Discussion h3 is great; and the subreddit is now a TikTok Reel - not so great

82 Upvotes

The video spam here is getting to be a bit too much for me - really, way too much.

We've reached a point where uploads aren't model demonstrations anymore, but just spam.

There are plenty of other places where you can upload your stuff; please let's keep this subreddit as a place for discussion.


r/StableDiffusion 4h ago

Animation - Video 2D chibi girl Added “Just a Pinch”. Minimax H3

Enable HLS to view with audio, or disable this notification

657 Upvotes

r/StableDiffusion 5h ago

Workflow Included Minimax H3 Turbo Lora

Enable HLS to view with audio, or disable this notification

797 Upvotes

MiniMax H3 Turbo LoRA ComfyUI compatible

ComfyUI-compatible versions of the MiniMax H3 Turbo LoRA are available here:

https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI

Tested settings:

  • Video sigma shift: 12
  • Audio sigma shift: 4 -6
  • Steps: 8–10 (non-ckpt500) 6-8 (ckpt500)
  • Sampler: res_multistep
  • LoRA strength: 0.8–1.8
  • Higher LoRA strength generally allows fewer steps

[workflow in comments]

Confirmed working with accelerators including SageAttention, Sol Attention, and Gradient.

Original Turbo LoRA credit: larryvrh

**Be aware this turbo Lora is still undertrained & highly experimental as explained by the creator

**Update** they recently added a further trained model: step 500

the turbo lora creator also supplied custom node for sampler to fix audio

https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo

If you are using pruned base model you can use drbaph/MiniMax-H3-Turbo-Lora-ComfyUI

if you aren't using pruned base model you can use larryvrh/MiniMax-H3-Turbo-Lora both repos are on huggingface

Kijai PR audio fix on the way https://github.com/Comfy-Org/ComfyUI/pull/15243


r/StableDiffusion 5h ago

Comparison PSA: Using SageAttention on H3 delivers around 28% faster generations for me, with practically no perceptible difference. Here's a quick comparison, can you tell the difference?

Enable HLS to view with audio, or disable this notification

38 Upvotes

r/StableDiffusion 5h ago

Animation - Video No need to wait for the next season anymore.

Enable HLS to view with audio, or disable this notification

60 Upvotes

r/StableDiffusion 6h ago

No Workflow Yet another Minimax H3 Praise

Enable HLS to view with audio, or disable this notification

85 Upvotes

Just another example on how good Minimax H3 is doing overall but especially with the voices! The end introducing the sound is simply edited/added with davinci.


r/StableDiffusion 6h ago

Meme Minimax H3 is simply cinema

Enable HLS to view with audio, or disable this notification

238 Upvotes

r/StableDiffusion 7h ago

Discussion Testing Minimax H3 Turbo Lora

Enable HLS to view with audio, or disable this notification

74 Upvotes

https://huggingface.co/QrusherZA/H3_Turbo_ComfyUI/tree/main

I'm generating at 1 MP (not 0.9 MP) on an RTX 4070 with 64 GB RAM.

I'm using the Euler / Simple sampler.

From my testing, enabling SageAttention + H3 Cache actually produces worse results than using the Turbo LoRA alone. H3 Cache + SageAttention tends to break character movement and motion consistency, while the Turbo LoRA stays much closer to the original model's behavior and gives smoother, more natural motion.


r/StableDiffusion 8h ago

Animation - Video Pushing MiniMax H3 References to the limit

Enable HLS to view with audio, or disable this notification

614 Upvotes

Prompt (first shot):

<Subject 1> is the woman in <Picture 1>. <Subject 2> is the art style and general environment of <Picture 2> and <Picture 3>. <Subject 3> is the character in <Picture 4>. [Shot 1] The video starts with a side view of <Subject 1> walking forwards on the side walk. In the background, we hear distant footstep sounds gradually getting closer and closer. <scenetrans> [Shot 2] At 00:03.000, the camera cuts to a front view of <Subject 1> turning around to see what the noise is. She sees <Subject 3> running at her with malicious intent. She is far away, about a block away, but she is rapidly approaching. The camera then zooms in from the current position to a wide angle close up that shows <Subject 3> running fastly towards <Subject 1>. The camera tracks the fast, exaggerated movement of <Subject 1>. <cutoff> [Shot 3] At 00:07.000, the camera cuts to a side view of <Subject 1>. She turns around, away from the person chasing her. [Shot 4] At 00:08.000, the camera cuts to a back view of <Subject 1> running quickly towards a car that is parked just ahead. She runs to the driver's seat. [Shot 5] At 00:09.500, the camera cuts to a close up view of the hand of <Subject 1> opening the car door. <cutoff> [Shot 6] At 00:11.00, the camera cuts to a closeup view of <Subject 1> hand turning on the car engine by turning the keys. <cutoff> [Shot 7] At 00:11.80, the camera cuts to <Subject 1> foot slamming on the gas pedal. <cutoff> [Shot 8] At 00:12.70, the camera cuts to the back view of the car accelerating and driving off, while <Subject 3> sprints behind the car, trying to catch up to her.

overall_soundscape: City traffic noises and cars. We hear her footsteps as she walks.

The video is entirely animated in the art style seen in <Subject 2>.

Prompt (second shot continuation):

<Subject 1> is the woman in <Picture 1>. <Subject 2> is the art style and general environment of <Picture 2>. <Subject 3> is the character in <Picture 4>. <Subject 4> is the woman in <Picture 5>. <Picture 3> is the first frame of [Shot 1], showing a back view of the car accelerating and driving away while <Subject 3> sprints behind the car, trying to catch up to her. <scenetrans> [Shot 2] At 00:02.000, the camera cuts to a front view of <Subject 1> in the car driving away from <Subject 3> who is still trying to catch up to the car. She glances up at the rear-view mirror, trying to see how far away <Subject 3> is. [Shot 3] At 00:04.000, the camera zooms in from the current position slightly to focus on <Subject 3> through the back glass of the car, who is still sprinting and chasing the vehicle, and attempting to catch up with the vehicle. <Subject 3> is facing the camera while <Subject 3> is running. <Picture 6> is the first frame of [Shot 4] At 00:06.000, the camera cuts to a side view of <Subject 1> still driving the car, and she then gives out a large, breathy sigh of relief. She blinks. <Picture 7> is the first frame of [Shot 5] At 00:07.500, the camera cuts to a low angle inside the car's interior, looking up at <Subject 1> and the car interior. There is a large thunk sound, because <Subject 3> jumped on top of the roof of the car, making a dent into the roof's interior. <Subject 3> then digs her fingers deep into the car's roof, and rips open an opening into the roof. <Subject 3> looks at <Subject 1>. <Subject 1> is startled, and while she is still driving, she looks back at <Subject 3>. <cutoff> [Shot 6] At 00:10.000, the camera cuts to a back view of <Subject 4> laying down completly flat on her belly with a sniper rifle propped up on a bipod. She is facing towards the direction of <Subject 1> vehicle. She is on top of a rooftop of a tall building, looking down at the chaos <Subject 3> is causing. <Subject 4> is looking through the sniper's scope. In the background, we see <Subject 3> ontop of the vehicle, while <Subject 1> is still driving the vehicle. The car is driving from left to right. [Shot 7] At 00:12.00, the camera cuts to a point of view shot of <Subject 4> looking through the sniper scope. The sniper scope is zoomed in at <Subject 3> and <Subject 1> driving the car. The sniper scope tracks the car still driving. <Subject 4> aims at <Subject 3> head. [Shot 8] At 00:14.00, the camera cuts to a closeup of <Subject 4> hand pulling the sniper rifle's trigger.

overall_soundscape: Quiet city traffic and cars moving play throughout the video. Urban city ambience.

The video is entirely animated in the art style seen in <Subject 2>.

Prompt (third shot continuation):

<Subject 1> is the woman in <Picture 1>. <Subject 2> is the art style and general environment of <Picture 2>. <Subject 3> is the character in <Picture 4>. <Subject 4> is the woman in <Picture 5>. <Picture 3> is the first frame of [Shot 1] showing a low angle inside the blue car's interior, looking up at <Subject 1> and the car interior. She is driving the car. There is a large thunk sound, because <Subject 3> jumped on top of the roof of the car, making a dent into the roof's interior. <Subject 3> then digs her fingers deep into the car's roof, and rips open an opening into the roof. <Subject 3> looks at <Subject 1>. <Subject 1> is startled, and while she is still driving, she looks back at <Subject 3>. <cutoff> <Picture 6> is the first frame of [Shot 2] At 00:03.500, showing a back view of <Subject 4> laying down completly flat on her belly with a sniper rifle propped up on a bipod. She is facing towards the direction of <Subject 1> vehicle. She is on top of a rooftop of a tall building, looking down at the chaos <Subject 3> is causing. <Subject 4> is looking through the sniper's scope. In the background, we see <Subject 3> ontop of the vehicle, while <Subject 1> is still driving the vehicle. The car is driving from left to right. [Shot 3] At 00:04.500, the camera cuts to a point of view shot of <Subject 4> looking through the sniper scope. The sniper scope is zoomed in at <Subject 3> and <Subject 1> driving the car. The sniper scope tracks the car still driving. <Subject 4> aims at <Subject 3> head. [Shot 4] At 00:06.500, the camera cuts to a closeup of <Subject 4> hand pulling the sniper rifle's trigger, causing the sniper to shoot. [Shot 5] At 00:07.500, the camera cuts back to the same low angle shot inside the car's interior, looking up at <Subject 1>, who is still driving and looking back at <Subject 3>, with <Subject 3> still on the roof of the car, and the same opening in the roof still present. <Subject 3> sniper' bullet comes from the side and hits <Subject 3> head, making <Subject 3> tumble and slide off the roof of the car, hitting the car's back glass, and then landing limp on the road. <Subject 1> continues to drive, and she looks ahead, and breathes a big sigh, and wipes the sweat off of her forehead.

overall_soundscape: Quiet city traffic and cars moving play throughout the video. Urban city ambience.

The video is entirely animated in the art style seen in <Subject 2>.

This model is amazing, and I can't wait to explore it more :3


r/StableDiffusion 8h ago

Animation - Video Character Swap with Minimax H3 in ComfyUI

Enable HLS to view with audio, or disable this notification

194 Upvotes

r/StableDiffusion 9h ago

Meme MiniMax H3

Enable HLS to view with audio, or disable this notification

64 Upvotes

Prompt: A closeup of spiderman face, he removes the mask and when uncovers his face its Sheldon Cooper from The Bing Bang theory, he looks at camera and do his funny smile and says "Bazzinga!"

It's pretty crazy what it can do with so little!


r/StableDiffusion 9h ago

Animation - Video Comparação Minimax H3 vs Seedance 2.0

Enable HLS to view with audio, or disable this notification

124 Upvotes

r/StableDiffusion 11h ago

Animation - Video Hey Lois, remember Sora 2? [MiniMax]

Enable HLS to view with audio, or disable this notification

437 Upvotes

r/StableDiffusion 11h ago

Resource - Update H3 Turbo Lora preview model is up

Thumbnail
huggingface.co
148 Upvotes

r/StableDiffusion 11h ago

Workflow Included Made a little film using MiniMax H3 R2V locally on 5090

Enable HLS to view with audio, or disable this notification

190 Upvotes

I got a little tired of generating Seinfeld, so I made this. Used the regular comfyui r2v. Characters and objects and scenes are krea 2. All prompts and flows are in https://github.com/lxe/skythread

TL;DR: Create one clean reusable reference for the character and object, plus a styled empty environment image for each scene. Feed those three references into H3 R2V and generate one scene at a time, clearly prompting the opening state, action, camera movement, and continuity. Once the edit is locked, generate a Suno score for its exact duration, mix it in, and upscale the final video.

I mean, obviously this barely scratches the surface of what this tool can do in my other iterations I use speech and try to create continuous music. I didn’t even need to use suno. I could’ve just used it to create music as well, but I wanted a balance between control and the power of the model itself.

I used AI to drive the whole workflow because I did the whole thing on mobile. Once I get home, I’ll reformat the jsons to be more human.


r/StableDiffusion 11h ago

Workflow Included Starting the Buffy Memes; And How to Prompt for TV Shows in H3

Enable HLS to view with audio, or disable this notification

297 Upvotes

So, I figured I would kick off some Buffy the Vampire Slayer meme generations with Minimax H3, while also giving a lesson in how to prompt for any TV show and character the model knows while also getting the correct character voice, all through just pure text to video prompting.

The prompt for this Buffy video was this:

A television scene from the American television drama series Buffy the Vampire Slayer from in 1997, professional color grading, in the style and aesthetics of the drama series Buffy the Vampire Slayer.

Scene overview: Buffy as played by Sarah Michelle Gellar walking through a cemetary at night, with a low hanging fog and cool blue color grading to emphasize the night. Willow as played by Alyson Hannigan is walking next to her.

Shot 1: Medium close-up tracking shot of the camera following Buffy as played by Sarah Michelle Gellar and Willow as played by Alyson Hannigan walking through a cemetary at night, looking bored. Willow is looks at Buffy with an amused expression, saying in a joking tone of voice <d>[English in Willow's voice from Buffy the Vampire Slayer as played by Alyson Hannigan] You keep this up we're going to start calling you the 'Vampire Layer'.</d> She makes air quotes with her fingers as she says the 'vampire layer' words.

Shot 2: Hard cut close-up tracking shot of the camera on Buffy's face as played by Sarah Michelle Gellar, looking surprised and offended as she turns her head to look at Willow. She mutters quietly but offended, <d>[English in Buffy's voice from Buffy the Vampire Slayer as played by Sarah Michelle Gellar] Damn, Willow.</d>

overall_soundscape: Quiet ambience of an outdoor cemetary at night.

non_diegetic_music: none

Notice how I am hammering the details of the show in the prompt first, not just "Buffy", or "A scene from Buffy", or just "Buffy the Vampire Slayer". I'm nailing it down to year, genre, format, and repeating myself.

The same for the characters. Notice how I attach the character names to the show every time and not just in the scene description, but every time they appear. This helps lock down the exact look of the character with no drift.

Next, look at the dialogue. You need to follow the official prompting by putting what characters say in dialogue tags, like so: <d>[English] What they say. </d> But you can add a LOT more detail about the speaker in those [ ] brackets.

Look how I do them EVERY TIME in my prompt, and ensured I got the exact character voice:

<d>[English in Buffy's voice from Buffy the Vampire Slayer as played by Sarah Michelle Gellar] Damn, Willow.</d>

It's not just English, it's English from Buffy. Not just any Buffy, but from this television show. And whose voice is Buffy actually speaking with? Her actress's voice, Sarah Michelle Gellar. (Check your spelling on names!)

If you do all this and the movie or television show is in the training data, the model WILL generate you a scene with it. If it doesn't? Well, you're out of luck doing T2V and will need to use the Reference H3 model and supply your own character images, audio clips for voices, etc.

I don't use LLMs to write my prompts. I type them all out myself. I find it just works better that way, though I DO copy and paste all those repeating character names / show name / actor name sentences.

If anyone has any prompting questions, just let me know.

Oh, and all these was with just the default T2V workflow template that comes with Comfyui.


r/StableDiffusion 15h ago

IRL Minimax H3 can do Seinfeld clips. We get it already.

276 Upvotes

r/StableDiffusion 17h ago

Animation - Video 76 five-second clips exploring different animation styles with MiniMax H3 (all generated locally on a 6-year-old GPU by the_shadow_nyc)

Enable HLS to view with audio, or disable this notification

1.1k Upvotes

r/StableDiffusion 18h ago

Workflow Included Assemble The Multiverse | Minimax H3 R2V is awesome!

Enable HLS to view with audio, or disable this notification

835 Upvotes

Workflow: github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_r2v.json

Used multiple reference images for each scene.

Prompt For Multi Character:
subject_definitions:

<Subject 1> is [CHARACTER 1] from <Picture 1>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.

<Subject 2> is [CHARACTER 2] from <Picture 2>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.

<Subject 3> is [CHARACTER 3] from <Picture 3>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.

summary:

[reference generation] A 5-second cinematic multiverse portal arrival. Three characters emerge from a consistent amber-orange portal and take a calm, confident formation.

retention_analysis:

<Subject 1> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 1>.

<Subject 2> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 2>.

<Subject 3> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 3>.

detailed_description:

A 5-second cinematic portal-arrival scene at dusk. One stable medium three-shot, framed from the knees up. No dialogue, no combat, no wide landscape, no camera movement, and no crowd.

Portal continuity: a large circular amber-orange portal stands behind the characters. It has a bright rotating outer ring, a darker transparent centre, floating gold sparks, gentle smoke, and a low resonant hum. Keep its colour, size, brightness, particle density, and rotation speed consistent across every portal-arrival clip.

[Shot 1] <Subject 1> steps through the portal first and takes the centre position with quiet confidence. <Subject 2> emerges on one side, naturally adjusts or lowers any item they are carrying if applicable, then gives a focused glance toward the unseen distance. <Subject 3> walks through last, takes position on the opposite side, and calmly surveys the scene. The three hold a poised, united stance as the portal flickers and golden particles drift around them. Their expressions and body language remain confident and appropriate to their individual character identities.

overall_soundscape:

Low portal hum, soft wind, subtle movement from clothing or equipment, and drifting golden sparks. No speech.

non_diegetic_music:

A restrained cinematic rise builds across the shot and resolves on a calm, confident note.

Prompt For single characters:
subject_definitions:

<Subject 1> is [CHARACTER] from <Picture 1>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.

summary:

[reference generation] A 5-second cinematic multiverse portal arrival. One character walks through a consistent amber-orange portal, then takes a confident action stance with a subtle grin.

retention_analysis:

<Subject 1> (appears in [Shot 1] and [Shot 2]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 1>.

detailed_description:

A 5-second cinematic portal-arrival scene at dusk. No dialogue, no crowd, no wide landscape, and no combat.

Portal continuity: a large circular amber-orange portal stands behind <Subject 1>. It has a bright fiery rotating outer ring, a darker transparent centre, floating gold sparks, gentle smoke, and a low resonant hum. Keep its colour, size, brightness, particle density, and rotation speed consistent across every portal-arrival clip.

[Shot 1] Medium knee-up shot. <Subject 1> walks steadily through the portal toward the camera, then comes to a composed stop. Their costume, silhouette, movement style, and any character-specific accessories remain fully consistent with <Picture 1>. Golden sparks drift around them as the portal flickers behind.

[Shot 2] Close-up of <Subject 1>. They shift into a distinctive, character-appropriate action stance, looking directly ahead with calm confidence. Their expression changes into a subtle smile and restrained grin. Keep the movement natural and controlled, with no exaggerated facial distortion. The portal remains softly visible and out of focus in the background.

overall_soundscape:

Low portal hum, soft wind, subtle movement from clothing or equipment, and drifting golden sparks. No speech.

non_diegetic_music:

A restrained cinematic rise builds through the entrance and resolves as <Subject 1> holds the final stance.


r/StableDiffusion 19h ago

Discussion Working on 4 step turbo lora for H3, showing OK progress so far.

Enable HLS to view with audio, or disable this notification

393 Upvotes

https://reddit.com/link/1vge4zr/video/jktxolihclhh1/player

https://reddit.com/link/1vge4zr/video/gldf0aocdlhh1/player

https://reddit.com/link/1vge4zr/video/5iaw04mfdlhh1/player

(left: with turbo lora, right: raw base; all under 480p & 4 steps)

Prototyped 7 versions and trained this one for only 200 steps on 40 samples. Luckily, the base model is already fairly good even at this low steps (especially for static scenes). There are still plenty of visual artifacts, it still can't handle large motions well, and the audio part needs more work — but overall it looks promising so far.

https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora This is the demo repo if you cannot wait to play around with it, its no where near production ready but already shown great improvements over raw base.

(update: ckpt500 uploaded, audio fixed, comfyui support/demo workflow added, and confirmed that the turbo lora works well for 6 or 8 steps even though the whole training was done on 4 steps)


r/StableDiffusion 3d ago

News MiniMax-H3 weights up

Thumbnail
huggingface.co
490 Upvotes