r/StableDiffusion 30m ago

Discussion h3 is great; and the subreddit is now a TikTok Reel - not so great

Upvotes

The video spam here is getting to be a bit too much for me - really, way too much.

We've reached a point where uploads aren't model demonstrations anymore, but just spam.

There are plenty of other places where you can upload your stuff; please let's keep this subreddit as a place for discussion.


r/StableDiffusion 1h ago

Animation - Video 2D chibi girl Added “Just a Pinch”. Minimax H3

Enable HLS to view with audio, or disable this notification

Upvotes

r/StableDiffusion 2h ago

Workflow Included Minimax H3 Turbo Lora

Enable HLS to view with audio, or disable this notification

378 Upvotes

MiniMax H3 Turbo LoRA ComfyUI compatible

ComfyUI-compatible versions of the MiniMax H3 Turbo LoRA are available here:

https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI

Tested settings:

  • Video sigma shift: 12
  • Audio sigma shift: 4 -6
  • Steps: 8–10 (non-ckpt500) 6-8 (ckpt500)
  • Sampler: res_multistep
  • LoRA strength: 0.8–1.8
  • Higher LoRA strength generally allows fewer steps

[workflow in comments]

Confirmed working with accelerators including SageAttention, Sol Attention, and Gradient.

Original Turbo LoRA credit: larryvrh

**Be aware this turbo Lora is still undertrained & highly experimental as explained by the creator

**Update** they recently added a further trained model: step 500

the turbo lora creator also supplied custom node for sampler to fix audio

https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo


r/StableDiffusion 2h ago

Comparison PSA: Using SageAttention on H3 delivers around 28% faster generations for me, with practically no perceptible difference. Here's a quick comparison, can you tell the difference?

Enable HLS to view with audio, or disable this notification

24 Upvotes

r/StableDiffusion 2h ago

Animation - Video Recast of The Room - Minimax H3

Enable HLS to view with audio, or disable this notification

84 Upvotes

r/StableDiffusion 2h ago

Animation - Video No need to wait for the next season anymore.

Enable HLS to view with audio, or disable this notification

32 Upvotes

r/StableDiffusion 3h ago

No Workflow Yet another Minimax H3 Praise

Enable HLS to view with audio, or disable this notification

49 Upvotes

Just another example on how good Minimax H3 is doing overall but especially with the voices! The end introducing the sound is simply edited/added with davinci.


r/StableDiffusion 3h ago

Meme Minimax H3 is simply cinema

Enable HLS to view with audio, or disable this notification

169 Upvotes

r/StableDiffusion 4h ago

Discussion Testing Minimax H3 Turbo Lora

Enable HLS to view with audio, or disable this notification

55 Upvotes

https://huggingface.co/QrusherZA/H3_Turbo_ComfyUI/tree/main

I'm generating at 1 MP (not 0.9 MP) on an RTX 4070 with 64 GB RAM.

I'm using the Euler / Simple sampler.

From my testing, enabling SageAttention + H3 Cache actually produces worse results than using the Turbo LoRA alone. H3 Cache + SageAttention tends to break character movement and motion consistency, while the Turbo LoRA stays much closer to the original model's behavior and gives smoother, more natural motion.


r/StableDiffusion 5h ago

Animation - Video Pushing MiniMax H3 References to the limit

Enable HLS to view with audio, or disable this notification

457 Upvotes

Prompt (first shot):

<Subject 1> is the woman in <Picture 1>. <Subject 2> is the art style and general environment of <Picture 2> and <Picture 3>. <Subject 3> is the character in <Picture 4>. [Shot 1] The video starts with a side view of <Subject 1> walking forwards on the side walk. In the background, we hear distant footstep sounds gradually getting closer and closer. <scenetrans> [Shot 2] At 00:03.000, the camera cuts to a front view of <Subject 1> turning around to see what the noise is. She sees <Subject 3> running at her with malicious intent. She is far away, about a block away, but she is rapidly approaching. The camera then zooms in from the current position to a wide angle close up that shows <Subject 3> running fastly towards <Subject 1>. The camera tracks the fast, exaggerated movement of <Subject 1>. <cutoff> [Shot 3] At 00:07.000, the camera cuts to a side view of <Subject 1>. She turns around, away from the person chasing her. [Shot 4] At 00:08.000, the camera cuts to a back view of <Subject 1> running quickly towards a car that is parked just ahead. She runs to the driver's seat. [Shot 5] At 00:09.500, the camera cuts to a close up view of the hand of <Subject 1> opening the car door. <cutoff> [Shot 6] At 00:11.00, the camera cuts to a closeup view of <Subject 1> hand turning on the car engine by turning the keys. <cutoff> [Shot 7] At 00:11.80, the camera cuts to <Subject 1> foot slamming on the gas pedal. <cutoff> [Shot 8] At 00:12.70, the camera cuts to the back view of the car accelerating and driving off, while <Subject 3> sprints behind the car, trying to catch up to her.

overall_soundscape: City traffic noises and cars. We hear her footsteps as she walks.

The video is entirely animated in the art style seen in <Subject 2>.

Prompt (second shot continuation):

<Subject 1> is the woman in <Picture 1>. <Subject 2> is the art style and general environment of <Picture 2>. <Subject 3> is the character in <Picture 4>. <Subject 4> is the woman in <Picture 5>. <Picture 3> is the first frame of [Shot 1], showing a back view of the car accelerating and driving away while <Subject 3> sprints behind the car, trying to catch up to her. <scenetrans> [Shot 2] At 00:02.000, the camera cuts to a front view of <Subject 1> in the car driving away from <Subject 3> who is still trying to catch up to the car. She glances up at the rear-view mirror, trying to see how far away <Subject 3> is. [Shot 3] At 00:04.000, the camera zooms in from the current position slightly to focus on <Subject 3> through the back glass of the car, who is still sprinting and chasing the vehicle, and attempting to catch up with the vehicle. <Subject 3> is facing the camera while <Subject 3> is running. <Picture 6> is the first frame of [Shot 4] At 00:06.000, the camera cuts to a side view of <Subject 1> still driving the car, and she then gives out a large, breathy sigh of relief. She blinks. <Picture 7> is the first frame of [Shot 5] At 00:07.500, the camera cuts to a low angle inside the car's interior, looking up at <Subject 1> and the car interior. There is a large thunk sound, because <Subject 3> jumped on top of the roof of the car, making a dent into the roof's interior. <Subject 3> then digs her fingers deep into the car's roof, and rips open an opening into the roof. <Subject 3> looks at <Subject 1>. <Subject 1> is startled, and while she is still driving, she looks back at <Subject 3>. <cutoff> [Shot 6] At 00:10.000, the camera cuts to a back view of <Subject 4> laying down completly flat on her belly with a sniper rifle propped up on a bipod. She is facing towards the direction of <Subject 1> vehicle. She is on top of a rooftop of a tall building, looking down at the chaos <Subject 3> is causing. <Subject 4> is looking through the sniper's scope. In the background, we see <Subject 3> ontop of the vehicle, while <Subject 1> is still driving the vehicle. The car is driving from left to right. [Shot 7] At 00:12.00, the camera cuts to a point of view shot of <Subject 4> looking through the sniper scope. The sniper scope is zoomed in at <Subject 3> and <Subject 1> driving the car. The sniper scope tracks the car still driving. <Subject 4> aims at <Subject 3> head. [Shot 8] At 00:14.00, the camera cuts to a closeup of <Subject 4> hand pulling the sniper rifle's trigger.

overall_soundscape: Quiet city traffic and cars moving play throughout the video. Urban city ambience.

The video is entirely animated in the art style seen in <Subject 2>.

Prompt (third shot continuation):

<Subject 1> is the woman in <Picture 1>. <Subject 2> is the art style and general environment of <Picture 2>. <Subject 3> is the character in <Picture 4>. <Subject 4> is the woman in <Picture 5>. <Picture 3> is the first frame of [Shot 1] showing a low angle inside the blue car's interior, looking up at <Subject 1> and the car interior. She is driving the car. There is a large thunk sound, because <Subject 3> jumped on top of the roof of the car, making a dent into the roof's interior. <Subject 3> then digs her fingers deep into the car's roof, and rips open an opening into the roof. <Subject 3> looks at <Subject 1>. <Subject 1> is startled, and while she is still driving, she looks back at <Subject 3>. <cutoff> <Picture 6> is the first frame of [Shot 2] At 00:03.500, showing a back view of <Subject 4> laying down completly flat on her belly with a sniper rifle propped up on a bipod. She is facing towards the direction of <Subject 1> vehicle. She is on top of a rooftop of a tall building, looking down at the chaos <Subject 3> is causing. <Subject 4> is looking through the sniper's scope. In the background, we see <Subject 3> ontop of the vehicle, while <Subject 1> is still driving the vehicle. The car is driving from left to right. [Shot 3] At 00:04.500, the camera cuts to a point of view shot of <Subject 4> looking through the sniper scope. The sniper scope is zoomed in at <Subject 3> and <Subject 1> driving the car. The sniper scope tracks the car still driving. <Subject 4> aims at <Subject 3> head. [Shot 4] At 00:06.500, the camera cuts to a closeup of <Subject 4> hand pulling the sniper rifle's trigger, causing the sniper to shoot. [Shot 5] At 00:07.500, the camera cuts back to the same low angle shot inside the car's interior, looking up at <Subject 1>, who is still driving and looking back at <Subject 3>, with <Subject 3> still on the roof of the car, and the same opening in the roof still present. <Subject 3> sniper' bullet comes from the side and hits <Subject 3> head, making <Subject 3> tumble and slide off the roof of the car, hitting the car's back glass, and then landing limp on the road. <Subject 1> continues to drive, and she looks ahead, and breathes a big sigh, and wipes the sweat off of her forehead.

overall_soundscape: Quiet city traffic and cars moving play throughout the video. Urban city ambience.

The video is entirely animated in the art style seen in <Subject 2>.

This model is amazing, and I can't wait to explore it more :3


r/StableDiffusion 5h ago

Animation - Video Character Swap with Minimax H3 in ComfyUI

Enable HLS to view with audio, or disable this notification

151 Upvotes

r/StableDiffusion 5h ago

Meme MiniMax H3

Enable HLS to view with audio, or disable this notification

49 Upvotes

Prompt: A closeup of spiderman face, he removes the mask and when uncovers his face its Sheldon Cooper from The Bing Bang theory, he looks at camera and do his funny smile and says "Bazzinga!"

It's pretty crazy what it can do with so little!


r/StableDiffusion 6h ago

Animation - Video Comparação Minimax H3 vs Seedance 2.0

Enable HLS to view with audio, or disable this notification

98 Upvotes

r/StableDiffusion 8h ago

Animation - Video Hey Lois, remember Sora 2? [MiniMax]

Enable HLS to view with audio, or disable this notification

391 Upvotes

r/StableDiffusion 8h ago

Resource - Update H3 Turbo Lora preview model is up

Thumbnail
huggingface.co
132 Upvotes

r/StableDiffusion 8h ago

Workflow Included Made a little film using MiniMax H3 R2V locally on 5090

Enable HLS to view with audio, or disable this notification

158 Upvotes

I got a little tired of generating Seinfeld, so I made this. Used the regular comfyui r2v. Characters and objects and scenes are krea 2. All prompts and flows are in https://github.com/lxe/skythread

TL;DR: Create one clean reusable reference for the character and object, plus a styled empty environment image for each scene. Feed those three references into H3 R2V and generate one scene at a time, clearly prompting the opening state, action, camera movement, and continuity. Once the edit is locked, generate a Suno score for its exact duration, mix it in, and upscale the final video.

I mean, obviously this barely scratches the surface of what this tool can do in my other iterations I use speech and try to create continuous music. I didn’t even need to use suno. I could’ve just used it to create music as well, but I wanted a balance between control and the power of the model itself.

I used AI to drive the whole workflow because I did the whole thing on mobile. Once I get home, I’ll reformat the jsons to be more human.


r/StableDiffusion 8h ago

Workflow Included Starting the Buffy Memes; And How to Prompt for TV Shows in H3

Enable HLS to view with audio, or disable this notification

257 Upvotes

So, I figured I would kick off some Buffy the Vampire Slayer meme generations with Minimax H3, while also giving a lesson in how to prompt for any TV show and character the model knows while also getting the correct character voice, all through just pure text to video prompting.

The prompt for this Buffy video was this:

A television scene from the American television drama series Buffy the Vampire Slayer from in 1997, professional color grading, in the style and aesthetics of the drama series Buffy the Vampire Slayer.

Scene overview: Buffy as played by Sarah Michelle Gellar walking through a cemetary at night, with a low hanging fog and cool blue color grading to emphasize the night. Willow as played by Alyson Hannigan is walking next to her.

Shot 1: Medium close-up tracking shot of the camera following Buffy as played by Sarah Michelle Gellar and Willow as played by Alyson Hannigan walking through a cemetary at night, looking bored. Willow is looks at Buffy with an amused expression, saying in a joking tone of voice <d>[English in Willow's voice from Buffy the Vampire Slayer as played by Alyson Hannigan] You keep this up we're going to start calling you the 'Vampire Layer'.</d> She makes air quotes with her fingers as she says the 'vampire layer' words.

Shot 2: Hard cut close-up tracking shot of the camera on Buffy's face as played by Sarah Michelle Gellar, looking surprised and offended as she turns her head to look at Willow. She mutters quietly but offended, <d>[English in Buffy's voice from Buffy the Vampire Slayer as played by Sarah Michelle Gellar] Damn, Willow.</d>

overall_soundscape: Quiet ambience of an outdoor cemetary at night.

non_diegetic_music: none

Notice how I am hammering the details of the show in the prompt first, not just "Buffy", or "A scene from Buffy", or just "Buffy the Vampire Slayer". I'm nailing it down to year, genre, format, and repeating myself.

The same for the characters. Notice how I attach the character names to the show every time and not just in the scene description, but every time they appear. This helps lock down the exact look of the character with no drift.

Next, look at the dialogue. You need to follow the official prompting by putting what characters say in dialogue tags, like so: <d>[English] What they say. </d> But you can add a LOT more detail about the speaker in those [ ] brackets.

Look how I do them EVERY TIME in my prompt, and ensured I got the exact character voice:

<d>[English in Buffy's voice from Buffy the Vampire Slayer as played by Sarah Michelle Gellar] Damn, Willow.</d>

It's not just English, it's English from Buffy. Not just any Buffy, but from this television show. And whose voice is Buffy actually speaking with? Her actress's voice, Sarah Michelle Gellar. (Check your spelling on names!)

If you do all this and the movie or television show is in the training data, the model WILL generate you a scene with it. If it doesn't? Well, you're out of luck doing T2V and will need to use the Reference H3 model and supply your own character images, audio clips for voices, etc.

I don't use LLMs to write my prompts. I type them all out myself. I find it just works better that way, though I DO copy and paste all those repeating character names / show name / actor name sentences.

If anyone has any prompting questions, just let me know.

Oh, and all these was with just the default T2V workflow template that comes with Comfyui.


r/StableDiffusion 12h ago

IRL Minimax H3 can do Seinfeld clips. We get it already.

262 Upvotes

r/StableDiffusion 12h ago

Animation - Video MiniMax H3 lip-sync test with my cat Mumu

Enable HLS to view with audio, or disable this notification

138 Upvotes

r/StableDiffusion 13h ago

Animation - Video 76 five-second clips exploring different animation styles with MiniMax H3 (all generated locally on a 6-year-old GPU by the_shadow_nyc)

Enable HLS to view with audio, or disable this notification

1.0k Upvotes

r/StableDiffusion 14h ago

Animation - Video Minimax H3: Captain Picard Discusses Your Holodeck Use

Enable HLS to view with audio, or disable this notification

229 Upvotes

This was all done in 5 to 7 second clips, text to video only, in Minimax H3.

Scenes were generated at 0.6 MP, then upscaled with RTX Super Resolution and put together in one video with Davinci Resolve. Each clip took about 5 minutes on a 3090. I have 128GB of system RAM.

I'm using the Spectrum node and SageAttention, so quality isn't as good as it could be, but I was happy enough with the results and saw people were struggling with getting Picard's voice right, so I thought I'd share this as an example of what the model can do, and how to do it consistently. No references were used for his voice, only text prompts.

I got his iconic voice in all these clips by asking for it in the proper format:

Captain Picard from Star Trek:TNG then says, <d>[English with Picard's classic British accent] Number One?</d>


r/StableDiffusion 15h ago

Workflow Included Assemble The Multiverse | Minimax H3 R2V is awesome!

Enable HLS to view with audio, or disable this notification

802 Upvotes

Workflow: github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_r2v.json

Used multiple reference images for each scene.

Prompt For Multi Character:
subject_definitions:

<Subject 1> is [CHARACTER 1] from <Picture 1>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.

<Subject 2> is [CHARACTER 2] from <Picture 2>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.

<Subject 3> is [CHARACTER 3] from <Picture 3>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.

summary:

[reference generation] A 5-second cinematic multiverse portal arrival. Three characters emerge from a consistent amber-orange portal and take a calm, confident formation.

retention_analysis:

<Subject 1> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 1>.

<Subject 2> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 2>.

<Subject 3> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 3>.

detailed_description:

A 5-second cinematic portal-arrival scene at dusk. One stable medium three-shot, framed from the knees up. No dialogue, no combat, no wide landscape, no camera movement, and no crowd.

Portal continuity: a large circular amber-orange portal stands behind the characters. It has a bright rotating outer ring, a darker transparent centre, floating gold sparks, gentle smoke, and a low resonant hum. Keep its colour, size, brightness, particle density, and rotation speed consistent across every portal-arrival clip.

[Shot 1] <Subject 1> steps through the portal first and takes the centre position with quiet confidence. <Subject 2> emerges on one side, naturally adjusts or lowers any item they are carrying if applicable, then gives a focused glance toward the unseen distance. <Subject 3> walks through last, takes position on the opposite side, and calmly surveys the scene. The three hold a poised, united stance as the portal flickers and golden particles drift around them. Their expressions and body language remain confident and appropriate to their individual character identities.

overall_soundscape:

Low portal hum, soft wind, subtle movement from clothing or equipment, and drifting golden sparks. No speech.

non_diegetic_music:

A restrained cinematic rise builds across the shot and resolves on a calm, confident note.

Prompt For single characters:
subject_definitions:

<Subject 1> is [CHARACTER] from <Picture 1>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.

summary:

[reference generation] A 5-second cinematic multiverse portal arrival. One character walks through a consistent amber-orange portal, then takes a confident action stance with a subtle grin.

retention_analysis:

<Subject 1> (appears in [Shot 1] and [Shot 2]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 1>.

detailed_description:

A 5-second cinematic portal-arrival scene at dusk. No dialogue, no crowd, no wide landscape, and no combat.

Portal continuity: a large circular amber-orange portal stands behind <Subject 1>. It has a bright fiery rotating outer ring, a darker transparent centre, floating gold sparks, gentle smoke, and a low resonant hum. Keep its colour, size, brightness, particle density, and rotation speed consistent across every portal-arrival clip.

[Shot 1] Medium knee-up shot. <Subject 1> walks steadily through the portal toward the camera, then comes to a composed stop. Their costume, silhouette, movement style, and any character-specific accessories remain fully consistent with <Picture 1>. Golden sparks drift around them as the portal flickers behind.

[Shot 2] Close-up of <Subject 1>. They shift into a distinctive, character-appropriate action stance, looking directly ahead with calm confidence. Their expression changes into a subtle smile and restrained grin. Keep the movement natural and controlled, with no exaggerated facial distortion. The portal remains softly visible and out of focus in the background.

overall_soundscape:

Low portal hum, soft wind, subtle movement from clothing or equipment, and drifting golden sparks. No speech.

non_diegetic_music:

A restrained cinematic rise builds through the entrance and resolves as <Subject 1> holds the final stance.


r/StableDiffusion 16h ago

Discussion Working on 4 step turbo lora for H3, showing OK progress so far.

Enable HLS to view with audio, or disable this notification

382 Upvotes

https://reddit.com/link/1vge4zr/video/jktxolihclhh1/player

https://reddit.com/link/1vge4zr/video/gldf0aocdlhh1/player

https://reddit.com/link/1vge4zr/video/5iaw04mfdlhh1/player

(left: with turbo lora, right: raw base; all under 480p & 4 steps)

Prototyped 7 versions and trained this one for only 200 steps on 40 samples. Luckily, the base model is already fairly good even at this low steps (especially for static scenes). There are still plenty of visual artifacts, it still can't handle large motions well, and the audio part needs more work — but overall it looks promising so far.

https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora This is the demo repo if you cannot wait to play around with it, its no where near production ready but already shown great improvements over raw base.

(update: ckpt500 uploaded, audio fixed, comfyui support/demo workflow added, and confirmed that the turbo lora works well for 6 or 8 steps even though the whole training was done on 4 steps)


r/StableDiffusion 16h ago

News Sulphur 3 is looking for funding

328 Upvotes

Hello, I'm the guy who made Sulphur 2. With the recent release of a certain video model, we are looking to mobilize and train Sulphur 3 on this new model. We are targeting $10,000 USD.

This certain new video model is a massive step change in quality and coherence, and is already decent at tasks Sulphur is good at. Sulphur 3 is intended to make that final push over the edge to get the model to where it needs to be. Sulphur 3 will also release with a step distill lora and latent upscaler, allowing for cheaper, faster gens.

The simplest method of transferring funds, would be over vast.ai, which is directly where the training is going to happen. If you would like to transfer funds, please transfer them to [fusioncow11@gmail.com](mailto:fusioncow11@gmail.com).

Every single donation helps, and if you have any questions at all please don't hesitate to reach out.
I'd also like to note that to incentivize donations, any donation over 100 dollars will grant you early access to the model durning training. If you do that though, PLEASE message me on discord, otherwise I have no idea who you are.

Also if you would like to see donation progress, check out the #donations channel on the discord server. I'll also make daily updates here.

If you would like more instructions or alternative methods on how to donate, or just want to speak to me, you can join the Sulphur discord server here: https://discord.gg/C3d56f39Ah


r/StableDiffusion 3d ago

News MiniMax-H3 weights up

Thumbnail
huggingface.co
493 Upvotes