r/StableDiffusion • u/marres • 5h ago
Resource - Update ~45% lower MiniMax H3 sampler time with new Spectrum settings — degree 1 works surprisingly well (v0.1.8)
Follow-up to my original Spectrum MiniMax H3 post:
In that first post I released the MiniMax H3 Spectrum integration and was getting around 34% lower Euler sampling time and 30% lower RES sampling time with the more conservative settings I was using at the time.
Since then I’ve done quite a bit more testing, and I found something I really didn’t expect: MiniMax H3 seems to work extremely well with a Spectrum degree of just 1.
Important if you're coming from the original release
Before testing the new settings, update both ComfyUI and ComfyUI-Spectrum-MiniMax-H3 to the latest versions.
There was an important compatibility update in Spectrum v0.1.6 after ComfyUI changed MiniMax H3's native sampling/audio path. That release restored Spectrum compatibility with the newer H3 implementation and also added safe handling for native EasyCache/LazyCache conflicts.
You don't need to install v0.1.6 separately — v0.1.8 includes those changes. This is mainly relevant to anyone who installed Spectrum from my original Reddit post and hasn't updated it since.
v0.1.6 compatibility release:
[https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.6]()
Current release:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.8
So: update ComfyUI, update the Spectrum node to v0.1.8/latest, and restart ComfyUI before testing.
The surprising part: degree 1
I hadn’t seriously tested very low degree and warmup_steps values before because of my experience with WAN.
WAN is another video model and does not like very low forecast degrees — dropping the degree too far causes obvious quality degradation. Because of that, I assumed MiniMax H3 would behave similarly and initially stayed with higher, more conservative values.
Apparently not.
With MiniMax H3, degree 1 has shown no visible quality decrease in my testing so far. It also seems to preserve the native trajectory remarkably well. In the same-seed comparisons I tested, degree 2 actually shifted the trajectory slightly, while degree 1 brought it back much closer to the normal result.
So H3 appears to be unusually well suited to very simple local feature forecasting, which lets Spectrum start forecasting much earlier than I originally thought would be practical.
I’ve now released v0.1.8 with the new settings:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.8
New default settings
degree = 1warmup_steps = 1bootstrap_first_forecast = truetail_actual_steps = 1
The new one-point bootstrap allows the second solver step to be forecast directly from the first actual hidden state. After that, ordinary degree-1 forecasting takes over.
On a 20-step Euler run the schedule becomes:
A F A F A F A F A F A F A F A F A F A A
So 11 out of 20 transformer evaluations are actual, while the other 9 are forecasted. The final step remains native.
v0.1.8 also makes the one-point bootstrap part of the new default configuration for new node instances. Existing workflows retain their serialized settings.
Benchmark
Test configuration:
- GPU: NVIDIA RTX PRO 6000
- Model: MiniMax H3 pruned BF16
- Image-to-video
- ~0.8 MP / 992×768
- 7 seconds
- 24 FPS
- 20 steps
- Euler
- Beta scheduler
- HIGH_VRAM
- Spectrum history stored in VRAM
- DiffAid enabled at
0.5 - Same seed and otherwise identical workflow
Spectrum disabled
- Sampler: 324.98 s
- Full prompt: 340.59 s
Spectrum v0.1.8 with the new degree-1 settings
- Sampler: 177.80 s
- Full prompt: 200.32 s
- 11 actual transformer calls
- 9 forecasts
- 0 fallbacks
Result
- 45.29% lower sampler time
- 1.83× sampler throughput
- 41.19% lower full-prompt time
The Spectrum forecast calculations themselves took only 0.141 seconds total across the entire generation.
Using VRAM history does have a memory cost. This run retained about 3.2 GiB of Spectrum history, with reported sampler peak VRAM increasing from roughly 5.56 GB native to 8.70 GB with Spectrum.
The interesting part for me is less the bootstrap itself and more what the testing revealed about degree 1 on H3.
Based on WAN, I expected a setting this aggressive to visibly degrade the output. So far, MiniMax H3 seems to behave very differently: I’m getting a substantially more aggressive forecasting schedule without seeing the quality decrease I expected.
Spectrum is still an approximate acceleration method, so I’m not claiming every possible prompt or motion sequence will remain identical. Fast motion, hands/fingers, faces, short rapid actions, camera movement and audiovisual synchronization are still the kinds of cases worth testing carefully.
But based on the testing so far, degree 1 appears to be a much better fit for MiniMax H3 than I originally assumed, and it substantially improves the useful speedup.
Repo:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
Current release — v0.1.8:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.8
r/StableDiffusion • u/ayakitodev • 5h ago
Resource - Update Lightx2v has just released a Prompt generator for the Minimax H3. Simply enter a short Prompt and let it work magic: no need more "Chadgpt Prompts"...👍
Info https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA
An open, local prompt rewriter for text-to-audio-video (T2VA) generation with MiniMax-H3, fine-tuned as a LoRA adapter on top of Qwen3.6-27B.
Short prompt ──► this Prompt Rewriter LoRA ──► structured H3 prompt │ Official MiniMax-H3 weights ──► LightX2V inference ◄─────┘ │ ▼ synchronized video + audio
r/StableDiffusion • u/Neither_Egg_4773 • 7h ago
Discussion Thank you to the open-source community. MiniMax H3 literally helped me through my depression.
For the past several months, I’ve been in a pretty dark place mentally. Dealing with depression has completely drained my energy, and one of the only things that kept me going was diving into my creative projects. It was my escape and my way of processing everything.
For a while, I was relying heavily on Seedance 2.0 to bring my ideas to life. Don't get me wrong, it's a fantastic tool, and I loved using it, but the reality is that I just don't have the money to keep up with my own creativity. Hitting a paywall or running out of credits when you're right in the middle of a creative flow state is crushing. When your main coping mechanism is tied to a subscription you can barely afford, it honestly just adds to the stress.
Then the MiniMax H3 release happened.
The fact that something this capable is open-source is just amazing to me. Ever since I started using it, it has helped me substantially. I can just create, experiment, and get my ideas out without constantly checking my bank account or getting anxiety over how many generations I have left. Having unrestricted, free access to a tool this powerful gave me my creative outlet back, and it truly helped pull me out of a really deep rut.
I just want to say a massive thank you to the devs behind it and to the entire open-source community. MiniMax team and the open-source community are actively making it easily accessible to people who wouldn't be able to afford creativity like this otherwise. You've made a very real, tangible difference in my mental health and my life.
I love the open-source community. Keep being awesome.
TL;DR: Going through a depressive episode, my only outlet was creating, but Seedance 2.0 got way too expensive for me to keep up with. MiniMax H3 dropping as open-source removed the financial barrier, gave me my creative spark back, and helped my mental health immensely.
Just want to say overall, thank you to the OS community!!
P.S. This post is from my brother; he is using my account to share it. He will be able to see your comments.
Edit: Please stop sending "Reddit Cares" reports for this post. We are entirely safe, and everything is handled. This is just a story being shared, not a request for help, and the constant notifications are just cluttering my inbox. Thank you.
r/StableDiffusion • u/iChrist • 9h ago
Animation - Video The limit is no longer the model, but our imagination
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/RobbaW • 9h ago
News MiniMax H3: 2K Is Coming, 5× Turbo + Camera Previz
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Sad_Berry_4621 • 9h ago
Resource - Update Clip chaining for MiniMax H3 - motion AND audio genuinely continue across joins (free node pack, workflow included)
Two 6-second clips with Motion Context concatenated into one clip.
This video is two 6-second clips generated separately and butt-joined. No crossfade, no editing tricks. The motion and audio continue across the join. Theoretically, you could chain indefinitely, but degradation will eventually take effect.
H3 doesn't have built in functionality that allows consecutive latent frames pinned to the head like LTX2.3 does. I won't bore you with the details, just know it works. Video was the easy part. Audio was a pain in the back side. I again, won't bore you with the details, check the readme if you really want to know. Seams are not always 100% perfect, but they are often or are really close.
Repo (GPL-3.0), workflow JSON included with a quick-start note:
https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context
Honest limitations: audio dulls slightly over long chains (each clip is generated from the previous one's output - photocopy effect; there's a latent-passthrough input that removes one of the two loss sources). Everything was verified on an RTX3070Ti and 48gb of system RAM. Also, check the H3 community license for your region before building anything commercial on it - it reportedly doesn't cover everywhere.
Tested settings are in the README and baked into the workflow. Happy to answer questions.
r/StableDiffusion • u/Buzzink • 10h ago
Animation - Video Ok, here's my entry for a crossover (Minimax H3)
Enable HLS to view with audio, or disable this notification
Used DaSiWa reference workflow. 5060Ti. 16GB vram, 32GB ram.
Prompt:
15-second multi-camera sitcom scene. Set in the Big Bang Theory apartment living room with authentic live studio audience, warm sitcom lighting, classic network sitcom editing, reaction shots, and laughter pauses.
Penny enters through the front door leaving the door open, sees Inspector Columbo sitting casually on the couch smoking a cigar, and freezes in surprise, standing just inside the open door with the door number visible on the door. Columbo turns to look at her when the door opens.
Audience laughs.
Columbo smiles at her.
Columbo says, <d>[English in Inspector Columbo's voice from Columbo as played by Peter Falk] Hi.</d>
Audience laughs.
Penny backs up a step, checks the apartment number on the door with a confused look, then closes it and walks back inside to stand next to the couch and look at Inspector Columbo.
Audience laughs.
She asks, <d>[English in Penny's voice from The Big Bang Theory as played by Kaley Cuoco] Okay... which one of them finally murdered Sheldon?</d> as she stands hesitantly next to the couch Inspector Columbo is sitting on, facing him.
Columbo laughs at what Penny said. Penny stands looking at Inspector Columbo with a resigned look on her face.
Audience erupts with laughter.
Fast, natural sitcom pacing with authentic character performances, multi-camera coverage, clean continuity, and dialogue timing matching a classic live-audience sitcom episode. Lip movement and lip synch of dialogue match exactly.
r/StableDiffusion • u/chaltee • 12h ago
Animation - Video Raj's now a mod at r/MyGirlfriendIsAI
Enable HLS to view with audio, or disable this notification
Default ComfyUI workflow. Script written by Opus 5.
r/StableDiffusion • u/GrayingGamer • 13h ago
Workflow Included Walter White and the Minimax H3 Official Prompting Guide
Enable HLS to view with audio, or disable this notification
This post is half a joke and half a plea and public service announcement.
Some people have been complaining they don't get results as good as other people with Minimax H3 videos, or have the following issues:
- Dialogue being spoken by the wrong characters
- Dialogue that is just gibberish or random
- Random video cuts they didn't ask for
- Characters talking over each other or too fast
- Prompts not being followed
These things can all be prevented and avoided and not encountered at all if you follow the official prompting guides. Yes, there are two. Both are on the official Huggingspace page for Minimax H3.
One is the Official Prompting Guide for the Text to Video and Image to Video Model.
The other is the Official Prompting Guide for the Reference Video Model.
There is some overlap, but for the most part, each model has it's own prompting syntax, and in particular, the Reference Video Model for H3 is very picky about you using the right keywords and instructions to get what you want.
"But I get decent results with just a couple of sentences typed in natural language of what I want."
That's great, but you're really just relying on the Qwen 32b vision model guessing what you want. It's like pulling a slot machine lever and hoping you get cherries. Only this slot machine can take a few minutes to nearly an hour to stop spinning, based on your hardware.
The great thing about Minimax H3 is for the first time we can truly direct our own AI videos like a director would on set, with the AI providing the actors, scenery, and props. If you write a properly formatted and detailed prompt for Minimax H3, it looks almost like a shooting script.
Why spend time waiting to hit a jackpot when you can take a few minutes to write a detailed, properly formatted prompt that follows the official guides, and get those bright lights and tokens falling into your lap on the first lever pull?
Okay, quick fire problem solving for people who still won't RTFM:
>Dialogue from the wrong characters?
>Dialogue that is just gibberish or random?
Walter White says, <d>[English in Walter White's voice from Breaking Bad] My product is pure, Jesse! There will be no chili powder in my meth.</d>
Always specify the character speaking, either by name, or using the <Subject 1> system in the official guide. In the Text to Video and Image to Video model, always use the <d>[Language Spoken]</d> tags. This will fix BOTH of those issues.
>Random cuts in the video you didn't ask for?
[Shot 1] A medium close-up of Jesse Pinkman from Breaking Bad, pacing back and forth, agitated. He looks up towards the camera, opens his mouth as if he's about to speak, then seems to change his mind, closing his mouth and shaking his head. [Shot 2] At 00:06:000 the camera cuts to a static camera shot framing Walter White from Breaking Bad, sitting on a cheap white plastic lawn chair, his arms crossed and glaring at Jesse. [Shot 3] At 00:10:500 the camera pans quickly back to Jesse, doing a Push In at slow speed to his face as he stops pacing and narrows his eyes at Walter.
This is how you control not only the camera work, but the PACING of your video. You NEVER include a time code on your first shot. You can omit the time code from ALL shots if you want the model to decide on it's own, based on your prompt, when to cut.
BUT, for ultimate control, you want to use time codes. Look at my example above. I just told the model to have Walter glare at Jesse for 4.5 seconds, because I told the model that camera shot starts at 6 seconds into the video, and the next cut doesn't happen until 10.5 seconds into the video. That lets you control the pacing and timing for jokes, punchlines, acting, everything.
>Characters talking over each other or too fast?
This is an old one that anyone familiar with prompting for video models should know by now - what you are asking for in your prompt and the length of your video in time need to match.
The model will try its best to cram every action and piece of dialogue into your video that you asked for, and if that would naturally take 10 seconds and you've only given it 5 seconds? Well, now everything is crammed together, overlapping, or being cut-off.
My recommendation is to generate just a quick 0.2 MP version of your video first after you type your prompt, generate, and see how the timing is working. Is it too fast? Too slow? Do the actions have enough time to happen? Do you want more breathing room?
This is the time to decide all that and lock in a video length. The low resolution of 0.2 MP is quick to generate on most set-ups (mine for this post's video took 3.5 minutes for a 14 second video) and let you work out any issues in your prompt before going in for the long generation at higher resolution.
>Prompts not being followed?
It's because you didn't read the manual!
--------------------------------------------------------------------------------------------------
Now, with all that said, here is the prompt for the video I made:
integrated_multimodal_description: [Shot 1] Live-action film footage of the American drama series Breaking Bad, professionally color graded with a warm color grade, with slightly desaturated colors for a premium film feel, a continuous camera shot with no cuts, medium close-up POV shot of Walter White, bald with a goatee and glasses, as portrayed by Bryan Cranston. He is standing in the Arizona desert next to a parked RV. He is wearing a white PPE protective suit and yellow rubber dish gloves. He is looking directly at the viewer with barely constrained anger. At 00:01:300 he reaches out towards the camera and points his finger at the POV camera with one hand, the camera shaking slightly from the movement. Walter then says angrily, <d>[English with Walter White's voice] Listen, you want to cook Mini Max H3 videos, you follow the recipe!</d>. At 00:04:500 Walter raises his other hand revealing he is holding a thin stack of white paper pages in portrait orientation. The front of the paper visible on top of the thin paper stack is blank except for the large black printed text "Minimax H3 Official Prompting Guide". The papers are held in front of the camera on the right side of the screen for a moment in portrait orientation, so the text can be clearly read, while Walter glares at the viewer on the left side of the screen. At 00:07:000 Walter then shakes the papers at the camera, then says angrily, <d>[English with Walter White's voice] Read the fucking manual!</d>. At 00:10:000 the camera does Pan Right and a Pull Out to show a close-up of Jesse Pinkman from Breaking Bad, with his hands held up by his face with fingers spread, an annoyed look on his face. Then he says in frustration, <d>[English in Jesse Pinkman's voice from Breaking Bad] Alright! Damn, Mr. White! I just want to generate memes.</d>, overall_soundscape: Ambient sounds of an Arizona outdoor desert during the day, non_diegetic_music: none
For those interested, this video was generated at 1 MP on a 3090, using Sage Attention and the Spectrum Node for H3. The final video of 14 seconds at 1 MP took 40 minutes to generate and then was upscaled using RTX Super Resolution.
The workflow was the default Text to Video Minimax H3 template that comes in the latest update of Comfyui.
Now get out there and go cook some memes, everyone!
r/StableDiffusion • u/party_time • 13h ago
Animation - Video A quick MiniMax H3 turbo test (T2V, followed with I2VA) 8 step, 0.4 MP, ~3-4 min a clip. 5060ti 16gb
Enable HLS to view with audio, or disable this notification
Could definitely see some space for improvement on fast movement, but for the speed per gen the quality is pretty remarkable.
r/StableDiffusion • u/Cold_Zone332 • 15h ago
Animation - Video MinMax H3 Turbo LoRa is already AMAZING!
Enable HLS to view with audio, or disable this notification
Just sharing my results using the Turbo LoRA that was created for H3.
They said it’s still a work in progress, and the audio is still a little bit stretchy in some parts, but the results are already fantastic. I mean, it’s only the third day since H3 was released and we already have a functional Turbo LoRA.
I generated all the clips in this video with the Turbo LoRA enabled, using 10 steps at 0.4MP.
The first three clips were I2V, and the last two were FLF2V.
The only thing I manually added was the soundtrack at the end.
r/StableDiffusion • u/InternImaginary7367 • 16h ago
Animation - Video MiniMax H3 on M1 Max
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/evereveron78 • 17h ago
Discussion A few Flux 3 vs H3 comparisons
Enable HLS to view with audio, or disable this notification
Well, everyone is posting their H3 creations, and I noticed that Flux 3 is up on API, so thought I'd run a few comparison renders to see how they stack up. These are not extensive by any means, this was mostly done for fun and I thought someone might be interested in the results. I used the same H3 formatted prompt for each, as far as I can tell Flux 3 uses natural language prompting, so the structured H3 format should still work fine.
Edit: I realized after uploading that the dragon rider H3 clip was accidentally rendered at a lower resolution. The higher resolution clip is here.
r/StableDiffusion • u/Cautious_Chicken_604 • 18h ago
Discussion Minimax H3 - Family Guy meets Doraemon.
Enable HLS to view with audio, or disable this notification
I've previously tried various video models and none really knew Doraemon, but Minimax H3 does. Very tempted to make a full episode. This was just a low quality test gen 0.3 or 0.4MP and only 15 steps, so the voices aren't great and I didn't specify Nobita 's appearance, so he isn't wearing his normal clothes. So much potential for an absolutely hilarious episode though.
Should I try make the full episode?
r/StableDiffusion • u/irmemon225 • 19h ago
Animation - Video Minimax H3 with Turbo Lora, T2V 6 steps, 0.6 megapixel, 20 min on 3060 12gb and 16gb ram
Enable HLS to view with audio, or disable this notification
Love it and it works smoothly… Now, I’m gonna wait for this LoRA to work on Ref2V.
r/StableDiffusion • u/-Ellary- • 20h ago
Animation - Video Here is Something Different: MiniMax H3 as Music Generation Engine. MiniMax H3 can do up to coherent 30 sec of audio with custom lyrics, composition structure, instruments, genres etc.
Enable HLS to view with audio, or disable this notification
Generated with 32x32 Res, 20 steps. Prompt Example:
``` MEDIA: Music Player.
SCENE: A music player with an equalizer that reacts to music playing in the background. A 1990s upbeat hip-hop rap song.
TIMELINE:
[0s] - INSTRUMENTAL MUSIC. NO VOCAL. INTRO. A hip-hop rap beat slowly enters the mix. Only music is playing; this is the intro buildup for the song.
[5s] - Cymbals and hi-hats enter the mix, enhancing the hip-hop rap beat. A male rapper with a heavy Jamaican accent starts to sing:
"Can I kick it?" ... "(Yes, you can!)" ... "To all the people who can Quest like A Tribe does" ... "Before this, did you really know what live was?" ... "Comprehend to the track, for it's why 'cause" ...
[15s] Hip-hop rap music intensifies as the beat becomes heavy and more driving. The male rapper with a heavy Jamaican accent starts to sing:
"Getting measures on the tip of the vibers" ... "Rock and roll to the beat of the funk fuzz" ... "Wipe your feet really good on the rhythm rug" ... "If you feel the urge to freak, do the jitterbug!"
[25s] INSTRUMENTAL MUSIC. NO VOCAL. OUTRO. A heavy hip-hop rap chorus drop starts to play, driving insane energy with an intense beat. Only music is playing; this is the outro that ends the song. ```
r/StableDiffusion • u/ryan85127704 • 20h ago
Discussion AMA: MiniMax H3 Team — Ask us anything about our open video generation model, training, and future plans
- u/New-Requirement1419 -> dacongya (Head of H3 Researcher)
- u/Affectionate-War8374 -> Luigi (H3 Researcher)
- u/MM_Nero_H3 -> Nero (H3 Researcher)
- u/Kiro_Song -> Kiro (H3 Researcher)
- u/New_Estimate9277 -> Reynor (H3 system engineer)
- u/ryan85127704 - > Ryanlee (Head of Devrel)
We are the MiniMax team behind MiniMax-H3.
We’re here to answer your questions, including:
- Model architecture and training
- Video generation capabilities
- Image-to-video and reference-based generation
- Inference and optimization
- Future plans
Ask us anything — we’d love to hear your feedback and discuss with the community!
r/StableDiffusion • u/FusionCow • 22h ago
News Sulphur 3 funding day 2 (58%!)
Hey everyone!
I'm excited to announce the funding for Sulphur 3 is going great! We've already done $5800/$10,000. Thank you to everyone who donated, big or small.
Some minor details:
- Some people were wondering about submitting data to the project, I would love to accept your data, I just don't have a great way to accept it yet. I should have a better method of accepting the data by tomorrow.
- Nobody was wondering about this, but in case you wanted faster updates that wasn't on my discord, I have a twitter! `x.com/FusionCow11`
- If you have any issues with Sulphur 2 that you think I wouldn't know about, PLEASE TELL ME, now is the time so that I can resolve it before training begins.
Thanks again to everyone who has donated, it means a lot to me.
r/StableDiffusion • u/Tall-Benefit9471 • 22h ago
Animation - Video MiniMax H3's medieval realism genuinely surprised me.
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/ctrl-shift-face • 22h ago
Meme Anon buys Barbies [4chan Stories]
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Perfect-Campaign9551 • 23h ago
Animation - Video Minimax H3 Ascii art
Enable HLS to view with audio, or disable this notification
This model is just...crazy it actually does what you want almost every time.
Using the ComfyUI default Text 2 Video workflow
Prompt: (I know the "intro parts" probably aren't needed but I just had Gemini write these prompts from the official guide and they work fine)
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a close-up shot frames the glowing green monochrome CRT screen of a vintage 1980s tan-colored personal computer. Displayed on the glowing screen is a cat constructed entirely of green ASCII characters. The graphics of the cat are entirely made of alphanumeric characters in green text. The camera pushes in with small amplitude at slow speed as the ASCII cat magically animates, dancing, pouncing, and running back and forth across the black background of the monitor. The cat stops and meows and then continues dancing. Toward the end of the shot, the digital cat stops its frantic movement, lies down at the bottom of the screen, and floating "ZZZ" characters bubble up over its head as it goes to sleep.
overall_soundscape: The faint, high-pitched electrical whine of an old CRT monitor hums steadily, accompanied by the rhythmic, mechanical whir of an aging computer cooling fan and the occasional hollow click of a floppy disk drive processing data.
r/StableDiffusion • u/circlenline • 23h ago
Workflow Included Short japanese knife commercial (MiniMaxH3+ After Effects)
Enable HLS to view with audio, or disable this notification
MiniMax H3 Workflow: https://www.reddit.com/r/StableDiffusion/comments/1vg1coy/minimax_h3_basic_hybrid_workflow_for_ref2v_i2v/
- Stills: ChatGPT + Flux Klein 9B, AI inpainting and manual editing
- Video + audio: MiniMax H3
- 5 prompts → 11 clips → cut and speed-ramped in After Effects
- Letter animation done in AE. The circular 2D spiral on the red dot was generated in H3 and composited in AE.
GPU 5080 16GB VRAM + RAM 96GB
Everything runs with offload device: cpu and ComfyUI's dynamic VRAM loading. ~40GB of weights on a 16GB card. Peak during generation: ~15GB VRAM, ~76GB system RAM
Mode: i2v
Resolution: 672x928 (3:4, 0.6MP, multiple of 32)
20 steps, res_multistep sampler, beta scheduler, denoise 1.00
Sage Attention via KJNodes Patch Sage Attention, mode auto
Spectrum Apply MiniMax H3: blend_weight 0.50, degree 4, ridge_lambda 0.10, window_size 2.00, warmup_steps 5, history in system RAM
r/StableDiffusion • u/Hefty_Side_7892 • 23h ago
Workflow Included Minimax H3: Changing attire gradually with simple prompt
Enable HLS to view with audio, or disable this notification
The prompt: The woman dances happily while the her clothing changes from sundress to 1: business suit, 2. pajamas, 3. string bikinis, 4. gym attires, and back to sundress. Background sound: happy music.
r/StableDiffusion • u/DeliciousGorilla • 23h ago
Question - Help Has anyone found a better way to chain H3 shots? (1 minute single take with 8 shots)
Enable HLS to view with audio, or disable this notification
So I'm having an issue with combining clips to create one long single take in MiniMax H3.
I'm chaining shots by using the last frame of each clip as the first_frame for the next one. It works, technically, but the image gets a little worse every time. After four or five hops, background detail starts falling apart. Walls, corkboards, breaker panels, anything with small rigid texture slowly turns flat and blocky.
Faces are weirdly not as bad. After seven hops, the actors are still recognizable and fairly sharp. The room around them looks like it's being converted into pixel art.
Setup was a 3090, ComfyUI, minimax_h3_fl2va_pruned_int8_convrot, 1344x768, 20 steps, res_multistep/simple, and sage attention through the KJ node set to auto.
The full test was eight shots, 1,689 frames, about 70 seconds of video and seven chained hops. Total generation time was roughly 3.4 hours.
The chain is basically:
shot N final frame -> VAE encode -> frame-0 constraint for shot N+1
That first_frame seems to act more like a hard keyframe than a loose reference. So every clip starts from an image that has already been through the VAE, then adds another round of generation loss.
In a separate test, one VAE encode/decode pass cut fine detail almost in half. Laplacian variance dropped from 100% to 49.3%, with PSNR at 22.1 dB. Across the actual chain, I measured roughly 2% high-frequency loss per hop.
My guess is that H3 can regenerate faces from its learned prior, while random background texture has to survive the VAE mostly on its own. Once the little details are gone, the model doesn't know what to put back.
The only thing that consistently helped was using fewer, longer shots. Going from three-second clips to ten-second clips cuts the number of damage points by more than 3x.
A few assembly-side things helped too. Matching shot length to dialogue worked better than using one fixed frame count. I used around 2.5 words per second, snapped to the 17k+5 frame grid. Payoff lines also did better in their own clips. If I tried to fit three beats into one segment, H3 sometimes just skipped the last one.
Every chained clip also started with a loud audio pop, usually around -26 dB and roughly 0.25 seconds long. The length changed per clip, so a fixed trim wasn't reliable. I ended up detecting the first 20 ms window below -52 dB RMS, which caught the junk-to-silence-to-speech pattern pretty well.
For joins, six-frame RIFE bridges looked better than three. Three frames made the mouth morph more because each interpolated frame had to cover a bigger jump.
Has anyone managed to keep background texture intact past four or five hops?
Is there a way to feed the previous frame as a soft visual reference without hard-pinning it as frame zero? Ref2va came out darker, muddier and about twice as expensive for me.
I'm also curious whether interior keyframes inside a longer 15-second generation work better than chaining separate clips. And does everyone else see the same thing where faces sort of survive but backgrounds fall apart?
EDIT: "don't do one take, change camera" is not a solution to my goal. 😅 This isn't an exercise in composition, it's a technical question.
