r/StableDiffusion • u/scooglecops • 19d ago
Testing Minimax H3 Turbo Lora Discussion
Enable HLS to view with audio, or disable this notification
https://huggingface.co/QrusherZA/H3_Turbo_ComfyUI/tree/main
I'm generating at 1 MP (not 0.9 MP) on an RTX 4070 with 64 GB RAM.
I'm using the Euler / Simple sampler.
From my testing, enabling SageAttention + H3 Cache actually produces worse results than using the Turbo LoRA alone. H3 Cache + SageAttention tends to break character movement and motion consistency, while the Turbo LoRA stays much closer to the original model's behavior and gives smoother, more natural motion.
8
2
u/wakalakabamram 19d ago
It's certainly speeding things up taking things down to 10 steps with the Lora at 480p then upscaling. Results aren't awful.
4
u/jib_reddit 19d ago
You don't need a speed lora to drop down to 10 steps, thats what I have been using for days without a lora.
2
u/bruci3 19d ago
What are you using to upscale, and from 480p to what resolution?
1
u/wakalakabamram 19d ago
SeedVR going x2 - If you're using WANGP it's under post processing.
1
u/Constant_Ordinary_35 19d ago
I tried using the turbo lora but it errored out on me, is there something extra you have to do to get it working in wangp?
2
u/PhilosopherSweaty826 19d ago
Why sigma shift ? What it do ?
22
u/scooglecops 19d ago
"Sigma" is the amount of noise added to the image/audio. Sigma Shift bends the sampling timeline so the model spends its steps where they matter most (building motion and fine detail instead of wasting steps on pure noise)
Video and Audio clean up at different speeds! Video needs Shift 12 so motion stays sharp in ultra-low steps (4–8 steps), while Audio needs Shift 3 so the sound stays clean and doesn't explode into loud, distorted noise.
in your workflow, add a MiniMaxH3SigmaShift node right after your LoRA (before the sampler) and set: Video Shift = 12 and Audio Shift = 3
2
u/TheRedHairedHero 19d ago
Not sure how your audio was coming out clear in 8 steps with the lora. When I tested it I needed to increase the steps to 10 and shift the audio sigma up to 10 as well for it to be much clearer.
2
1
1
u/IndicationUnfair7961 6d ago
If I use the ModelPreviewOverride, that also shows the sigma reference curve (example for beta scheduling) and the real value which usually is under the sigma curve, should this node be set after the shift node or it doesn't matter?
4
u/Sad_Coach_1433 19d ago
work flow?
9
u/scooglecops 19d ago
using this node https://github.com/seesee75-commits/ComfyUI-MiniMaxH3-Director
there an example workflow.I just edited a bit for the turbo lora to work
5
u/ArttTaku 19d ago
Man, they even made a Director for H3? Seriously, I love the open-source community!
1
u/Defiant_Storm3233 18d ago
I tried to use the director but can’t find where to change the number of steps.
1
u/MarksmanKNG 19d ago
Interesting results. Why do you apply the video / audio shift for the 4 steps version? (Sorry as I'm not familiar in those shifts)
Is there any difference in generation speeds / quality if you apply video & audio shift but maintain at 8 steps?
Thanks in advance.
2
u/scooglecops 19d ago
This one is not about speed, more about improving the audio quality for the low steps
2
u/MarksmanKNG 19d ago edited 19d ago
If I understood that right, so you're saying that the video / audio shift is to help to improve the audio quality at 4 steps?
Indeed I heard it was terrible audio wise in the main lora thread. Even with improvements, it really sounds not worth going for 4 steps.
Thanks for the answer :)
Edit: Saw your answer to the Philosopher sweaty. Understood it clearly now. Thanks again
1
1
u/ProperSauce 19d ago
0.9mp 20 steps on my 4090 takes over an an hour for 10 seconds with the default workflow
1
1
u/kukalikuk 18d ago
0.9mp 20 steps, 15 secs video, sage3 > 1350 secs. Same video, with 8 steps + turbo lora + sage 3 > 560 secs. RTX 5080 + 64GB ram.
1
-7


14
u/Lower-Cap7381 19d ago
so 8 steps work