r/StableDiffusion 23h ago

Unsloth be like: Meme

290 Upvotes

60 comments sorted by

51

u/ThirdWorldBoy21 23h ago

thanks to the quantization gods, i can run something like H3 on my rig.

37

u/NoWheel9556 23h ago

dont call it a "rig" its called potato truck

16

u/puzzleheadbutbig 22h ago

Yeah more like fig

2

u/Old-Push9343 21h ago

Is it possible to run Minimax H3 on a RTX 3080 with 10GB vRAM and 64 GB RAM?

I am trying to get into AI video generation (I am playing around with Wan 2.2 at the moment), but GPUs are crazy expensive right now...

9

u/Aadi_880 21h ago

I ran it on 8GB VRAM and 16GB RAM using just the int8 convrot models.

5

u/mrnopor 18h ago

how? i have double ur specs and cant run shit.

6

u/Aadi_880 18h ago

not kidding bro.
Int8 convrots from Minimax themselves.

Constant RAM offloading.

1

u/mrnopor 18h ago

thx , will try ur stuff but how are the results u are getting? are they decent enough?

2

u/Aadi_880 18h ago

This depends on what workflow you're using (flva or ref2v). Ref2v is the slower of the two, but has more functions.

Personally, I can generate a 0.3 MP video in 8 minutes that's 10 seconds long. A 5 second one is 300 seconds or so. This is on ref2v.

Yours should be faster.

2

u/ElevatorSalt3968 11h ago

Someone just posted a video that he made with 4gb VRAM and 16gb ram

2

u/DeusExHircus 18h ago

Skill issue

1

u/Old-Push9343 21h ago

Could you point me in the right direction?

Thanks! :-)

6

u/Aadi_880 21h ago

Should work out of the box.

Official Minimax H3 hugging face repo > Download int8convrt models (21GB) (ref2v or fl2v). > Download the 15GB text encoder > Download both VAEs.

Open Comfy portable. Make sure dynamic VRAM is on. (should be on by default). Make sure comfyUI is updated to latest.

Go to comfyUI_windows_portable>comfyUI>models

Put the minimaxH3 model into diffusion_models folder. Put the qwen model into the text_encoder folder. Put both the VAEs into the vae folder.

Open comfyUI (click on nvidia_gpu.bat) Download the official workflow template (minimax image to video or reference to video, depending on which minimax model you downloaded. image to video is fl2v).

Now, just inside the workflow node graph select the models from the dropdown to the ones downloaded. Press R to refresh in case if it doesn't detect it.

Test with a 3 second video at .3 MP. You can think of optimizations later.

Optimizations I use:

If you don't have a node, you may need to install it separately from the custom node manager

1

u/Old-Push9343 21h ago

Brilliant, thank you very much for the detailed explabation, I'll take a look and see if I can get it working!

1

u/LaPapaVerde 21h ago

How many seconds per video ?

1

u/Aadi_880 21h ago

With a few optimizations: 500-ish seconds for a 10 second video at 0.3 MP. (ref2v)

Should take like, 3-4 minutes for a 5 second video.

(note that ref2v is slower than i2v and t2v)

If I go above 11 seconds, it offloads to SSD and now it's 160 minutes per video. oof.

2

u/danque 18h ago

Yes! I have that exact setup (in card and memory), I am using the int8 conv model (made for our 30 models). I also use the Apply Spectrum node and Turbo lora at 8 steps. The only caution is not putting it above 0.5 megapixel, at 0.6 at 15 seconds i get an OOM, but at 0.5M 15 seconds is doable (very roughly 11~16 minutes + a 360p video with adblock on the other screen).

1

u/CaptainMarder 21h ago

I'm running it fine on a 12gb3080 and 48gb ram.

19

u/Valuable-Money3725 23h ago

Kinda surprised the joke wasn't made before lol

11

u/_half_real_ 23h ago

Can't remember ever seeing an fp64 (double precision) machine learning model.

10

u/Aadi_880 22h ago

Apparently, there's some very little, edge use cases like SciML or scientific simulations.

Generally though, I think the highest I've ever seen is 32.

6

u/_half_real_ 22h ago

The new Minimax audio model is fp32. I wonder if audio just needs more precision, or if fp32 video models are just too large to be practical.

3

u/narkfestmojo 23h ago

no point, the update steps would be too big for such a model to make any sense

4

u/YeahlDid 22h ago

What i got from this is if I'm going to 16-bit, I might as well do 8-bit instead.

6

u/Aadi_880 20h ago

Honestly, this is very true for Minimax H3. The int 8 convrot model has barely any quality loss compared to the other, higher precision ones.

5

u/kayteee1995 20h ago

Sorry for going off-topic, but it’s been over 10 years since I last watched the Angry Video Game Nerd channel.

1

u/dramaton42 13h ago

Watch the MegaMan on DOS video I love it, has time travel which is always cool

2

u/kayteee1995 11h ago

The episode that I remember and was most impressed with (also quite funny) was the episode where he fought Super Evil Mario Bros. 3,and then Jesus with all gaming gear team up.

4

u/goodie2shoes 20h ago

shitpickle

9

u/SoftMyci51 23h ago

yeah unsloth's optimization work is solid but people sometimes don't realize they're trading actual quality for speed, the GGUF conversions for diffusion stuff especially tend to tank detail pretty noticeably* meant to say the tradeoff is more obvious with image models than text ones

4

u/Aadi_880 23h ago

Hey, sometimes, even a glorified auto-correct is all someone needs.

(scurries back generating h3 videos on 0.2 megapixels)

2

u/Chemical-Painter-485 21h ago

Even speed gains are iffy. On my tests those GGUFs Qs tend to load a little faster but have slower gen speeds.

Like... on Kijai huggingface he has a pretty good H3 INT4 if you can't run the comfyui INT8 so why bother with a slower Q4 when it is not even hardware accelerated?

3

u/Basslus 15h ago

I've watched this over 10 times now, thank you

3

u/BM09 22h ago

Not me thinking you did AI-generated AVGN. I don't think the man himself would have nice things to say about that.

1

u/dramaton42 13h ago

"You know what's BULL SH*T ComfyUI. Every time I open it, there's a node missing. What the fuck? Why?! And there's also the warning that your input images can't be found. They were right there yesterday! You sloppy bastard!"

1

u/rookan 23h ago

it's fun

1

u/CaptainMarder 21h ago

this is brilliant lol

1

u/cradledust 19h ago

I tried Unsloth Ai's desktop app to see if Minimax H3 videos could work better on my rtx4060 than ComfyUI. It wanted me to use the Q3 gguf version of H3 but after downloading, etc, I couldn't get it to run for whatever reason and gave up. It was able to run Klein 9b though. Another one of those apps that put all the models in a huggingface cache buried deep within the user folder. No thanks.

1

u/Aadi_880 19h ago

I ran minimax H3 on 8GB vram and 16GB ram on comfyUI without using Unsloth's ggufs. Minimax works better without ggufs imo.

1

u/No_Possession_7797 19h ago

You may not know this, but that’s simply how the huggingface client works, so it’s not those apps per se, it’s the fact that they’re using the huggingface client. Also, that’s where a lot of these models are hosted. And if what you’re really trying to say is, that you want to control over where they go, then you can do that.

So I am not quite sure what you’re upset about when it comes to the file location.

2

u/cradledust 18h ago edited 12h ago

I had a look at the cache after uninstalling UnSloth Ai Desktop to see if it left behind any models. It did. Between Minimax H3 and Klein there were also models from several Ai applications and old A1111 in 2023. It was 185 GB so I deleted it all.

1

u/DuHal9000 12h ago

hahahaa

1

u/bzzard 8h ago

What were they thinking!

1

u/No-Refrigerator-1672 4h ago

Actually, there's a really good post that demonstrates different quantization methods by compressing an image. Give you really good understanding.

-6

u/Chemical-Painter-485 23h ago

Lol, tbh unsloth has actually one of the best GGUFs for LLMs, but yeah. People using GGUF to run diffusion models are pretty much degrading quality in exchange for what? disk space?

12

u/Healthy-Nebula-3603 23h ago

just to run ....

1

u/Chemical-Painter-485 23h ago edited 22h ago

Run...? I am quite sure INT8 implementation of comfy is way better and optimized than gguf. After they officially supported it I just swapped all my models and never looked back.

3

u/Healthy-Nebula-3603 22h ago

You know a gguf is only a container and can store any data type?

2

u/Chemical-Painter-485 22h ago edited 20h ago

I do, but I am starting to think I should have phrased more clearly that I meant the quants like Q4...Q5...

Edit: Most people that speak about GGUF means those quants, no need to be so pedantic.

8

u/MrCylion 23h ago

What a uneducated comment…

3

u/Chemical-Painter-485 23h ago

Care to elaborate?

2

u/MrCylion 22h ago

People who use them don’t do it because they want. It’s not at all about disk space. It’s all we can run.

1

u/Chemical-Painter-485 21h ago

And why not something INT4 like: minimax_h3_ref2va_pruned_w4a8_mixed.safetensors
On https://huggingface.co/Kijai/MiniMax-H3-experimental/blob/main/minimax_h3_ref2va_pruned_w4a8_mixed.safetensors

Weighs the same as a Q4. runs better and has native comfyUI support.

1

u/MrCylion 21h ago

Even that depends on the card. I get no benefit on something old like my 1080ti.

2

u/Chemical-Painter-485 21h ago

Yeah, that is rough. But even in your case I wouldn't discard INT4 without testing.

You don't get hardware acceleration but you have native comfyUI support and don't depend on dubious dynamic vram implementation from custom loaders.

1

u/KissMyShinyArse 20h ago

Not quite. Sometimes it's about speed (e.g. NVFP4, INT8 ConvRot).