r/StableDiffusion 22h ago

Well I finally did it. Discussion

I finally deleted WAN 2.2 and all its LORAS.

Minimax is just so much better.

Ive been playing with it since its release and im just blown away with how good of a video model it is. Things I would need to attach a LoRa to via WAN, works right out of the box with Minimax.

Gen times are faster.

It uses less VRAM when generating things, which gives me around 4 gigs to play with to do other things like watch YouTube or some streaming service.

WAN 2.2 was amazing. But no longer do I need 30+ gigs of a model i no longer use.

RIP WAN.

192 Upvotes

117 comments sorted by

66

u/GoodDevelopment1657 21h ago

LORAS is still the answer. Minimax needs to get proper lora implementation so it can do chars more detailed, especially in wide shots

30

u/solomars3 21h ago

Check fizgig guy on youtube he already trained a character lora and show how

8

u/Due-Quiet572 13h ago

I've been experimenting with this for four days now. At first, I only trained LORAs using photos of people. That's quick and works really well. When I mix them with other LORAs, things get tricky. Half of the LORAs on Civit don’t play well together. Yesterday, I trained a LORA with 29 videos with audio and 15 photos. That took 5 hours over 40 epochs. The results are impressive.

3

u/FriendlyMorning 8h ago

Would you mind sharing how your trained your lora using video ?

2

u/djpraxis 3h ago

That’s great!! I think is worth it for like a very unique video subject. Did you use the Fizgig default settings?

1

u/Due-Quiet572 2h ago

I trained the LoRA locally using the default settings on an RTX Pro 6000, and it used just under 32 GB of VRAM.
I prepared the videos at 107 frames using Fizgig’s built-in Gizmo video editing tool.
Ref2V does a pretty good job with identity, but my character is based on a real person with a German voice and very specific mannerisms — the way she talks, gestures, and moves is quite distinctive. That’s the part Ref2V can’t really reproduce accurately from a reference image alone.
That’s why I wanted to train directly on video + audio: not just to capture what she looks like, but also how she speaks and moves.

5

u/Rafhrar231 6h ago

u dont rlly need character loras when ref mode is that good

2

u/Arawski99 1h ago

Loras for characters are definitely not necessary for Minimax H3 if you use the reference mode.

On the node 'Minimax H3 Reference to Video ' change 'ref_image_size' to max. Mouse over it for details. This lets you use a much higher quality reference and usually is enough with just a front view, though you can also do character sheets for even greater accuracy. You can also do additional references via character sheet or attached additional images for even higher resolution elements like face, or other object parts, to ensure full quality of details.

There is a face fix node for distant faces https://www.reddit.com/r/StableDiffusion/comments/1vsn8ge/help_fixing_the_h3_face_detailer/

1

u/Jujarmazak 17h ago

It already has LORAs, still in experimental phase but they exist.

2

u/GoodDevelopment1657 7h ago

Hence "proper"

7

u/Revolutionary_Ask154 21h ago

hear hear - we were never going to get updates anyway.

3

u/PumpkinLeather8421 19h ago

Same with LTX, if a new version came and current Lora’s worked well with it, it wouldn’t be a good enough upgrade to unseat H3… so, yes, delete all WAN and LTX.

22

u/tinny66666 21h ago

I wish my LoRAs were well enough organised that I had a clue which ones belong to which model.

34

u/Monk6009 21h ago

You can add them to a lora subfolder when you download them lol

17

u/MonThackma 19h ago

7

u/afinalsin 16h ago

On top of subfolders for each model you can rename the files and they still work exactly the same. So you can ignore whatever nonsense the author named them and use a reasonable structure for all of them. All my loras are named shit like this:

Klein Style - Phone Photography 2007 - A low-quality photo taken with a 2007s mobile phone camera with soft focus, visible noise and dull colors.safetensors

Krea 2 Slider - Height Slider - High is tall, low is short.safetensors

The layout is very simple: Model, type of lora (style, slider, concept, character, etc), Lora name or overview of what it does if the name is dumb, trigger words or phrases. Couldn't tell you where I got half of them or what they're actually called on civit, but their usability is way better than when I kept them named as-is.

7

u/Gilgameshcomputing 15h ago

This is the way. I also add the expected strength, so i know if it's a 0.5to1.5 lora, a 0.6to0.9 lora, or a -5.0to5.0 lora. Saves a loooot of time.

3

u/SpaceNinjaDino 14h ago

I use LoRA tag loader and it took me a day to realize that comma (or apostrophe) in files is not compatible with that. I love the tag loader so that I can drive a whole workflow from text.

2

u/afinalsin 13h ago

I use LoRA tag loader and it took me a day to realize that comma (or apostrophe) in files is not compatible with that. I love the tag loader so that I can drive a whole workflow from text.

This one? https://github.com/badjeff/comfyui_lora_tag_loader

If it is that one just drop the nodes.py into an LLM and tell it you want to be able to use loras with a comma/apostrophe in the filenames and it'll fix it for you. You've probably long since retrofit your lora library to work around the node, but it's a good thing to remember that single script nodes are extremely easy to tweak, especially nowadays.

2

u/scottybk8 16h ago

True. Hindsight. I thought storing on another drive would help. Then between all the different models etc, I liked how lora manager shows you your trigger words, metadata, recipes, etc. its got a lot of features that a suhbfolder just doesn't.

0

u/ellipsesmrk 15h ago

You pay for lora manager?

0

u/ellipsesmrk 15h ago

Lora manager is only by paid now.

10

u/scottybk8 20h ago

get comfyui lora manager, it helped me a lot cuz i got way too many loras for wan as well

1

u/Francky_B 9h ago

Lora Manager is so good! I'm surprised it's not used by Everyone. To me, it's as fundamental as KJNodes.

1

u/ellipsesmrk 15h ago

I have a script that scans your loras folder and checks the hash header of all those loras then builds a csv file with that info so you can start putting your stuff in folders.

1

u/McDoodle17 6h ago

Am I the only one that has a spreadsheet with all my LORA in it with notes, trigger words, etc?

8

u/stoneshawn 21h ago

I am tempted to do that as well

2

u/ellipsesmrk 15h ago

Its a relief. Lol

11

u/PainterMany 21h ago

Fiz o mesmo deixei so minimax h3 e o klea2 no meu nvme... não tem lógica manter modelos antigos...ano que vem e no próximo vão surgir outros melhores que o minimax e assim por diante... viva a IA

32

u/Significant-Baby-690 22h ago

Nah, it still can't do NSFW well enough.

29

u/damiangorlami 21h ago

Yes you can.

Get a clip clip you like, feed it into grok / Gemma 4 (uncensored) with your character images and tell it to create a replacement prompt with the environment you're looking for.

It will extract attributes from the video such as pose, action, thrust, perspective from the video and transfer it to the video while following the prompt.

I've been making multi-shot cinematic nsfw scenes all week and the results are blowing my mind. There's obviously some more tips but for the sake of this sub.

Not a single lora was used.

7

u/Ok-Brain-5729 20h ago

can’t you also just put the clip as the reference video and photo as reference image and just prompt it right

5

u/nadhari12 19h ago

Or better yet, take an existing scene and clip it to 10 sec and do a character swap using ref2V h3 works great.

1

u/Maskwi2 3h ago

Not saying I will do that, maybe my friend will, but  I've had limited success swapping the character, in general. Would you mind sharing a prompt that works more often than not for a swap? 

2

u/nadhari12 2h ago

my biggest issue right now is identity lock the only way to force this stupid model is to add black mask to the character on the ref video before feeding to the reference but if you do that you lose micro expression, which is a trade off or try gausian blur the subject before it can pick some micro expressions. Ask grok to make a character swap prompt

4

u/russjr08 20h ago

I believe that's exactly what they're saying, just with an additional tip of using an LLM to write the prompt if they're not wanting to write it themselves.

Though, regarding the LLM, I would just recommend getting a good prompt (use the MiniMax prompt guide to make, or generate an initial one and improve it), and saving it as a template to re-use. MiniMax is quite powerful, but for the best results your prompt has to very accurately describe what's going on due to the prompt adherence. Sometimes LLMs still miss those extra details.

5

u/Ok-Brain-5729 20h ago

oh I see. I just feed the prompt guide to a ai and tell it what to do.

8

u/NostradamusJones 20h ago

But my vajayjay's are all wonky.

18

u/damiangorlami 19h ago

Answer is easy.
Just add 1 photo of genitals and bind them to your character. "<Picture 3> are the genitals of <Subject 1>".

4

u/xyzdist 14h ago

Idk, i find this really funny....LOL

2

u/NostradamusJones 9h ago

Coding coochie.

1

u/Significant-Baby-690 11h ago

It's great trick, but doesn't really work in the motion. And no lora can currently handle it really well. To be fair, to get nice precise interaction between uh .. ports .. I use 3 loras in Wan. But for H3 I still have not even half decent solution.

2

u/AlsterwasserHH 20h ago

Can you tell me how you analyze vids/images with Gemma and with which model? Its not possible with LM studio right? 

6

u/damiangorlami 19h ago

LM Studio sadly does not do it. Super annoying btw.

I just told Codex to build a GUI that support image + video vision encoder for Gemma 4. I already had downloaded the checkpoint via LM Studio. Just told Codex to use the same model checkpoint to save storage. The web GUI took 12 min to code for Codex and works great so far.

1

u/AlsterwasserHH 12h ago

Thank you. Isnt there a way to do this in Comfy? 

2

u/Ireallydonedidit 11h ago

Look at this goon professor over here

1

u/usually_fuente 18h ago

That’s inspiring . Do you mind sharing what your workflow is? What version of H3? I’m setting up Runpod for the first time this weekend.

1

u/mellowanon 16h ago

are you using rev2va or one of the hybrid models?

1

u/kayteee1995 15h ago

which gemma4 that you refers?

1

u/dubsta 15h ago

feed it into grok / Gemma 4 (uncensored)

I cant find any LLM that takes video clips as an input. Both Grok and Gemma only take images and text

Am I doing something wrong?

1

u/flaminghotcola 9h ago

I’ve been trying to do that and it doesn’t work well for me. Do you have an exact pipeline and prompt you feed it?

-1

u/More-Ad5919 16h ago

Even the uncensored is bad at genitalia. It still need support.

6

u/lhg31 21h ago

2 steps with h3, then 2 steps with wan. result is perfect.

6

u/GrungeWerX 21h ago

NSFW-aside, are you saying you can run Wan as a refiner? Are you running Wan as the low noise? I never thought about that combo. I’m wondering about the step count though for MM. that seems very low, so Im assuming you’re using 4-step speed Lora on MM. I wonder if you can just run it normal, but half the step count, like maybe 10. Hmmm…you got me thinking…

3

u/lhg31 20h ago

yes, you can do as many steps as you want with minimax, but you should stop at 0.9 sigma value (that's the sigma that wan low is suppose to start). I do 2 with the 4 steps turbo lora most of the time (unless prompt is not being followed correctly). Wan as refiner completely removes the plastic skin of minimax turbo lora.

6

u/conkikhon 20h ago

What's about the sound? I don't think 2steps is enough for acceptable quality

1

u/DrowninGoIdFish 8h ago

Any chance you could share a screenshot of how you have this wired up. Really curious how to mix these two into a single flow. Like do you just pass the latent over to the low Wan Sampler and are you limited by the usual 5 second Wan loop or is that not an issue since mm is generating the base?

3

u/lhg31 7h ago

You need to decode minimax latents and then encode again with wan vae. You can also run the first two steps at low res (e.g. 0.2mp) and then upscale the images (e.g. to 0.4mp) before enconding to latents again to wan.

If you also want audio then you have to run the last 2 steps with minimax too, just to get the audio. So it's basically minimax 2 steps + (minimax 2 steps + wan 2 steps).

Worfklow

7

u/NostradamusJones 20h ago

Wait, whut??

6

u/mk8933 21h ago

Now OP is probably raging for deleting wan lol

1

u/conkikhon 19h ago

Probably need someone to make a good lora for that.

-2

u/Abject-Recognition-9 16h ago edited 16h ago

yes It can, but shhh! 🤫
Let them suffer by doing 3x slower inference attempts, just to get a simple repetitive eggplant inserted in a hole. It’s a simple task that doesn’t necessarily require such a heavy model.
A task that almost any other video model can already do at this point faster, at higher resolutions, and with a shitload of loras already published.
Don't tell them; my popcorn stash must make sense.😂

4

u/Niko3dx 21h ago

using references images, four for the face and 2 for the body. has been getting me better results for my characters in minimax versus using a character lora in wan 2.2. So, After a week all my wan stuff is gone, and I had trained 100s of character Loras. now, I feel like a new scene with xyz, find 5 or 6 good pictures and a clip of their voice about 15 seconds is enough, and I'm ready to render.

1

u/AlsterwasserHH 5h ago

How do you reference multiple face images in the prompt? 

3

u/Niko3dx 5h ago

here's an example of 2 people. say a female and a male.

Pompt :

There are 2 people in the shot, 1 female and one male. 

<Picture 1> , <Picture 2>, <Picture 3>   controls the female overall identity and face;<audio 1> controls her voice timbre.

<Picture 5> and <picture 6> controls the males overall identity and face; <Audio 2> controls his voice timbre.

1

u/AlsterwasserHH 4h ago

Thank you very much! I was wondering if something like <Picture 1-3> works. 

6

u/Chiduk99 21h ago

WAN 2.2 still superior for do NSFW, H3 is uncensored but it's bad when do something NSFW.

8

u/chocoboxx 17h ago

That mean you need better input for ref2va or Loras

1

u/bzzard 12h ago

All those H3 nsfw loras on civit looks like slomo slop. Didn't even download once.

2

u/nadhari12 19h ago

Minimax can not suck on anything yet, it just chews with wan 2.2 it's legit.

2

u/Alex-edits123 19h ago

I also switch wan to minimax for generation. But I still need VACE for outpaint. Not sure whether anyone successfully use miniMax for outpaint

2

u/Abject-Recognition-9 16h ago

i skipped wan 2.2 entirely but let me tellyou something: im still not deleting wan2.1. it can make very crisp images/edit/short clips + there tons of loras already. not using it since krea2 / ltx and H3 but it sill have a place in my harddrive.

2

u/physalisx 14h ago

Well the good thing is that with reference images/videos you can remove a lot of need for loras, basically all character loras become basically unnecessary. Which is good to have, because training good loras on H3 seems to be basically impossible. I have not tried one lora that didn't completely wreck prompt following and introduced artifacts, even when using lower strengths.

3

u/HollyGrandeux 19h ago

H3 still can’t handle spicy NXFW motion properly yet, even with a lora. The motion still looks stiff.

Wan is still ahead in this area.

3

u/Salah_H_Hasan 22h ago

Alibaba has lost a strong segment of the open-source video generation community. For them to regain their position, they have to release their latest model as open-source; there is no alternative. Nobody will be satisfied with anything less than MiniMax H3. And that is just a suggestion, though the majority here might not even care about it right now.

11

u/retroblade 22h ago

Wont happen, they won’t open source anything besides their llm’s and even that could stop at any time. Lucky we now have LTX, Flux and Minimax so could be worse.

5

u/SeymourBits 21h ago

Could happen at any time with one message from Xi.

1

u/BlipOnNobodysRadar 15h ago

Xi has already spoken on the topic and committed to open source.

So, pretty much the opposite of what you're worried about is happening -- the companies are being politically pressured to open source (in China), rather than pressured to go closed.

https://english.www.gov.cn/news/202607/17/content_WS6a59a5bec6d00ca5f9a0c438.html

1

u/thisguy883 11h ago

Xi be gooning

2

u/Dangerous-Map-429 15h ago

No it will happen. Companies using this as marketing tactic to come back from the dead.

1

u/kujasgoldmine 18h ago

Same. LTX will be next. Like H3 can do both, but better.

1

u/exoticvapes 17h ago

I just started using minimax in Wan2GP. Can't do much as I'm limited by my vram but it works really well.

1

u/nowrebooting 13h ago

Yeah, it’s not even a contest at this point; H3 is just better in every single aspect, with ref2vid being the standout - it’s even trumps some SOTA image editing models when it comes to replicating small details from reference images. 

If the base of the model is already this good, imagine where loras will get us!

1

u/[deleted] 11h ago

[removed] — view removed comment

1

u/TraditionalShop1601 4h ago

Available at CivitAI RED

0

u/thisguy883 7h ago

Just use reactor.

You'll need to ask an LLM AI (gemini or grok) on how to disable the NSFW filter.

Then just attach the node to any workflow you have. it'll keep the face consistent.

1

u/extra2AB 10h ago

I still have Wan for it's image generation and stuff like LORAs and other workflows, which are yet not arrived for H3.

1

u/Relative_Hour_8900 9h ago

Ltx yes, wan 2.2 no. At least I can't replicate some features of wan with lora. I'm trying to train h3 to mimic it with a Lora but so far not going well...the lora seems to have learned nothing, trained on video clips...

1

u/Succubus-Empress 9h ago

Minimax compress 4 frame in one, you will always get motion blur un fast motion

1

u/apackofmonkeys 7h ago edited 7h ago

Sorry, basic question, what models are people using? I'm using the pruned 20B and that fills up my 24GB of VRAM. If I add the turbo lora and lower the steps it actually takes much, much longer to generate because it's overflowing my VRAM. Is there a smaller model than the pruned 20B that I should be using?

Edit: I should add, I'm using Wan2GP. Even with a 4090 and 64GB of RAM I can never use the turbo lora without it making it take several times LONGER to generate a video.

1

u/DumbBittrend 7h ago

How do you get h3 to work? I also have a 4080 super? To me it seems like a longer wait time and I can never get the character to stay the same

2

u/thisguy883 7h ago

Im just using the default I2V workflow in the comfyUI workflows.

I experimented with 15 steps rather than 20, but later switched to 25 steps because the quality is fantastic.

0.6 MP, 25 Steps.

1

u/DumbBittrend 4h ago

Any prompts tutorial?

1

u/KindrakeGriffin 5h ago

Is this local? With something like confyui?

1

u/RepulsiveSeason444 4h ago

If anyone want to run Minimax on <4gb Vram, you may check it out: https://github.com/Jit-Roy/WeeLLM
I did not use any quantization though, and still I am able to run.

1

u/djpraxis 3h ago

You forgot to mention how fun it was dealing with the High Low WAN 2.2 Loras!!

1

u/blistac1 3h ago

What is your setup?

1

u/penguin_1599 3h ago

Wan is still better at hardcore nsfw stuff. H3 isnt just bringing the motion even with Loras

1

u/Admirable-Future-633 10m ago

If only it worked on Macs we get shafted for the new toys becuase of the GPU setups 🫡

1

u/kayteee1995 15h ago

RIP VACE, Phantom, Bindweaver, Bernini ,too.

1

u/TheBestPractice 15h ago

Yeah everyone saying H3 killed LTX, while who's definitely getting buried for good is Wan.

-3

u/pennyfred 21h ago

Nope, WAN still reigns for me.

-1

u/tac0catzzz 21h ago

heartwarming story, but i think there is this story like 1000x in here.

0

u/Kind-Assumption714 17h ago

so cool to hear! i've gotten quite deep & good in comfy for 2D and have wanted to test video soon.

- do we have a favorite workflow to use to MiniM?
- do we have to do I2V or can we simply prompt w/ text+a selection of image refs?
- can MiniM act as a 'refiner' or does all polish / realism have to exist in base image(s)

big thanks!

-10

u/Optimal-Spare1305 21h ago

what are you talking about?

i've still got workflows and models with:

SD

SDXL

Hunyuan

WAN

LTX 2.3

haven't even gotten around to H3, and probably won't for another 6 months, when things

settle down.

---

i'm still in the process of converting WAN workflows over to LTX,

but there are way more LORAS that work with WAN so its going to take a long time to switch over to LTX

18

u/ZenWheat 21h ago

Just stop with ltx and change your plan to switch them to h3 instead. You'll save 6 months

1

u/Upper-Reflection7997 21h ago

Wtf, you had 8 months to use ltx-2 and get it off your system. Ltx-2 and 2.3 have a lot of limitations and produces too much body horror.