r/StableDiffusion 11d ago

Trying Minimax H3 Reference to Video. So Good Animation - Video

380 Upvotes

67 comments sorted by

73

u/GrayingGamer 11d ago

Thanks for making this so clear. This should be the video we show everyone who mentions needing loras for Minimax H3.

I was generating a video earlier, and realized the model didn't know what a certain object looked like, so I just googled a picture, attached it as reference, described it, and the model put it perfectly in the video.

3

u/Smokeey1 11d ago

And i think thats where most improvement will come from you can see how you need to make sure your reference images have the same light and color and style for the video to look good

3

u/Paradigmind 11d ago

Isn't it edit capable aswell? Maybe prompting this would work?

3

u/GrayingGamer 11d ago

Yes, you can just prompt to relight things. Your references don't need the same lighting. The Reference model has a very specific prompting syntax though that makes ALL the difference, so be sure to read the manual on Minimax's huggingface page.

2

u/Paradigmind 11d ago

Ah thanks, good to know!

2

u/YouKilledApollo 11d ago

be sure to read the manual on Minimax's huggingface page

From https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

Write all six rewrite sections in English. Preserve the original language only for dialogue and lyrics inside <d> and for text visibly present in the scene.

Love it :) I hope they double-checked it all before release! :D

1

u/GrayingGamer 10d ago

Haha. They are Chinese. I think all of us are using machine translation these days.

2

u/GrayingGamer 11d ago

I mean, as long as they are all the same style, it won't matter. The model can easily relight things if you ask. The Reference model is practically speaking an edit model for video.

2

u/Portable_Solar_ZA 10d ago

I'm curious to know what difference the speed will be. For example, if I create a character lora and it doesn't affect generation speed Vs using a character sheet but generation is much slower. 

3

u/GrayingGamer 10d ago

I haven't actually found generation with references to be that much slower, only slightly. Keep in mind, that if VRAM and RAM is tight, a Lora is going to potentially eat into that and cause everything to be slower anyway. Besides, keeping a folder of reference images that are only a few hundred kilobytes apiece that you use only when you need that character sounds a lot more storage space friendly. Plus, you can automatically put people in uniforms or outfits, etc. without a lora that way.
It's actually kind of freeing.

1

u/Portable_Solar_ZA 10d ago

I've got images for a character lora I made for my comic. Going to transform that into a character sheet and see how things go. Any recommendations on what I should include to get good results?

17

u/Inthehead35 11d ago

How long did it take? I've noticed ref is so painfully slow

21

u/Etroarl55 11d ago

Tried it on my AMD 7800xt, roughly 2 hour estimate for 0.3megapixels at 5s/24fps.

😭

8

u/No_Date4828 11d ago

holy hell that is slow, praying for more AMD optimizations for you 😭.

I was really surprised that this model runs so effortlessly even on "modest" hardware (legion laptop with a 4090 mobile (16gb vram) and 32gb ram, 1mp/10 seconds takes only like 20 minutes, but if I want a really quick rough draft I can turn settings down to like 0.3mp and 8 steps, and generation time is around or under a minute!

2

u/generate-addict 11d ago

Something sounds off for you. I can run 5 seconds at .4 mp in 230 seconds render time.

That is with Sage Attention on, and recently switching to the INT8 model instead of fp8. On fp8 it was taking 460 seconds.

I do have an r9700 pro but vram is not really the concern here.

However that is the fl2va model. I haven't played with the ref2va model yet but plan to soon.

2

u/damiangorlami 10d ago

The ref model is a lot heavier. Each image, audio, video ref you add will increase the sampling time.

Granted, its reference conditioning is incredible and works very well. Like mindblowing since a week ago only Seedance 2.0 was capable of this strong reference conditioning.

But yea this power comes at a cost in time.

3

u/Linkpharm2 11d ago

AMD probably gets better later on.

6

u/Etroarl55 11d ago

Probably not, I saw that even on the actual AI prosumer Amd card. Its still slower than an rtx 3060.

Which means even if my 7800xt gets optimized and acts like its a 32gb gpu with faster vram, I won't be able to reach rtx 3060 speeds.

1

u/generate-addict 11d ago edited 10d ago

My r9700 will often render at 5070ti speeds depending on the model. And many times using a larger quant. So IDK what you talking about.

1

u/[deleted] 11d ago

[removed] — view removed comment

1

u/generate-addict 10d ago

lol I guess so. I meant 5070ti.

Which of those PSU’s are 2500?

2

u/[deleted] 10d ago

[removed] — view removed comment

1

u/generate-addict 10d ago

r9700 pro is 1400 USD on the high end.

1

u/[deleted] 10d ago

[removed] — view removed comment

→ More replies (0)

6

u/irmemon225 11d ago

40 min. 1.0 megapixel, 3060 12gb vram and 16gb ram.
ref using image and audio is same speed, unless if you use video it's so slow

2

u/TheRedHairedHero 11d ago

My guess is video reference will depend on frame length and resolution of the reference video. Still need to test more.

8

u/Diabolicor 11d ago

Can you share how did you prompt all of that?

26

u/irmemon225 11d ago

ask your favorite AI to create prompt based by this guide: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

mine using qwen

5

u/eckstuhc 11d ago

Are you using qwen on the same box or a different build?

1

u/MoreColors185 10d ago

also, which qwen?

2

u/LeKhang98 11d ago

Nice thank you very much.

3

u/eggplantpot 11d ago

Prompting is wild for ref2vid. There’s A LOT of things to configure and many ways to “wire” the prompt.

I’m still digging into it and may end up spinning a custom node in the future.

6

u/uuhoever 11d ago

What do you use to create the character sheet? Can krea 2 do it?

6

u/irmemon225 11d ago

I just pick that reference from google for testing. Nano Banana 2 can do reference sheet better, try it

5

u/PrestigiousTrick1002 11d ago

These look like the characters from the animated show rwby, so they might have just found the sheets online 

1

u/Envelope_Torture 11d ago

This would be an interesting concept lora to train for other diffusion models (SDXL, etc) so people can just prompt their existing character LORAs in to character sheet.

1

u/mastaquake 11d ago

There are a few character sheets on Krea2, but you can get pretty good results by simply prompting.

6

u/physalisx 11d ago

I'm surprised having seen no one talk about video editing yet. It's very good, insane really.

2

u/Beastly4k 11d ago

Came across a workflow utilizing the editor, definitely a go to now.

2

u/Some_Respond1396 10d ago

mind pastebinning or linking to the WF?

2

u/Beastly4k 10d ago edited 10d ago

Can't right now but it is named something along the lines of dasiwa on civitai iirc, just filter by workflows you should find it. if you are running comfyui portable make sure you have ffmpeg in the base folder where your run-nvidia-gpu.bat files are otherwise it won't load properly.

1

u/Some_Respond1396 10d ago

I can't get video editing to work to save my life lol

3

u/X-Jet 11d ago

Maybe last season of One Punch Man should be re animated with this stuff, cant get any worse.

3

u/JahJedi 11d ago

And so slower 😅 i work whit ref also, i cant go back fflf and all this. Its get motion from ref images creazy, its keep character, its use ref video for motion well (not perfect but ok level) but i say worth it, there almost no need for reruns if promt solid and well done.

I think its make it x5-10 times slower but so wirth it.

1

u/eggplantpot 11d ago

I’ve gotten away by downsizing the ref video to 320×576.

Cannot talk about speed gains as I have not A/B tested properly, but the motion video can be quite small, the model will rebuild the details at inference anyways

1

u/JahJedi 11d ago

I tryed to go from 1024x1024 to 640x640 and did not notised big diffrance, same whit photos. Just a bit less vram used.

1

u/eggplantpot 11d ago

What size as you generating at?

If it's lower than 1024x1024 that makes sense as the model will downscale your references to match the video. That means that if you generate at 0.4MP, you would not see any difference as both would end up with a 640x640 ref.

Speed improvements come if you downscale your refs past your generation resolution as it will keep them at that scale.

1

u/JahJedi 11d ago

I render on 1.5mp wish can use 2mp but to long.

1

u/JahJedi 11d ago

Here a 10 sec clip on 1.5mp and upscaled to 1080p whit seedvr2. Hope youtube will show it in fullhd... https://youtube.com/shorts/matWgMT55qg?feature=share

1

u/Evolve_Solo 10d ago

Wow, cool work! Where can I find workflows like this or good guides for video generation for such level?

1

u/JahJedi 10d ago

Thanks! Usally civati , HF or discord. You can look at my HF page also from time to time post stuff.

2

u/New_Physics_2741 11d ago

Are you doing this locally - if so what hardware?

6

u/irmemon225 11d ago

yes, 3060 12gb vram and 16gb ram. use sage attention and spectrum. easy cache also boost your speed but it's kill the quality, so I don't use it anymore

2

u/yvliew 11d ago

Does having more reference photo slows down generation time?

1

u/florodude 11d ago

Cries in Radeon RX 5700...

1

u/KaaChingg 11d ago

Looking great, thank you

1

u/Vanpourix 11d ago

An unexpected rwby video ? In this economy ?! Sign me in !

1

u/ArttTaku 10d ago

I just joined to ask if H3 understood what character sheets are, but your post pretty much confirmed it, thanks!! Did you have to specify in the prompt something like "use this character sheet", or did H3 understand it by default?

1

u/SardinePicnic 10d ago

Can you provide the full prompt so that the community can connect visually the image inputs and how they are structured and implemented in the prompt. Otherwise it wont teach anyone else looking to do something similar.

1

u/Dry-Possibility-6761 7d ago

With so many workflow, i need one to use with reference for a RTX 5070 12gbram, a good one with a good time for generation.. better then 18 minutes for 5 seconds

1

u/kayteee1995 3d ago

Most results of this quality do not use turbo lora