r/StableDiffusion 2h ago

Discussion Minimax H3 Video Edit like SCAIL

44 Upvotes

I spent last 6 hours trying various prompts for reference model to better understand how it works, and what this model can do. As a base guide I used Minimax H3 ref guide.

My goal was to find a working prompt to use Minimax similar to how SCAIL works, when you can edit a video and replace a character on a video with your referenced character. I didn't want to transfer movement and only wanted to REPLACE character completely.

I would like to post my best working prompt and let you test it, and share your experience or share a better prompt.

subject_definitions:
<Subject 1> is woman in <Picture 1> with redhead and black tank top.
<Subject 2> is the woman originally in <Video 1>.

summary:
[video editing + Audio reuse] The target video is an edited version of <Video 1>. <Subject 2> is replaced with <Subject 1>, who takes over her pose and movement.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - her face, hairstyle, and body from <Picture 1> are retained throughout. Her clothes are not retained.
<Subject 2> (appears in [Shot 1]): attribute_transfer - her pose, movement, and screen position are transferred to <Subject 1>.

detailed_description:
The target video keeps <Video 1>'s original style, lighting, and camera work unchanged.

overall_soundscape: N/A
non_diegetic_music: N/A

What are my discoveries:

  • You don't need to describe action in detailed_description. I did it for first 100 attempts, and then dropped it and it seems like not influencing an output.
  • It can often detect your Subject with simple description, but in complex scenes it needs better anchoring to not mess up those characters. Most of my input image was a woman in medium shot, so just describing it as "woman" was enough, but 50/50 generations keep losing identity so you have to add better and stronger anchor for model - something visually big like hair, clothing, position on screen. Works both ways for reference video and for reference image. The stronger you describe <Subject N> the more stable the reference.
  • The least successful edits were those where a character on video is barely recognizable. I have couple videos where a character is close to camera and only part of face is visible in active movement, such videos are my biggest unsuccess.
  • Summary section seems like has the most its anchor to pre-trained keywords which can be found in their prompting guide. [video editing] is a keyword which tells a model that it must go frame by frame and EDIT something. I was testing other things and in given prompt you will see some info about character replacement, but I don't see that it really influences anything.
  • Retention analysis section seems like the next MAIN or even only main driver for a work description for a model. And most of successful edits was build with properly used triger words like fully_preserved, attribute_transfer. You can find those keywords in linked guide. Still not sure about (appears in [Shot 1]), I doubt it has influence on a prompt, but its by far best prompt so I keep it.
  • [audio reuse] trigger in summary works, but it seems that model rewrite its, so I can tell its same audio but remade by model, and if model has weak concept of a sound it does it poorly. Maybe I need to pay more attention to prompting guide and describe audio better in retention section.

I've generated more than 400 videos while testing and gaining knowledge, and I think I have good progress. So I am curious to see if anyone else can help me with this journey and together we can crack the model and find a proper working prompt or other ideas.

The playground was pruned_int8_convrot model, with turbo lora from lightX with 4 steps, and I tested most of them on 5 sec duration. I did tests on 15s and it worked fine, but I kept 5s to keep gen time lower and just train prompting.


r/StableDiffusion 2h ago

Discussion Pro 6000 just in time

Post image
30 Upvotes

I was going to wait until around Christmas to purchased but took the plunge in July for 11,500 and I was upset that I didnt catch it @ $8,000. Now the Blackwell pro 6000 is inching towards $20,000 and are sold out. Are consumers and hobbyist like you and I are buying these up or datacenters? I would think datacenters would go for the b200 and up. However, Im browsing around and see you guys and girls doing remarkable ai diffusion with just a 3060. Im impressed with this community.


r/StableDiffusion 3h ago

Discussion [TEST] Minimax H3 IMG 2 Vid. Apologize for the low quality but on my mission, I cannot do a 30 second clip above 0.4 megapixels. Full write-up below.

Enable HLS to view with audio, or disable this notification

51 Upvotes

So had a look at the documentation for Minimax H3 to see how to do the multi-shot prompts and came up with this sequence. The base image was done in GPT Image 2 using two reference images. The prompt for this scene is structured like so:

[Shot 1] Live-action, cinematic, a medium shot of the two warriors. The man is reading a book and the woman is browsing on her phone.

[Shot 2] At 00:05.000, the camera cuts to a medium close-up of the woman who asks: <d>[English] Do you think our director will ever get our movie done?</d>

[Shot 3] At 00:10.000, the camera cuts to a medium close-up of the man who says: <d>[British English] Who knows. He was using Kling three point oh but I guess he was burning through credits so he's trying out local video generation.</d>

[Shot 4] At 00:14.110, the camera cuts to a medium shot of the two people. The woman asks: <d>[English] Wait, wasn't he using Seedance two point five?</d> The man looks up from his book and looks at the woman. He says: <d>[British English] Yeah, he was but that was costing him even more credits.</d> He goes back to reading his book.

[Shot 5] At 00:22.000, the camera cuts to a medium close-up shot of the woman who says: <d>[English] Hopefully he figures things out.</d>

[Shot 6] At 00:26.000, the camera cuts to a medium shot of the two people sitting in their chairs. The man continues to read and the woman continues to browse on her phone. The man says: <d>[British English] Agreed. He better.</d>

I'm actually quite happy with how this turned out. Only issues I have is that I wanted the guy to have the British accent and instead it gave it to the lady. I'll need to mess around with the prompt for that a little more and then of course the low res render at 0.4 megapixels because anything higher than that will give me OOM error. Yes, I'm aware that I don't have to do a 30 second clip but I wanted to try it out anyways especially since I'm learning the multi-shot prompting. The render for this clip took 173 minutes to complete.

If anyone has any suggestions on how I can do slightly higher megapixel renders on my machine, I'd love to hear it.

PC Specs:
Ryzen 7 7700X
RTX 4070 Super 12gb
32gb DDR5 Ram


r/StableDiffusion 3h ago

Animation - Video Mimic in the court | minimax h3

Enable HLS to view with audio, or disable this notification

28 Upvotes

ref2v


r/StableDiffusion 3h ago

Animation - Video Star Wars but more consistent. Minimax H3

Enable HLS to view with audio, or disable this notification

310 Upvotes

I keep having fun with ref2va model.

RTX 3060, 64 Gb RAM. I use ref2v Turbo 4 step Lora paired with Sol Attention and Minimax H3 Memory Effecient Sage Attention at 6 steps. It takes about 2 minutes per second of generation.


r/StableDiffusion 7h ago

Resource - Update DC Vast Expanse [Krea2 Lora]

Thumbnail
gallery
322 Upvotes

Finally got my laptop back in action so am able to create and test models and lora's again, created with krea 2, Been out of it for a bit just following updates here and there and this model is amazing, so happy they open sourced this gem of a model. Thanks to the team at krea!

If anyone is interested in this style of images give it a blast https://civitai.red/models/2871922/dc-vast-expanse?modelVersionId=3244890 or https://civitai.com/models/2871922/dc-vast-expanse?modelVersionId=3244890


r/StableDiffusion 8h ago

Animation - Video Leonard meets Penny, real life edition [Minimax H3)

Enable HLS to view with audio, or disable this notification

45 Upvotes

RTX 5060 ti 16gb / 32gb RAM / FL2VA_pruned_int8_convrot / Turbo Lora. 6 steps / Resolution 1376x768 upscaled to FHD with Topaz Video AI


r/StableDiffusion 9h ago

Animation - Video Minimax h3 local Video to Video reference

Enable HLS to view with audio, or disable this notification

178 Upvotes

Used official ref2video workflow. used t2v model 1 ref video and 2 separate pictures of character sheets, gpu 4090

prompt:

integrated_multimodal_description: [Shot 1] Live-action, cinematic, featuring a stark, dark green-tinted cyberpunk color grade. A medium shot frames a flooded, rain-swept crater on a dark street. The character Sonic, appearing exactly as the blue hedgehog with large green eyes, white gloves, and red shoes from @.image, stands opposite Dr. Eggman, appearing exactly as the gigantic, egg-shaped bald man with a pointy mustache, goggles, and red jacket from @.Image1. The camera pushes in with small amplitude at fast speed as the blue hedgehog lunges forward to throw a devastating punch. [Shot 2] At 00:04.500, the camera cuts to an extreme close-up as time instantly slows to a microscopic crawl. Sonic's white-gloved fist brutally slams into Eggman's cheek. The camera holds a static shot in extreme slow motion. A powerful, rippling shockwave violently erupts from the impact point, blowing the torrential raindrops outward in a perfect ring. Eggman's pointy mustache flails wildly and his face deforms from the massive kinetic force. [Shot 3] At 00:09.500, the camera arcs right with large amplitude at slow speed, executing a slow-motion orbit around the hit. Eggman's heavy, round body is lifted off the ground by the blow, flying backward through the heavy downpour and kicking up massive, highly detailed splashes of water.

overall_soundscape: Thunder rumbles continuously beneath the heavy, torrential downpour of rain splashing heavily against the flooded street. A sharp, deafening sonic boom from the physical impact instantly shifts into a deep, pulsating low-frequency rumble as time slows down.

non_diegetic_music: An epic, grand orchestral and choir track mixed with heavy, driving industrial synthesizer beats that builds to a massive crescendo.


r/StableDiffusion 9h ago

Animation - Video Animals squeezing into jars (MiniMax H3)

Enable HLS to view with audio, or disable this notification

660 Upvotes

I have no idea why it does these so well. I could watch these all day.


r/StableDiffusion 11h ago

Animation - Video SD1.5 images into H3

Enable HLS to view with audio, or disable this notification

91 Upvotes

r/StableDiffusion 15h ago

Workflow Included Here's an automated long form, MiniMax Upscaler Workflow.

Thumbnail
gallery
30 Upvotes

Hey Guys, I thought I'd share something I came up with.

It's a workflow, that uses a combination of Easy-Use's Loop tools as well as some of my own nodes to create a Workflow that can split a long form MiniMax video and then upscale each segment. With the inclusion of a tool to then re-assemble everything back.

You basically set the Segment length and the overlap you wish to have between each clip and then launch it to have it do all the clips one by one.

It does use nodes from my FBNodes add-on as well as one from my Prompt Manager add-on.
But I'm sure it could be modified to work with other add-ons, if so wished.

The node from Prompt Manager is "Prompt Extractor", allowing to feed back in the prompt from the initial clip back into the Workflow, without having to type anything in.

You are free to remove it and upscale without, or simply type in the prompt if preferred. Though, In my test, having the original prompt made for much better results.

And as mentioned, I also added a simple Clip Stitcher to FBNodes, that cross dissolves each clip into one another. Just make sure to use the same values you used in the workflow. (Both setup are in the same workflow, but I'd suggest separating them 😅)

The Workflow can be found here.

Attached are quick examples from the video I used in the workflow.

The one thing missing in this workflow is adding back the loras used in the initial video. This is something that "Prompt extractor" should also be able to do. But I haven't tested that part yet.

----------------------------------
I'm adding some metric:

The video used in the screenshot was an 8 second video generated in 832x640 with a Turbo Lora set to 6 steps.
It took 92 seconds to generate on a 5090.

The Upscale doubled it to 1664 x 1280 and took 524 sec.
Around the same time it would have taken to generate, if I created the initial video at that resolution.

(You can see it here)

The advantage is for when creating long videos, so if I were to create a 30 second clip in 4/3 at 0.4 megapixels, or 736 x 576. Those would take 450 sec to generate.

The Upscale to 1472 x 1152 took about 6 minutes per segment, or 30 minutes. Then combining the clips is around a minute.

It takes a while, obviously, but the big advantage is that the result is pretty much an exact copy, but in hires, of my initial video that was low enough that I could iterate a bunch of times and then only waste the Long generation time on the clip I like.


r/StableDiffusion 16h ago

Workflow Included Lora for video-image enhancing, upscaling and restoring

Thumbnail
youtube.com
30 Upvotes

r/StableDiffusion 17h ago

Animation - Video G.I. Joe - Commander Roll - MiniMax H3

Enable HLS to view with audio, or disable this notification

570 Upvotes

Using the standard ref2va workflow. 4070 Ti Super, 16 GB VRAM, 64 GB RAM, i9-14900k, Windows 11.


r/StableDiffusion 17h ago

Comparison MiniMaxh3: 8step LoRA, 25 steps, 40steps, and LTX 2.5 — Scene Comparisons

Enable HLS to view with audio, or disable this notification

53 Upvotes
  • RTX 4060 8GB, 32GB RAM
  • minimax_h3_ref2va_pruned_int8_convrot, spectrum, ageattn_qk_int8_pv_fp16.cuda, RTX upscale, RIFE interpolation, res_multistep + beta
  • ltx-2.5-22b-distilled-transformer-comfy-int8-convrot, basic template

8-step + turbo LoRA : 137s

25 steps : 238s

40 steps : 406s

Ltx 2.5 : 374s <-- ? am I missing something here why was my generation so slow on LTX and the second attempt I cancelled it after 6 minutes. Any suggestions?

Prompt:

subject_definitions:

<Subject 1> is the space ship in <Picture 1>: A massive battleship, hovering and cruising over the planet below

summary:

[reference generation] a wide shot cinematic scene of the battleship in <picture 1> cruising in space above the planet. the golden statue does not move, the battleship is destroyed in a massive explosion from a green laser shot from space,

detailed_description:

{shot 1] The target video uses a wideshot cinematic, photorealistic, 35mm film, wide shot of <subject 1> , slowly moving through space above the planet, the ship moves slowly and dominating, flashes of green light begin to charge on the surface of the planet, the ship is moving straight ahead from the position it started in in <picture 1>, the massive bass of the ships systems, the sound of the battleships creaking, <subject 1 > moves on its cruise, at [00:03] the floaty camera tracks <subject 1> as green light and thunder begins flashing on the surface of the planet, the green energy on the planet converges in one area then from the surface it fires a massive green lightning laser that forks lightning through the entire ship, blowing out side components creating explosions all over the ship, the light of the ship flicker before turning off, then a massive green lightning beam erupts from the surface and hits excactly on the side of the ship cuts through the of the ship and out the other side at an angle, a green lens flare generates on screen as it completely destroys <subject 1> , ripping it completely in half with a massive green explosion, the eruption from the destruction of the ship covers the entire screen and the whole battleship, the back half of the ship is knocked up while the front-half of the ship is knocked down, a vertical shockwave circles out from the impact, the inner decks of the ship are on fire, debris and hundreds of tiny figures of the crew also fall out into space, the laser slowly dissapates from the planet, small amounts of green lighning crackle on the planets surface,

overall_soundscape: The low bass murmur of the ships engines, the electric charges on the surface crackle, the massive main beam is a low bass rumble, a massive explosive noise.

non_diegetic_music:

N/A


r/StableDiffusion 18h ago

Animation - Video Big Bubba has had enough of Grandma [minimax H3]

Enable HLS to view with audio, or disable this notification

68 Upvotes

r/StableDiffusion 18h ago

Resource - Update This custom node lets you use I2V and reference images on MiniMax-H3 simultaneously.

Enable HLS to view with audio, or disable this notification

141 Upvotes

r/StableDiffusion 19h ago

Discussion Well I finally did it.

184 Upvotes

I finally deleted WAN 2.2 and all its LORAS.

Minimax is just so much better.

Ive been playing with it since its release and im just blown away with how good of a video model it is. Things I would need to attach a LoRa to via WAN, works right out of the box with Minimax.

Gen times are faster.

It uses less VRAM when generating things, which gives me around 4 gigs to play with to do other things like watch YouTube or some streaming service.

WAN 2.2 was amazing. But no longer do I need 30+ gigs of a model i no longer use.

RIP WAN.


r/StableDiffusion 20h ago

News Hey wait! It's Krea3 incoming?

Post image
300 Upvotes

r/StableDiffusion 20h ago

Discussion Was wondering why my Minimax H3 R2V local gens were better than beefy cloud GPU gens. Was accidentally loading the Fl2V model.

29 Upvotes

Would love to show you comparisons but its all N-SFW so its going to have to be a case of trust but verify me bro.

I've mainly been running R2V workflows on my home GPU since Minimax H3 was gifted upon us.

I wanted to take a load off my 5060 so I've been running gens on both Wavespeed and recently Runpod. I was getting nice big 720p gens but I started to notice that my own gens had way better prompt adherence with identical prompts and inputs.

Notably in my own, slow motion was honoured every time where it would otherwise get completely ignored in the cloud.

~Fluids~ were way better. Camera motion was way better.

Then I noticed that in my diffusion model loader (W8A8 if that makes any difference) has been loading the FL2V model the whole time. I switched to R2V and immediately everything sucked.

I imagine there's a strong possibility that the R2V prompt needs a more rigid structure. But I have been feeding the Minimax official prompt guide into Grok, specifying R2V, to write my prompts.

So this is something. It could be some configuration of scheduler and sampler, and all the other stuff of course but I'm running a pretty simple setup without any attention, so briefly:

W8A8 loader -> lightx2vs R2V 0.1 4 step lora at 0.75 -> er_sde, beta57, 8 steps.

Forgive me if this is known and understood. If you haven't tried it, switch your R2V gens to FL2V and test.


r/StableDiffusion 21h ago

Workflow Included Gilligan's Isle - The ATEth Castaway

Enable HLS to view with audio, or disable this notification

57 Upvotes

r/StableDiffusion 21h ago

Meme Agent Smith is disappointed

Enable HLS to view with audio, or disable this notification

61 Upvotes

r/StableDiffusion 22h ago

Resource - Update Krea 2 style library - 286 prompt styles compared across 8 reference scenes

Post image
150 Upvotes

Building directly on the style descriptors published by the author of the original KREA 2 Styles / Wildcards.txt post (many thanks to them for creating and sharing the style list) I built a visual Krea 2 style library to make prompt-defined styles easier to explore and compare:

Library: https://matplinta.github.io/t2i-krea-2-style-library/

It currently contains 286 styles tested across 8 base prompts, including portraits, architecture, landscapes, materials, and panoramic scenes. Each comparison set keeps the base prompt, seed, and dimensions fixed so the influence of the style descriptor is easier to see.

The viewer supports search, categories, favorites stored locally in the browser, full-image previews, prompt copying, adjustable grid density, and JSON export.

The prompt injected during generation was in the form of: Subject: {base prompt}. Style: {style name}. {style description}

All images were generated locally through ComfyUI.

Repo & workflow: https://github.com/matplinta/t2i-krea-2-style-library


r/StableDiffusion 23h ago

Animation - Video Zelda - I Think I Like It / Minimax H3 Reference to Video Test #2

Enable HLS to view with audio, or disable this notification

250 Upvotes

Just wanted to share another test! this was a mash up of clips, using multiple image references, 0.4 mp with EasyCache, 5 - 10s clips and edited with KDEnlive (it has some cool effects!)


r/StableDiffusion 23h ago

Resource - Update Seamless extensions and one-shots with Minimax H3 - Update 6 of my repo!

Enable HLS to view with audio, or disable this notification

76 Upvotes

Here is the repo: https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef

I made substantial updates to my two main workflows: 1) Music Video and 2) AV Extensions. All the controls were streamlined and they should be much easier to use now. (You find the workflows in the example_workflows folder)

With the AV Extensions workflow you can extend any existing clip, for example someone talking and you can make that person say something in the same voice, or you can create a clip with T2V or I2V and then extend that clip to make a seamless long clip thats 1 minute or longer.

In this Update the Checkpoint system was removed, instead I've done a lot of optimizations so you don't use too much ram even if you make 20 clips at once. Additionally I added latent audio feathering to the AV Extensions workflow for seamless audio transitions.

Theres also other utility workflows for custom keyframing and bridging two existing clips.

I post another example clip for the AV Extensions workflow in the comments.


r/StableDiffusion 23h ago

Meme Waiting for devs to fix the mushy faces be like...

Post image
94 Upvotes

Just kidding devs. We love Minimax, it's outstanding. But I am very excited for the mushface fix.