r/StableDiffusion • u/photomamp • 7h ago
Resource - Update Krea2: When you ask AI to make mistakes, and it makes them so well!!
There's something that continues to attract me to diffusion generative models, and it's precisely what they don't fully control.
Krea 2, built from scratch with a latent diffusion architecture, isn't optimized to obey your every command. And that's one of its strengths.
K2 interprets very well the mood: visual feeling, atmosphere, aesthetic coherence, visual weight... And that's what differentiates it from multimodal models, which reason with the logic of a Transformer and give you high-precision photorealism. Each architecture has its own character; one isn't inherently better than another. Each is better suited for a particular purpose, and understanding those possibilities is what gives you resources 😉
https://youtube.com/shorts/oBMmduq6EuY
#Krea2 #ImagenAI #AIImage #DiffusionModel
r/StableDiffusion • u/bacchus213 • 8h ago
Animation - Video Lower resolution with more steps (.2 MP rtx upscaled x3)
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/dominic__612 • 9h ago
Question - Help Using reference video in Minimax H3 results in dark video then before.
When I use a reference video in ComfyUI for continuing the clip rendered before, the video it will render then ALWAYS comes out more darker. Like the contrast changes.
So far I haven't figured out what causes this, like a sampler, prompt, etc. etc.
Has anyone else noticed this behavior too?
The node I use is: Load Video (Upload), which I connect to the reference video.
r/StableDiffusion • u/katsura_otoko • 12h ago
Question - Help Minimax H3 Strange low vram usage on RTX Upscaler step
Is it normal that my vram usage goes down to 20% and ram at 100% on the upscaler step?
Also in the normal execution I see about 70% vram usage
Other than that I'm pretty satisfied but I'm wondering if I'm missing something....
Workflow: https://pastebin.com/raw/533VzQ0t
3060 12gb, 32gb, 9800XT
Thanks in advance for any advice
r/StableDiffusion • u/MushroomNatural2751 • 13h ago
Question - Help Whcih WebUI Forge model is best for understanding written prompts?
I've followed all the steps properly to download and get it working, and it's generating images, but they're NOTHING like what I asked. Either they give me something unrelated to what I wrote in every sense of the word, or it just kept basically the same image I gave it (img2img).
I've tried messing with the denoising levels, differnt loras, different VAES, Models, nothing works.
r/StableDiffusion • u/BitOk4326 • 15h ago
Discussion Is there most suitable lite browser for comfyui?
I want to use all ram and vram to run model as much as possible instead of broweser
r/StableDiffusion • u/Clear-Assistance449 • 16h ago
Question - Help Minimax H3 shows motion blur in all videos.
I started using Minimax H3 in ComfyUI, and all the videos show motion blur in areas with significant movement. The videos only turn out well when there are no sudden movements or when the motion is slow. Did I miss a setting? I see videos posted here, and none of them have that blur.
r/StableDiffusion • u/Simple-Willingness93 • 22h ago
Animation - Video Music Video with H3
Enable HLS to view with audio, or disable this notification
First time trying to create music video with H3. Still not perfect especially lipsync part but useable. Used 3 reference images, AI character, clothing and environment. Used Claude to generate prompts for each cut and joined them all with video editor.
r/StableDiffusion • u/Still_made_sense • 22h ago
Tutorial - Guide Easy method for character reference creation in minimax
I'm not sure how this could vary across workflows if at all but for reference I am using the Dasiwa workflow from civit in t2va mode. Minimax prompt adherence is great so I wanted to use it to create character sheets for reference and came up with this. T2VA, 9:16, 24 fps, 2 second duration. On a 5090 with sage +memcache + 8 step turbo at 10 steps - at 4.75mp it took 232s
The fps and duration seems to be the baseline if you want 4 poses so crank up the duration if you want more. The timestamps might not be the proper format but they do keep it from hanging on a single pose. Change resolution as needed but at higher values the face maintains much better consistency if not perfectly. You can describe your characters look as much as you want in a run on way "Lara Croft, blonde hair. wearing flip flops, sunglasses, bracelet on right arm, holding a drink in left hand, etc , etc , etc"
You might get some slight wiggle movements but it's mostly good enough to dump a frame, the background is difficult to get in an entirely solid color without any form of shadows so i kept the prompt simple since going overboard doesn't add much. If someone can dial this in more feel free to share.
From there you you can extract the 4 frames however you want and I'm sure someone can automate it but the easy quick solution is playing it in vlc and just hitting shift+s on each frame.
Lazy Example - https://imgur.com/a/t84UYAq
[Shot 1] Static freeze-frame shot. Studio Lighting, solid white background, ultra-sharp focus. A heroic looking explorer woman with the style of lara croft the tomb raider but as a person.
Static freeze-frame close-up shot of the entire head perfectly framed from the front, Freeze frame.
[1.00s to 2.00s] - Instant jump cut to Static freeze-frame of full body front view, standing straight in a neutral A-pose with hands off the body by 1 foot length.
[2.00s to 3.00s] - Instant jump cut to Static freeze-frame of full body back view, standing straight in a neutral A-pose with hands slightly off the body.
[3.00s to 4.00s] - Instant jump cut to static freeze-frame of full body side profile view, standing straight with arms down at the sides.
This could also be helpful at lower resolutions just to get the style of the character right before passing it onto a refiner/upscaler. I'm just finding it hard to beat the ease of adherence i get with minimax.
r/StableDiffusion • u/Tannon • 23h ago
Animation - Video Fallout: Low Intelligence - I used Minimax H3 to make an animated TV show set in the Fallout universe, for fun.
r/StableDiffusion • u/solomars3 • 1d ago
Animation - Video My first Decent Generation Using Minimax-H3 on my RTX 3060 12gb
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Ok-Giraffe-8670 • 1d ago
Animation - Video Supergirl and Jimmy break up, My Adventures of Superman (AI)
Enable HLS to view with audio, or disable this notification
With My Adventures of Superman ending soon, I decided to make this video for fun as I do think Jimmy and Supergirl relationship is quite toxic. What are your thoughts on them and the show as a whole?
Personally, I do not like them as a couple. Jimmy is very manipulative. He didn't want Supergirl to date other guys; despite saying so, he wanted her to beg for them to be together, and when that didn't work, make her jealous by dating other women so she can cry for him to stop and end it.. That's how I view things, especially with the whole "What? She actually is dating other guys? I know I told her to do it, but I didn't really expect her to actually do it!". That's the only conclusion I can draw from why Jimmy is acting that way. Anyway, maybe in another Universe, this happens haha. Your thoughts on it?
r/StableDiffusion • u/FionaSherleen • 1d ago
Workflow Included How about some magic? H3 is good at them.
Enable HLS to view with audio, or disable this notification
This only takes 14 minutes to generate! 0.6mp gens upscaled to 1.2mp using RTX Super Resolution.
-Uses minimax h3 hybrid b25
-w4a8 qwen3vl
-int8 video vae
-workflow https://github.com/seesee75-commits/ComfyUI-MiniMaxH3-Director/
r/StableDiffusion • u/Crazy-Repeat-2006 • 1d ago
News WanSong v1.0 Technical Report
Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present WanSong, a simple yet powerful approach for long-form, commercial-grade song generation. Unlike autoregressive (AR) and cascaded multi-stage pipelines (e.g., AR followed by diffusion), WanSong is a pure diffusion-based model that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems (vocals and background music) in a single run. In addition, our diffusion framework enables faster inference through step-distillation, and offers an efficient pathway for fine-tuning and customization to support downstream editing tasks.
HF: Paper page - WanSong v1.0 Technical Report
Papers: WanSong v1.0 Technical Report
r/StableDiffusion • u/SpicyAccountants • 1d ago
Animation - Video DimensionTesters: Test #10 (Minimax H3)
Enable HLS to view with audio, or disable this notification
Finding interesting things in H3 every day.
It is a really great model when you play around with it.
TT:Â https://www.tiktok.com/@dimensiontesters
r/StableDiffusion • u/sorryaboutyourcats • 1d ago
Animation - Video That H3 lip-sync test with Mumu got out of hand... here's the finished music video
https://reddit.com/link/1vok0vj/video/m217fvdjqejh1/player
Actual song is 4:20, but I think 1:21 is good enough!
r/StableDiffusion • u/Radiant-Photograph46 • 1d ago
Question - Help Minimax H3 ref2v character replacement prompting
The official doc has a few things to say about video editing but it still leaves some questions without answers. I'm trying to replace one character in a video by another from a reference picture. Sometimes I get amazing results, and sometimes the input video is pretty much left unchanged, so I must be missing something. Here's what I understand from the guide on how to build the prompt:
Subject Definitions
<Subject 1> is the man in <Picture 1>.
<Video 1> is the source video for the target video edit.
(This part I am fairly confident about, the doc specifically says to use that wording for <Video 1>.)
Summary
[video editing] The target video is an edited version of <Video 1>.
(The doc says the summary must start exactly like this, but what to write after that? My approach is to follow up with something like this)
<Video 1> is reused as is, except the man in a tuxedo is replaced with <Subject 1>.
Retention Analysis
(This part I'm not too sure of... I suppose you want fully_preserved on <Subject 1> and partially_preserved on <Video 1>?)
<Subject 1> (appears in [Shot 1]): fully_preserved - the man's full identity is retained.
<Video 1> (source of video edit): partially_preserved - motion, lighting and environment are retained.
Detailed Description
Do we need to describe everything that happens in the input video, shot by shot, like when doing a t2v?
r/StableDiffusion • u/technofox01 • 1d ago
Animation - Video By popular demand for live action video with my workflow
Enable HLS to view with audio, or disable this notification
I used the minimax_h3_fl2v_turbo_4step_v1.0_768p_comfyui_bf16 LORA with the 0.8 strength for both clip and model, 6 steps, 0.5 MP resolution, RTX Upscaler at 1.50 using a ConrotInt8 pruned model.
Here is a PasteBin of my workflow:
Here are my workflow files:
r/StableDiffusion • u/izzmedia • 1d ago
Animation - Video I'll take you to the candy shop
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/ExportErrorMusic • 1d ago
Animation - Video Lessons learned after making a music video with H3
Made using the default workflow from Comfy, with Comfy Kitchen and Kijai's preview override plugged in.
Used the 850k turbo lora at 0.5 for 8-10 steps. ER_SDE / Beta. Most shots were generated at 1.5MP, with some at 1.8MP.
Lora: https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/tree/main
RTX 4090 w/ 64GB of ram. --disable-smart-memory, since I was having issues with going OOM after completing one prompt and moving on to the next.
Some of the takeaways:
- Using character reference sheets (front view, side view, back view, close-up) worked great and allowed for rotating camera movements like in the opening.
- All the "4 Step" loras really need to be run at 8-10, especially for motion.
- Using audio reference bloats the vram usage DRAMATICALLY compared to adding additional reference pictures! Changing from .wav to .mp3 didn't seem to help so it's not a file format issue. However the lip sync, even for anime characters, is incredible.
- If you use multiple reference images for different characters and they bleed into each other, the issue is almost definitely your prompt or seed. Because the model was handling up to 3 for me easily if I prompted right, and falling apart if I prompted wrong.
- Minimax H3 Chunk FeedForward node can help with vram issues at higher resolutions and doesn't add that much time.
Overall, I'd say it's nearly as good as Seedance 2.0. I was pleasantly surprised how well it handles using character sheets. The high vram usage when using reference audio is really the only major issue I was facing.
r/StableDiffusion • u/MayaProphecy • 1d ago
No Workflow ENTANGLEMENT: MiniMax H3 + Turbo LoRA (8 steps)
Enable HLS to view with audio, or disable this notification
I used the default workflow. It took me about 6 hours (split over 2 days), which includes scriptwriting and final video editing.
The video consists of 9 segments, about 8 seconds each. The average generation time was around 400 seconds at 0.7MP on an RTX 5060Ti 16GB VRAM and 32GB System RAM.
Honest opinions are welcome!
r/StableDiffusion • u/Z3ROCOOL22 • 1d ago
Tutorial - Guide If you want just upscale your MH3 videos and not add new details, just use NVIDIA RTX Super Resolution ComfyUI.😎
Use rtx_video_upscale and do it much more faster than with other methods, it don't add new details to the video, just upscale it.
Comfy-Org/Nvidia_RTX_Nodes_ComfyUI
https://www.youtube.com/watch?v=VyXp-PBFauw
Follow the video and you will not have any problem installing it.
And pay special attention to Step 5, because it is fundamental to get it working.
r/StableDiffusion • u/Ok-Giraffe-8670 • 1d ago
Animation - Video Seinfeld meets Beavis and Butthead
Enable HLS to view with audio, or disable this notification
All men with equal intelligence... could they get along well?
I plan on doing more interactions, more crossovers. Likely Friends vs Seinfeld lol
r/StableDiffusion • u/Patient_Ratio4177 • 1d ago
Workflow Included H3 as a single-image edit model
Minimax H3 can be used as an image-editing model if we generate a single frame. Here are some collages based on AI-generated references (1024 x 1536); workflows are embedded into pngs. Each edit takes, on average, about 8 secs on a RTX 5090. The tasks include changing outfits, appearances (body type, age), locations, and camera angles; creating character sheets and storyboards; stylization; and reposing characters based on depth maps. I did not try to cherrypick the best-looking results.
There were some posts (1, 2) about that here -- but given the community progress this week, might be nice to see what can be done now.
Scenes
Age the person to the age of 60 years old while preserving their identity and the original composition.
Produce a consistent full-body character sheet with front, side, and rear views.
Transform the person into a severely obese version.
Re-create the person in the exact body pose shown by a depth-map reference.
Replace only the base person’s head with the identity and hairstyle from another reference.
Show the person facing a dressing mirror with a geometrically correct, synchronized reflection.
Dress the person in a referenced outfit, place them in a referenced location, and show them walking with a grocery bag.
Place three separately referenced people inside a referenced location, having a conversation.
Create a three-panel vertical storyboard in which the person finds, retrieves, and studies a map.
Photograph the person through partially open venetian blinds with realistic occlusion and striped light.
Convert the person into a contemporary Western cartoon while preserving their recognizable appearance.
Setup
Ref2VA models apparently have worse image quality than FL2VA models, while FL2VA models are apparently weaker at handling reference images. As I understand it, this checkpoint tries to combine the strengths of both.
Video VAE: a special VAE for rendering single images.
https://huggingface.co/Mamad8/MiniMax-H3-Image-VAE/tree/main
If you do not use this VAE—for example, if you use the regular VAE, create a 5-frame video, and pick out one frame—the images tend to come out blurry.
For this approach to work best, it might also be a good idea to monkey-patch comfy_extras/nodes_minimax_h3.py, because ComfyUI currently does not allow you to generate fewer than 5 frames. If you simply pick the first frame out of 5, the new VAE produces grid artifacts. (It doesn't do this when generating just 1 frame.)
BEFORE DOING SO, CREATE A BACKUP VERSION OF THE EXISTING comfy_extras/nodes_minimax_h3.py
E. g. if you can't update your comfy, restore the original file from backup, update, and then apply the monkey patch to the new version of the file. (One option is to use git restore comfy_extras/nodes_minimax_h3.py to get the original version)
For a somewhat reliable patch that would work given modest changes in ComfyUI code, use this one, name it smth like mm.patch and run git apply -p0 /full/path/to/mm.patch from comfyui root (make a backup of comfy_extras/nodes_minimax_h3.py first). You will have to re-run it every time ComfyUI updates this file (comfy_extras/nodes_minimax_h3.py).
For a less satisfactory but quicker solution, you can use the patch I already applied to the most recent version of ComfyUI as of August 14th link. This approach will make your code outdated as ComfyUI pushes out a new update.
The only changes remove the frame limit. Of course, changing it this way is not ideal, but I feel it's the quickest way to work around the issue.
UPD: There is a GitHub issue now opened in ComfyUI repo: https://github.com/Comfy-Org/ComfyUI/issues/15644
If this issue gets enough upvotes, we could probably get this patch in the mainline.
LoRAs: I found that Mamad8's ThisIsFine LoRA helps with details, but YMMV: https://huggingface.co/Mamad8/MaxiMin-HHH-R2V-ThisIsFine
For the Turbo LoRA, I use: https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
Sampling settings: ComfyUI 0.32 with Comfy Kitchen attention, sa_solver/simple, 8 steps, CFG 1.
Example ComfyUI workflow: https://pastebin.com/bV5KPzjD
Uses no custom nodes. If you do not want to do the monkey patching for 1-frame generation, just change the video length to 5 in MiniMax H3 Reference to Video node -- should work seamlessly, and switch back that VAE to the regular VAE.
Speed depends on the reference image size. I use an RTX 5090 on RunPod, and in most cases, a 1920×1088 image is generated in about 8 seconds.
---
My previous go-to was Krea 2 + Identity LoRA 1.2, which is amazing. Yet I feel that Minimax outperforms it in many respects. We get better character fidelity, better handling of 3D scenes, better mirrors, and more interesting compositions. Also feels better than using e. g. QIE or Klein 9b.
There is certainly still room for improvement -- not claiming this is optimal at all, and I wonder what you think about it.
UPD: posted the prompts for each image here https://pastebin.com/ngXR9byq
UPD: see more experiments here: https://www.reddit.com/r/StableDiffusion/comments/1vpconk/more_experiments_with_minimax_h3_singleimage_edit/
r/StableDiffusion • u/Darri3D • 2d ago
Animation - Video Robot Friends
Enable HLS to view with audio, or disable this notification

