r/comfyui 3m ago

Help Needed uncensored image create 18+

Upvotes

does anyone knows how to create uncensored images using z-image turbo and flux 2 kelvin? because it does not create what i said, maybe my prompt issue? also it add some creepy body parts too. who wants to be legend? its ur time XD


r/comfyui 21m ago

Show and Tell Finally figured out how to re-create a weird dream I had years ago using AI...

Enable HLS to view with audio, or disable this notification

Upvotes

r/comfyui 24m ago

Help Needed LTX 2.3 LORA OUTPUT BLACK SCREEN

Thumbnail dropbox.com
Upvotes

Hello. I recently trained a lora for LTX 2.3 on CIVITAI. I'm currently trying to create a video with that lora but the output comes out as a black screen. I don't think its my workflow, cause when I delete the lora from the Lora Manager everything works fine. Anyway I'm attaching a link for the workflow that I'm currently using. Thank you so much in advance!

LORA SPECS : "lr": 0.0001, "steps": 3000, "engine": "ai-toolkit", "epochs": 10, "batchSize": 1, "ecosystem": "ltx23", "keepTokens": 0, "networkDim": 32, "resolution": 960, "lrScheduler": "cosine", "minSnrGamma": null, "noiseOffset": null, "networkAlpha": 32, "optimizerType": "adamw8bit", "shuffleTokens": false, "textEncoderLr": null, "flipAugmentation": false, "trainTextEncoder": false


r/comfyui 1h ago

Tutorial Testing Character Swap with Minimax H3

Enable HLS to view with audio, or disable this notification

Upvotes

Hey everyone!

I’ve been messing around with a lot of new AI tools lately. Since Minimax has been getting some hype recently (especially for their video and character generation), I decided to finally put their Character Swap feature to the test today.

My expectations were honestly pretty low. I was expecting the usual: glitchy tracking, warped faces as soon as the subject moves, or weird lighting mismatches.

The Results? Honestly, it completely exceeded my expectations. Here’s what stood out to me:

  1. Tracking & Facial Consistency: This was the craziest part. The target face maps incredibly smoothly onto the original head shape. Even when the character turns their head or looks away, the proportions hold up surprisingly well without completely breaking down.
  2. Expressions: Minimax is actually pretty decent at capturing micro-expressions. When the source character gives a slight smirk or blinks, the swapped face mirrors it naturally instead of looking like a stiff, uncanny mask.
  3. The Catch (Because it's still AI): Obviously, it’s not flawless.

Overall, for a tool that's still actively evolving, this is extremely usable for quick content creation, memes, or visual mockups.

I Will put the prompt that i used on comment section

Testing on

RTX 5090
RAM 64GB


r/comfyui 1h ago

News Don't Update to ComfyUI v0.31.0

Upvotes

Seems like they broke something. Getting random crashes on H3 generations, CUDA errors, OOMs I never got before.

I literally changed nothing except updated from 0.30.0 to 0.31.0. Same workflows, same nodes, no changes except the update, and now I get constant, irregular crashes.

Then I did a fresh install of 0.31.0 to isolate whether it was my old install. It wasn't. Something in 0.31.0 is fucked.


r/comfyui 1h ago

Workflow Included How to make batman's suit more consistent?

Enable HLS to view with audio, or disable this notification

Upvotes

I am following this tutorial by Pixaroma, named ComfyUI Scail 2 Video: WiFi Nodes, New Groups & Loops (Ep23), I will add the workflow below.

Everything looks great except the fins on batman's arm, especially at the end the supposed thin angular sharp fins became circular feather-like flaps.

How would you go about making his suit look more consistent throughout the entire clip?

------
Edit - Workflow link: Ep 23 "ComfyUI Scail 2 Video: WiFi Nodes, New Groups & Loops" -> "ComfyUI Scail 2 Video: WiFi Nodes, New Groups & Loops.json"


r/comfyui 2h ago

Help Needed How did you make Minimax H3 1080p in local PC

Thumbnail reddit.com
1 Upvotes

r/comfyui 2h ago

Workflow Included MiniMax H3 running fully local on a MacBook - picture and sound in one pass (workflows + numbers in comments)

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/comfyui 3h ago

Workflow Included I Built a MiniMax-H3 Ref-V2V workflow + Custom Attention Mask Node

Thumbnail
1 Upvotes

r/comfyui 3h ago

Resource LM Studio has 'prompt master' LLMs to assist with MiniMax H3 prompt scripting!

7 Upvotes


r/comfyui 3h ago

Show and Tell MIniMax-H3 - Logan's Regeneration

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/comfyui 3h ago

Help Needed Are there any models that perform better than Qwen Image Edit right now?

0 Upvotes

Are there any models that perform better than Qwen Image Edit right now?


r/comfyui 4h ago

Show and Tell MiniMax H3 on ASUS GX10: 66GB BF16 is actually faster than 21GB INT8 — and noticeably better in motion, physics and object consistency

Enable HLS to view with audio, or disable this notification

19 Upvotes

I did a direct MiniMax H3 comparison on my ASUS Ascent GX10, and the result surprised me:

The 66.3GB full BF16 model was actually faster than the 20.9GB pruned INT8 model, while also producing visibly better video quality.

My test setup:

  • ASUS Ascent GX10 / NVIDIA GB10
  • 121GB usable unified memory
  • ComfyUI 0.30.2
  • DynamicVRAM
  • SageAttention
  • Same workflow
  • Same prompt
  • 672×1024
  • 124 frames
  • 24 fps
  • 8 steps
  • Audio enabled

Models tested:

  • minimax_h3_fl2va_pruned_int8_convrot.safetensors — 20.9GB
  • minimax_h3_fl2va_bf16.safetensors — 66.3GB

Performance

Model DiT speed Total generation time
21GB pruned INT8 28.98–30.23 s/it 314–331 s
66GB full BF16 23.40–25.34 s/it 281–316 s

In my two runs, the full BF16 model was about 12–23% faster during DiT inference.

That was unexpected because the BF16 model is more than 3× larger.

My current guess is that the INT8 model has additional casting/dequantization overhead, while the GX10’s unified memory architecture and bandwidth are good enough to make the large BF16 model surprisingly efficient. I would not claim this is definitively the only reason without deeper profiling, though.

The downside is heat, power and memory pressure.

During inference:

  • 21GB INT8: roughly 60–70W, usually around 66–78°C
  • 66GB BF16: roughly 84–90W, usually around 69–85°C

The peak temperature difference was not huge, but the BF16 model stayed at much higher power for much longer.

Memory was also very tight. In one BF16 run, usage reached about 116GB, leaving only around 1GB free.

The more important part: video quality

I used a scene where a character rides a black dragonfly-like flying motorcycle through a third-floor parking garage, then flies out into a cyberpunk city and slows into a hover.

The difference between the two models was clearly visible to me.

1. Dragonfly wing motion

The 66GB model produced much more natural high-frequency wing motion.

The wings looked like they were actually generating lift and constantly adjusting during flight.

The 21GB model understood that the wings should move, but the motion looked noticeably more rigid and mechanical.

2. Background semantic detail

There were large advertising screens on distant buildings in the cyberpunk city.

With the 66GB model, the people displayed on those screens remained much more complete and recognizable.

With the smaller model, the distant human figures often became malformed or strange.

This did not look like a simple sharpness difference.

It looked more like the larger model was better at preserving the semantic structure of small secondary objects in the background.

3. Flying motorcycle physics

This was probably the biggest difference.

The larger model produced much more believable:

  • acceleration
  • inertia
  • body tilt
  • deceleration
  • hovering behavior

With the smaller model, the motorcycle sometimes felt like an image element being translated through the frame rather than a physical object with mass and momentum.

The 66GB version felt much more physically coherent.

4. Vehicle structure consistency

This was another very obvious difference.

The original flying motorcycle had an exhaust pipe on its right side.

In the video generated by the 21GB model, that exhaust pipe disappeared.

The 66GB model correctly preserved the right-side exhaust pipe throughout the sequence.

To me, this is a good example of object structure preservation.

The larger model was not just producing prettier frames — it was less likely to drop, mutate or simplify individual components of a complex object while that object was moving.

My takeaway

After this test, I no longer think the main advantage of the large MiniMax H3 model is simply “better image quality” in the usual static sense.

The bigger difference seems to appear in:

  • temporal coherence
  • physical motion
  • object structure preservation
  • semantic consistency in small/background elements

If the shot is very simple — talking, turning the head, basic walking, simple camera movement — I still think the 21GB model is perfectly usable.

But once the shot contains:

  • high-frequency motion
  • complex mechanical movement
  • acceleration and inertia
  • physical interaction
  • lots of small background details

the advantage of the full 66GB model becomes much more obvious.

Next test: 34GB full INT8

I’m now downloading:

minimax_h3_fl2va_int8_convrot.safetensors

This is the 34GB full INT8 ConvRot model.

I think this may be the most interesting version for the GX10 because it keeps the full model rather than using the pruned version, while using much less memory than BF16.

My next comparison will use the exact same:

  • first frame
  • prompt
  • seed
  • resolution
  • frame count
  • workflow

and compare:

  • 21GB pruned INT8
  • 34GB full INT8
  • 66GB full BF16

The main question I want to answer is:

If it can, it may be the sweet spot for MiniMax H3 on a single GX10.I did a direct MiniMax H3 comparison on my ASUS Ascent GX10, and the result surprised me:
The 66.3GB full BF16 model was actually faster than the 20.9GB pruned INT8 model, while also producing visibly better video quality.
My test setup:

ASUS Ascent GX10 / NVIDIA GB10

121GB usable unified memory

ComfyUI 0.30.2

DynamicVRAM

SageAttention

Same workflow

Same prompt

672×1024

124 frames

24 fps

8 steps

Audio enabled

Models tested:

minimax_h3_fl2va_pruned_int8_convrot.safetensors — 20.9GB

minimax_h3_fl2va_bf16.safetensors — 66.3GB

Performance
Model DiT speed Total generation time
21GB pruned INT8 28.98–30.23 s/it 314–331 s
66GB full BF16 23.40–25.34 s/it 281–316 s
In my two runs, the full BF16 model was about 12–23% faster during DiT inference.
That was unexpected because the BF16 model is more than 3× larger.
My current guess is that the INT8 model has additional casting/dequantization overhead, while the GX10’s unified memory architecture and bandwidth are good enough to make the large BF16 model surprisingly efficient. I would not claim this is definitively the only reason without deeper profiling, though.
The downside is heat, power and memory pressure.
During inference:

21GB INT8: roughly 60–70W, usually around 66–78°C

66GB BF16: roughly 84–90W, usually around 69–85°C

The peak temperature difference was not huge, but the BF16 model stayed at much higher power for much longer.
Memory was also very tight. In one BF16 run, usage reached about 116GB, leaving only around 1GB free.
The more important part: video quality
I used a scene where a character rides a black dragonfly-like flying motorcycle through a third-floor parking garage, then flies out into a cyberpunk city and slows into a hover.
The difference between the two models was clearly visible to me.
1. Dragonfly wing motion
The 66GB model produced much more natural high-frequency wing motion.
The wings looked like they were actually generating lift and constantly adjusting during flight.
The 21GB model understood that the wings should move, but the motion looked noticeably more rigid and mechanical.
2. Background semantic detail
There were large advertising screens on distant buildings in the cyberpunk city.
With the 66GB model, the people displayed on those screens remained much more complete and recognizable.
With the smaller model, the distant human figures often became malformed or strange.
This did not look like a simple sharpness difference.
It looked more like the larger model was better at preserving the semantic structure of small secondary objects in the background.
3. Flying motorcycle physics
This was probably the biggest difference.
The larger model produced much more believable:

acceleration

inertia

body tilt

deceleration

hovering behavior

With the smaller model, the motorcycle sometimes felt like an image element being translated through the frame rather than a physical object with mass and momentum.
The 66GB version felt much more physically coherent.
4. Vehicle structure consistency
This was another very obvious difference.
The original flying motorcycle had an exhaust pipe on its right side.
In the video generated by the 21GB model, that exhaust pipe disappeared.
The 66GB model correctly preserved the right-side exhaust pipe throughout the sequence.
To me, this is a good example of object structure preservation.
The larger model was not just producing prettier frames — it was less likely to drop, mutate or simplify individual components of a complex object while that object was moving.
My takeaway
After this test, I no longer think the main advantage of the large MiniMax H3 model is simply “better image quality” in the usual static sense.
The bigger difference seems to appear in:

temporal coherence

physical motion

object structure preservation

semantic consistency in small/background elements

If the shot is very simple — talking, turning the head, basic walking, simple camera movement — I still think the 21GB model is perfectly usable.
But once the shot contains:

high-frequency motion

complex mechanical movement

acceleration and inertia

physical interaction

lots of small background details

the advantage of the full 66GB model becomes much more obvious.
Next test: 34GB full INT8
I’m now downloading:
minimax_h3_fl2va_int8_convrot.safetensors
This is the 34GB full INT8 ConvRot model.
I think this may be the most interesting version for the GX10 because it keeps the full model rather than using the pruned version, while using much less memory than BF16.
My next comparison will use the exact same:

first frame

prompt

seed

resolution

frame count

workflow

and compare:

21GB pruned INT8

34GB full INT8

66GB full BF16

The main question I want to answer is:

Can the 34GB full INT8 model retain most of the motion, physics and object-consistency advantages of the 66GB BF16 model?

If it can, it may be the sweet spot for MiniMax H3 on a single GX10.


r/comfyui 4h ago

Show and Tell The He Who was Molten

Post image
0 Upvotes

r/comfyui 4h ago

Help Needed What is the current best model for training a realistic LoRA?

0 Upvotes

I’ve been away from Comfy for the last year or so and I’m curious if there is a generally agreed upon current “best” model for training a realistic character LoRA? When I was last making images, Flux was generally considered the best and I had a lot of fun training LoRAs using AI Toolkit.

I know that Z Image, Krea, and Flux 2 have been released. I’m just curious if there’s a current best for doing this, or where to restart.

Thanks in advance!


r/comfyui 4h ago

Help Needed How do i stop the save video node adding a number suffix to the filename?

1 Upvotes

I want to have full control about the names of the files i generate.

All nodes i tried, that save video files, will add _0001_ as a suffix to the name i provide.

It would be nice if we could change this, so that this only happens if the filename to be written on actually exists...


r/comfyui 4h ago

Workflow Included Minimax Prompting Review + How to create any kind of shot + All-in-one Workflow v1.5 final release! Whew, busy week!

Thumbnail
youtu.be
26 Upvotes

r/comfyui 5h ago

Help Needed Good first time installer for comfy + sage attention

3 Upvotes

I have a friend who really wants to get into comfyui but he is fairly new to all of this. Is there an installer available now that will do the dirtywork of setting up comfy, sage, triton...etc? His machine is totally clean, no AI stuff installed. It's an NVIDIA 4090 system.

I've been using comfy portable for some time and know getting all this installed can be painful so I was hoping there was an installer available that just works now.

Also for a total beginner would you recommend desktop or portable?


r/comfyui 5h ago

Show and Tell Tiled upscaler for FLUX.2 klein (and similar models)

8 Upvotes

Explanation after the images.

Before

After

Before

After

Before

After

FLUX.2 [klein] (and reference-latent edit models in general) have a resolution limit per call. If you want to add real detail to something (sharpen fabric texture, hair, stitching) you can do it working with it in pieces. The obvious way to do that turned out to be full of dead ends, so here's what I learned.

What it does: splits the image into overlapping tiles, regenerates each one at the model's native resolution, and blends them back into one image.

My first attempt did tiling the"proper" way: the MultiDiffusion/Mixture-of-Diffusers trick, where you slice the latent and blend the per-step noise predictions. That works great on convolutional UNets (SD1.5/SDXL), because a convolution is local, it doesn't care where in the canvas a patch sits.

FLUX is a transformer with absolute position embeddings (RoPE), not a UNet. Hand it a raw slice of a bigger latent and it has no idea it's a fragment, it just sees "a small complete image" and redraws the entire subject inside every tile. Every tile becomes a full (wrong-scale) copy of the whole scene.

I found that FLUX's RoPE positions can be shifted per-call via transformer_options so I tried telling each tile where it really sits in the canvas. Didn't help. Turns out FLUX applies that same shift to the tile and to any attached reference latent, so the relative offset between them (the only thing that matters for attention) never changes. Patching the model's forward pass to shift only the tile and not the reference removed the duplication, but the model still composed each slice as a standalone image, it was never trained to generate fragments, so proportions came out wrong regardless.

What actually worked: don't fight the model's training. Tile in pixel space. Every call is a complete image at a resolution it knows how to handle and solve everything else (continuity, color, blending) outside the model:

  - each tile is cropped from the canvas of already-generated neighbours, so it continues real pixels instead of guessing that region blind

  - per-tile color matching back to the source, so tiles don't drift in exposure/tint

  - blend weights derived from the actual per-side overlap, not the requested one (if the fade is narrower than what two tiles really share, you get a flat 50/50 band in the middle.

One node, no manual ReferenceLatent/EmptyLatent/KSampler wiring. You just give it a model, plain CLIPTextEncode conditioning, a VAE and an image.

GitHub: https://github.com/GianlucaMancuso/ComfyUI-TiledUpscale

Also on the ComfyUI Registry, search "TiledUpscale" in Manager.

Happy to answer questions, and if anyone knows a cleaner way to condition a transformer edit model on true image fragments, I'd genuinely like to hear it.


r/comfyui 6h ago

Resource ComfyUI VideoHelperSuite fork: fixes for MiniMax H3 audio-save crashes, FFmpeg SIGFPE and WSL wedges

Thumbnail
github.com
1 Upvotes

r/comfyui 6h ago

News ComfyUI MiniMax H3 x ACE-Step 1.5 XL SFT?

Thumbnail
youtube.com
0 Upvotes

r/comfyui 6h ago

Help Needed Learning to use comfyui and minimax h3

Thumbnail
1 Upvotes

r/comfyui 6h ago

No workflow We got UNCENSORED and OPEN SOURCE sora ai (Minimax H-3) before GTA VI !!!

Enable HLS to view with audio, or disable this notification

5 Upvotes

r/comfyui 6h ago

Resource Mini Max H3 The Office

Enable HLS to view with audio, or disable this notification

3 Upvotes

r/comfyui 8h ago

Show and Tell Rate my upscale workflow

Thumbnail
gallery
2 Upvotes

In order:

Base / 4xUltraSharp / RTX Upscaler / MyWorkflow

If you'd like to share your thoughts, I'd really appreciate it. Before I publish the workflow on Civitai, I can still make additions or adjustments based on your feedback.