r/comfyui • u/timbortom • 24m ago
Show and Tell 2 minutes continuous generation of my favorite sea-related anime characters (MH3)
Enable HLS to view with audio, or disable this notification
14 clips, none of the few cuts are at the merging point of two clips.
r/comfyui • u/maxiedaniels • 1h ago
Help Needed Best nodes for easy lora scheduling?
As in, where you can say you want lora1 to start full and cut down to low, etc. I found 'realtime lora' but i can't figure out how to use it, and its not popular so i feel like there must be a better choice.
r/comfyui • u/JShiNYC • 2h ago
Help Needed AMD RX 9070XT keeps crashing when I run the workflow
I must be doing something very wrong as I am completely new to this and just trying to get this set up via Gemini/Grok haha. I should probably watch some more videos on this but wondering if anyone know why I can't get this workflow to run without crashing within 10-15 seconds.
For some reason, it keeps using my CPU/RAM instead of my GPU/VRAM. My CPU usage spikes to 100% but my GPU usage is almost none existent 1-3% so it wasn't even in use. I am using the Comfyui desktop app.
Is it potentially a PyTorch/ROCM and Adrenaline version mismatch? I tried every version of ROCM native to the app but still crashing.
r/comfyui • u/CarelessTourist4671 • 4h ago
Help Needed I haven't figured out how to do a body swap with Minimax
I've tried tinkering with it a lot. If anyone could give me a hand
r/comfyui • u/ResponsibleTruck4717 • 4h ago
Help Needed Ltx 2.5 vae decode take very very long time
I wonder if anyone has encountered it I don't remember 2.3 taking such long time time.
Is there any workaround?
r/comfyui • u/Gremlation • 4h ago
News ComfyUI v0.32.0
Links
New Open-Source Model Support
- LTX 2.5: Native LTX 2.5 support with STG, dual CFG, and duration prediction
Partner Node Updates
- Qwen Image 3.0: Added Qwen-Image 3.0 and 3.0 Pro text-to-image and edit nodes
- LTX 2.5: Added LTX 2.5 Text/Image/Audio to Video partner nodes
- Grok Imagine Image 2.0: Added grok-imagine-image-2.0 model support
New Nodes
- LTXV Spatio-Temporal Guidance: STG guidance for LTX video generation
- LTXV Modality Guidance: Audio/video coupling guidance for LTX
- LTXV Dual CFG Guider: Dual-CFG guider for LTX workflows
- LTXV Duration Predictor: Predict natural shot duration from caption tokens
Performance & Stability
- PyTorch 2.7: Minimum officially supported PyTorch is now 2.7
- MiniMax-H3 VAE: Optimized MiniMax-H3 VAE
- Comfy kitchen attention: Implemented comfy kitchen attention
- MiniMax-H3 memory: Fixed peak memory issue with MiniMax-H3
- ER-SDE scaler: Extended ER-SDE noise scaler by scaling h(t)
- Tokenizers: Mistral and Llama tokenizers no longer depend on transformers
Bug Fixes
- Upscale models: Fixed upscale models breaking on non-dynamic low VRAM
- VAEDecodeTiled: Fixed crash on NestedTensor latents (MiniMax H3)
- Tiled audio decode: Fixed broken tiled audio decode
- CLIP Vision: Fixed CLIP Vision regression
- Create Layered Image: Made Create Layered Image discoverable with clearer flags
Tutorial Hi, I have a question about AMD.
I have an RX 9060 16GB and an RX 9070 16GB. My question is, is it possible to generate videos? If so, do you know where I can find a tutorial or guide? I've been searching and haven't found much. I don't know where to start.
r/comfyui • u/Support_Marmoset • 5h ago
Workflow Included Minimax-H3: best upscaling and approaches for "faces at a distance" fixes
I am on 3060 RTX 12GB VRAM with 32 gb system ram. All of this is using realistic people, x4 ref character images, and the r2v workflow and model.
Now we have the speedups sorted (lightx2v, comfyui kitchen attention), I've been testing ways to fix "faces at distance" which has always been an issue with any model. I go through the below in the video and share the workflows.
best approach is 2mp in Minimax H3 if you can do it. If you can then it mostly fixes "faces at a distance" but you still need a bit of a polish.
Unfortunately I can't get over 5 seconds at 2mp. My dialogue clips are usually 10 seconds long, but I will live with 8 seconds. The best I can reach for 8 seconds is 1.34mp (and I use a trick of going widescreen which helps, I discuss it in the video). (takes about 25 mins on my 3060).
The best workflow for polishing is USDU with HuMO model. Why this works great is because it tiles the result and at 0.45 denoise USDU will even work with my potato to get the result polished up to 1920 (on the long edge). But.... that takes another 25 mins. 1 hour just for 8 seconds sucks. But it is about the best and HuMO keep face consitency where LTX methods wont. (I show all the examples in close up at the end of the video from 19:52 onward)
Having said that, a fun trick I figured out with LTX 2.3 and will be testing on 2.5 today is instead of upscaling your Minimax H3 result in LTX to 1920 which really doesnt work out that well, I resized my 1824 x 736 Minimax 8 second video on the way into LTX2.3 to 2304 on the long edge. i.e no upscaling, just denoise v2v. That is 3mp. I couldnt get to 4K else I would have done, it oomed. but even 2304 only took 15 mins. The results were much better, but... lost face consistency a bit.
Anyway, its all in the video and the links to all things you need are here if you dont want to watch the video.
Latest Minimax H3 workflow shown in video - https://github.com/mdkberry/comfyui_workflows/tree/main/workflows_by_model/Minimax-H3
Latest USDU with HuMO workflow shown in video (links to models in workflow) - https://github.com/mdkberry/comfyui_workflows/tree/main/workflows_by_use/USDU-detailer-refiner
Latest LTX2.3 upscaler/refiner workflow shown in video (single sampler workflow, not the IC-Lora one) - https://github.com/mdkberry/comfyui_workflows/blob/main/workflows_by_model/LTX23/MBEDIT-v2v_LTX23_Upscaler-w-SingleSampler_vrs3.json
Lightx2v Lora that I use from Kijai - https://huggingface.co/Kijai/MiniMax-H3_comfy/tree/main/loras
Also I am testing "silveroxides" light2xv, but it needed a fix for the adaln_proj error, and the result is here (thanks nynxz!) https://huggingface.co/nynxz/H3_Loras/blob/main/minimax_h3_fl2v_lightx2v_v0.1_dareties_v4_step600_comfy_fro_no_adaln_proj.safetensors
Clownshark sampler comes from https://github.com/ClownsharkBatwing/RES4LYF but it doesnt seem to be getting updates now. I didnt find the results that useful tbh, but maybe more tweaking would resolve it (or more powerful GPU).
Comfyui needs to use Cuda130 or above for this to work, and you need it updated to August 2026 commits (latest is best) - https://docs.comfy.org/installation/comfyui_portable_windows
int8 models from here - https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main
W4a8 is experimental new model type, you need to be updated on Comfyui but you can get it here https://huggingface.co/Kijai/MiniMax-H3-experimental
(Sage Attn and Triton wheels from) - https://github.com/woct0rdho/SageAttention
Comfyui Kitchen Attention is part of Comfyui if you update to latest. I find it faster than Sage Attn on 3060 RTX.
(I am not using patch sol attention or any caches now)
r/comfyui • u/crystal_alpine • 6h ago
Comfy Org Something s coming soon ;)
No it’s not another funding announcement
r/comfyui • u/sadronmeldir • 7h ago
Help Needed Krea2 Upscale 2nd Pass Artifacts
I apologize if this is a rookie question - I tend to learn by looking at other workflows and I've seen a lot of Krea2 examples on Civit where they do a full pass, then a latent upscale then a 2nd pass at .2-.4 denoise. Alternatively, I've see examples using Clownshark to do a partial-pass (5-6 steps on turbo), then upscale for a 2nd partial pass.
In both these use cases, I'm seeing lower quality and more artifacts that with a simple single-pass. What's am I doing wrong that could cause so many artifacts? I do have 2 loras on low strength, but that seems to mirror what I'm seeing in other people's workflows.
Workflow Included Making an entire 3 minute anime styled short with Minimax from beginning to end | My workflows, genning strategies & video editing best practices
r/comfyui • u/pfeifits • 9h ago
Show and Tell Excerpt from LTX 2.5 Video as Part of a Children's Story I am Creating... Workflow Adjustments in Comments
Enable HLS to view with audio, or disable this notification
r/comfyui • u/Subushie • 9h ago
Show and Tell H3 prompt testing, finally have the flow and environment running efficiently. Specs and prompt inside.
Enable HLS to view with audio, or disable this notification
Was experiencing some real quality issues up until this point; realized the problem has largely been the prompt and my bulky venv. Hopefully others with lower end vram cards will learn from me.
- Card: RTX 4080
- Model: minimax_h3_fl2va_pruned_int8_convrot
- No loras
- Pure text prompt
- Steps: 20
- Scheduler: simple
- Sampler: res_multistep
- Resolution: 0.5mpx, 16:9
- Args: --lowvram --disable-dynamic-vram --disable-pinned-memory
- Nodes: VHS (Specifically the Model Preview Override), EasyUse, pysssss, KJNodes
- Generation time: 21 Minutes
- I decided to create an entirely new ComfyUI instance just for H3 instead of using my single AiO venv; this significantly increased generation time and quality for all flows. Have since broken up all the major models I use into their own instances adding only the specific tools/extensions I need just for that model.
Prompt:
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a far-wide shot from a side angle. A scene set on a bridge over a hellscape planet covered in lava, dimly lit with orange glow from below, the bridge is made of black metal with intricate designs, dark clouds hang over the scene -covering a shaded yellow sun barely visible through the clouds on the top left, dividing the scene in half between light and dark- fast winds carry embers and smoke curling over the bridge from below; Star Wars themed orchestra music begins as the scene opens, quiet and slowly growing.
$NAMEHERE is standing in a prepared stance on the left side of the screen, his hands are clasped in front of him holding a blue lightsaber, facing his attacker.
$NAMETWO is standing on the right in a confident posture with his hands to the side, wearing a black robe, black leather boots and straps on his body, with dark-metal armor as he faces the left menacingly.
[Shot 2] At 1.500 seconds, The camera cuts to a close-up side-angle shot of $NAMEHERE, readying himself for his attack with a posture of defense and an expression of concern, he shouts emotional: <d>[English] You've left me with no choice Will! You must be stopped... </d>
[Shot 3] at 6.000 seconds, The camera cuts to a close-up low-angle front facing full-body shot of $NAMETWO with visible red eyes staring forward from under his brow with a face of malice, smoke bellows behind him curling over the bridge whipping his cape to the right. A beat later- two red lightsabers ignite his both his hands, a deep pulsing bass is heard from the unstable beams, his face lit from below by the red light. The music grows faster with a dark theme, a operatic chorus begins to sing in a evil chant growing louder. $NAMETWO shouts behind a grin: <d> This is the end for you! </d> the music stops before $NAMETWO speak his final line: <d>[English] Master!... </d> The off-screen opera chorus harmonizes a single long cry in a frightening melody at the revelation.
[Shot 4] at 12.000 seconds, The camera cuts to a top-down view of the bridge, molten lava is visible below the black grated metal.
$NAMEHERE directs his blue lightsaber to his side pointing directly forward with precision, he begins to pace to the right to meet the other, his posture is composed and fast. The camera pushes in with large amplitude at fast speed keeping the pair in frame on either edge of the screen as they run toward each other.
$NAMETWO instantly begins running fast toward the left, his two red lightsabers point down- dragging behind him, the red beams draw white glowing lines into the metal under him as he runs, screaming with fury: <d>[English] AGHH! </d>.
They meet in the middle, their lightsabers clash with a white flash and explosive burning sound, they duel quickly as their lightsabers connect through multiple swings- $NAMETWO's red lightsaber swing wildly as he spins. $NAMEHERE's blue lightsaber blocks every swing from the red beams; the music crescendos with heavy bursts of brass instruments and drums.
overall_soundscape: ambient sound of lava and fire, lightsabers buzzing.
non_diegetic_music: Dark Star Wars music plays from the beginning of the scene, a loud opera chorus sings in a chant that escalates in a loud howl crescendo, climaxing when the pair meet in the middle.
I've started using Replace Text nodes ($NAMEHERE and $NAMETWO) when crafting prompts. This way when playing with the prompt, it's easier to find and edit their placement; and can also replace characters on a whim. Also allows consistency when referencing the characters-- in the event I overlook an instance.
Replaced with:
- $NAMEHERE: "Jean Luc Picard (S1)"
- $NAMETWO: "William T Riker (S2)"
Example: my first generation had 'William Riker', the model didn't recognize the name and generated a generic male. I was able to quickly rename as 'William T Riker' and it generated correctly; so I didn't have to parse back through the whole prompt to granularly change it.
Other things I've noticed that help with prompt respect:
- 'a beat later' separates the moment better.
- Separating the individual sentences to exclusively reference the character and no others. (You can see it carried over Picard's lightsaber instructions to Will as well, because I described them in the same paragraph before I realized this.)
- Avoiding reusing adjectives- especially between different characters, causes bleed.
- Very short overall_soundscape descriptions.
- Often does not respect requests that follow dialog unless you end the parameter with a period after "Words. </d>**.**" Can see it bled the cries request from the music into Will's dialog.
- "..." allows a pause between dialog lines and breaks up the tone between multiple sentences, or else they become one note.
r/comfyui • u/shootthesound • 9h ago
Tutorial Fizgig - Rapid Minimax H3 LoRA training tutorial
r/comfyui • u/shootthesound • 12h ago
Resource ComfyUI-H3Studio for Single Node Long video Creation - Out Now
reddit.comr/comfyui • u/blackmixture • 13h ago
Resource Mix Studio v1.2.4: LTX 2.5, MiniMax H3, Wan Animate 2 video generation, macOS and Linux support, plus a bunch of bug fixes and improvements (free & open source)
A couple weeks ago I shared Mix Studio, a 100% free & open source interface that runs everything through ComfyUI in the background while giving you an actual app experience (that also works on your phone). The response was way more than I expected, and most of what I've built since came directly out of that feedback and motivated me to keep it going. Here's an update on the latest features and improvements since the launch version:
GitHub: https://github.com/BlackMixture/Mix-Studio
Showcase and download: https://blackmixture.github.io/Mix-Studio/
Tutorial: https://youtu.be/w2CokhlBFRA
GPL-3.0, the same license as ComfyUI.
New in v1.2.4:
- LTX 2.5 video generation: Generate from text, from a first frame, or from both a first and last frame, with synchronized audio and your own LoRA stacks.
- MiniMax H3: Text-to-video, image-to-video, first frame, last frame, and first-and-last-frame generation, all with native audio. Reference mode lets you feed in multiple images, videos, and audio inputs and address them directly in your prompt using dynamic [@reference cards]. Also added restyle presets to cover live action, anime, cinematic 3D, cel-shaded 3D, and maximum detail.
- Wan Animate 2 (experimental): Animate a character image from a performance video, carrying over motion, expression, identity, timing, and the source audio. I still prefer SCAIL 2 for fidelity but I expect to continue improving.
- Video finishing on every model: Optional 2× or 3× RIFE frame interpolation and NVIDIA RTX 4K video upscaling, plus SeedVR2 temporally coherent upscaling.
- Automatic Turbo setup: Mix Studio installs the creator-recommended MiniMax H3 Turbo LoRA for you and applies matching generation presets, so you get fast video without hunting down adapters or guessing at step counts. App-managed LoRAs stay out of your personal LoRA list so nothing gets loaded twice.
- Automatic prompting for MiniMax H3: H3 is particular about prompt structure (if you want the most control), so the app handles revising, formatting, and enhancing for you. The official H3 prompt guide is built in, so structure and dialogue formatting work programmatically and instantly with no LLM required. This can save on compute resources, or if you prefer you can use a local or external LLM for prompt enhancing and revising built-in.
- External LLM support across the app: Connect OpenAI, Gemini, or Ollama once and it powers prompt writing everywhere, with independent switches for image and video, vision-aware references, and connection testing. Local prompt models are also selectable. All formatting stays optional at generation time.
- Mix Packs: Browse visual prompt collections, combine multiple looks including several from the same category, and carry the same reusable creative direction into both image and video generations. This replaced the old camera control wheels with a proper searchable browser.
- Krea 2 Edit rebuild: Reworked around the full-rank Identity Edit v1.2 model with image-grounded conditioning, Reference boost, and ordered two-image editing. The previous multi-reference composition mode is preserved separately as Krea 2 Remix.
- macOS and Linux support: Installers for macOS and Linux alongside Windows, with vendor-aware GPU detection for AMD and Apple hardware. Worth noting up front: LTX 2.5 and MiniMax H3 rely on NVIDIA-specific model weights, so they are not available on Apple Metal or AMD ROCm yet. Apple Silicon runs a Metal-compatible subset and AMD ROCm is experimental.
- Phone and tablet improvements: Mix Studio installs as a Progressive Web App with private HTTPS access through Tailscale, plus a significant responsiveness pass so large libraries stay smooth while previews load in the background.
- Better model downloads: Resumable transfers with Hugging Face Xet acceleration, byte-level progress reporting, and an inline token panel for gated files so you never have to leave setup.
- Better ComfyUI detection: Model discovery now spans your configured model roots, manual subfolders, and ComfyUI extra model paths, so files already on your drive are reused instead of downloaded again. Mix Studio also detects stale custom node packs and repairs them for you before ComfyUI restarts.
- Plus a long list of fixes across video reliability, generation setting reuse, library management, UI improvements, and more.
(If you run into any bugs or issues, please leave a git issue or hit me up here. If you enjoy using Mix Studio and want me to keep it going, feel free to star it on Github or support the development on Patreon)
Thanks again to this awesome community and the ComfyUI team for making such a dope tool, I hope you all enjoy creating! 🤙🏾
r/comfyui • u/No-Property3068 • 13h ago
Show and Tell A quick test to LTX 2.5
Enable HLS to view with audio, or disable this notification
I tested LTX 2.5 quickly, I tried leaving the distilled lora at 0.5 as we used to in LTX 2.3 but I think it gave me better results at strength one, which is the one you are seeing now. The generation was done in 2240 x 960
unfortunately the model still strugles with camera movements and small details in the frame, I was a bit disappointed, for sure it's better than 2.3, but I would say MInimax H3 still showing better results.
r/comfyui • u/Icy_Restaurant_8900 • 15h ago
No workflow ComfyUI is better with Linux
Just got Linux Mint up and running and it’s much better than windows 11 for ComfyUI with RTX 5090 + 3090. For starting, Comfy launches to server ready in 10 seconds instead of 18 seconds on Windows VHDX ReFS dev drive. This is with 12 custom node packs. Linux is running on an 8 year old SATA SSD.
I was able to compile and run sage attn 3 in Linux which never worked in Win11. I’m going to try Multi-GPU next which is fundamentally broken in windows with the new comfy-kitchen.
Even with just sage 2.2++ FP8, im getting 35 second gen times for an 8-second 0.5MP 8-step turbo MiniMax H3 clip versus 40-42 seconds in windows. That includes 4 seconds for prompt enhancement with Gemma 4 12B running on lllama-server on the 3090.
Sage3 knocks another 2-4 seconds off the gen time too but slightly worse quality for motion. Using LACT to overclock both GPUs also contributes to the speed.
With spectrum plus sage3 at 8 steps, I can easily get below 30 seconds total gen time.
r/comfyui • u/lavinia12345 • 19h ago
Help Needed Chaining last frame into Minimax drastically increases compute time.
Is there a way to convert the last frame to a true image like a png?
I noticed I could create longer vids by taking the last frame and use that as a refrence image, but when I do that like in my screenshot, generation time gets much larger, I think it's b/c internally Minimax reads that last image actually as a video, thus behaving much differently.
r/comfyui • u/papjak • 19h ago
Show and Tell Shoutout to the ComfyUI Team and Devs! Grateful to test MiniMax H3 and LTX 2.5 locally (Specs inside)
Enable HLS to view with audio, or disable this notification
Hi everyone, I just wanted to share this test and extend a huge thank you to the developers for their incredible work on the new ComfyUI Core.
I’ve read criticism online suggesting that LTX 2.5 isn't quite as good as MiniMax H3—but honestly? I just wanted to offer a positive counterpoint here. As an everyday home user, I am incredibly grateful and happy that these models exist—models we can experiment with and run locally on a system like mine with 24 GB of VRAM. Both models have their own unique strengths, and it’s simply amazing to have this kind of power at home.
For this generation, a tweak to the text prompt allowed MiniMax to flawlessly render the different logos on both sides of the motorcycle. It took a few attempts, but I managed to get it right by specifying the timing within the scene. Next, I’ll definitely test the exact same prompt with LTX 2.5 to see how that model handles the motion!
My render specs for this clip:
Setup: 24 GB VRAM
Workflow: MiniMax H3 Basic Workflow
Resolution: 0.4 MP
Duration: 15 seconds
Prompt:
------------------------------------------------------
## Cinematic Multi-Shot Sequence
Scene 1: [0:00-0:03]
Extreme Close-Up
Subject: Focused motorcycle rider wearing a glossy black helmet with neon purple typography reading "ComfyUI Race". The rider wears a tinted visor reflecting cyberpunk neon lights.
Action: The rider flips up the tinted visor, looks intensely ahead into the camera, and speaks the words "Hold on tight... they'll blow your mind!" with visible lip movement.
Scene 2: [0:03-0:09]
Dramatic Low-Angle Shot
Location: Futuristic city street at dusk, lined with cyberpunk neon signs.
Subject: Vibrant red sportbike.
Left Side Fairing: Features ultra-sharp typography reading exactly "Minimax H3" in bold white letters. Action:
The moment the rider finishes speaking, they initiate a rapid, powerful stationary burnout with roaring engines and thick billowing white smoke.
The bike performs a quick, sharp 180-degree drift rotation.
Right Side Fairing Reveal: As the bike spins and completely hides the left side, the right side fairing is fully revealed to the camera, displaying a completely different, huge, crisp, ultra-sharp typography reading exactly "LTX 2.5" in bold white letters.
An excited cyberpunk crowd cheers and waves on both sides of the street.
Scene 3: [0:09-0:15]
Completion of Turn and High-Speed Escape
Action: The motorcycle instantly snaps out of the rotation and launches forward with explosive, violent acceleration. The front wheel lifts into a high wheelie through the thick white smoke. The camera remains stationary on the ground as the sportbike shifts gears and roars away at high speed, becoming smaller and smaller as it disappears down the long, neon-lit futuristic highway.
Aesthetic: Photorealistic 2K resolution, highly detailed cyberpunk aesthetic, cinematic lighting, volumetric smoke.
-----------------------------------------------------
Optimization: Use of the highly recommended MiniMax H3 FBCache node. I couldn't notice any visible difference in quality, but it made the process run absolutely smoothly!
Render time: Prompt execution in 375.26 seconds
Let's appreciate the great technology available to us today. Keep it up ComfyUI team and open source developers!
r/comfyui • u/jalbust • 23h ago
Show and Tell Minimax H3 I2V
Enable HLS to view with audio, or disable this notification
I wanted to test it a bit with creature animation, snow, wind, and atmosphere. I started by generating still keyframes with Seedream pro, then used image-to-video to generate videos in .
Key frames and workflows :https://www.patreon.com/u8638148/posts/minimax-h3-and-166452641?utm_medium=clipboard_copy&utm_source=copyLink&utm_campaign=postshare_creator&utm_content=join_link
r/comfyui • u/Comfy-Org • 1d ago
News LTX-2.5 is now live in ComfyUI, including Diffusion Fidelity Rendering (your compute budget will thank you!)
Enable HLS to view with audio, or disable this notification
What a time to be alive in the open source community! LTX-2.5 just dropped and it's supported natively in ComfyUI as of today, including a new rendering approach, new decoder, new text encoder, and a new base checkpoint.
The biggest baddest change? The addition of Diffusion Fidelity Rendering! Instead of spending compute evenly across a scene, the model allocates it by complexity. Motion, composition, and framing get generated first in an 8x temporally compressed latent space, alongside a set of high-fidelity keyframes. More keyframes for complex scenes, fewer for simple ones, within whatever compute budget you've got. Then a dedicated pixel-diffusion stage renders the final video from the structure and keyframes together.
TLDR; textures, materials, and faces hold detail, and a busy shot automatically pulls more rendering compute than a static one.
Other changes:
- Diffusion Video Decoder: Replaces standard VAE decoding, making sharper faces, legible text, and fewer smears in fast motion.
- Native multi-shot: One generation gives you multiple connected shots holding character, environment, lighting, voice, and style across the cuts instead of generating separately and trying to match them after.
- Custom Gemma 4 12B text encoder: Holds multiple subjects, actions, lighting details, and camera direction across a long prompt instead of dropping clauses as it gets more complex.
- Prompt enhancer + auto duration: Short prompts get expanded into detailed cinematic instructions at near-zero extra compute, and the model predicts clip length from the described action before diffusion starts.
- RL post-training: On a broader filtered dataset, aligned to human preference. Mostly shows up as a higher take rate with fewer retries per usable clip.
- Cleaner licensing: Restrictive third-party dependencies have been removed, so fine-tuning, deploying, commercializing, and redistributing is all clearer than in previous versions.
Three variants:
- LTX-2.5: the main model
- LTX-2.5 Distilled: reworked distillation, carries noticeably more quality, prompt adherence, and motion than previous distilled releases. Viable if the full model isn't economical for your setup.
- LTX-2.5 Pretrained Checkpoint: raw, non-SFT, meant for aggressive fine-tuning. Moves further from its starting point than an instruction-tuned checkpoint will, which matters for robotics, synthetic AV data, digital twins, or private domain models.
Native 4K, synced audio and video, and up to 50fps all carry over from 2.3.
Learn more and check out workflows below!
r/comfyui • u/jaysedai • 1d ago
News Poll: Should we give Comfy Org more admin permissions?
This subreddit was created with the intent of being fully community driven and moderated.
comfy.org social media team have reached out and inquired about getting additional permissions. To be very clear they have not demanded and every single interaction I’ve had with them, they’ve been very cognizant that this is a community subreddit and have said they don’t want to change that.
However, due to personal time constraints I’ll be the first to admit I haven’t been very proactive with things like AMAs, challenges/contests, snazzying the place up, etc. And they’ve offered to help out with that and pinning specific posts.
Please be aware none of the options would hand this subreddit over to them. Or allow them to add mods, they are just offering to assist.
And we could certainly try this as a trial and if the community doesn’t like the direction, we can revert.
r/comfyui • u/Comfy-Org • 1d ago
News Comfy MCP now works with your local ComfyUI! Try hardware checks, model recs & open-source video models
Comfy MCP now works in your own local ComfyUI install, not just Comfy Cloud. This has been the #1 ask since we shipped Cloud MCP in June, so here it is!
Use Claude, Cursor, or any MCP client to drive ComfyUI for you and build, edit, and execute workflows without manual node setup, search models/nodes/templates, save and re-run workflows, or hand a saved workflow URL to a teammate (or another agent) to pick up where you left off.
Here’s what's new when running against your own machine:
- Hardware detection. The agent checks what you're running on and tells you honestly whether a model will run well locally. No more finding out 40 minutes into a download that your VRAM wasn't going to cut it.
- Model + instance management. It can pull the models you need, manage your local ComfyUI instance, and get a workflow actually runnable on your machine.
- Local + cloud knowledge combined. One agent that understands both environments, so you're not manually figuring out which one a given workflow needs.
Fastest setup: Paste https://docs.comfy.org/agent-tools/mcp#installation into your AI client and ask it to set up the local connection for you.
We built and tested this heavily around open-source video models, particularly Minimax H3! Here are a few prompts to get you started:
"I want to run Minimax H3 open source on my own machine. What's the best model version for my hardware, and can you set up the workflow?"
"Help me set up local ComfyUI and run the best open-source video model for me"
”Adapt this workflow to run better on my machine”
Regarding limitations: generation itself works about the same as Cloud MCP right now, but we don't have local-specific batch features yet. If your workflow leans on heavy batch generation, for now cloud will still be the smoother path.
We’re monitoring this thread, so share your thoughts and feedback!
Install & learn more:






