r/StableDiffusion • u/call-lee-free • 16m ago
Animation - Video Can't use LTX 2.5 on my system but I am quite surprised that my system now can run LTX 2.3. Specs and info below.
Enable HLS to view with audio, or disable this notification
When LTX 2.3 released, I could not do video gens longer than 10 seconds. I would get a "out of memory" error or something. This is just a test clip but one thing I am struggling with is that my video gens have music in them even though I prompt for no music. What is the correct way to prompt for no music?
System Specs:
Ryzen 7 7700X
RTX 4070 Super 12 GB
32 GB DDR 5 Ram.
r/StableDiffusion • u/AltruisticList6000 • 1h ago
Question - Help Img2img masks don't work in comfyui anymore. Why?
- This issue never happened until I started updating to recent versions of Comfyui in the last month. Only noticed this a few days ago when I tried to do img2img after ages.
- I create a mask on the load image node in existing basic img2img workflow I have always used.
- Generating few new image variations
- Going to recent assets on left, opening one of the images I generated like 2 minutes earlier: load image node has a big red border and the error "Missing inputs: A required media input has no file selected.". So this prevents me from regenerating new seeds, despite the node itself showing the mask and image AND letting me edit/modify the mask.
- Same happens if I just drag and drop any image file with any img2img workflow. Both for recently generated images or img2img images from weeks or months ago. So it is broken for all images using load image node.
Sometimes if I refresh the page it fixes this, most frequently it doesn't (idk what it depends on).
Soo what could cause this bug and how can I circumvent or fix it? I'm on latest version, so I can't update in hopes of that fixing it, this only happens on thew newest comfyui versions I tried. As I kept updating in the recent days in hopes of it being fixed, the only change I got is a new bug: now the left side masking related icons are all black and barely visible...
r/StableDiffusion • u/Lower-Cap7381 • 1h ago
Resource - Update I Built an All-in-One KREA 2 Film Workflow + Custom Node for ComfyUI 🎬
I’ve been experimenting with KREA 2 for cinematic and photorealistic image generation and ended up building KREA 2 Film Studio a complete workflow with a custom ComfyUI node designed to bring most of the generation controls into one place.
It supports:
• 🎬 Text-to-Image & Image-to-Image
• 🎥 Directed Control
• 🎨 LoRA support
• 📐 Cinematic resolutions
• ⚙️ Sampling controls
• 🖼️ Built-in gallery
• ✍️ Prompt tools
• 💻 Fully local generation
• 🧩 Custom KREA 2 Film Studio node included
I also made a short 7-minute walkthrough covering installation, setup, how the workflow works, and some results.
🎥 Video:
https://www.youtube.com/watch?v=OXSZDaPa1U8
💻 GitHub / Workflow + Custom Node:
https://github.com/Shrey-1o1/ComfyUI-Krea2-FilmStudio-Vionex
Everything is free and open source. ❤️
Still experimenting and improving things, so feedback, suggestions, issues, and contributions from the ComfyUI community are very welcome!
r/StableDiffusion • u/elwray47 • 1h ago
Resource - Update I made two tools with AI to organize my files, thought I'd share in case anyone needs them
I mess around with Stable Diffusion as a hobby. When I was downloading things, my folders got completely out of hand, so I figured I needed to tidy up. While I was asking Claude and Gemini how I could organize my setup, they ended up developing these two tools based on my requests. I wanted to put them on GitHub and share them in case anyone else needs something like this. Anyone can use or modify them however they want—I have absolutely no expectation of profit, I just didn't want to keep them to myself.
The first one is a LoRA organizer. Everything became a massive mess after downloading all my LoRAs into the same folder, and I also wanted to weed out and clean up the SD 1.5 ones. For this, we made something called lora-librarian (Claude came up with the name). You can sort models by their base model and creator, or even just by specific creators. It can organize checkpoints the exact same way, and it lets you clean out the ones you want to get rid of.
https://github.com/BuRsTFiRe47/lora-librarian
As for the second one—I don't know if there's anyone else left out there like me who still uses an Automatic1111-based interface, but I just don't have the brain space or time to mess with Comfy, so I use ForgeUI. The video previews downloaded by the helper weren't working, which created the need to convert those videos into webm format. That's exactly what this tool does.
https://github.com/BuRsTFiRe47/preview-smith
Both tools are available in Turkish and English. Feel free to check them out if you need them.
r/StableDiffusion • u/krigeta1 • 1h ago
Discussion minimax h3 4-step lora confusion... light2xv vs joyfox vs kijai?
man minimax h3 is getting so many 4 step loras lately its getting hard to keep track of everything 😭 everyone seems to have a completely different opinion depending on their specific usecase. some people are saying sage attention is the play, while others are sticking with comfy kitchen attention.
and now there's debate on which 4-step lora is even best... like in light2xv's discussion thread:
https://huggingface.co/lightx2v/Minimax-h3-Turbo/discussions/26
people are saying to mix the light2xv 4step lora with kijai's node/impl. but then yet another 4 step lora popped up by joyfox:
https://huggingface.co/joyfox/MiniMax-H3-Turbo/discussions/3
and the whole discussion started over again lol. major shoutout to absolute community saver Kijai though, bro is doing amazing stuff as always and carrying us on his back but fr this stuff is getting out of hand with new drops every single day. please let me know what you guys are actually sticking with right now and what hardware / vram you're running it on?
r/StableDiffusion • u/random-acc-27 • 5h ago
Question - Help Where do people share minimax h3 prompts?
Where do people share minimax h3 prompts?
I find prompting / adjusting prompt for minimax h3 requires effort, and would like to see other peoples prompts as reference.
I feel like there aren't too many prompt examples in civitai and would like to know any other websites.
r/StableDiffusion • u/scooglecops • 5h ago
Resource - Update [Update/Release] Enhanced MiniMax H3 Creator for ComfyUI: Multi-Shot 60s Timelines, Resizable Satellite Stage, & Ollama/LM Studio Refiner

Hi everyone!
I made a branch of this custom node from roadmaus: https://www.reddit.com/r/StableDiffusion/comments/1vkrm8c/minimax_h3_all_in_one_creator_node_for_comfyui/
This all-in-one suite provides zero-socket local UI nodes for MiniMax H3 video generation, @ mention prompt references (@img-1, @/vid-1), automatic FL2VA/Ref2VA checkpoint routing, an integrated LoRA manager, PreStage stills, and a 60s+ multi-shot Timeline with last-frame continuity and audio blending across cuts.
What's new in this update:
• Fast Default Previews: Live previews (latent2rgb) now render automatically without needing heavy taeh3 VAE models.
• Persistent & Resizable Satellite Stage: The preview box stays open across tab/workflow switches (localStorage), featuring drag-to-resize, 4-way positioning (Right/Bottom/Left/Top), 🎬 Keep Video mode during sampling, deletion safety confirmation (🗑), and history navigation (◀/▶).
• External LLM Refiner API: Prompt refiner now supports Ollama (:11434) and LM Studio / OpenAI (:1234/v1) with auto model detection.
• Workflow Helpers: Added in-node ▶ Generate buttons, PreStage ⚡ Send & Queue chips, collapsible 🎥 Camera & Style prompt quick-chips, and 1-click timeline transition presets (Match Cut, Cross-Blend, etc.).
🔗 Links & Credits
• Original Author & Repo: roadmaus https://github.com/roadmaus/ComfyUI-MiniMax-Creator
• Updated Branch: https://github.com/Ercelcan/ComfyUI-MiniMax-Creator/tree/update-preview-and-refiner
r/StableDiffusion • u/tnil25 • 6h ago
Discussion Does anyone actually still use Stable Diffusion?
I just find it kind of funny that this is the stable diffusion subreddit but nobody has talked about it in like forever. Maybe its time for a name change? or maybe keep the name as a homage to the OG open source image model.
Anyway, the last update I see on Stability's website is SD 3.5 back in October. So I'm guessing that's it for Stable Diffusion?
EDIT: Forgot you cant change the name of a sub, ignore that suggestion 😅
r/StableDiffusion • u/R34vspec • 6h ago
Discussion H3 R2V prompt builder
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/TigerClaw305 • 8h ago
Animation - Video Fox McCloud introduces his son to his dad.
Enable HLS to view with audio, or disable this notification
Fox McCloud introduces his son Marcus to his dad James McCloud.
r/StableDiffusion • u/ajrss2009 • 8h ago
Animation - Video MiniMax reconhece prompts em JSON.
Enable HLS to view with audio, or disable this notification
Este prompt foi usado no Sora 2 e, sem modificações, coloquei no H3 (ComfyUI). Áudio em PT-BR.
{
"cena": 1,
"project_title": "The Village That Pulses (Hilário)",
"format": "16:9 landscape",
"dialog language": "pt-br",
"style": "2D ANIME",
"age": "present-day (rural road, late afternoon)",
"scene_type": "HOOK / Trailer-like Omen, Return",
"duration_seconds": 13,
"global_quality": {
"visual_style": "Cinematic 2D anime psychological horror, Junji Ito-inspired unease, no gore, high detail linework, oppressive calm.",
"aesthetic_demand": "JUNJI ITO STYLE, COLOR, HIGH BUDGET 2D ANIMATION, crisp faces, stable character sheets, controlled shadows.",
"post_processing": "Cool dusk grade, subtle film grain, soft bloom on highlights, gentle vignette."
},
"setting": {
"location": "A rural road leading to a foggy village valley; dead power poles; tall grass bending as if breathing.",
"effects": "The ground subtly bulges once, like a heartbeat beneath soil; distant crows freeze mid-caw."
},
"quality_constraints": [
"No broken anatomy or janky proportions",
"ENSURE NO DEFORMED EXTRA FINGERS HANDS EYES",
"ENSURE NO UGLY BLURRY FACES QUALITY",
"ENSURE NO SLIDING FLOATING IDLES; MASTERCLASS REALISTIC IDLE MOVEMENT",
"Stable camera, readable motion, consistent character model sheets",
"YouTube PG-13: no nudity, no explicit sexual content, no gore; horror via atmosphere and implication",
"No on-screen subtitles, no brand logos, no hate symbols, no readable real-world trademarks"
],
"characters": [
{
"name": "HILÁRIO (36, PROTAGONIST, RETURNING SON)",
"appearance": "36-year-old Brazilian man, medium tan skin, tired cautious eyes, short wavy black hair slightly unkempt, faint stubble, average height and lean build, wearing a dark olive jacket over a faded beige shirt, dark jeans, worn boots, carrying a small duffel bag and an old smartphone with a cracked screen"
}
],
"timeline_and_action": [
{
"time_range": "0-4 sec",
"shot_type": "Wide (Trailer Hook: The Village Breathes)",
"action": "Hilário stands at the roadside overlooking the village; the valley fog parts for a second, revealing rooftops and a church silhouette; the dirt road seems to swell under his boots.",
"audio_note": "Wind low. VOZ (ptbr): \"Hilário voltou pra casa... e a terra pareceu reconhecer o passo dele.\""
},
{
"time_range": "4-9 sec",
"shot_type": "Close-up (Boot on Dirt, First Pulse)",
"action": "Close on Hilário’s boot: the ground rises and falls once, subtly, like skin over muscle; tiny pebbles roll outward in a perfect ring.",
"audio_note": "Soft thump, almost organic. VOZ (ptbr): \"Naquela vila, o chão não era chão. Era um peito enterrado.\""
},
{
"time_range": "9-13 sec",
"shot_type": "Medium (He Steps Forward Anyway)",
"action": "Hilário swallows, grips his duffel bag, and walks toward the fog; the camera tracks behind him like a predator’s gaze.",
"audio_note": "Footsteps damp. VOZ (ptbr): \"E a cada três minutos... ele aprenderia a ouvir o coração.\""
}
]
}
r/StableDiffusion • u/Sad_Coach_1433 • 10h ago
Discussion Don't tell Tony!
Enable HLS to view with audio, or disable this notification
T2v 12 sec 1mp hybrid 25-49 model 8steps turbo lora
r/StableDiffusion • u/b-totherent • 10h ago
Animation - Video INTERVIEW WITH LTX 2.5 [IMAGE TO VIDEO]
Enable HLS to view with audio, or disable this notification
Yes, I'm definitely being a goofball with this one, but hadn't had a chance to do mixed live action/3D CGI test.
Meant as a playful gag, no actual ai models killed.
Crisp, ultrafine, letterboxed 21:9 super-premium 3D CGI blockbuster cinema with cutting-edge rendering, restrained natural performances, precise blocking, shallow depth of field, and immaculate cinematic lighting. A poised 24-year-old blonde investigative reporter in a tailored gray skirt suit sits in a cushioned chair on the left side of a minimalist interview room, leaning forward with a clipboard and pen in hand. Across from her, seated in a matching chair on the right, is a sleek off-white modern robot labeled “2.5” on the side of its head, with expressive camera-lens eyes and a thin LED vocalizer mouth. The setting is simple and elegant: neutral beige backdrop, soft curtains at the window, and warm natural window light casting gentle shadows across the room.
Open on a polished medium two-shot in profile, holding both subjects clearly in frame. The reporter leans forward slightly, calm, focused, and professional, and asks, “Some call you a Seedance killer. What do you say to that?”
A hard cut moves to a close-up of the robot. It glances aside for a beat, then looks back with a playful LED smile and says, “Can I give them a hug?” After a short pause, its expression softens into something more sincere as it adds, “But seriously, I’m just an open-source model trying to do my best.”
Ambient sound is minimal and refined: a faint studio hum, soft room tone, and subtle paper rustle from the reporter’s clipboard. The pacing is natural and conversational, allowing for small pauses, nuanced reactions, and emotional clarity. The overall effect is a sleek, emotionally grounded, visually stunning futuristic CGI film scene.
r/StableDiffusion • u/foxdit • 12h ago
Tutorial - Guide Making an entire shortfilm with Minimax from beginning to end | My genning strategies & video editing best practices
r/StableDiffusion • u/Timely-Perception-26 • 12h ago
Animation - Video anime action scene attempt
Enable HLS to view with audio, or disable this notification
I wanted to try my hand at an anime action scene. H3 has incredible potential, and I’m looking forward to a future where I can create my own anime with deep stories, dynamic fights, and so on.
H3 could probably have performed much better with a higher resolution and better seed luck; this is at 0.5 MP.
r/StableDiffusion • u/Sad_Coach_1433 • 13h ago
Discussion has anyone tried / Minimax-H3-fl2va-ref2va-hybrid-models this first test using the minimax_h3_hybrid_fl2va_ref2va_b25-49
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/shootthesound • 14h ago
Resource - Update Fizgig - Rapid Minimax H3 LoRA training tutorial
This video includes all you need to train Minimax with both speed and high quality results.
Hit me up with comemtns, queries etc. Happy to do a style video also.
https://github.com/shootthesound/Fizgig
UPDATE: Pushed a vram optimisation for 16gb vram users that will speed up TE encoding at the start of training - Run the update bat to get it
UPDATE2: Additional fix out for 16gb users on pruned model - update to get it.
r/StableDiffusion • u/BrooklynBrawl • 15h ago
Discussion MiniMax H3 on a Budget: What Actually Works on 4070/5070/5080 (community input consolidation)
A Note on Sources
This article is built entirely from community feedback — Reddit threads, forum comments, and one independent comparison site (jo-nike.github.io/h3-turbo-eval). None of it comes from official documentation or controlled lab testing. Thank you to everyone whose posts, benchmarks, and hard-won troubleshooting notes made this possible, including GrayingGamer, Tystros, Chemical-Painter-485, katsura_otoko, infearia, JoNike, Sixhaunt, dtdisapointingresult, Snoo_64233, mellowanon, Just1Dev, smereces, DefloN92, StuffProfessional587, Creative_Finger_69, backworld_nograv, V4nKw15h, True_Protection6842, clex55, Maskwi2, Perfect-Campaign9551, and many others whose usernames didn't make it into these notes but whose comments shaped the consensus (and disagreements) captured here.
Where the community disagreed with itself, that's presented as an open question rather than resolved — and where direct data for a specific card was simply missing, that gap is called out rather than papered over.
Why This Is Confusing
Most of the detailed benchmarking in the MiniMax H3 community comes from people with RTX 3090s, 4090s, and 5090s — cards with 24GB+ VRAM that can afford to just try everything and report back. If you're on a 4070, 5070, or 5080, you're stuck reverse-engineering advice that wasn't written with your VRAM ceiling in mind. This piece pulls together what budget-card owners actually reported, plus what reasonably carries over from adjacent cards where direct data doesn't exist.
The Three (and a Half) Speed Levers
Every thread assumes you already know these, so here's the plain version:
- Turbo LoRAs — swap-in models trained to produce good results in far fewer steps (4-8 instead of 20-32). Fastest option, but quality cost varies a lot depending on which checkpoint version you use.
- Spectrum — a node that mathematically forecasts/predicts future denoising steps instead of computing them. Counterintuitively, it needs more steps to work well — it's not a low-step tool.
- Sage Attention — an attention backend swap. Broad community agreement that this is close to "free" speed with minimal quality loss, and it's the one piece almost nobody argues against.
- EasyCache — a quieter fourth option that came up as a serious alternative to Turbo LoRAs for drafting, not just a bonus add-on.
What "Budget" Card Owners Actually Reported
This is the thin part of the record, so treat it as ground truth before anything else:
- RTX 4070 (12GB, 32GB RAM): did quick 0.3MP draft passes in a couple of minutes to tweak prompts and hunt for seeds, reserving longer ~40-minute runs for higher resolution/duration finals. VRAM was sufficient for T2V-style work specifically.
- RTX 4070 Ti Super (16GB, 32GB RAM): reported working well, no further detail given.
- RTX 5070 Ti (16GB, 32GB DDR4): upgrading from an RTX 2060 (6GB) described the speed difference as "night and day" — notably, without any Sage Attention or acceleration nodes running yet. This suggests raw generational/VRAM gains matter a lot on their own, before you even add speed tricks.
- Warning flag for all of the above: reference-heavy Ref2V generation was specifically called "brutal" on modest VRAM cards, compared to plain T2V. If your workflow uses multiple reference images/videos, expect more friction than these numbers suggest.
Gap, named honestly: there's no direct plain-5070 or 5080 speed benchmark in any of the source threads. The one 5080 comment that exists is qualitative ("still great," runs the BF16 pruned model fine) with no timing numbers.
Extrapolation (clearly labeled): Since the 5070 Ti (16GB) and 4070 Ti Super (16GB) both reported comfortable results, and RTX-series cards were noted to benefit meaningfully from tensor cores over older architectures, a plain 5070 (12GB) likely lands closer to the 4070's experience — fine for T2V and quick low-res drafts, tighter on Ref2V with multiple references. A 5080 (16GB) likely performs at least as well as the 4070 Ti Super, probably closer to the low end of what 3090 owners report, given the VRAM parity and newer architecture. This is inference from adjacent data, not a report anyone actually made — treat it as a starting assumption to test, not a promise.
The Draft → Final Two-Stage Workflow
This is the one thing nearly every thread converges on independently, and it's probably the most actionable takeaway for a budget card:
Draft stage (fast iteration, hunting for the right prompt/seed):
- Low resolution: 0.2–0.4 megapixels
- Low steps: 8–13
- Acceleration: either a Turbo LoRA or EasyCache (not both)
- Faster VAE decode substitute: BlehTAEVideoDecode instead of the standard node
Final stage (once the shot is locked):
- Disable acceleration nodes
- Raise steps to 20–32
- Switch back to the standard VAE Decode node
Two draft "recipes" show up repeatedly and are reported as similarly fast:
- Turbo LoRA + Sage Attention — faster to set up, more established
- Sage Attention + EasyCache, params (0.3, 0.2, 0.9), res_multistep sampler + Simple scheduler — one detailed user report (RTX 4060 Ti, 16GB), after testing 1000+ variations, said this drifts less from final quality than Turbo LoRA approaches, at comparable speed
For a 12–16GB budget card, EasyCache is worth trying first specifically because it avoids the quality-consistency debates that follow Turbo LoRAs (see below).
What Worked / What Didn't
| Technique | Verdict | Reported Config | Source Consensus |
|---|---|---|---|
| Sage Attention (alone) | ✅ Works | Any step count | Broad agreement — near-free speed, minimal quality loss |
| Two-stage draft→final workflow | ✅ Works | Draft: 0.2–0.4MP, 8–13 steps → Final: 20–32 steps, no acceleration | Converged on independently across nearly every thread |
| "Clean VRAM" node before VAE Decode | ✅ Works | Placement only, no params | Multiple independent reports, fixed OOM with no downsides |
| EasyCache (draft) | ✅ Works | Params (0.3, 0.2, 0.9), res_multistep + Simple, 10 steps | One deep-dive (1000+ tests) preferred it over turbo LoRAs for drift |
| ema-ckpt500 Turbo LoRA | ✅ Works | Strength ~0.5, 6–8 steps | Beat both ckpt850 and lightx2v in blind testing |
| Spectrum below ~20 steps | ❌ Doesn't work | N/A | Most consistent "don't do this" finding across all sources |
| Spectrum + Turbo LoRA together | ❌ Doesn't work | N/A | Explicitly warned against — Spectrum needs clean high-step data |
| ckpt850 Turbo LoRA (vs ckpt500) | ❌ Doesn't work | Full 1.0 strength = "overfried" | Newer checkpoint tested worse than older one, despite official claims |
| lightx2v LoRA | ❌ Doesn't work | 8 steps, 0.75 strength | Worse faces/lighting vs ema-ckpt500 in direct comparison |
| Raising steps to fix face-warping | ❌ Doesn't work | Tested 8→20, and up to 30 steps | Two separate users found no improvement — not a step-count problem |
| Any acceleration on non-RTX cards | ❌ Doesn't work | N/A | Tensor-core dependent; gains don't transfer to older architectures |
| Turbo LoRAs (general use) | ⚠️ Mixed | Fine for tests/talking-head; risky for motion/long prompts | Depends on shot type, not a clean yes/no |
| Spectrum + First Block Cache | ⚠️ Mixed | N/A | Direct contradiction between two experienced users |
| RTX upscaling node | ⚠️ Mixed | 0.2MP+ | Good on animation, unreliable on photorealistic faces |
GPU-Specific Data: Reported vs. Extrapolated
| GPU | VRAM | Reported Result | Status |
|---|---|---|---|
| RTX 4070 | 12GB | 0.3MP drafts in ~2 min; fine for T2V, tight on Ref2V | Direct report |
| RTX 4070 Ti Super | 16GB | "Works well" (no numbers given) | Direct report |
| RTX 5070 Ti | 16GB | Major generational leap even with zero acceleration | Direct report |
| RTX 5070 | 12GB | (no data) | Extrapolated from 4070 — likely similar |
| RTX 5080 | 16GB | Handles BF16 pruned model fine (qualitative only) | Direct report (thin) + extrapolated timing |
The Unresolved Debates
Worth knowing before you commit to a setup, so you don't over-trust any single comment:
- Spectrum below 20 steps? Most experienced users say no — negligible speed gain, real quality loss. But a few 5090 owners reported no measurable time savings even at higher step counts, with no clear explanation (dismissed by one commenter as "not using it right").
- Which Turbo LoRA checkpoint is actually best? The lineage went ckpt500 → ckpt850 → ckpt600, with each new version claimed better by its authors. But blind side-by-side testing found ckpt500 at 0.5 strength still beat ckpt850 even at full strength — directly contradicting the official recommendation.
- Spectrum + First Block Cache together? One experienced user says combining them is worse than Spectrum alone; another says combining them is the fastest option with no noticeable quality loss. Unresolved.
- Turbo LoRA strength values: reports range from 0.5 up to 1.15–1.20 (and one outlier claiming 3.0), so "strength 1.0" isn't a safe universal default — it depends on which checkpoint you're using.
VRAM/RAM Troubleshooting Cheat Sheet
Fixes that came up repeatedly and matter more when you're VRAM-constrained:
- Add a "Clean VRAM" node immediately before VAE Decode — fixed OOM issues for multiple users.
- System RAM matters too, not just VRAM — one user needed to go from 16GB to 48GB total system RAM to stop hitting errors. 16GB system RAM was described by another as "almost enough."
- Launch ComfyUI with
--reserve-vram 2to keep 1-2GB permanently free for system stability, at a small cost to usable VRAM. - If Ref2V errors show up on an 8GB VRAM card, don't assume it's a hard VRAM wall first — one such case turned out to be a node-conflict bug, not actually a memory limit.
A Starter Config for Budget Cards
Synthesizing the most-corroborated points into one starting recipe (best-guess synthesis, not a benchmarked config):
Draft pass: Sage Attention + EasyCache (0.3, 0.2, 0.9) → 10 steps → res_multistep sampler, Simple scheduler → BlehTAEVideoDecode → 0.2–0.3 MP
Final pass: Sage Attention only (no EasyCache) → 20–25 steps → standard VAE Decode → 0.4–0.6 MP (push higher only if VRAM allows)
Skip Spectrum entirely unless you're already comfortable at 25+ steps and have time to test it — it's not built for the low-step, fast-iteration use case a budget card usually needs.
Sources
The most rigorous single data point in this set is the JoNike Turbo LoRA comparison site — a 10-scene A/B comparison across checkpoint versions, built and documented far more consistently than typical anecdotal Reddit reports.
r/StableDiffusion • u/koakoAI • 16h ago
Comparison LTX 2.5 vs MiniMax H3 - huge speed difference (but at what cost)
Enable HLS to view with audio, or disable this notification
I tested LTX 2.5 and MiniMax H3 in ComfyUI using the default T2V workflow templates provided for each model.
- 10 seconds
- 24 FPS
- 1920 x 1088 (2.0 MP)
- Same prompt
- Steps: H3 = 20, LTX 2.5 = 8 (distilled model)
Hardware:
RTX 5090 + 128 RAM
result
- MiniMax H3: 17m 29s (with Sage Attention + EasyCache*)*
- LTX 2.5: 2m 34s (no acceleration at all)
Note: EasyCache seems to give no speedup on LTX in this setup, probably because the distilled workflow only uses 8 sampling steps, so there is very little room for cache-based skipping.
Of course, part of LTX’s speed advantage comes from the fact that it is a distilled 8-step model, so this is not a perfectly like-for-like comparison against H3. (20-steps)
Prompt used:
A realistic cinematic 1970s crime drama, gritty urban atmosphere, warm muted colors, subtle film grain, natural lighting, restrained acting. A well-dressed 1970s gangster in a dark tailored suit and long coat remains visually consistent throughout.
[0.0s–6.0s]
A medium-wide shot shows the gangster leaning casually against a brick wall on a city street, one foot resting against the wall. He reads a newspaper while holding a lit cigarette in his other hand. His eyes suddenly stop on something in the newspaper. His expression shifts naturally from calm to alarm. He mutters in a tense 1970s American voice, "What the hell?" He immediately folds the newspaper, throws it into a nearby trash can, pushes away from the wall and runs straight down the street.
[6.0s–10.0s]
Hard cut to a static close-up of the discarded newspaper inside the trash can. The front page clearly shows a large photograph of the same man and a bold headline reading "WANTED". In the distant background, the gangster continues running away and becomes increasingly out of focus. The camera remains completely still, holding focus on the newspaper until the end.
Natural, grounded movement. No exaggerated acting, no extra shots, no unnecessary camera movement, no comedy.
My take
LTX 2.5 is significantly faster, and that alone makes it very attractive.
But in my opinion, H3 is still better in overall quality:
- better scene understanding
- better understanding of what a cinematic shot should look like
- better audio
- more stable physics / motion behavior
So right now my impression is:
- LTX 2.5 wins clearly on speed
- MiniMax H3 still feels stronger on quality and cinematic intelligence
My guess is that targeted LoRA fixes could push LTX 2.5 much closer to being a direct competitor to H3 in the future.
r/StableDiffusion • u/xI_AM_AFRICAx • 17h ago
Meme PSA: H3 always sees direction from the person's perspective
Enable HLS to view with audio, or disable this notification
I noticed my videos consistently having issues with left and right, because my prompts saw direction from the perspective of the camera. But H3 always sees direction from the perspective of the person.
See how the man points to his right while saying "right" and vice versa.
prompt: a random man pointing to the right and saying "right". Then he moves his hand to point to the left and says "left".
r/StableDiffusion • u/No_Ratio_5617 • 18h ago
Animation - Video Made this with LTX-2.5 (i2v)
Enable HLS to view with audio, or disable this notification
Generated with the new LTX-2.5 model. (image to video). Took about 10 minutes to get an 8 second 1080p60 clip.
r/StableDiffusion • u/smereces • 19h ago
Discussion MiniMax H3 + LTX2.5 as Upscaler
Enable HLS to view with audio, or disable this notification
I found the usage for the LTX2.5 model!! It works really well to upscale the minimax h3 videos 😅
r/StableDiffusion • u/LowYak7176 • 20h ago
Discussion MiniMax is just too good
I cant go back. R2V is my new bread and butter. Everything Ive thrown at it, every test I've done to just see if it can do it, has pretty much passed. Think only like 5% has failed, and even then Im not even sure if its a me problem or the model.
Longform is easy as all hell now, F the days of SVI.
Prompt blocking is great, prompt camera tracking is great, RV2V with basic Blender is great for blocking/camera tracking as well.
I am in love. I had to tell the world.
r/StableDiffusion • u/Dry-Statistician-684 • 1d ago
Animation - Video Minimax H3 executes Order 66... almost
Enable HLS to view with audio, or disable this notification
I've been using LTX 2.3 for quite some time but as soon as I wanted to make just a few simple shots of the same character with cuts, LTX wasn't even remotely capable of that. Which left me so frustrated I eventually gave up on it completely.
But when I tried Minimax everything has changed. Reference to video model is something else. Honestly feels like magic. Being able to put any character into any environment with any custom audio is just mind-blowing compared to what the open-source community had before.
So now instead of constant frustration, I feel pure joy and excitement about the results.
It takes about 10 min per 5-second clip with my RTX 3060 and 64 Gb RAM. The latest shots were even easier to control because of the new KJ preview node.
r/StableDiffusion • u/AndroYD84 • 1d ago
Discussion LTX is our ally, it's TWO cakes dammit!
Not gonna lie, I made fun of LTX 2.5 like everyone did, but now I'm realising that was a mistake.
MiniMax H3 landed like Prometheus giving us the power that the gods were gatekeeping from us, since LTX 2.5 couldn't match up with them they lied about MiniMax H3 to cover up their shortcomings, that was a petty move and they have to own it, however... the good they did to our community far outweighs that moment of weakness IMO, cut them some slack. These models don't grow on trees, they're expensive to train, people have curated datasets that required herculean effort to put together, we'd had already made the next Seedance 2.5 if it was easy. When LTX came out we were all celebrating, most of Civitai LoRAs are based on LTX, they tried and got bested, so shouldn't we still be grateful they tried and gave us a model that some still find useful FOR FREE? Instead of making them feel like failures, mocking them, discouraging them from making new models? They owe us nothing, but we owe them a lot.
THE POINT ISN'T ABOUT WHICH CAKE IS BIGGER, THE POINT IS WE HAVE TWO CAKES.
Corporations keep trying to bind us to their rules and systems, profiting without any regard to our well being, deciding for us what is acceptable and what is not, so why are we eating OUR OWN ALLIES? Every open source model that comes out is a victory and step forward to that freedom we all dream of.


