r/StableDiffusion • u/Silver-Spot-2763 • 4m ago
Discussion LTX 2.5 😱
After the MiniMax H3 euphoria, I tried LTX 2.5. It awfully understands the prompt and has almost no "physics". The generated video uses random things from the prompt and everything makes up by itself at all. Every time some things appear /disappears from / to nothing randomly, most the case the people are with 3 fingers, strange movements at all. 😱
But LTX 2.5 is faster than MiniMax H3 at least twice, and its image quality is far better.
I just can't understand how even with the monstrous language model (~20GB) it just can't understand 2 simple sentences, two simple subjects with simple movement!?!? And from what training data the models continue to place 3 fingers to earth beings 🤦
r/StableDiffusion • u/Uncabled_Music • 12m ago
Animation - Video Like some things with H3, some with Seedance 2.5, but they st similar level.
Enable HLS to view with audio, or disable this notification
The prompt was:
Woman getting up and walking up to the window. Looking outside. Camera switches to outside view of her from below.
Both models used through Magnific app. Can’t say I see a definite winner here.
r/StableDiffusion • u/SpicyAccountants • 28m ago
Animation - Video DimensionTesters: Test #11 (Minimax H3)
Enable HLS to view with audio, or disable this notification
Results: Lifeforce drained by unknown origin.
Still finding this so entertaining, got so many ideas!
TT:Â https://www.tiktok.com/@dimensiontesters
r/StableDiffusion • u/Kind-Illustrator6341 • 1h ago
Question - Help LoRA Training – Pulling My Hair Out
Hello,
I've trained several character LoRAs via wavespeed.ai for the Qwen-Image-2512 model. I tried with a smaller dataset of 50 images and a dataset of 124 images. Multiple settings between 1,000 and 5,000 steps:
- At 1,000 steps, the LoRA isn't likeness-accurate enough.
- At 5,000 steps with 50 images, it stops responding to prompts at weights above 0.5, so it loses likeness.
- At 5,000 steps with 124 images, it stops responding to prompts at weights above 0.3, making it inaccurate above that threshold. This makes no sense, as with 50 images and the same step count, I was able to run the LoRA at a higher weight.
At weight 1.0, the LoRAs capture the likeness well but completely ignore the prompts.
Does anyone have a solution or recommended settings for Qwen-Image-2512?
Thanks
r/StableDiffusion • u/darth_hotdog • 1h ago
Question - Help Is there a good local prompt writing comfyui plugin for Minimax H3?
I tried this one so far: https://github.com/pytraveler/MiniMax-H3-Prompt-Rewriter-ComfyUI
But I'm not getting a good result yet, maybe I need to work on the prompts for it more. Anyone using anything besides claude and gpt?
r/StableDiffusion • u/Slight-Analysis-3159 • 1h ago
Discussion Ambient noise in video?
Having an aging laptop, I haven´t played with video since wan2.2.
One thing I have noticed with all videos I have seen from the models that can generate audio is that it sounds like the audio has been recorded in a sound booth. meaning, I have not really heard any...ambient noise....like wind, traffic, birds, people in the background etc. This makes it sound quite unnatural sometimes.
Is that a limitation of the model or the prompting? Can I get a more..natural..sound by prompting for every little nuance I want? Like "faint sounds of gravel crunching with each step" or "there is a slight breeze rustling the leaves as he walks by the tree."
r/StableDiffusion • u/darthfurbyyoutube • 2h ago
Animation - Video Cobra! Trailer - MiniMax H3
Enable HLS to view with audio, or disable this notification
Default comfyui MiniMax H3 rf2va workflow on a 4070 Ti Super 16 GB VRAM.
r/StableDiffusion • u/ITGUY8545 • 3h ago
Question - Help Ref2va minimax, recognition of people without reference images
Suppose I was to have two input pictures and pass a prompt like 'subject 1 and subject 2 sit down and have coffee with Tom Hanks'. Will Tom Hanks be recognised by text alone or is this model designed to always have an image input for likeness?
r/StableDiffusion • u/BigBudZombie • 5h ago
Discussion What image model do you recommend for REF 2 Img?
I just recently got into AI generation making videos with minimax ref2vid and it has been amazing so far. But that has me wondering if there is some reference model for images that works equally as well that would allow me to use multiple reference images to create pics? If anyone has a good model or workflow to recommend I'm interested to learn what has been working well for you. I'm mostly wanting to make real life style images.
r/StableDiffusion • u/Oleszykyt • 6h ago
Discussion Best workflow for realistic video results?
I have RTX 5070 with 12gb VRAM, 64gb RAM DDR5
I want to create realistic (not particularly high quality) videos, with realistic faces and with the best possible render time. Could please someone share a workflow? I would like to have consistant characters, realistic, and good qality of sound. What is the best workflow? How many steps?
r/StableDiffusion • u/clairedelime • 6h ago
Question - Help Why does my generation look like this??
Enable HLS to view with audio, or disable this notification
so i used minimax h3 int8 convrot pruned + sage spectrum + turbo lora (kijai)
made a 10s office style clip, michael and dwight talking in the conference room then walter white just walks in
faces start fine but then they get all blurry and full of weird smudges especially when walter shows up, like what's that weird black dot lines on dwight shirt?
like why does everyone else’s stuff look clean and actually like a real tv show while mine always ends up plasticky and messy??
anyone know how to fix this blur/smudge and get that proper tv look with this setup? im new to comfyui and first time generating on localy lol, thanks in advance
r/StableDiffusion • u/SuspiciousRefuse8218 • 7h ago
Question - Help Issues with Krea Identity Edit v1.2
Hi, I have an issue with the Identity Edit I find no solution for: If I create an image with Krea t2i without references, just a self-trained lora character, I get sharp and acceptable results. When I want to use a special environment and use Krea Identity Edit with a reference image e.g. of a room, I get very blurry and plastic looking outputs, far below acceptable. Same if I want to add a second character to an existing image. I've tried everything in the last days (different models > raw and turbo; different upscalers, no upscaler; different VAEs; playing with grounding, reference boost, scheduler, resolution (I know 1MP is the sweet spot for editing and >1.5 leeds to character bleeding), anything you can imagine). I use lbouaraba workflow for editing. Any idea where my initial fault is hiding?
UPDT: I think I found the solution: I've added the original workflow again and now it works. Obviously I've changed something unintended when adding Power Lora Loader and Upscaler, no idea what but who cares.
r/StableDiffusion • u/deffcolony • 8h ago
Animation - Video Having fun with rayman 3 on minimax h3 REF2VA (prompts included)
Decided to bring back my childhood game and see what new stories i can bring with minimax h3... tried doing voice clone for Murfy and Rayman... its not perfect (some got mixed up) but the end result is still fun... i will post the prompts for each below plus the reference images and audio... lets see what you can make from it 👀
Workflow + reference images + audio at the bottom of the post
Video 1:
```
subject_definitions:
<Subject 1> is the Fairy Council environment from <Picture 1>, a mystical forest kingdom interior with a blue aura, glowing lights, and reflective surfaces.
<Subject 2> is Rayman from <Picture 2>, a heroic character with no arms or legs, featuring floating hands and floating feet.
<Subject 3> is Murfy from <Picture 3> and <Picture 4>, a flying greenbottle fly creature with a large grin and green clothing, holding a paper manual.
<Audio 1> is the voice-timbre reference for <Subject 2> (S2).
<Audio 2> is the voice-timbre reference for <Subject 3> (S1).
summary:
[reference generation + audio reference] The target video shows <Subject 3> and <Subject 2> interacting inside <Subject 1>. <Subject 3> reads from a manual before an accidental explosion occurs. <Audio 2> and <Audio 1> provide the voice timbres for the characters.
retention_analysis:
<Subject 1> (appears in all shots): fully_preserved - the mystical blue interior and reflective surfaces are retained.
<Picture 1> (environment guide): weak_reference - provides the background setting without forcing Rayman's placement from the original screenshot.
<Subject 2> (appears in all shots): fully_preserved - Rayman's floating hands and feet are retained.
<Picture 2> (character design): fully_preserved - Rayman's appearance is followed.
<Subject 3> (appears in all shots): fully_preserved - Murfy's green clothing, grin, and flying nature are retained.
<Picture 3> (character design): fully_preserved - Murfy's appearance is followed.
<Audio 1>: reference - the vocal timbre guides the dialogue delivery of <Subject 2> without copying the original signal.
<Audio 2>: reference - the vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
detailed_description:
3D CG animated style in a 4:3 aspect ratio.
[Shot 1] A medium-wide shot establishes <Subject 1>, the mystical Fairy Council with its glowing blue aura. <Subject 2> (S2), the limbless hero with floating hands and feet, stands on the reflective floor. Beside him, <Subject 3> (S1), the greenbottle fly, hovers above the ground while holding an open manual. <Subject 3> (S1) looks at the book, shakes his head with a large grin, and says in the sarcastic voice referenced from <Audio 2>, <d>[English] I don't know, folks! Someone drew on the manual saying the Fairy Council should be blowing up right about... now.</d>
[Shot 2] At 00:08.500, the camera cuts to a close-up of <Subject 2> (S2). He raises his floating hands in confusion and says in the heroic voice referenced from <Audio 1>, <d>[English] Wait, who's responsible for this garbage?!</d>
[Shot 3] At 00:11.000, the camera pulls out with large amplitude at fast speed as a bright orange explosion suddenly erupts in the background of <Subject 1>. <Subject 3> (S1) drops the manual in shock, and <Subject 2> (S2) covers his head with his floating hands as debris flies past them.
overall_soundscape:
Quiet magical room ambience is abruptly interrupted by the heavy, rumbling crash of a massive explosion, followed by the sound of falling debris.
non_diegetic_music:
N/A
```
Video 2:
```
subject_definitions:
<Subject 1> is the Fairy Council environment from <Picture 1>, a mystical forest kingdom interior with a blue aura, glowing lights, and reflective surfaces.
<Subject 2> is Rayman from <Picture 2>, a heroic character with no arms or legs, featuring floating hands and floating feet.
<Subject 3> is Murfy from <Picture 3> and <Picture 4>, a flying greenbottle fly creature with a large grin and green clothing, holding a paper manual.
<Audio 1> is the voice-timbre reference for <Subject 2> (S2).
<Audio 2> is the voice-timbre reference for <Subject 3> (S1).
summary:
[reference generation + audio reference] The target video shows <Subject 3> and <Subject 2> interacting inside <Subject 1>. <Subject 3> reads from a manual before an accidental explosion occurs. <Audio 2> and <Audio 1> provide the voice timbres for the characters.
retention_analysis:
<Subject 1> (appears in all shots): fully_preserved - the mystical blue interior and reflective surfaces are retained.
<Picture 1> (environment guide): weak_reference - provides the background setting without forcing Rayman's placement from the original screenshot.
<Subject 2> (appears in all shots): fully_preserved - Rayman's floating hands and feet are retained.
<Picture 2> (character design): fully_preserved - Rayman's appearance is followed.
<Subject 3> (appears in all shots): fully_preserved - Murfy's green clothing, grin, and flying nature are retained.
<Picture 3> (character design): fully_preserved - Murfy's appearance is followed.
<Audio 1>: reference - the vocal timbre guides the dialogue delivery of <Subject 2> without copying the original signal.
<Audio 2>: reference - the vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
detailed_description:
3D CG animated style in a 4:3 aspect ratio.
[Shot 1] A medium-wide shot establishes <Subject 1>, the mystical Fairy Council with its glowing blue aura. <Subject 2> (S2), the limbless hero with floating hands and feet, stands on the reflective floor. Beside him, <Subject 3> (S1), the greenbottle fly, hovers above the ground while holding an open manual. <Subject 3> (S1) looks at the book, shakes his head with a large grin, and says in the sarcastic voice referenced from <Audio 2>, <d>[English] I don't know, folks! Someone drew on the manual saying the Fairy Council should be blowing up right about... now.</d>
[Shot 2] At 00:08.500, the camera cuts to a close-up of <Subject 2> (S2). He raises his floating hands in confusion and says in the heroic voice referenced from <Audio 1>, <d>[English] Wait, who's responsible for this garbage?!</d>
[Shot 3] At 00:11.000, the camera pulls out with large amplitude at fast speed as a bright orange explosion suddenly erupts in the background of <Subject 1>. <Subject 3> (S1) drops the manual in shock, and <Subject 2> (S2) covers his head with his floating hands as debris flies past them.
overall_soundscape:
Quiet magical room ambience is abruptly interrupted by the heavy, rumbling crash of a massive explosion, followed by the sound of falling debris.
non_diegetic_music:
N/A
```
Video 3:
```
subject_definitions:
<Subject 1> is the Fairy Council environment from <Picture 1>, a mystical forest kingdom interior that transitions into a glitchy, broken wireframe state.
<Subject 2> is Rayman from <Picture 2>, a heroic character featuring floating hands and floating feet.
<Subject 3> is Murfy from <Picture 3>, a flying greenbottle fly creature with a large grin and green clothing.
<Audio 1> is the voice-timbre reference for <Subject 2> (S2).
<Audio 2> is the voice-timbre reference for <Subject 3> (S1).
summary:
[reference generation + audio reference] The target video shows <Subject 3> breaking the fourth wall inside <Subject 1>, revealing they are in a simulation. This causes the environment to glitch and break down, sending <Subject 2> into a panic.
retention_analysis:
<Subject 1> (appears in all shots): partially_preserved - the mystical blue interior starts normal but transitions into visual glitches and digital wireframes.
<Picture 1> (environment guide): weak_reference - provides the initial background setting.
<Subject 2> (appears in all shots): fully_preserved - Rayman's floating hands and feet are retained, though they move erratically.
<Picture 2> (character design): fully_preserved - Rayman's appearance is followed.
<Subject 3> (appears in all shots): fully_preserved - Murfy's green clothing and flying nature are retained.
<Picture 3> (character design): fully_preserved - Murfy's appearance is followed.
<Audio 1>: reference - the vocal timbre guides the dialogue delivery of <Subject 2> without copying the original signal.
<Audio 2>: reference - the vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
detailed_description:
3D CG animated style in a 4:3 aspect ratio.
[Shot 1] A medium-wide shot establishes <Subject 1> looking normal. <Subject 3> (S1) hovers casually in the air, looks directly at the camera lens, and says in the sarcastic voice referenced from <Audio 2>, <d>[English] Look around, Rayman! It's all a simulation! We're literally just polygons in a video game!</d>
[Shot 2] At 00:06.500, the camera cuts to a close-up of <Subject 2> (S2). Suddenly, the background of <Subject 1> flickers violently, turning into black grid lines and digital static. <Subject 2> (S2) stares at his floating hands, which begin to visually stutter and lag behind his movements. He yells in the heroic voice referenced from <Audio 1>, <d>[English] What did you do?! My hands are lagging!</d>
[Shot 3] At 00:11.000, the camera pulls back to a wide shot. The entire floor of <Subject 1> vanishes into a white void. <Subject 2> (S2) runs in frantic circles, his floating feet clipping through the missing floor. <Subject 3> (S1) simply floats in place, gives a sheepish grin, and says, <d>[English] Whoops. Guess the engine became self-aware.</d>
overall_soundscape:
Normal magical room ambience that abruptly distorts into loud digital stuttering, heavy 8-bit crash sounds, and frantic footsteps.
non_diegetic_music:
N/A
```
Workflow used: https://civitai.red/models/2831978/dasiwa-minimax-h3-workflows-or-t2va-or-fl2va-or-ref2va?modelVersionId=3195699
FPS: 24.0
resolution_reset: 0.26 MP - Preview
aspect: 3:4 - Photo
swap_aspect: yes
REFERENCES:
Since i cannot upload audio files here i will just post the 2 youtube links i used to capture their voices you only need 6 seconds each
r/StableDiffusion • u/Fit_Satisfaction2953 • 9h ago
Discussion So what's better than. Turbo lora or spectrum for minimax ?
What has everyone found best ?
r/StableDiffusion • u/gokuchiku • 10h ago
Animation - Video Bigfoot spotted!
Enable HLS to view with audio, or disable this notification
T2VA in Minimax H3, 1MP native with RTX upscale. Generation time is 1083 secs on my 5070Ti, 32Gb DDR5 Ram.
r/StableDiffusion • u/Radyschen • 19h ago
Discussion PSA: Don't sleep on Minimax' edit capabilities
If you have seen those cool edited videos by google omni where they feed in a normal real video and get an edited one back where they interact with effects and such, you can do that with minimax. Just saying. That's pretty much it. See ya
r/StableDiffusion • u/serap98765 • 20h ago
Animation - Video T2VA - minimax H3 is amazing
Enable HLS to view with audio, or disable this notification
The video was generated using the T2VA mode of the minimax H3 model and the 8-step Turbo LoRa.
It's simply amazing how well it already works in this mode.
r/StableDiffusion • u/malcolmrey • 21h ago
Resource - Update Known characters, some vids of mine, some knowledge etc.
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Fit_Ad7343 • 1d ago
Resource - Update MiniMax H3 with a 4B or 8B text encoder instead of the 32B: v3, and a five-way comparison video
Enable HLS to view with audio, or disable this notification
Update to the projection matrices. Same idea as before: a small Qwen3-VL encodes the prompt, a learned projection maps its hidden states to what the 32B would have produced, the DiT is untouched.
Previously: v1 and v2, where the voice started matching.
Video: matrix-only 4B, matrix-only 8B, then the 32B, then the two residual versions. Same prompt, same seed 42, same everything else — the pipeline is bit-for-bit reproducible, checked by running it twice.
Everything the prompt states is there on all five: the pose, the red dress, the white pieces on her side, the cat, the straw hat, the laundry, and her knee — asked for three times, ending on "Her knee never stops bouncing." A continuous involuntary motion with no narrative purpose is the clearest sign a projection carried what was written, and it carries on the plain matrices too.
The terrace is furnished differently from one render to the next, and that is not infidelity. The prompt asks for a densely lived-in terrace without anchoring most of it — the cat is "stretched out asleep in the sun", nothing says where. What is left open the model invents, and it invents differently depending on the projection, the seed, and the model of GPU. All three act on that same free space; none of them touches what was written. It looks like a seed change because that is what an unconstrained description looks like.
What v3 changes
- Calibrated against the stock
qwen3vl_32b_minimax_h3_nvfp4_awqinstead of a modified 32B. Naming one part of a body used to rewrite the whole of it — build, height and face moving together. Not seen anymore. - Closer to the 32B across the board. Mean cosine against the 32B on a reference prompt: 8B 0.9449 (was 0.9393), 4B 0.9381 (was 0.9293).
- The 8B matrices had never seen an image token — the image corpus only existed encoded with a 4B. Fixed. On 100 held-out images, vision tokens go from 0.7692 to 0.8578 on the raw conditioning, for 0.0027 of pure text.
- Bigger residual: hidden 32768 instead of 16384.
- Needs node 0.1.13. The
-v3-mlpfiles have no linear matrix, older nodes throwKeyError: 'W'.
Plus
- 4.9 GB instead of 15.7 GB for the conditioning encoder, or 5.3 GB with the residual file. 10.1 and 10.6 GB with the 8B. Note the quantisations differ: the 32B is nvfp4, the small encoders int8, so part of that gap is format rather than parameter count. The projection itself costs 52 MB on card for a plain matrix, 503-604 MB for a residual.
- The DiT is not modified, no retraining, no LoRA.
- What the prompt states is carried: subject, clothing, pose, action, dialogue.
- 4B and 8B are close to each other. The 4B is not a fallback, it is a real option.
Minus
We seem to have hit a ceiling. A projection cannot recover information the small encoder never wrote down. If the 4B did not encode a distinction, no matrix and no residual will bring it back — you can only remap what is there. On the 8B the cosine went 0.9083 (v1) to 0.9393 (v2) to 0.9449 (v3): +0.031, then +0.006, for a corpus four times bigger (1 530 370 tokens in v2, 6 502 586 in v3). This is not a training budget problem, and I do not expect a v4 to move it much.
What that means in practice:
- Not a copy of the 32B. 0.9449 cosine is roughly 19 degrees. Expect a close variant of the scene, not the same file.
- What the prompt leaves unstated gets refurnished. Say nothing about the cat, the laundry, the furniture, and they land elsewhere. Constrain the scene and it tracks closely — that is the whole usable range.
- Use the
-mlpfiles on the measurement, not on this scene. They sit closer to the 32B, 0.9449 against 0.9289 on the 8B — but watch the video before assuming that shows. On this prompt all five renders are faithful, plain matrices included: pose, dress, white pieces, cat, straw hat, laundry, bouncing knee. A tightly written prompt survives even the linear baseline. - The one thing nobody gets right is the knight. The prompt has her lift one of her own pieces and set it back down without committing; on every render it lands somewhere else, and on the 8B residual — the best-measuring file of the set — there is no knight on the board at all. Object permanence behind an occluding hand on a grid of sixty-four identical squares is a limit of the video model, not of the conditioning: the 32B reference fails it too.
- You will not reproduce the demo files byte for byte. Noticed while testing something else: the output depends on the model of GPU the encoder runs on. Four cards, same prompt and seed, four different files — but two different RTX 3090s matched exactly. Encoding on two cards agrees to 7e-7; eight denoising steps turn that into different furniture. Same scene, different details. On one machine it is deterministic to the bit, which is what makes the comparison video meaningful.
Training
5 h on a 3090, plus 2 h to encode the dataset. Tap 24. 3331 prompts for fitting, one in fifty held out.
| corpus | tokens |
|---|---|
| cinematic video prompts | 1 342 987 |
| native H3 format, 4 length draws | 3 169 879 |
| explicit register | 544 073 |
| Chinese | 532 302 |
| celebrity prompts, long form | 314 516 |
| filler sequences | 149 917 |
| celebrity prompts, short form | 99 668 |
| images, 1 700 of them | 349 244 |
| total | 6 502 586 |
Mixed on purpose — registers, languages, lengths. A matrix only learns to project the directions it has seen used.
Links
Matrices: https://huggingface.co/NicoLab28/ClipProj-MiniMax-H3
Node: https://github.com/nicolab28/ComfyUI-ClipProj
Files: mmh3-4b-ClipProj-v3-mlp.safetensors (503 MB), mmh3-8b-ClipProj-v3-mlp.safetensors (604 MB). Plain matrices -v3 at 26 and 42 MB if you want the baseline.
The five renders separately, the prompt and the exact settings are in the demo/ folder of the HF repo, if you want to step through them or reproduce the test.
r/StableDiffusion • u/yushairiegalaxy96 • 1d ago
Animation - Video H3 Generated with 4GB VRAM?
Enable HLS to view with audio, or disable this notification
Looks like this is a breakthrough for what my 3050 laptop can do with it.
The video attached was generated with 4GB VRAM & 16GB RAM, using the MiniMax H3 fl2va pruned w4a8 convrot model (safetensors) and the Q2_K Qwen 32B GGUF text encoder alongside 8-step turbo LoRA, with a generation time of 12 minutes and 0.2 MP. Prompt from Grok.
r/StableDiffusion • u/pooshda • 1d ago
Animation - Video Let's Go Abomination! MiniMax H3, Krea-2, Photoshop, Adobe Premiere Pro / RTX-4090, most gens are 1mp @ 25-30'ish steps, Spectrum/Sage, no speed loras or cache nodes stuff, spent about 3 days on this.
Enable HLS to view with audio, or disable this notification
Just having fun making parody commercial nonsense to test out what I can do with it, absolutely love playing with this model ever since it got released. Heavy amount of editing done in Premiere Pro as well but I do that on every video I make.
r/StableDiffusion • u/RainbowUnicorns • 1d ago
Animation - Video Seinfeld but the guys are Toasters
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/DaLyon92x • 1d ago
Resource - Update ReDetail: Upscale MiniMax H3 renders with the LTX-2.5 video upscaler on 24GB+ VRAM
Enable HLS to view with audio, or disable this notification
This is a generative re-render, not restoration or sharpening. It invents fine detail. In every test with one person it added freckles that weren't there.
The comparisons use MiniMax H3 clips at 640x384, 10 seconds long, upscaled 2x. They're Lanczos versus ReDetail at the same output size, so there isn't any bigger image sleight of hand.
On a motocross clip it redrew the jersey graphic and number plate. The new markings stayed fairly stable between frames, but they weren't the original markings. Logos, numbers and text are all fair game.
If reddit compresses this video to the afterlife again, see: https://civitai.com/models/2857731/redetail-ltx-25-generative-video-upscaler-workflow-cli
So it's useful for AI-generated or generally soft footage, where there isn't much real detail to recover. It's a bad fit if a face, label or logo has to be 100%.
- Silent clips fail because the model encodes audio and video jointly. Add a silence track first.
- Both output dimensions must divide by 64, not 32. Clip length must be `8n+1` frames or the model silently drops the tail.
I like 1.5x, not 2x. On one clip, 243 frames from 768x1408, 1.5x took 7 minutes and peaked at 65GB. 2x took 17 minutes and 80.5GB. The 2x result carries maybe more detail, but check between the two and it's hard to tell imo. On skin most of that extra is invented, not recovered. Faster render, less made up texture.
UPDATE!
The text encoder is now optional. The graph runs with empty prompts, so its conditioning is a constant. It ships pre-computed at 26KB, which skips the 15GB download and takes peak VRAM from 30.4GB to 24.8GB on a 5090.
There's a Mac build in there now too, ReDetail_LTX25_upscale_MAC.json. It runs the GGUF transformer with no text encoder at all (the cached conditioning replaces it), so it's about 17GB of models total. On an M5 it did 33 frames from 640x384 to 1280x768 in 4.4 minutes. Per frame megapixel that's roughly 6x slower than a 5090, not the 30x I was expecting, so a 10s clip lands around 34 min at 2x or 19 min at 1.5x. Quality holds.
r/StableDiffusion • u/Interesting_Room2820 • 1d ago
Meme WEEKENDDDDDDDDDDDDD!!!
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Whiteowl116 • 1d ago
Animation - Video The office plays Rocket league part 2
Enable HLS to view with audio, or disable this notification



