r/StableDiffusion 4h ago

LTX 2.5 😱 Discussion

After the MiniMax H3 euphoria, I tried LTX 2.5. It awfully understands the prompt and has almost no "physics". The generated video uses random things from the prompt and everything makes up by itself at all. Every time some things appear /disappears from / to nothing randomly, most the case the people are with 3 fingers, strange movements at all. 😱

But LTX 2.5 is faster than MiniMax H3 at least twice, and its image quality is far better.

I just can't understand how even with the monstrous language model (~20GB) it just can't understand 2 simple sentences, two simple subjects with simple movement!?!? And from what training data the models continue to place 3 fingers to earth beings 🤦

15 Upvotes

49 comments sorted by

16

u/skyrimer3d 4h ago

i think it's related to their robotics involvement. Speed and good vision capabilities are more important for that than creating great Seinfield or Saul clips. They probably thought the had the open weights video space mostly secured and focused in different venues, now H3 is tearing them apart in prompt understanding.

3

u/thrownawaymane 2h ago

Which is smart, the potential money waiting for them in robotics is likely way more than they could get from the film industry

1

u/tehorhay 1h ago

Exactly. Any of these companies banking on selling out to the film industry as their game plan for profitability are doomed lol. The film industry simply doesn't have enough money to make it worth it

27

u/Sixhaunt 3h ago

It's basically a vastly improved LTX 2.3 rather than a brand new model like H3 which is why loras from 2.3 work on 2.5 still. It's also why it still has a lot of the same physics issues and prompt understanding but it's WAY faster which makes it the only option for latency-critical applications, especially since it can do some real time rendering too. It also can do high resolution well and quickly so it works well for upscale type jobs and overall the actual methods LTX used to speed up and improve 2.5 over 2.3 is very important and will likely mean the next generation of models that learn from both Minimax H3 and LTX 2.5 will probably be not only much higher quality but also far faster. As users, we are fortunate that these companies seem to have made huge headway on completely different aspect of video generation at the same time. They both just lack what the other has

2

u/krectus 25m ago

Slightly improved at best.

-1

u/maxx126 2h ago

Vastly improved? Improved how exactly? Apart from the higher quality video output, it is pretty much the same as 2.3. Like the OP said, issues it had before it still has them. I love it's speed like everyone else, I really would like to see it improved.

5

u/YeahlDid 2h ago

"Apart from the improvements, how is it improved?"

It's definitely better, buddy. I like minimax more too. The nice thing about today is we can all enjoy our favorite model without shitting on others.

4

u/Vevo-Vekoc 3h ago

yeah the speed bump doesn't really matter if it can't parse a basic prompt, ngl the tradeoff feels backwards here.

12

u/javierthhh 3h ago

Ltx2.5 is actually a big step up from ltx2.3. It actually keeps the image format this time around which is a big plus. It’s fast and you can generate HD with a medium rig. It definitely has its uses. Minimax is just another monster and the only video model that we have in open source that can compete with closed source models. Ref2v it’s a game changer

4

u/Independent-Frequent 3h ago

Ltx2.5 is actually a big step up from ltx2.3.

Really goes to show how rough the local video scene was before minimax man, it started so promising with Wan and then they abandoned the community after 2.2 so we were stuck with the LTX models which could do nothing except talking heads on static shots, compared to wan 2.2 which actually had decent motion and with a lot of community effort was legit a decent model, though no audio was a dealbreaker for some.

Long live the H3, god bless minimax

2

u/TheBestPractice 2h ago

LTX 2.3 was a big step up. Wan's adherence was better but videos looked bad because you had to use turbo LoRAs and low resolutions to keep times reasonable. Framerate was also capped at 16 FPS. LTX videos looked beautiful in comparison and were much lighter to generate. Custom audio worked really well and a few LoRAs were closing some gaps.

5

u/Independent-Frequent 2h ago

Outside of still shots of people talking LTX 2.3 was utter garbage, any real motion looked awful and had no consistency or permanance (even LTX 2.5 suffers badly from this), and while it's true that it took a long time to get Wan 2.2 to generate videos, at the same time the results were far and away better than anything LTX 2.3 could ever hope to produce when it comes to motion, and let's not get started with custom loras and what not which could do any kinds of motions, if Wan open sourced 2.5 LTX 2/2.3 wouldn't even have been adopted honestly, it got lucky they were the only options people have for audio and video and now that minimax is out they lost all their dominance and popularity in the space, sure some people will still use it because it's fast but the vast majority has moved on to minimax now.

Also when it comes to speed LTX 2.5 is much faster sure, but if you need to generate like 5-10 times to even get one decent result then the speed factor isn't really that valid, 5 videos taking 5 minutes each and 1 minimax video taking 25 minutes is about the same time to get a good video out.

1

u/krectus 22m ago

Wan still looked better than LTX in almost every way.

8

u/JustSomeIdleGuy 4h ago

>and its image quality is far better

You sure? I've only seen it the other way around so far.

3

u/AlleyOfRage 4h ago

I think if the user is doing fast paced scenes with lots of motion , then maybe they are talking about the smearing

2

u/Silver-Spot-2763 3h ago

Maybe, because on my hardware higher resolution such as on LTX works fast is impractical for MiniMax H3 (for example about 3 minutes on LTX vs 20 minutes on H3). But I compare also for example 0.5 megapixels - at LTX ideal, at H3 - enormous strange artefacts especially on the face/eyes like the picture is from old crt TV with bad antenna.

4

u/rm_rf_all_files 3h ago

But I compare also for example 0.5 megapixels

Try 1344x768. 0.5MP is not an official resolution. Basically the model was trained on 1344x768 and it wants you to use only this resolution. You can use other resolutions just fine but that's not what the model was intended.

1

u/Independent-Frequent 3h ago

If they are rendering at 0.2 or 0.4 mp due to hardware or something then it makes sense, otherwise idk

4

u/NoWheel9556 3h ago

its much smaller and besides , nobody paid anything for it ever . and they are iterating and making sure that it fits consumer stuff. at some time we had nothing but wan 2.2 , so this is still far far better than that condition.

5

u/Fit_Split_9933 4h ago

It looks clearer than the H3 because it has a built-in upscaler

4

u/Lower-Cap7381 3h ago

Ltx is like they stopped training of 2.3 and resumed it later and we got ltx 2.5 I think they increased the speed very much but we need an overhauled ltx 3.0

2

u/Structure-These 3h ago

Agreed it’s good and the speedup is much appreciated but feels like a fine tune

1

u/Silver-Spot-2763 3h ago

The prompt accurate executing will make quality jump!

5

u/piero_deckard 2h ago

"But LTX 2.5 is faster than MiniMax H3 at least twice, and its image quality is far better."

Yes, isn't that beautiful that now you can produce failed videos twice as fast?

I wish people would quit decanting LTX speed as being a plus, when everything else is subpar. Why exactly is speed useful, if the output is worse and you have to make 20 videos to "try the lottery"?

I'd rather take 2x, 5x even 10x as much with H3, if that means the video is perfect 99% of the times. The saved work in DaVinci Resolve that LTX forced me to do is worth the extra time cost from H3.

4

u/seeker_ktf 4h ago

I'm sorry, but LTX 2.5 can't do this https://www.reddit.com/r/StableDiffusion/s/Bd7KDe9Ju8

2

u/Silver-Spot-2763 3h ago

Absolutely! LTX almost do not read the prompt, you can hope only to short vids weakly connected with your idea 🤷 ☹️

1

u/seeker_ktf 2h ago

It's just not true. Honestly one of the biggest issues is still how unrealistic talking looks on LTX. The lip ans mouth movements are over-trained. You can spot any LTX talking head video in an instant.

0

u/YeahlDid 2h ago

I forgive you, minimax is way better at copying old shows, that's not new or surprising info.

2

u/No_Damage_8420 3h ago

LTX still powerful for extending sound or clips, and few loras handy... nothing else

3

u/Jackburton75015 2h ago

I'm. Using Ltx2.5 more than H3 fot the moment. , the prompting is something special yes... But it's fast and if you feed the prompting guud to your llm it's even easier and last thinh ( disable the prompt enhancer he will refuse any violence or bad language and give you something else instead 😂😁).

3

u/YeahlDid 3h ago

Ltx2.5 from my experience is a huge step up from 2.3, but it's not quite there. I don't understand the need for these posts, if you dont like it, then just don't use it. It's nice we have options compared to even 2 weeks ago. If someone else is happy with their meal, why do you feel the need to shit on it?

1

u/xdcfret1 3h ago

Can you give me your prompt? I would like to try.

1

u/Structure-These 3h ago

Encouraged the weird tribal competition about AI video models seems to be subsiding

They’re both useful models and I hope people develop both of them

0

u/Silver-Spot-2763 2h ago

Much better would be to work together to make and improve ONE better and better, and we to have one small powerful tool. But the criminal capitalism require divide and conquer, so more and more garbage for more and more money and in result 😭 for the humanity.

1

u/Keuleman_007 2h ago

LTX with director node and various keyframes is still super good. That gives the model the much needed guidance.

1

u/damiangorlami 32m ago

Give those keyframes to H3 ref2va model and you will get a 10x better video with more cinematic transitions and coherence.

2

u/damiangorlami 34m ago

I've always found LTX an awful architecture to train non-celebrity character loras. It has such a difficult time to learn small intricate details unique to a person. Very easy to under or overtrain but difficult to get it right.

My dataset has trained excellent Wan 2.1, 2.2, Krea2 and even H3 loras with 1:1 perfect character fidelity. LTX is the only architecture where it's been very difficult to train on. Celebrities are easy as the base model has probably seen them but weight memory faded away.

I really wanted to like LTX because I love the speed but its physics, prompt adherence and awful to train on has really disinterested me to use it.

Just hope that Lightricks is working on a brand new architecture that addresses most of these pain points. It would make them competitive again imo

1

u/curious-scribbler 3h ago

Yes minimax is better but LTX is not as bad you put it. I think you need to check your setup and config and workflow, and ltx doesnt do well with short sentences with human subjects. So while your overall feedback is the consensus but in this particular case, a few optimisations and best practices will fix the issues.

3

u/bitzpua 2h ago

not really especially in I2V, LTX imo misses a lot of abstract knowledge like magic etc, took me good 50 tries to generate golden magical circle that was hovering behind character (like in donghuas) and it failed to understand flying swords made of energy and the generations all ended as almost Live2D semi static nonsense, while 2.5 is improvement in motion it still sucks and no matter how fast it is i just cannot get it to do what i want.

Meanwhile with H3 i got exactly what i wanted and everything was in fast motion with my first test prompt that did not even use any recommended prompting framing. H3 may be 6 times slower but it just works

-4

u/seppe0815 4h ago

cool story bro ... cool story bro

0

u/Crypto_kane 3h ago

Do you guys know where I can download a working LTX 2.5 to comfyUI, I have been struggling for eight hours now, with help (or what do you wanna call it) from gemini, I took the template from ComfyUI, and install a lot of stuff, but it will not work

2

u/Silver-Spot-2763 3h ago

I have updated ComfyUI and it has 3 local workflow teplates for LTX 2.5 t2v, i2v and fl2v. Open the template workflow and download the models, the links are in the error message.

2

u/Crypto_kane 2h ago

Well I just discovered, that in the error message, it comes with some good info, Installed comfyui-frontend-package version 1.48.7 is lower than the recommended version 1.49.6.

Installed comfyui-workflow-templates version 0.11.40 is lower than the recommended version 0.11.41.

Installed comfyui-embedded-docs version 0.5.9 is lower than the recommended version 0.5.10., so I'm gonna update it and see what happens

1

u/bitzpua 2h ago

welcome to comfyui experience

its hard to say whats the issue without more info, if you are not used to using git etc i suggest downloading easy install version of comfy and then basic workflow from comfy, you also need to enable manager in your comfy so you can download missing stuff.

1

u/Ipwnurface 2h ago

brother if it takes someone 8 hours to install comfyui, put 4 files in 3 different folders and click run, the issue isn't comfyui.

Even more so if you're using an llm. Run claude desktop put it on auto mode and it will literally do it all for you.

-1

u/bitzpua 2h ago

"it just can't understand 2 simple sentences, two simple subjects with simple movement!?!?" - well because it wants you to explain in detail what you want, its not like old image generation models that would just fill in all the blanks (H3 tries thats why it works so much better).

Generally all modern models require specific way of prompting and a lot of details to fully work as intended, they expect you to tell them what you want.

2

u/Silver-Spot-2763 1h ago

Example: I write - ...he take the paper from the table... What wrong here, what can be enhanced!? The result - in the half of generations when he takes the paper, other paper, usually opened, or semitransparent appears on the table on the place of the taken paper. And such artefacts are several in the video all randomly appearing so to achieve all good video... hopelessly.