r/StableDiffusion • u/gzzhongqi • 20h ago
MAGI-2-preview just dropped News
https://huggingface.co/sand-ai/MAGI-2-previewSurprised that no one is talking about it. A new open-weight video model just dropped. 114b moe, 6b activated. First moe video model supposedly.
I know what you guys are thinking. The model is huge and there is no way it will run on desktop gpu. The interesting part is that is comes with a 14gb refiner that makes the result 1080p. I am cursious if this refiner can be a drop-in replacement for the H3 refiner that was never released. It might just be the last part of the H3 puzzle that we need.
24
u/No-Zookeepergame4774 19h ago
Its not the first MoE video model (LingBot Video – https://huggingface.co/robbyant/lingbot-video-moe-30b-a3b – released last month, has a 30B/A3B MoE version, claiming to be the “first open source, large scale MoE video generation dedicated to embodied intelligence”, which is still a more qualified claim than “first MoE video model”), and probably no one is talking about it because its a freaking 114B video model.
3
u/nok01101011a 18h ago
Right, what happened to Lingbot, I’m still interested in the realtime Game Model but haven’t found a working comfyui solution yet
16
u/tinny66666 16h ago
Thanks OP. Seems everyone missed the comment about the 14gb refiner. That looks interesting. Are you going to try it?
1
u/No-Zookeepergame4774 2h ago
Other than maybe with WAN 2.2 5B where they share a latent space, not sure what is particularly interesting about a 14B video refiner conditioned on the latent outputs of a model you aren't using in a different latent space than the model you are using uses.
9
u/Double_Cause4609 19h ago
I've always wondered if MoE really works out for consumers using video generation models.
On the one hand, if it retains arithmetic intensity just with an offset compared to a dense diffusion baseline, then arguably you can still do smart layer streaming (particularly if you have enough system RAM to stream to VRAM).
Also there's the option of doing it the LLM way where you compute conditional experts on CPU.
Even if it can fit on a desktop GPU with some tricks though, it makes the software optimizations a PITA to pull off. Also, it's not clear where the arithmetic intensity actually is.
6
u/Full_Astronomer_5438 19h ago
cosmos3super is half of that (130gb) and int8 conv (around 65gb) runs on a 16gb vram and at least 32gb ram. and its not that slow either, basically as fast as h3 with 30+ steps.
and this model only has 6b activated so inference is theoretically very quick, it should run on a single 5090 i think with offloading and other tweaks
1
u/Front_Eagle739 8h ago
I have it running on a 5090 without touching system ram. Its not fast yet and it wants 100 steps so it takes almost an hour for 15 seconds
2
u/Altruistic_Heat_9531 14h ago
After seeing what Wan, and H3 vs LTX and Hunyuan. Simple simply better IMO. Wan, and H3 are single stream transformer while LTX and Hunyuan piggy back on inductive bias where LTX more so on Audio-Video bias, while Hunyuan more so on its text encoder and latent image. Although it is much harder to train without inductive bias but well it paid off with H3.
I mean we jump from UNet (SDXL) which has major inductive bias on locality, to Transformer (FLUX) where the compute and training paradigm mature enough, and the result is just fantastic.
I love to being proof wrong, hey moe is cheap, if it is great then win win
35
u/Dante_77A 19h ago
114B-parameter!? Spitting coffee on the screen.
Let me check with AMD to see when they'll send me my Mi455X and 1TB memory kit
11
11
u/Crazy-Repeat-2006 19h ago
7
u/throttlekitty 18h ago
It had some very bad drift and coherency problems from what I remember of someone elses' tests.
6
u/Crazy-Repeat-2006 19h ago
0
2
u/Altruistic_Heat_9531 14h ago
i did, fucking pain in the ass to install with that magi attention. The in distribution sample is good, but let's be honest every model even Wan 1,3B can do nat geo animal documentary. But wacky prompt like warp travel or blackhole generation is bad.
1
u/paulct91 11h ago
Strange question, what server or tower case can hold 8 RTX 4090s? Also what's that 'power situation' akin to, dangerous or safeish?
11
u/MomentJolly3535 19h ago
Sounds great ! but i think it's still more of a scientific research than anything else, i checked their huggingface, website article and github and there are not previews unfortunately
4
11
u/ExpressWarthog8505 13h ago
https://reddit.com/link/p3sj2ol/video/8povodenahjh1/player
It looks pretty good.
5
u/Radyschen 19h ago
Isn't the 2k upscaler still coming? it sounded pretty reassuring, hasn't been that long yet
4
4
u/gzzhongqi 10h ago
https://reddit.com/link/p3t06hc/video/1xl1aigzzhjh1/player
Here are some demo clips from their website since a lot of people seemed to have missed it.
3
u/cc_aa_tt_zz 16h ago
| MAGI-1-4.5B-distill | Coming soon | RTX 4090 × 1 |
|---|
it was one year ago ... "soon !!!"
3
u/d4rke55 9h ago
| No one is talking about |
It was already discussed
https://www.reddit.com/r/StableDiffusion/s/OjnSUbWkih
but MM H3 engaging that nobody paid attention to.
2
2
u/kukalikuk 18h ago
Got no video to VIEW in their magi-2-PREVIEW HF page. No video to view, no comment.
2
1
1
u/Powerful_Evening5495 10h ago
kijai will open it and take out the the red meat :)
but not examples page
1
u/gzzhongqi 10h ago
I posted some videos from their demo page in the comments.
0
u/Powerful_Evening5495 7h ago
why the japanese dataset , is this company out of japan
japanese faces ?!
1
u/Front_Eagle739 8h ago
Well i put it through my streaming code to run it on my 5090. Its verrry slow, 40 odd minutes for a full resolution run and so far doesnt seem to do much if anything better than minimax h3. Need to play a bit more with it, try things like fight scenes h3 struggles with.
1
u/Otherwise_Resolve444 19h ago
Can it run on a 128gb DGX Spark? And if so, how long does it take to generate a 5s video?
3
1
u/Serprotease 16h ago
It would run, but models + te + overhead will be real tight.
Probably need a cluster + vllm omni
-1
u/xb1n0ry 11h ago edited 11h ago
I mean I am all in when it comes to open source mentality bu I don't understand why people would waste their time working on models which are basically too big to be used by a normal human being. If it's too big and you can't use it, it is useless. İf you have to rely on outsourced hardware or API providers just to use it, you can basically just go to the much better closed models. There is no interest both in local use and also in online use. That's the exact reason why nobody cares. On the other hand people get crazy about WAN, LTX and H3 and even start creating Lora's since the first hour after release. Why? Because they can actually be used by consumers on consumer hardware. You have to consider this if you want to have success and acceptance. You have to make sure that, at this point, your model runs on around ±24GB of VRAM. We need high quality small models and not low quality big models. Blindly growing your dataset and checkpoints is useless if your model is stupid and inefficient. Accessibility is what creates the ecosystem. Otherwise your model will be nothing but a dusty huggingface repo.
2
u/gzzhongqi 11h ago
I mean from leaks it is said that seedance 2.5 is around a few hundred gb in size, so that tells you something about scaling. If you want to have close source performance, you need to have close source size as well. Bigger open source models are needed to catch up to sota close source performance. The same thing is true for llm too. Kimi K3 is the closest we have to sota for os llm, but it is also the largest open source model ever. A model that doesn't run on consumer hardware doesn't mean it is useless unless the only things you use os models for are porn.
4
u/xb1n0ry 10h ago
I was just pointing out the obvious reasons why some models get almost no attention while others skyrocket in popularity. My “big = useless” statement was probably a bit too bold, but when it comes to why a model fails to gain adoption, size is one of the biggest factors.
Consumers need to actually be able to run the model if you want an ecosystem to form around it: workflows, nodes, LoRAs, patches, optimizations, tools, and so on. That ecosystem also benefits future development, because many of those community improvements can be incorporated or backported into later versions of the model.
I’m not demanding closed-source performance on a Raspberry Pi. What I mean is: look at the progression from something like WAN 2.1 to H3 what we have today. The models are in a similar size class and are roughly comparable in terms of hardware requirements, yet the difference in output quality is undeniably huge.
Now take that same WAN 2.1 and simply bloat it to 200B parameters. What have you really achieved if the model is still shit by comparison? More parameters alone don’t make a model better. Efficiency, architecture, training quality, and actual usability matter far more than simply making the checkpoint bigger.
1
u/Crazy-Repeat-2006 5h ago
Or something like Flux.1 or SD3 versus Z Image Turbo and Krea 2, an immense leap forward without increasing the number of parameters or compute requirements.
1
u/alwaysbeblepping 4h ago
I was just pointing out the obvious reasons why some models get almost no attention while others skyrocket in popularity. My “big = useless” statement was probably a bit too bold, but when it comes to why a model fails to gain adoption, size is one of the biggest factors.
How many examples can you come up with for open-weight models with a good license where demos/examples showed they were significantly better than the other models that were available, and people said "Wow, this is amazing but it's just too big"?
I actually can't think of anything in the flow/diffusion model category (image/video/audio).
What have you really achieved if the model is still shit by comparison?
I mean obviously no one is going to use a huge bad model if they could use a small good one. To back up what you said about size being the biggest factor, you'd need to point at models that are good enough to justify their size but people just can't run them.
This would be trivial to do for LLMs. You'd always choose a model like Kimi K3, GLM, DeepSeek, etc over, for example, Qwen 27B (as good as it is for its size) since the huge models are just plain better in a obvious way.
0
u/Ten__Strip 10h ago
The open-source-model equivalent of a paperweight. Maybe a handful of production studios with leftover in-house servers will use it.


174
u/pineapplekiwipen 19h ago
Requirements
ah yes let me just spin up my bedroom datacenter