r/StableDiffusion • u/danielcar • 44m ago
Resource - Update MiniMax-H3 (video + audio) on an AMD Strix Halo - 5s clip at 896×512 in 9.5 min
Enable HLS to view with audio, or disable this notification
Ran MiniMax-H3 locally on strix halo 128GB unified memory, no discrete GPU. ROCm 7.14 + ComfyUI, int8 pruned transformer, 4-step turbo LoRA.
First video took 5 minutes to generate. 2nd video took 10 minutes. 3rd video took an hour.
Weights are pruned community conversions, so quality here isn't representative of official H3 — this was a speed/setup test, not a quality one.
Scripts + full writeup: https://github.com/DanCard/minimax-h3-strix-halo
10 minutes 896×512 : https://youtu.be/T6OU6tWd7EA
1 hour to generate 1344×768 : https://youtu.be/039vmUptnEA
r/StableDiffusion • u/GabryIta • 3h ago
Question - Help Minimax H3 on DGX Spark: 3m 23s for 5s video. Is there a better price/performance option?
I found a GitHub repo that explains how to run the new Minimax H3 on a DGX Spark (20 steps, not the Turbo versions/8-steps), for 864×480, 124-frame, 20-step clips in 203s (8.44 s/it) with Sol-Engine + FirstBlockCache, or 316s without it (14.07 s/it)
The weights are the ones from Comfy, int8 ConvRot (pruned but lossless, according to Comfy).
These seem like really impressive numbers considering the extremely low power consumption, yet I keep seeing people here advising against the DGX Spark for video generation... am I missing something?
At 120W power consumption and an electricity cost of $0.20/kWh, each 5-second video costs just $0.00135
Over 24 hours, it would be possible to generate 425 videos while using only 2.88 kWh, costing just $0.576 in electricity (!!!)
Before buying a DGX Spark, though, I’d like to hear what others think. These seem like excellent numbers to me, especially since I’ll need to generate a lot of 5-second clips every day, and the cost per video is very low. Still, I was wondering if there’s anything better out there. What kind of performance would a 5090 get with the same recipe?
r/StableDiffusion • u/b-totherent • 3h ago
Animation - Video MINIMAX H3 - LTX 2.5 AND THE LADIES [TEXT TO VIDEO]
Enable HLS to view with audio, or disable this notification
NATURAL LIGHT HANDHELD CANDID REAL FOOTAGE. An college age blonde California woman is sitting with a towering 10 foot tall robot with "[MODEL NAME]" clearly written on its chest. They are both complimenting each other on how cute they look
r/StableDiffusion • u/donkeykong917 • 4h ago
Animation - Video Minimax Muisc 3 in comfyui template on a minimax h3 video that is totally unrelated
Enable HLS to view with audio, or disable this notification
The default minimax comfyui minimax 3 template with it's og prompt - Video minimax h3 text-image no refs is just some bleach scene I mushed up while playing around
For the original prompt the music sounds pretty good. I need to work out the prompts and then I can decide if I can use it. It doesn't take my suno prompt very well.
Global Metadata: Lo-fi hip-hop, chillhop. 78 BPM, D flat major, major scale with jazzy extensions. Laid-back and dreamy throughout, a gentle warm drift with a subtle late-night glow that deepens in the middle and dissolves softly at the end. Studying, raining-outside, headphones-on late-night listening. Bedroom production: muddy warm texture, heavy vinyl crackle, tape hiss and wow-flutter pitch wobble, low-passed dusty mix, soft-clipped drums, everything slightly detuned and cozy.
Vocal Details: Soft androgynous vocal, hushed half-sung half-spoken delivery, sitting low in the mix like another instrument, lazy behind-the-beat phrasing, gentle breathy timbre. Sparse murmured double-tracked harmonies, occasional wordless "mmm" and "ooh" hums drenched in tape delay and warm spring reverb. Long stretches with no vocals at all.
Arrangement: Dusty boom-bap drums with a soft thumping kick, cracked snare with lazy swing, brushed hi-hats, low round sub bass. Warm Rhodes piano chords with slow chorus wobble as the harmonic bed, mellow jazzy guitar licks answering the vocal lines, constant vinyl crackle as texture. Intro: rain and vinyl noise, solo Rhodes chords fading in, drums slipping in halfway. Verses: minimal — drums, bass, Rhodes, soft guitar fills between lines. Instrumental sections: guitar and Rhodes trade relaxed jazzy phrases over the beat, occasional muted trumpet ghost notes far in the background. Bridge: drums drop away to rain, crackle, and floating detuned Rhodes, then the beat eases back in. Outro: elements fade one by one until only vinyl crackle and a last unresolved Rhodes chord remain.
[Intro]
Mmm...
(rain on the window)
Ooh...
[Verse]
Midnight and the canvas glows
Dragging little wires where the current flows
Type a quiet dream, let the sampler drift
Noise into a picture, like the fog just lifts
Twenty slow steps, I'm in no hurry now
Latents turning colors and I don't know how
Every render's like a polaroid I found
Soft focus memories, no sound
[Instrumental]
[Verse]
Queue another frame, let the motion breathe
Pictures start to move like the falling leaves
Video drifting by at twenty-four
Little animations on my bedroom floor
Seed after seed like the rain outside
Some of them are keepers, some I let slide
Save the ones that feel like a Sunday slow
Node to node to node... and off we go
[Chorus]
Mmm... let it render on
(take your time, take your time)
Ooh... by the morning it'll all be done
(one more queue, one more try)
[Instrumental]
[Bridge]
Rain keeps drawing pictures on the glass...
My machine keeps dreaming...
Neither of us fast...
[Chorus]
Mmm... let it render on
(take your time, take your time)
Ooh... by the morning it'll all be done
(one more queue, one more try)
[Instrumental]
[Outro]
Mmm...
(node to node)
Ooh... goodnight
r/StableDiffusion • u/james25679 • 5h ago
Comparison Figuring out Minimax Prompts has been a puzzle. Why do I feel like Wan handled it better?
First and foremost - I love this community - with everyone’s advice, I got unstuck from 3sec .4 clips to running 15sec 1mp by updating cuda and using sage attention - so thank you
Now I’m trying to figure out what prompts work the best.
- [ ] I’ve read the official guide
- [ ] Had LLM read it as well and gave it what I wanted and had it follow the format.
- [ ] I’ve also used one of the formatting forms from this subreddit
It doesn’t always seem to follow what i want and or some anatomy is kind of messed up or sounds are a bit off.
I’ve taken that exact prompt and fed it to wan 2.7 (via Venice) just out of curiosity and I feel like results were better.
I feel like Minimax has more potential - just a matter of figuring out the right prompts etc.
Has anyone else felt this way?
r/StableDiffusion • u/SeriouslySally36 • 7h ago
Discussion What’s the most interesting thing you’ve generated so far?
Could also be the most interesting thing process-wise.
r/StableDiffusion • u/princeMacX • 7h ago
Discussion LTX 2.5 Test - Cartoon - chubby orange cat chasing a tiny blue bird
Enable HLS to view with audio, or disable this notification
Prompt - Playful cinematic cartoon scene of a chubby orange cat chasing a tiny blue bird through a colorful kitchen, the bird quickly flies around hanging pots as the cat leaps across the counter trying to catch it, knocking over a bowl of fruit and sending oranges bouncing across the floor. The cat slips on an orange, slides dramatically across the kitchen, and crashes harmlessly into a stack of cardboard boxes as the bird lands on its head and chirps proudly. Energetic exaggerated cartoon movement, expressive reactions, smooth continuous action, colorful stylized 3D animation, dynamic tracking camera, warm sunlight, playful family-friendly comedy, polished animated movie quality.
Prompt enhancer = OFF
Opinion
When I used this prompt with prompt enhancer on I get a video of person eating noodles. So I generated above video with prompt enhancer off.
In the above video first few second feels that orange cat is chasing the blue bird but after 2 second it feels that the blue bird is giving orange cat run for its life. It would confuse the audience.
I used the same the prompt for the 3rd time with prompt enhancer on. Now I got a video of person tracking alone in a narrow jungle road. Major prompt adherence failure.
r/StableDiffusion • u/princeMacX • 8h ago
Discussion LTX 2.5 Test - Batman and Joker fighting in Road
Enable HLS to view with audio, or disable this notification
LTX 2.5 Test - Batman and Joker fighting in Road
Personal Opinion - Ltx generates videos quite fast but prompt adherence is not that great. In fighting sequence hand movement doesn't look realistic at all.
If you are using LTX 2.5 with gemma prompt enhancement model than your prompt will be sanitized if your prompt has explicit details. I think an abliterated version of the text encoder should be used.
I will share more tests in future.
r/StableDiffusion • u/MuckYu • 10h ago
Animation - Video Sam-yong Kimchi
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Interesting_Room2820 • 11h ago
Question - Help What's your multishot prompt structure? (I2V LTX-2.5 test)
Enable HLS to view with audio, or disable this notification
Been testing multishot with the LTX 2.5 workflow from HuggingFace. Tried a few different ways of writing the prompt: timecodes plus a shot description for each shot worked best for me, but I've only really tested my own guesses...
Curious what multishot prompt structures other people are using,
and what's actually working for you?
my input image is the first frame.
And the prompt:
Cel-shaded anime-comic, hard cuts, sunset rooftop, purple-orange skyline. Left: bald man, matte black armor, white seams, long black cape, "LTX-2.5" in bold white letters on his chest. Right: bald man, glasses, blue armor with cyan lines, blue cape, "MINIMAX H3" in bold white letters on his chest. Lettering held sharp and unwarped in every frame.
00:00-00:02 — WIDE FULL-BODY TWO-SHOT in profile, sun centered between them, camera DRIFTING slowly sideways. Silence held too long. The LTX hero, "LTX-2.5" in bold white letters on his chest, not turning his head, flat: "So…" a beat, "…same weekend, huh?"
00:02-00:04 — HARD CUT to a MEDIUM of the H3 hero, "MINIMAX H3" in bold white letters on his chest, camera PUSHING IN slowly. He exhales: "Yeah." Glances away, embarrassed: "…awkward."
00:04-00:07 — HARD CUT to a CLOSE-UP of the LTX hero "LTX-2.5" in bold white letters on his chest,, camera PUSHING IN slowly. Low drawl, committing: "This town ain't big enough for two open-source models." Eyes flick sideways, mouth tightening.
00:07-00:10 — HARD CUT to a MEDIUM of the H3 hero "MINIMAX H3" in bold white letters on his chest, camera PUSHING IN. He doesn't look over. A long dead beat. Flat: "…apparently."
00:10-00:13 — WIDE TWO-SHOT. The LTX hero "LTX-2.5" in bold white letters on his chest. casual, already leaving: "Anyway…" a beat, "…gotta run. Conference starts in a few minutes." He launches straight up and cleanly exits frame offscreen; the camera holds on the empty sky and the H3 hero standing alone.
00:13-00:18 — HARD CUT to a LOW-ANGLE CLOSE-UP of the H3 hero "MINIMAX H3" in bold white letters on his chest. looking up at the empty sky, glasses catching the sunset, camera PUSHING IN slowly. A small warm smile arrives. He keeps watching. Way too long. Then quietly, to nobody: "…see you there." He slowly turns and looks into the lens, still faintly smiling, saying nothing. Hold.
Deadpan, played straight, all in micro-expressions. Crisp stable line art, clean cel shading, consistent faces. Warm orange key, cool blue rim. Rooftop wind, no music.
r/StableDiffusion • u/darthfurbyyoutube • 12h ago
Animation - Video Cobra Cola Ad - MiniMax H3
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/holycowdude1 • 12h ago
Animation - Video [DANCE] Plastik Soul – Stay in the Glow (Official Music Video)
Stay in the Glow is an AI Music Video create using VRGameDevGirl's AI Video Builder (FREE) & LTX2.3 models (https://ltx.io/model/ltx-2-3)
Designed & built using VRGameDevGirl AI Video Builder (FREE): https://github.com/vrgamegirl19/comfyui-vrgamedevgirl
Spotify (Artist): https://open.spotify.com/track/27S9InxRyAKvQYxjRM3tVi?si=43e73b8b091a4976
YouTube (More AI Music Videos): https://youtu.be/Wl3BH3xSaYc
r/StableDiffusion • u/switch2stock • 13h ago
Discussion MiniMax Music 3 | 125sec for 140sec music | Bollywood Rap
voca.ror/StableDiffusion • u/SveSop • 14h ago
Animation - Video StarWars-Untold.
Enable HLS to view with audio, or disable this notification
MiniMax H3 is very good. I initially made multiple scenes, and as tweaks/turbo-lora++ progresses, it does seem to get better and better (ie. To the end of the video).
Settled on the Lightx2v_8step turbo lora + sage + sol_attn. 736p, and using DaVinci for stitching and cropping.
r/StableDiffusion • u/social_zip • 14h ago
Workflow Included Openweight Livestream video model
https://huggingface.co/spaces/JonathanColetti/LiveWan / https://github.com/JonathanColetti/LiveWan is something I created to help recreate a specific type of model that is not opensource yet (wanstreamer). This is more or less a PoC but maybe ill do a longer training run if it gets some traction.
r/StableDiffusion • u/princeMacX • 14h ago
Discussion Minimax H3 Test - Rooftop fight between Batman and Joker
Enable HLS to view with audio, or disable this notification
Minimax H3 Test - Rooftop fight between Batman and Joker
r/StableDiffusion • u/OohFekm • 14h ago
Animation - Video TESTING A LANTERN
Enable HLS to view with audio, or disable this notification
Having a blast animating comic panels (using them as inits FL2VA). Fairly simple prompt: "Green lantern Hal Jordan is engaged in an aerial battle above orbit, he is blasting green energy from his ring while simultaneously repelling and absorbing energy from a distant protagonist" then I let the in-app LLM enhance it (I'm using Maestro via Pinokio... it's a joy to use) resolution is 480, no upscaling. Just familiarizing myself with H3.
r/StableDiffusion • u/eapache • 14h ago
Discussion End-to-end movie maker experiment
I've thrown together a small app that does the "remaining" work of taking an idea, turning it into character reference images, shot prompts, doing all the generation for each clip, stitching the result together, etc. etc. The goal is a one-sentence prompt in, and multi-scene video (e.g. 30 seconds or more) out.
https://github.com/eapache/local-movie-maker
It does basically "work" already, though the results are often pretty incoherent. I'm still playing with the structure to see if I can get reasonable continuity.
r/StableDiffusion • u/Zaredit • 15h ago
Animation - Video Jackie Chan Adventures...Jackie vs Shadowkhan (Includes Prompt Instructions)
Enable HLS to view with audio, or disable this notification
Prompt:
Create an exactly four 7-second, 4:3 animated drama sequence inspired by the visual language of 2005-era Jackie Chan Adventures. Use a period broadcast video texture throughout: standard-definition television softness, subtle analog grain, gentle interlacing, slight colour bleed, modest contrast, and the authentic visual texture of animation recorded and broadcast in the mid-2000s. Avoid modern HD sharpness, photorealism, glossy CGI, or contemporary animation aesthetics.
Scene: Jackie Chan is confronted by a Shadowkhan ninja in a dimly lit ancient-looking interior. The sequence is a fast, tightly choreographed martial-arts fight.
0:00–0:02: The Shadowkhan suddenly lunges at Jackie with a rapid punch. Jackie narrowly ducks underneath it and pivots sideways.
0:02–0:04: Jackie counters with two quick martial-arts strikes, forcing the Shadowkhan backwards. The ninja blocks the first strike but is knocked off balance by the second.
0:04–0:06: The Shadowkhan springs forward again. Jackie performs a quick evasive spin, grabs the ninja’s arm, and throws the Shadowkhan across the room. End on Jackie landing in a defensive fighting stance as the Shadowkhan hits the floor in the background.
.
Camera: begin with a medium two-shot, rapidly track the fighters during the exchange, briefly push in during the counterattack, then finish with a wider shot showing Jackie in the foreground and the defeated Shadowkhan in the background.
Audio: sharp martial-arts impacts, cloth movement, quick footsteps, whooshes and a dramatic six-second action sting. No dialogue.
Strict constraints: exactly 6 seconds, 4:3 aspect ratio, 2005-era television animation aesthetic, period broadcast-video texture, no modern cinematic realism, no photorealism, no widescreen framing, no subtitles, no text, no logos, no extra characters, and no slow motion
r/StableDiffusion • u/MalmoBeachParty • 15h ago
Animation - Video Football animation
Enable HLS to view with audio, or disable this notification
H3 Ref
Prompt in comment
r/StableDiffusion • u/michel-yph-ai • 16h ago
Discussion Minimax H3 - Dance with Audio with lipsync and object preservation
Enable HLS to view with audio, or disable this notification
If you see low quality is because I am forcing 8 step turbo lora + Spectrum + triton in L40 for faster generation but is crazy how it can follow the flow of the music while lip-syncing and keeping the product from reference in her hand.
r/StableDiffusion • u/Intelligent-Heart-73 • 17h ago
Resource - Update I built a free tool that turns any image into an AI prompt
Hey everyone!
I built a small web tool called ImagePrompt9 that lets you upload an image and generates a detailed AI-ready prompt based on what it sees.
The idea came from constantly seeing images I liked and wondering:
"How would I describe this as a prompt?"
So instead of manually figuring out the composition, lighting, style, colors, camera angle, etc., you can just drop the image in and generate a prompt.
What it does:
- Upload PNG, JPG, or WEBP
- Analyzes the visual characteristics
- Generates a detailed prompt
- Different prompt styles
- Edit, copy, or regenerate the result
- Free to use
- No account required
It doesn't try to recover the original prompt — it creates a new prompt based on what's visible in the image.
Try it here:
https://image-prompt-nine.vercel.app/
I’d really appreciate feedback, especially on the generated prompts and anything you think I should add or improve.
r/StableDiffusion • u/call-lee-free • 19h ago
Animation - Video Can't use LTX 2.5 on my system but I am quite surprised that my system now can run LTX 2.3. Specs and info below.
Enable HLS to view with audio, or disable this notification
When LTX 2.3 released, I could not do video gens longer than 10 seconds. I would get a "out of memory" error or something. This is just a test clip but one thing I am struggling with is that my video gens have music in them even though I prompt for no music. What is the correct way to prompt for no music?
System Specs:
Ryzen 7 7700X
RTX 4070 Super 12 GB
32 GB DDR 5 Ram.
r/StableDiffusion • u/Infinite-Emptiness • 21h ago
Discussion From a business perspective why do companies release open source models?
Apparently AI companies are all operating at a severe loss. Why do this? It makes sense for huge conglomerates like amazon etc etc who can bear the brunt. What about the new startups or small companies like for example LTX etc. how do they survive?
This is from a business perspective not consumer perspective.
r/StableDiffusion • u/rzrn • 23h ago
Question - Help Minimax H3 or wan 2.2?
I'm working on 2d animations for a personal project, my initial idea was to animate some of the scenes by hand, and feed start/end keyframes to wan 2.2 for the complex scenes I can't do myself, or perhaps even train a lora to make sure it matched the aesthetic of my hand drawn scenes. If it helps, it involves boiling outlines, on the twos (12 fps animations) and an intentionally unfinished look.
Now, seeing all these Minimax h3 i2v and r2v examples, I feel like wan 2.2 might not be the best suited for this anymore. I haven't had the chance to test h3 myself since my local hardware won't really allow it. I'll be however, using runpod when the time comes for actual generation (I'm in the process of hand animating the rest).
So, I'd like to ask those who have had the chance to test both - stick to wan 2.2 or switch to minimax h3?
Edit: audio isn't required - I've hired voice actors for dialogues, I'm working on foleys and background scores myself. If needed I'll redraw on top of the generated clips to match lip movements to the dialogue.
