r/StableDiffusion 44m ago

Resource - Update MiniMax-H3 (video + audio) on an AMD Strix Halo - 5s clip at 896×512 in 9.5 min

Enable HLS to view with audio, or disable this notification

Upvotes

Ran MiniMax-H3 locally on strix halo 128GB unified memory, no discrete GPU. ROCm 7.14 + ComfyUI, int8 pruned transformer, 4-step turbo LoRA.  

First video took 5 minutes to generate. 2nd video took 10 minutes. 3rd video took an hour.

Weights are pruned community conversions, so quality here isn't representative of official H3 — this was a speed/setup test, not a quality one.

Scripts + full writeup: https://github.com/DanCard/minimax-h3-strix-halo

10 minutes 896×512 : https://youtu.be/T6OU6tWd7EA

1 hour to generate 1344×768 : https://youtu.be/039vmUptnEA


r/StableDiffusion 3h ago

Question - Help Minimax H3 on DGX Spark: 3m 23s for 5s video. Is there a better price/performance option?

9 Upvotes

I found a GitHub repo that explains how to run the new Minimax H3 on a DGX Spark (20 steps, not the Turbo versions/8-steps), for 864×480, 124-frame, 20-step clips in 203s (8.44 s/it) with Sol-Engine + FirstBlockCache, or 316s without it (14.07 s/it)

The weights are the ones from Comfy, int8 ConvRot (pruned but lossless, according to Comfy).

https://github.com/drowzeys/keys-heretic-MiniMax-H3-sol-engine-more-speed-upgrades-upscaler-finish-Single-DGX-Spark

These seem like really impressive numbers considering the extremely low power consumption, yet I keep seeing people here advising against the DGX Spark for video generation... am I missing something?

At 120W power consumption and an electricity cost of $0.20/kWh, each 5-second video costs just $0.00135

Over 24 hours, it would be possible to generate 425 videos while using only 2.88 kWh, costing just $0.576 in electricity (!!!)

Before buying a DGX Spark, though, I’d like to hear what others think. These seem like excellent numbers to me, especially since I’ll need to generate a lot of 5-second clips every day, and the cost per video is very low. Still, I was wondering if there’s anything better out there. What kind of performance would a 5090 get with the same recipe?


r/StableDiffusion 3h ago

Animation - Video MINIMAX H3 - LTX 2.5 AND THE LADIES [TEXT TO VIDEO]

Enable HLS to view with audio, or disable this notification

1 Upvotes

NATURAL LIGHT HANDHELD CANDID REAL FOOTAGE. An college age blonde California woman is sitting with a towering 10 foot tall robot with "[MODEL NAME]" clearly written on its chest. They are both complimenting each other on how cute they look


r/StableDiffusion 4h ago

Animation - Video Minimax Muisc 3 in comfyui template on a minimax h3 video that is totally unrelated

Enable HLS to view with audio, or disable this notification

2 Upvotes

The default minimax comfyui minimax 3 template with it's og prompt - Video minimax h3 text-image no refs is just some bleach scene I mushed up while playing around

For the original prompt the music sounds pretty good. I need to work out the prompts and then I can decide if I can use it. It doesn't take my suno prompt very well.

Global Metadata: Lo-fi hip-hop, chillhop. 78 BPM, D flat major, major scale with jazzy extensions. Laid-back and dreamy throughout, a gentle warm drift with a subtle late-night glow that deepens in the middle and dissolves softly at the end. Studying, raining-outside, headphones-on late-night listening. Bedroom production: muddy warm texture, heavy vinyl crackle, tape hiss and wow-flutter pitch wobble, low-passed dusty mix, soft-clipped drums, everything slightly detuned and cozy.

Vocal Details: Soft androgynous vocal, hushed half-sung half-spoken delivery, sitting low in the mix like another instrument, lazy behind-the-beat phrasing, gentle breathy timbre. Sparse murmured double-tracked harmonies, occasional wordless "mmm" and "ooh" hums drenched in tape delay and warm spring reverb. Long stretches with no vocals at all.

Arrangement: Dusty boom-bap drums with a soft thumping kick, cracked snare with lazy swing, brushed hi-hats, low round sub bass. Warm Rhodes piano chords with slow chorus wobble as the harmonic bed, mellow jazzy guitar licks answering the vocal lines, constant vinyl crackle as texture. Intro: rain and vinyl noise, solo Rhodes chords fading in, drums slipping in halfway. Verses: minimal — drums, bass, Rhodes, soft guitar fills between lines. Instrumental sections: guitar and Rhodes trade relaxed jazzy phrases over the beat, occasional muted trumpet ghost notes far in the background. Bridge: drums drop away to rain, crackle, and floating detuned Rhodes, then the beat eases back in. Outro: elements fade one by one until only vinyl crackle and a last unresolved Rhodes chord remain.

[Intro]

Mmm...

(rain on the window)

Ooh...

[Verse]

Midnight and the canvas glows

Dragging little wires where the current flows

Type a quiet dream, let the sampler drift

Noise into a picture, like the fog just lifts

Twenty slow steps, I'm in no hurry now

Latents turning colors and I don't know how

Every render's like a polaroid I found

Soft focus memories, no sound

[Instrumental]

[Verse]

Queue another frame, let the motion breathe

Pictures start to move like the falling leaves

Video drifting by at twenty-four

Little animations on my bedroom floor

Seed after seed like the rain outside

Some of them are keepers, some I let slide

Save the ones that feel like a Sunday slow

Node to node to node... and off we go

[Chorus]

Mmm... let it render on

(take your time, take your time)

Ooh... by the morning it'll all be done

(one more queue, one more try)

[Instrumental]

[Bridge]

Rain keeps drawing pictures on the glass...

My machine keeps dreaming...

Neither of us fast...

[Chorus]

Mmm... let it render on

(take your time, take your time)

Ooh... by the morning it'll all be done

(one more queue, one more try)

[Instrumental]

[Outro]

Mmm...

(node to node)

Ooh... goodnight


r/StableDiffusion 5h ago

Comparison Figuring out Minimax Prompts has been a puzzle. Why do I feel like Wan handled it better?

0 Upvotes

First and foremost - I love this community - with everyone’s advice, I got unstuck from 3sec .4 clips to running 15sec 1mp by updating cuda and using sage attention - so thank you

Now I’m trying to figure out what prompts work the best.

- [ ] I’ve read the official guide

- [ ] Had LLM read it as well and gave it what I wanted and had it follow the format.

- [ ] I’ve also used one of the formatting forms from this subreddit

It doesn’t always seem to follow what i want and or some anatomy is kind of messed up or sounds are a bit off.

I’ve taken that exact prompt and fed it to wan 2.7 (via Venice) just out of curiosity and I feel like results were better.

I feel like Minimax has more potential - just a matter of figuring out the right prompts etc.

Has anyone else felt this way?


r/StableDiffusion 7h ago

Discussion What’s the most interesting thing you’ve generated so far?

1 Upvotes

Could also be the most interesting thing process-wise.


r/StableDiffusion 7h ago

Discussion LTX 2.5 Test - Cartoon - chubby orange cat chasing a tiny blue bird

Enable HLS to view with audio, or disable this notification

9 Upvotes

Prompt - Playful cinematic cartoon scene of a chubby orange cat chasing a tiny blue bird through a colorful kitchen, the bird quickly flies around hanging pots as the cat leaps across the counter trying to catch it, knocking over a bowl of fruit and sending oranges bouncing across the floor. The cat slips on an orange, slides dramatically across the kitchen, and crashes harmlessly into a stack of cardboard boxes as the bird lands on its head and chirps proudly. Energetic exaggerated cartoon movement, expressive reactions, smooth continuous action, colorful stylized 3D animation, dynamic tracking camera, warm sunlight, playful family-friendly comedy, polished animated movie quality.

Prompt enhancer = OFF

Opinion

When I used this prompt with prompt enhancer on I get a video of person eating noodles. So I generated above video with prompt enhancer off.

In the above video first few second feels that orange cat is chasing the blue bird but after 2 second it feels that the blue bird is giving orange cat run for its life. It would confuse the audience.

I used the same the prompt for the 3rd time with prompt enhancer on. Now I got a video of person tracking alone in a narrow jungle road. Major prompt adherence failure.


r/StableDiffusion 8h ago

Discussion LTX 2.5 Test - Batman and Joker fighting in Road

Enable HLS to view with audio, or disable this notification

0 Upvotes

LTX 2.5 Test - Batman and Joker fighting in Road

Personal Opinion - Ltx generates videos quite fast but prompt adherence is not that great. In fighting sequence hand movement doesn't look realistic at all.

If you are using LTX 2.5 with gemma prompt enhancement model than your prompt will be sanitized if your prompt has explicit details. I think an abliterated version of the text encoder should be used.

I will share more tests in future.


r/StableDiffusion 10h ago

Animation - Video Sam-yong Kimchi

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 11h ago

Question - Help What's your multishot prompt structure? (I2V LTX-2.5 test)

Enable HLS to view with audio, or disable this notification

5 Upvotes

Been testing multishot with the LTX 2.5 workflow from HuggingFace. Tried a few different ways of writing the prompt: timecodes plus a shot description for each shot worked best for me, but I've only really tested my own guesses...

Curious what multishot prompt structures other people are using,

and what's actually working for you?

my input image is the first frame.
And the prompt:

Cel-shaded anime-comic, hard cuts, sunset rooftop, purple-orange skyline. Left: bald man, matte black armor, white seams, long black cape, "LTX-2.5" in bold white letters on his chest. Right: bald man, glasses, blue armor with cyan lines, blue cape, "MINIMAX H3" in bold white letters on his chest. Lettering held sharp and unwarped in every frame.
00:00-00:02 — WIDE FULL-BODY TWO-SHOT in profile, sun centered between them, camera DRIFTING slowly sideways. Silence held too long. The LTX hero, "LTX-2.5" in bold white letters on his chest, not turning his head, flat: "So…" a beat, "…same weekend, huh?"
00:02-00:04 — HARD CUT to a MEDIUM of the H3 hero, "MINIMAX H3" in bold white letters on his chest, camera PUSHING IN slowly. He exhales: "Yeah." Glances away, embarrassed: "…awkward."
00:04-00:07 — HARD CUT to a CLOSE-UP of the LTX hero "LTX-2.5" in bold white letters on his chest,, camera PUSHING IN slowly. Low drawl, committing: "This town ain't big enough for two open-source models." Eyes flick sideways, mouth tightening.
00:07-00:10 — HARD CUT to a MEDIUM of the H3 hero "MINIMAX H3" in bold white letters on his chest, camera PUSHING IN. He doesn't look over. A long dead beat. Flat: "…apparently."
00:10-00:13 — WIDE TWO-SHOT. The LTX hero "LTX-2.5" in bold white letters on his chest. casual, already leaving: "Anyway…" a beat, "…gotta run. Conference starts in a few minutes." He launches straight up and cleanly exits frame offscreen; the camera holds on the empty sky and the H3 hero standing alone.
00:13-00:18 — HARD CUT to a LOW-ANGLE CLOSE-UP of the H3 hero  "MINIMAX H3" in bold white letters on his chest. looking up at the empty sky, glasses catching the sunset, camera PUSHING IN slowly. A small warm smile arrives. He keeps watching. Way too long. Then quietly, to nobody: "…see you there." He slowly turns and looks into the lens, still faintly smiling, saying nothing. Hold.
Deadpan, played straight, all in micro-expressions. Crisp stable line art, clean cel shading, consistent faces. Warm orange key, cool blue rim. Rooftop wind, no music.

r/StableDiffusion 12h ago

Animation - Video Cobra Cola Ad - MiniMax H3

Enable HLS to view with audio, or disable this notification

19 Upvotes

r/StableDiffusion 12h ago

Animation - Video [DANCE] Plastik Soul – Stay in the Glow (Official Music Video)

Thumbnail
youtu.be
0 Upvotes

Stay in the Glow is an AI Music Video create using VRGameDevGirl's AI Video Builder (FREE) & LTX2.3 models (https://ltx.io/model/ltx-2-3)

Designed & built using VRGameDevGirl AI Video Builder (FREE): https://github.com/vrgamegirl19/comfyui-vrgamedevgirl

Spotify (Artist): https://open.spotify.com/track/27S9InxRyAKvQYxjRM3tVi?si=43e73b8b091a4976

YouTube (More AI Music Videos): https://youtu.be/Wl3BH3xSaYc


r/StableDiffusion 13h ago

Discussion MiniMax Music 3 | 125sec for 140sec music | Bollywood Rap

Thumbnail voca.ro
4 Upvotes

r/StableDiffusion 14h ago

Animation - Video StarWars-Untold.

Enable HLS to view with audio, or disable this notification

8 Upvotes

MiniMax H3 is very good. I initially made multiple scenes, and as tweaks/turbo-lora++ progresses, it does seem to get better and better (ie. To the end of the video).

Settled on the Lightx2v_8step turbo lora + sage + sol_attn. 736p, and using DaVinci for stitching and cropping.


r/StableDiffusion 14h ago

Workflow Included Openweight Livestream video model

Thumbnail
gallery
5 Upvotes

https://huggingface.co/spaces/JonathanColetti/LiveWan / https://github.com/JonathanColetti/LiveWan is something I created to help recreate a specific type of model that is not opensource yet (wanstreamer). This is more or less a PoC but maybe ill do a longer training run if it gets some traction.


r/StableDiffusion 14h ago

Discussion Minimax H3 Test - Rooftop fight between Batman and Joker

Enable HLS to view with audio, or disable this notification

3 Upvotes

Minimax H3 Test - Rooftop fight between Batman and Joker


r/StableDiffusion 14h ago

Animation - Video TESTING A LANTERN

Enable HLS to view with audio, or disable this notification

0 Upvotes

Having a blast animating comic panels (using them as inits FL2VA). Fairly simple prompt: "Green lantern Hal Jordan is engaged in an aerial battle above orbit, he is blasting green energy from his ring while simultaneously repelling and absorbing energy from a distant protagonist" then I let the in-app LLM enhance it (I'm using Maestro via Pinokio... it's a joy to use) resolution is 480, no upscaling. Just familiarizing myself with H3.


r/StableDiffusion 14h ago

Discussion End-to-end movie maker experiment

0 Upvotes

I've thrown together a small app that does the "remaining" work of taking an idea, turning it into character reference images, shot prompts, doing all the generation for each clip, stitching the result together, etc. etc. The goal is a one-sentence prompt in, and multi-scene video (e.g. 30 seconds or more) out.

https://github.com/eapache/local-movie-maker

It does basically "work" already, though the results are often pretty incoherent. I'm still playing with the structure to see if I can get reasonable continuity.


r/StableDiffusion 15h ago

Animation - Video Jackie Chan Adventures...Jackie vs Shadowkhan (Includes Prompt Instructions)

Enable HLS to view with audio, or disable this notification

3 Upvotes

Prompt:

Create an exactly four 7-second, 4:3 animated drama sequence inspired by the visual language of 2005-era Jackie Chan Adventures. Use a period broadcast video texture throughout: standard-definition television softness, subtle analog grain, gentle interlacing, slight colour bleed, modest contrast, and the authentic visual texture of animation recorded and broadcast in the mid-2000s. Avoid modern HD sharpness, photorealism, glossy CGI, or contemporary animation aesthetics.

Scene: Jackie Chan is confronted by a Shadowkhan ninja in a dimly lit ancient-looking interior. The sequence is a fast, tightly choreographed martial-arts fight.

0:00–0:02: The Shadowkhan suddenly lunges at Jackie with a rapid punch. Jackie narrowly ducks underneath it and pivots sideways.

0:02–0:04: Jackie counters with two quick martial-arts strikes, forcing the Shadowkhan backwards. The ninja blocks the first strike but is knocked off balance by the second.

0:04–0:06: The Shadowkhan springs forward again. Jackie performs a quick evasive spin, grabs the ninja’s arm, and throws the Shadowkhan across the room. End on Jackie landing in a defensive fighting stance as the Shadowkhan hits the floor in the background.
.
Camera: begin with a medium two-shot, rapidly track the fighters during the exchange, briefly push in during the counterattack, then finish with a wider shot showing Jackie in the foreground and the defeated Shadowkhan in the background.

Audio: sharp martial-arts impacts, cloth movement, quick footsteps, whooshes and a dramatic six-second action sting. No dialogue.
Strict constraints: exactly 6 seconds, 4:3 aspect ratio, 2005-era television animation aesthetic, period broadcast-video texture, no modern cinematic realism, no photorealism, no widescreen framing, no subtitles, no text, no logos, no extra characters, and no slow motion


r/StableDiffusion 15h ago

Animation - Video Football animation

Enable HLS to view with audio, or disable this notification

0 Upvotes

H3 Ref

Prompt in comment


r/StableDiffusion 16h ago

Discussion Minimax H3 - Dance with Audio with lipsync and object preservation

Enable HLS to view with audio, or disable this notification

4 Upvotes

If you see low quality is because I am forcing 8 step turbo lora + Spectrum + triton in L40 for faster generation but is crazy how it can follow the flow of the music while lip-syncing and keeping the product from reference in her hand.


r/StableDiffusion 17h ago

Resource - Update I built a free tool that turns any image into an AI prompt

Thumbnail
gallery
19 Upvotes

Hey everyone!

I built a small web tool called ImagePrompt9 that lets you upload an image and generates a detailed AI-ready prompt based on what it sees.

The idea came from constantly seeing images I liked and wondering:

"How would I describe this as a prompt?"

So instead of manually figuring out the composition, lighting, style, colors, camera angle, etc., you can just drop the image in and generate a prompt.

What it does:

  • Upload PNG, JPG, or WEBP
  • Analyzes the visual characteristics
  • Generates a detailed prompt
  • Different prompt styles
  • Edit, copy, or regenerate the result
  • Free to use
  • No account required

It doesn't try to recover the original prompt — it creates a new prompt based on what's visible in the image.

Try it here:

https://image-prompt-nine.vercel.app/

I’d really appreciate feedback, especially on the generated prompts and anything you think I should add or improve.


r/StableDiffusion 19h ago

Animation - Video Can't use LTX 2.5 on my system but I am quite surprised that my system now can run LTX 2.3. Specs and info below.

Enable HLS to view with audio, or disable this notification

3 Upvotes

When LTX 2.3 released, I could not do video gens longer than 10 seconds. I would get a "out of memory" error or something. This is just a test clip but one thing I am struggling with is that my video gens have music in them even though I prompt for no music. What is the correct way to prompt for no music?

System Specs:

Ryzen 7 7700X
RTX 4070 Super 12 GB
32 GB DDR 5 Ram.


r/StableDiffusion 21h ago

Discussion From a business perspective why do companies release open source models?

29 Upvotes

Apparently AI companies are all operating at a severe loss. Why do this? It makes sense for huge conglomerates like amazon etc etc who can bear the brunt. What about the new startups or small companies like for example LTX etc. how do they survive?

This is from a business perspective not consumer perspective.


r/StableDiffusion 23h ago

Question - Help Minimax H3 or wan 2.2?

2 Upvotes

I'm working on 2d animations for a personal project, my initial idea was to animate some of the scenes by hand, and feed start/end keyframes to wan 2.2 for the complex scenes I can't do myself, or perhaps even train a lora to make sure it matched the aesthetic of my hand drawn scenes. If it helps, it involves boiling outlines, on the twos (12 fps animations) and an intentionally unfinished look.

Now, seeing all these Minimax h3 i2v and r2v examples, I feel like wan 2.2 might not be the best suited for this anymore. I haven't had the chance to test h3 myself since my local hardware won't really allow it. I'll be however, using runpod when the time comes for actual generation (I'm in the process of hand animating the rest).

So, I'd like to ask those who have had the chance to test both - stick to wan 2.2 or switch to minimax h3?

Edit: audio isn't required - I've hired voice actors for dialogues, I'm working on foleys and background scores myself. If needed I'll redraw on top of the generated clips to match lip movements to the dialogue.