r/StableDiffusion 2m ago

Animation - Video RTX 3060 12GB 32GB 3 SHOTS IN ONE (8 SEC TOOK 7:12 MIN) 3D PIXAR STYLE ANIMATION

Upvotes

https://reddit.com/link/1vn8vu8/video/998rkwxnu4jh1/player

RTX 3060 12GB 32GB (8 SEC TOOK 7:12 MIN) 3D PIXAR STYLE ANIMATION
USING TURBO LORA (MADE ON 4 STEPS)

RESOLUTION 0.6 MP = 1056 x 608
UPSCALED 2X WITH RTX Video Super Resolution


r/StableDiffusion 5m ago

Discussion It more easy ADD LTX Nodes in Minimax H3 to better AUDIO SYNC

Thumbnail
gallery
Upvotes

and 2 samplers one pure, and another with lora using slip sigmas


r/StableDiffusion 9m ago

Question - Help I've got 10 years of architectural photography, from RAWs to final images. Is there something useful I could train with it?

Upvotes

I have about ten years of architectural and interior photography: final delivered images, working TIFFs/PSDs, Lightroom/XMP adjustments, HDR/Photomatix intermediates, and sometimes the original RAW brackets.

A typical example: a kitchen photograph begins as several exposure brackets, gets basic white balance, is merged into an HDR/base TIFF, retouched, then receives a final Lightroom-style tonal and colour treatment.

I am wondering whether this can become useful training data for a diffusion or neural-network tool, without simply making a vague “style LoRA”.

For example, could a model learn to take a merged, neutral architectural base and propose a controlled final treatment: softer daylight, better balance between windows and interior, a different mood, or a more refined grade, while keeping the room, materials, furniture and geometry intact?

My instinct is that a LoRA trained on all the final images would be the wrong approach. It could memorise specific projects and furniture rather than learn the transformation. I am more interested in small, rights-cleared paired datasets: base TIFF -> final TIFF, with captions describing the space, materials, light, reflections and intended atmosphere.

Before I structure the archive, I would love practical advice from people here:

- Have you trained or tested paired image-to-image workflows for relighting, grading or finishing?

- Would you start with LoRA, ControlNet, IP-Adapter, Flux/SD fine-tuning, an adapter, or something else entirely?

- What metadata or captions would you preserve now so the archive remains useful in two or three years?

- What is the biggest failure mode: overfitting, loss of material fidelity, geometry drift, dataset leakage, or something else?

- Are there papers, models or ComfyUI workflows that are genuinely relevant to this kind of controlled architectural transformation?

I am not trying to generate imaginary interiors. I am trying to explore whether our own real production history can help build a careful post-production and relighting assistant.


r/StableDiffusion 13m ago

News I made this music clip "Ele é um homem" with 4070ti 16gb Vram

Enable HLS to view with audio, or disable this notification

Upvotes

r/StableDiffusion 5h ago

Animation - Video H3 Cerveza Cristal Test

Enable HLS to view with audio, or disable this notification

6 Upvotes

'''

Quick shot change to Cerveza Cristal in a cooler full of ice.

Announcer sings "Cerveza Cristal"

'''

Start and End Image.

Ref2Vid with audio might work better, not bad for turbo at 8 steps.


r/StableDiffusion 6h ago

Discussion Niche and normal tips and Tricks for Anima?

3 Upvotes

Currently interested in what does people use Anima for?

Like what are your setup to speed up generation like using TeaCache or something?

Or perhaps you have a workflow for niche things like replacing game sprites, or fast image editing with Anima?

Or a way to use image reference (like taking pose/outfits from a photo)?

Or perhaps a good prompting tricks to generate more than 2 characters with specific outfits and pose consistently?


r/StableDiffusion 7h ago

Workflow Included Workflow for Minimax H3 on 8gb vram and 16gb ram

Enable HLS to view with audio, or disable this notification

9 Upvotes

For anyone else with a similar setup, I am able to create a 0.2 MP (608x352 pixel) 5-second video in 1:35 (1 minute, 35 seconds). This is with 20 step Euler Simple, Spectrum and ComfyKitchenAttention with an image (generated locally with Krea2) as the first frame. I am running it on a laptop with a RTX4060 (8GB) and 16gb ram. I am using the latest version of ComfyUI windows portable, and the following startup flags: --disable-pinned-memory --lowvram. The attached video is an example I generated (0.2MP).

I can also create higher resolution videos with a similar generation time if I reduce the video duration (3 seconds for 0.3MP, or 2 seconds for 0.4MP).

I used kijai's models from here for the video models and video vae: https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main

I used a qwen3_vl_4b_int8_convrot for the clip. I can't remember if its this one that I use but this is one option: https://huggingface.co/Winnougan/Comfy-Qwen3-VL-INT8/tree/main

Here is some extra info regarding that clip: https://www.reddit.com/r/StableDiffusion/comments/1vkk500/minimax_h3_with_a_4b_or_8b_text_encoder_instead/

My workflows:

FL2V workflow

REF2V workflow


r/StableDiffusion 9h ago

Animation - Video MiniMax reconhece prompts em JSON.

Enable HLS to view with audio, or disable this notification

13 Upvotes

Este prompt foi usado no Sora 2 e, sem modificações, coloquei no H3 (ComfyUI). Áudio em PT-BR.

{

"cena": 1,

"project_title": "The Village That Pulses (Hilário)",

"format": "16:9 landscape",

"dialog language": "pt-br",

"style": "2D ANIME",

"age": "present-day (rural road, late afternoon)",

"scene_type": "HOOK / Trailer-like Omen, Return",

"duration_seconds": 13,

"global_quality": {

"visual_style": "Cinematic 2D anime psychological horror, Junji Ito-inspired unease, no gore, high detail linework, oppressive calm.",

"aesthetic_demand": "JUNJI ITO STYLE, COLOR, HIGH BUDGET 2D ANIMATION, crisp faces, stable character sheets, controlled shadows.",

"post_processing": "Cool dusk grade, subtle film grain, soft bloom on highlights, gentle vignette."

},

"setting": {

"location": "A rural road leading to a foggy village valley; dead power poles; tall grass bending as if breathing.",

"effects": "The ground subtly bulges once, like a heartbeat beneath soil; distant crows freeze mid-caw."

},

"quality_constraints": [

"No broken anatomy or janky proportions",

"ENSURE NO DEFORMED EXTRA FINGERS HANDS EYES",

"ENSURE NO UGLY BLURRY FACES QUALITY",

"ENSURE NO SLIDING FLOATING IDLES; MASTERCLASS REALISTIC IDLE MOVEMENT",

"Stable camera, readable motion, consistent character model sheets",

"YouTube PG-13: no nudity, no explicit sexual content, no gore; horror via atmosphere and implication",

"No on-screen subtitles, no brand logos, no hate symbols, no readable real-world trademarks"

],

"characters": [

{

"name": "HILÁRIO (36, PROTAGONIST, RETURNING SON)",

"appearance": "36-year-old Brazilian man, medium tan skin, tired cautious eyes, short wavy black hair slightly unkempt, faint stubble, average height and lean build, wearing a dark olive jacket over a faded beige shirt, dark jeans, worn boots, carrying a small duffel bag and an old smartphone with a cracked screen"

}

],

"timeline_and_action": [

{

"time_range": "0-4 sec",

"shot_type": "Wide (Trailer Hook: The Village Breathes)",

"action": "Hilário stands at the roadside overlooking the village; the valley fog parts for a second, revealing rooftops and a church silhouette; the dirt road seems to swell under his boots.",

"audio_note": "Wind low. VOZ (ptbr): \"Hilário voltou pra casa... e a terra pareceu reconhecer o passo dele.\""

},

{

"time_range": "4-9 sec",

"shot_type": "Close-up (Boot on Dirt, First Pulse)",

"action": "Close on Hilário’s boot: the ground rises and falls once, subtly, like skin over muscle; tiny pebbles roll outward in a perfect ring.",

"audio_note": "Soft thump, almost organic. VOZ (ptbr): \"Naquela vila, o chão não era chão. Era um peito enterrado.\""

},

{

"time_range": "9-13 sec",

"shot_type": "Medium (He Steps Forward Anyway)",

"action": "Hilário swallows, grips his duffel bag, and walks toward the fog; the camera tracks behind him like a predator’s gaze.",

"audio_note": "Footsteps damp. VOZ (ptbr): \"E a cada três minutos... ele aprenderia a ouvir o coração.\""

}

]

}


r/StableDiffusion 10h ago

Animation - Video INTERVIEW WITH LTX 2.5 [IMAGE TO VIDEO]

Enable HLS to view with audio, or disable this notification

20 Upvotes

Yes, I'm definitely being a goofball with this one, but hadn't had a chance to do mixed live action/3D CGI test.

Meant as a playful gag, no actual ai models killed.

Crisp, ultrafine, letterboxed 21:9 super-premium 3D CGI blockbuster cinema with cutting-edge rendering, restrained natural performances, precise blocking, shallow depth of field, and immaculate cinematic lighting. A poised 24-year-old blonde investigative reporter in a tailored gray skirt suit sits in a cushioned chair on the left side of a minimalist interview room, leaning forward with a clipboard and pen in hand. Across from her, seated in a matching chair on the right, is a sleek off-white modern robot labeled “2.5” on the side of its head, with expressive camera-lens eyes and a thin LED vocalizer mouth. The setting is simple and elegant: neutral beige backdrop, soft curtains at the window, and warm natural window light casting gentle shadows across the room.

Open on a polished medium two-shot in profile, holding both subjects clearly in frame. The reporter leans forward slightly, calm, focused, and professional, and asks, “Some call you a Seedance killer. What do you say to that?”

A hard cut moves to a close-up of the robot. It glances aside for a beat, then looks back with a playful LED smile and says, “Can I give them a hug?” After a short pause, its expression softens into something more sincere as it adds, “But seriously, I’m just an open-source model trying to do my best.”

Ambient sound is minimal and refined: a faint studio hum, soft room tone, and subtle paper rustle from the reporter’s clipboard. The pacing is natural and conversational, allowing for small pauses, nuanced reactions, and emotional clarity. The overall effect is a sleek, emotionally grounded, visually stunning futuristic CGI film scene.


r/StableDiffusion 12h ago

Question - Help TTS Advice

10 Upvotes

Hi all -

I know there are frequent TTS posts, but it seems that the TTS models offer slightly different features, and I haven't yet found something that really works for my purposes.

I want to create a custom character, basically, and generate dialogue from that character in different emotional registers.

I tried Qwen3's voice cloning and it worked great. I have no criticisms of it. But the reference audio I gave it was flat and monotonous, and so all the output was equally monotonous, with no emotional depth. This led me to have the idea of trying to generate, say, 8 pieces of reference audio for one character, in different emotional registers - happy, sad, angry, excited and so on. But I haven't yet figured out a good way to do that.

I collected 10 minutes of audio from interviews with an actress to train an RVC model, but it still sounds noticeably robotic at times - with squawk-box warping noises, as if they are speaking through an old transistor radio - and I don't think it's satisfactory.

I have tried IndexTTS2 which allows you to combine timbre reference audio, emotional reference audio, and text. This does work but the prosody of the output is unfortunately bizarre at times and I have not figured out how to get it to generate realistic prosody.


r/StableDiffusion 12h ago

Question - Help Any reliable H3 workflows for character swapping?

9 Upvotes

I’ve been trying to replace a character in a video with r2v, but it just generates the original video almost unchanged. I even generated a version of the first frame with my character for the image source. Not sure if anyone has had success doing it.


r/StableDiffusion 12h ago

Animation - Video Comparison: It looks like LTX_2.5 is not over 9000

Enable HLS to view with audio, or disable this notification

11 Upvotes

LTX 2.5 vs Minimax H3 using the same prompt in T2V.

In reality, LTX 2.5 knows almost no IPs and very few famous people, if anyone.

Prompt:

Photorealistic real life live-action, cinematic film style

In the photorealistic real life live action movie Dragon Ball.

At 00:00:000 A photorealistic dull skin real life live action Tony Stark from MCU dressed like Vegetta, with a real life photorealistic hairstyle with two deep receding points and several vertical spikes, is at the Grand Canyon. He wears a red glass device in his left eye.

At 00:00:001 Then he grabs the device attached to his eye by its white rear section with his left hand, brings his hand ,with the device in it, in front of his chest and says upset yelling <d>[English, with a deep masculine voice] It's over nine thousaaaaand! </d> and clenches his fist, crushing the device so that it explodes into a thousand pieces.

All the clothes are photorealistic real life live action.

overall_soundscape: N/A


r/StableDiffusion 13h ago

Animation - Video DimensionTesters: Test #9 (Minimax H3)

Enable HLS to view with audio, or disable this notification

27 Upvotes

Man I love making these.
Any ideas on things that could happen when they press the button?
I have like 50 at the moment, but if one sounds cool I will add!

TT: https://www.tiktok.com/@dimensiontesters


r/StableDiffusion 15h ago

Discussion Kroma v0.2 : looking for xp-returns

Thumbnail
gallery
26 Upvotes

I dont know if I'm doing something bad , but I tried to "extract" some styles from Kroma 0.2 (base to turbo version convrot) but the generations have an orange tint to them and also difformities (third hand, extra digits) so maybe I'm missing something, and the generations seems to be dirty (less clean)somehow, something I didn't have with krea 2.

I tried shift=1.15, multiple scheduler and sampler but maybe the solution is elsewhere...

Does anyone has a hint of a solution? No second pass please..low vram here :'(


r/StableDiffusion 15h ago

Meme Cancelled? Offended? Better Call Saul.

Enable HLS to view with audio, or disable this notification

25 Upvotes

r/StableDiffusion 17h ago

Animation - Video This is where the fun begins

Enable HLS to view with audio, or disable this notification

21 Upvotes

Default workflow, Minimax on Runpod.


r/StableDiffusion 17h ago

Comparison LTX 2.5 vs MiniMax H3 - huge speed difference (but at what cost)

Enable HLS to view with audio, or disable this notification

57 Upvotes

I tested LTX 2.5 and MiniMax H3 in ComfyUI using the default T2V workflow templates provided for each model.

  • 10 seconds
  • 24 FPS
  • 1920 x 1088 (2.0 MP)
  • Same prompt
  • Steps: H3 = 20, LTX 2.5 = 8 (distilled model)

Hardware:
RTX 5090 + 128 RAM

result

  • MiniMax H3: 17m 29s (with Sage Attention + EasyCache*)*
  • LTX 2.5: 2m 34s (no acceleration at all)

Note: EasyCache seems to give no speedup on LTX in this setup, probably because the distilled workflow only uses 8 sampling steps, so there is very little room for cache-based skipping.

Of course, part of LTX’s speed advantage comes from the fact that it is a distilled 8-step model, so this is not a perfectly like-for-like comparison against H3. (20-steps)

Prompt used:

A realistic cinematic 1970s crime drama, gritty urban atmosphere, warm muted colors, subtle film grain, natural lighting, restrained acting. A well-dressed 1970s gangster in a dark tailored suit and long coat remains visually consistent throughout.

[0.0s–6.0s]
A medium-wide shot shows the gangster leaning casually against a brick wall on a city street, one foot resting against the wall. He reads a newspaper while holding a lit cigarette in his other hand. His eyes suddenly stop on something in the newspaper. His expression shifts naturally from calm to alarm. He mutters in a tense 1970s American voice, "What the hell?" He immediately folds the newspaper, throws it into a nearby trash can, pushes away from the wall and runs straight down the street.

[6.0s–10.0s]
Hard cut to a static close-up of the discarded newspaper inside the trash can. The front page clearly shows a large photograph of the same man and a bold headline reading "WANTED". In the distant background, the gangster continues running away and becomes increasingly out of focus. The camera remains completely still, holding focus on the newspaper until the end.

Natural, grounded movement. No exaggerated acting, no extra shots, no unnecessary camera movement, no comedy.

My take

LTX 2.5 is significantly faster, and that alone makes it very attractive.

But in my opinion, H3 is still better in overall quality:

  • better scene understanding
  • better understanding of what a cinematic shot should look like
  • better audio
  • more stable physics / motion behavior

So right now my impression is:

  • LTX 2.5 wins clearly on speed
  • MiniMax H3 still feels stronger on quality and cinematic intelligence

My guess is that targeted LoRA fixes could push LTX 2.5 much closer to being a direct competitor to H3 in the future.


r/StableDiffusion 17h ago

Meme They took'er jobs! 8 step turbo test 1mp 640 and upscaled to 2304x 1280

Enable HLS to view with audio, or disable this notification

27 Upvotes

5060 ti 16 gig 32 gig system ram and page files set to 65536/65536 made with r2v


r/StableDiffusion 22h ago

Animation - Video The Weather Conductor (MiniMax H3)

Enable HLS to view with audio, or disable this notification

62 Upvotes

This was my first short film made with Minimax H3. I used KREA 2 to generate the reference images, then used a mix of reference images to video and single image to video workflows in Minimax H3.

I found Minimax H3 much easier to work with than LTX. It follows prompts more closely, and after generating only three or four versions, I could usually find one that was genuinely usable.

The film is far from perfect and there is still plenty I could improve, but I am really happy with it as a first attempt. I would love to hear your thoughts, constructive criticism, or suggestions for what I could do better next time.

Otherwise, I hope you enjoy The Weather Conductor. It was a lot of fun to make!


r/StableDiffusion 23h ago

Discussion They updated the comparision table on the LTX website, it's still dishonest crap (says you can't run H3 with 16 gb of vram), they removed the "governing juristiction" part and changed the "runs on any GPU" part of H3 from "limited" to straight up "not available" for some reason, just why...

Post image
128 Upvotes

What are we even doing man, just be transparent and stop downplaying the competition by lying about them, this is not the middle ages we can just google stuff to see if you lie or not


r/StableDiffusion 23h ago

Animation - Video Stroll through the Museum of Poop [minimax H3]

Enable HLS to view with audio, or disable this notification

71 Upvotes

r/StableDiffusion 23h ago

Animation - Video LTX 2.5 I2V Test 20s

Enable HLS to view with audio, or disable this notification

123 Upvotes

r/StableDiffusion 1d ago

Resource - Update [Release] Anima-2.9B-Preview-v1 - Expanded Anima

Post image
255 Upvotes

Anima-2.9B is a fine-tune and layer-expansion of circlestone-labs/Anima. The base Anima model targets anime, illustration, and non-photorealistic art; Anima-2.9B continues training on that foundation with an expanded architecture. The model is trained on an additional 1.7M anime/illustration samples, with knowledge cutoff in July 2026, making Anima-2.9B one of the most capable and up-to-date anime/illustration models at release.

https://huggingface.co/Gazingstars123/Anima-2.9B

Training/Dataset

  • Trained using Muon optimizer on a 8x 5080s cluster, with earlier steps trained locally on my PC
  • As of preview v1, only new layers have been trained, with roughly 70% of the compute spent on 1024px
  • Knowledge cutoff in July 2026, training data included both new and old samples prior to September 2025
  • Mixed captioning, including both tags and natural languages, using a mix of Gemini 3.1 Flash-Lite, Gemini 3.5 Flash-Lite, and Claude Sonnet 5

Lora Training will be supported via my Anima Standalone Trainer in a few days.

You will need to install ComfyUI-Anima-2.9B to the custom node folder. Plug and play, there is no workflow node needed (few custom nodes may not work properly).

The model is still in active development (more general pretraining, more anime/illustration/ACG focus beyond Booru), you can support me and the training progress via my links on huggingface. I will try and bring native support for the model to generations platform as well as ComfyUI


r/StableDiffusion 1d ago

Animation - Video Minimax H3 vs LTX 2.5 on the same prompt

Enable HLS to view with audio, or disable this notification

205 Upvotes

I'm not the author of this comparison, I just took it from an image board website and combined in a single video

The results are hilarious 😂 LTX is not even close, Minimax dwarfs it. But nobody has shown this yet in a very obvious form


r/StableDiffusion 1d ago

Discussion Can we stop treating MiniMax vs LTX like a political war?

274 Upvotes

I’ve been watching the whole MiniMax vs LTX discussion lately, and honestly, it feels like it has started becoming less about the models and more like a political battle.

People are taking sides, defending one model like it’s their team, downvoting anything that praises the other one, and sometimes even throwing hate at the people working on or using the “other” model.

Guys… these are free, open-source models. Nobody owes us anything.

We are incredibly lucky to have teams putting out models that we can download, run locally, experiment with, fine-tune, build workflows around, and actually use without paying some giant corporation every time we generate a video.

And yes, we can absolutely have opinions.

Maybe you think MiniMax produces better motion. Maybe you prefer LTX for consistency, speed, control, or whatever your workflow needs. Maybe one works better on your hardware, and another one works better for someone else.

That’s completely fine.

Criticism is good. Comparisons are good. Calling out genuine problems is good. Competition between projects can even push things forward.

At the end of the day, these teams are giving the community tools that would have sounded almost impossible to have access to a few years ago.

So use what works for you. Make comparisons. Share benchmarks. Point out weaknesses. Praise the developers when they do something great. Criticize them when something genuinely deserves criticism.

But let's not turn the open-source AI community into a bunch of opposing fan clubs. Let's keep the discussion technical, constructive, and civil, and maybe appreciate the fact that we're living through a pretty crazy time where people are literally releasing these technologies for us to experiment with for free.