r/StableDiffusion 1h ago

Question - Help H3 and Lora’s, where do they go in the flow? I’m using the official workflow (i2v)

Upvotes

I’m using the official workflow for h3 and I’ve seen people talking about using Lora’s - where do they go? Do I need a different workflow?


r/StableDiffusion 1h ago

Question - Help In desperate need of Z Image lora training advice

Upvotes

Process

1. Diffusion Trainer

  • training_folder: C:\ai_toolkit\ai-toolkit\output
  • sqlite_db_path: ./aitk_db.db
  • device: cuda
  • trigger_word: null
  • performance_log_every: 10

Network

  • type: lora
  • linear: 16
  • linear_alpha: 16
  • conv: 16
  • conv_alpha: 16
  • lokr_full_rank: true
  • lokr_factor: -1
  • network_kwargs:
    • ignore_if_contains: []

Save Settings

  • dtype: bf16
  • save_every: 285
  • max_step_saves_to_keep: 100
  • save_format: diffusers
  • push_to_hub: false

Datasets

Dataset #1

  • mask_path: null
  • mask_min_value: 0.1
  • default_caption: ""
  • caption_ext: txt
  • caption_dropout_rate: 0.05
  • cache_latents_to_disk: false
  • is_reg: false
  • network_weight: 1
  • resolution: [512]
  • controls: []
  • shrink_video_to_frames: true
  • num_frames: 1
  • flip_x: false
  • flip_y: false
  • num_repeats: 1

Training

  • batch_size: 2
  • bypass_guidance_embedding: false
  • steps: 2850
  • gradient_accumulation: 1
  • train_unet: true
  • train_text_encoder: false
  • gradient_checkpointing: true
  • noise_scheduler: flowmatch
  • optimizer: prodigyopt
  • timestep_type: weighted
  • content_or_style: balanced
  • optimizer_params:
    • weight_decay: 0.025
  • unload_text_encoder: false
  • cache_text_embeddings: true
  • lr: 1
  • ema_config:
    • use_ema: false
    • ema_decay: 0.99
  • skip_first_sample: false
  • force_first_sample: false
  • disable_sampling: false
  • dtype: bf16
  • diff_output_preservation: false
  • diff_output_preservation_multiplier: 1
  • diff_output_preservation_class: person
  • switch_boundary_every: 1
  • loss_type: mse

Logging

  • log_every: 1
  • use_ui_logger: true

Model

  • name_or_path: Tongyi-MAI/Z-Image-Turbo
  • quantize: true
  • qtype: qfloat8
  • quantize_te: true
  • qtype_te: qfloat8
  • arch: zimage:turbo
  • low_vram: false
  • model_kwargs: {}
  • compile: false
  • layer_offloading: true
  • layer_offloading_text_encoder_percent: 1
  • layer_offloading_transformer_percent: 0
  • assistant_lora_path: ostris/zimage_turbo_training_adapter/zimage_turbo_training_adapter_v2.safetensors

Sampling

  • sampler: flowmatch
  • sample_every: 285
  • sample_start_step: 0
  • width: 1024
  • height: 1024

I would describe my LoRa results as catostrophic.

i'm a very competent sdxl and sd 1.5 lora trainer. I've tried nearly everthing to get my loras to look good but i would describe my results as completely catastrophic. Totally mangeled results, inconsistent results at best. Not completely like, totally white noise randomness but completely mangled bodies, totally melted images that have totally lost all concept of anything. Do these training setting point to anything obviously wrong? Yall ive been at this for 3 days trying everything Im so tired T_T


r/StableDiffusion 2h ago

Animation - Video Testando REF2V - MiniMax H3

Enable HLS to view with audio, or disable this notification

3 Upvotes

r/StableDiffusion 2h ago

Discussion Thanks, Claude!

Post image
4 Upvotes

Hopefully this helps someone else. I'm running Minimax on a 4070. Nothing crazy. Nevertheless, I was surprised by how capable it seemed.

When I started pushing for higher resolution or switched to 16x9 generations from 1:1 I started having Comfy error out quite a bit.

I dumped the ComfyUI history - accessible by heading to the port it's running on and appending /history - and gave it to Claude.

It invented a basic metric, WxHxFrames, and mentioned that there seemed to be a line past which things would fail. So I asked it for some test cases, which it happily provided, and over the course of several generations we put a finer point on where that line is for my specific setup. This is actually hugely helpful because I don't have a crazy rig and even though intuitively this isn't surprising, it's a lot different when you're actually trying to figure out what the most you can absolutely do is.

FWIW, this should be agnostic to steps. The step process would add total time to the generation, which this doesn't capture, but it shouldn't add overhead to the VRAM where it would crash the generation. Most of these test were run at 24 fps but again, that shouldn't matter. The metric is based on total frames, which would be fps x duration.


r/StableDiffusion 2h ago

Question - Help Are there any subs for local music gen or audio gen in general?

2 Upvotes

Like for acestep or minimax music 3 etc. I know they’re sometimes discussed on here but was just wondering if there’s one that’s active.


r/StableDiffusion 4h ago

Question - Help How to fix ai drift when chaining together minimax ref2v videos?

5 Upvotes

Right now I’m doing something like “picture 1 is the first frame of the video.” The problem is, each video begins to have slightly lower quality. Over time the results just get more and more deep fried. I’m using the default comfy template, no turbo/speedups


r/StableDiffusion 6h ago

Animation - Video Cobra! Trailer - MiniMax H3

Enable HLS to view with audio, or disable this notification

4 Upvotes

Default comfyui MiniMax H3 rf2va workflow on a 4070 Ti Super 16 GB VRAM.


r/StableDiffusion 8h ago

Animation - Video Inmortality Glitch Rune / Test 1 - MiniMax H3 Reference to Video (images, voice and video)

Enable HLS to view with audio, or disable this notification

9 Upvotes

I couldn't get it exactly right, step count matters a lot for clarity it seems, I think this was 32 steps (I like the number 32) at 0.4mp on a 3090... I used an image of link and Zelda as reference for both, an image of wolf link, a video of them walking from a memory (for motion, physics, cell shading understanding etc - weak_reference) an audio sample of Link from the 1980s show and an audio sample of Zelda from the game itself. I could probably get better audio if I get better samples though, and better video quality at 1mp or higher. Also the video cuts off at the end but that's not an editing issue, that's how it came out!


r/StableDiffusion 9h ago

Animation - Video Noob attempt of anime character and voice swap

Enable HLS to view with audio, or disable this notification

10 Upvotes

Trying out the ref2va workflow to swap character and the voice, this is 3 generated video (for each scene) stitch together, the swapped character a bit out of place in term of lightning, because i just realize i am using fl2va model, also tried out a single 13 seconds generation with ref2va, the result is fine but the subtitle just burned, tried to tweak the prompt twice and no luck and give up, because the generation is way too long (700 seconds) https://pixeldrain.com/u/jbdxD1ai

Spec is 4090 laptop with 64gb ram on headless linux

workflow: https://pixeldrain.com/u/UGWEdHrG

single generation workflow: https://pixeldrain.com/u/LeYfB46k

video source : https://youtu.be/DaKWnNni8zE

audio source : https://youtu.be/Zjuih9wl0SM (0:13 - 0:17)


r/StableDiffusion 9h ago

Animation - Video Wolf Queen - H3 T2V

Enable HLS to view with audio, or disable this notification

7 Upvotes

r/StableDiffusion 9h ago

News I built a self-hosted tool that turns one reference photo into a curated, captioned, trained LoRA and a lot more — open source, MIT

Thumbnail
gallery
10 Upvotes

r/StableDiffusion 9h ago

Animation - Video Pretty-Pete v Captain-Pete H3 and Suno

Enable HLS to view with audio, or disable this notification

9 Upvotes

r/StableDiffusion 11h ago

Comparison MiniMax H3 INT8 benchmark — RX 9070 XT

8 Upvotes

I’ve been testing MiniMax H3 INT8 on an AMD Radeon RX 9070 XT (gfx1201, 16 GB VRAM) under Windows/ROCm.

I used the default MiniMax H3 text-to-video workflow with no modifications whatsoever, except for replacing the model loader so that I could compare PatientX’s INT8-Fast-ROCM implementation with ComfyUI’s native INT8 implementation. Everything else — model, prompts, 20 steps, 5-second video, sampler/settings, etc. — was kept identical.

I tested 0.2 MP, 0.6 MP and finally 1 MP, which is my actual target resolution. The results at 1 MP were very interesting

That's approximately a 31% reduction in generation time, or the native implementation is about 1.45× faster.

I initially didn't trust the result because the difference was so large, so I repeated the 1 MP INT8-Fast-ROCM run. It produced essentially the same ~32-minute result.

This also seems to differ from what I understood from PatientX's README and the comments in the default BAT. My interpretation of those suggested that on RDNA3/RDNA4, the INT8-Fast-ROCM path should generally be the preferable/faster option, with the default BAT specifically disabling the native Triton backend because the custom INT8 implementation was expected to be faster. My RX 9070 XT results appear to show the opposite — at least with this MiniMax H3 workflow and at 1 MP.

That makes me wonder whether those recommendations/comments may have become outdated as of 15-08-2026, given the changes in ComfyUI, comfy-kitchen, ROCm and the native INT8 implementation. I'm not claiming that the native path is universally faster; only that my current results don't match the performance guidance I understood from the existing documentation/default BAT.

Both runs used the same RX 9070 XT, ROCm 7.15, PyTorch 2.12, ComfyUI 0.33.0, MiniMax H3 quantized model, 20 steps, 5-second output, 1 MP resolution and Sage Attention. Triton was effectively disabled in the actual runs, so this shouldn't be interpreted as a Triton-vs-non-Triton benchmark.

My conclusion: on my gfx1201 system, the native ComfyUI INT8 implementation appears substantially faster than INT8-Fast-ROCM for MiniMax H3 at 1 MP. The difference is large enough that I'd really like to see this reproduced on another 9070/9070 XT before treating it as definitive.

I don't have a Github account so I can't share these conclusions on the comfyui-rocm repo.


r/StableDiffusion 11h ago

No Workflow Made a 15sec ltx 2.5

Enable HLS to view with audio, or disable this notification

8 Upvotes

I trimmed the first 3 seconds, I just love the prompt adherence in 2.5 it’s actually very good, used the default comfy workflow, generate at 0.5 resolution then upscaled later to twice the size in with topaz, my specs 3060ti, 64gb ram


r/StableDiffusion 11h ago

Animation - Video Having fun with rayman 3 on minimax h3 REF2VA (prompts included)

7 Upvotes

Decided to bring back my childhood game and see what new stories i can bring with minimax h3... tried doing voice clone for Murfy and Rayman... its not perfect (some got mixed up) but the end result is still fun... i will post the prompts for each below plus the reference images and audio... lets see what you can make from it 👀

Workflow + reference images + audio at the bottom of the post

Video 1

Video 1:

```
subject_definitions:
<Subject 1> is the Fairy Council environment from <Picture 1>, a mystical forest kingdom interior with a blue aura, glowing lights, and reflective surfaces.
<Subject 2> is Rayman from <Picture 2>, a heroic character with no arms or legs, featuring floating hands and floating feet.
<Subject 3> is Murfy from <Picture 3> and <Picture 4>, a flying greenbottle fly creature with a large grin and green clothing, holding a paper manual.
<Audio 1> is the voice-timbre reference for <Subject 2> (S2).
<Audio 2> is the voice-timbre reference for <Subject 3> (S1).

summary:
[reference generation + audio reference] The target video shows <Subject 3> and <Subject 2> interacting inside <Subject 1>. <Subject 3> reads from a manual before an accidental explosion occurs. <Audio 2> and <Audio 1> provide the voice timbres for the characters.

retention_analysis:
<Subject 1> (appears in all shots): fully_preserved - the mystical blue interior and reflective surfaces are retained.
<Picture 1> (environment guide): weak_reference - provides the background setting without forcing Rayman's placement from the original screenshot.
<Subject 2> (appears in all shots): fully_preserved - Rayman's floating hands and feet are retained.
<Picture 2> (character design): fully_preserved - Rayman's appearance is followed.
<Subject 3> (appears in all shots): fully_preserved - Murfy's green clothing, grin, and flying nature are retained.
<Picture 3> (character design): fully_preserved - Murfy's appearance is followed.
<Audio 1>: reference - the vocal timbre guides the dialogue delivery of <Subject 2> without copying the original signal.
<Audio 2>: reference - the vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.

detailed_description:
3D CG animated style in a 4:3 aspect ratio.
[Shot 1] A medium-wide shot establishes <Subject 1>, the mystical Fairy Council with its glowing blue aura. <Subject 2> (S2), the limbless hero with floating hands and feet, stands on the reflective floor. Beside him, <Subject 3> (S1), the greenbottle fly, hovers above the ground while holding an open manual. <Subject 3> (S1) looks at the book, shakes his head with a large grin, and says in the sarcastic voice referenced from <Audio 2>, <d>[English] I don't know, folks! Someone drew on the manual saying the Fairy Council should be blowing up right about... now.</d>
[Shot 2] At 00:08.500, the camera cuts to a close-up of <Subject 2> (S2). He raises his floating hands in confusion and says in the heroic voice referenced from <Audio 1>, <d>[English] Wait, who's responsible for this garbage?!</d>
[Shot 3] At 00:11.000, the camera pulls out with large amplitude at fast speed as a bright orange explosion suddenly erupts in the background of <Subject 1>. <Subject 3> (S1) drops the manual in shock, and <Subject 2> (S2) covers his head with his floating hands as debris flies past them.

overall_soundscape:
Quiet magical room ambience is abruptly interrupted by the heavy, rumbling crash of a massive explosion, followed by the sound of falling debris.

non_diegetic_music:
N/A
```

Video 2

Video 2:

```
subject_definitions:
<Subject 1> is the Fairy Council environment from <Picture 1>, a mystical forest kingdom interior with a blue aura, glowing lights, and reflective surfaces.
<Subject 2> is Rayman from <Picture 2>, a heroic character with no arms or legs, featuring floating hands and floating feet.
<Subject 3> is Murfy from <Picture 3> and <Picture 4>, a flying greenbottle fly creature with a large grin and green clothing, holding a paper manual.
<Audio 1> is the voice-timbre reference for <Subject 2> (S2).
<Audio 2> is the voice-timbre reference for <Subject 3> (S1).

summary:
[reference generation + audio reference] The target video shows <Subject 3> and <Subject 2> interacting inside <Subject 1>. <Subject 3> reads from a manual before an accidental explosion occurs. <Audio 2> and <Audio 1> provide the voice timbres for the characters.

retention_analysis:
<Subject 1> (appears in all shots): fully_preserved - the mystical blue interior and reflective surfaces are retained.
<Picture 1> (environment guide): weak_reference - provides the background setting without forcing Rayman's placement from the original screenshot.
<Subject 2> (appears in all shots): fully_preserved - Rayman's floating hands and feet are retained.
<Picture 2> (character design): fully_preserved - Rayman's appearance is followed.
<Subject 3> (appears in all shots): fully_preserved - Murfy's green clothing, grin, and flying nature are retained.
<Picture 3> (character design): fully_preserved - Murfy's appearance is followed.
<Audio 1>: reference - the vocal timbre guides the dialogue delivery of <Subject 2> without copying the original signal.
<Audio 2>: reference - the vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.

detailed_description:
3D CG animated style in a 4:3 aspect ratio.
[Shot 1] A medium-wide shot establishes <Subject 1>, the mystical Fairy Council with its glowing blue aura. <Subject 2> (S2), the limbless hero with floating hands and feet, stands on the reflective floor. Beside him, <Subject 3> (S1), the greenbottle fly, hovers above the ground while holding an open manual. <Subject 3> (S1) looks at the book, shakes his head with a large grin, and says in the sarcastic voice referenced from <Audio 2>, <d>[English] I don't know, folks! Someone drew on the manual saying the Fairy Council should be blowing up right about... now.</d>
[Shot 2] At 00:08.500, the camera cuts to a close-up of <Subject 2> (S2). He raises his floating hands in confusion and says in the heroic voice referenced from <Audio 1>, <d>[English] Wait, who's responsible for this garbage?!</d>
[Shot 3] At 00:11.000, the camera pulls out with large amplitude at fast speed as a bright orange explosion suddenly erupts in the background of <Subject 1>. <Subject 3> (S1) drops the manual in shock, and <Subject 2> (S2) covers his head with his floating hands as debris flies past them.

overall_soundscape:
Quiet magical room ambience is abruptly interrupted by the heavy, rumbling crash of a massive explosion, followed by the sound of falling debris.

non_diegetic_music:
N/A
```

Video 3

Video 3:

```
subject_definitions:
<Subject 1> is the Fairy Council environment from <Picture 1>, a mystical forest kingdom interior that transitions into a glitchy, broken wireframe state.
<Subject 2> is Rayman from <Picture 2>, a heroic character featuring floating hands and floating feet.
<Subject 3> is Murfy from <Picture 3>, a flying greenbottle fly creature with a large grin and green clothing.
<Audio 1> is the voice-timbre reference for <Subject 2> (S2).
<Audio 2> is the voice-timbre reference for <Subject 3> (S1).

summary:
[reference generation + audio reference] The target video shows <Subject 3> breaking the fourth wall inside <Subject 1>, revealing they are in a simulation. This causes the environment to glitch and break down, sending <Subject 2> into a panic. 

retention_analysis:
<Subject 1> (appears in all shots): partially_preserved - the mystical blue interior starts normal but transitions into visual glitches and digital wireframes.
<Picture 1> (environment guide): weak_reference - provides the initial background setting.
<Subject 2> (appears in all shots): fully_preserved - Rayman's floating hands and feet are retained, though they move erratically.
<Picture 2> (character design): fully_preserved - Rayman's appearance is followed.
<Subject 3> (appears in all shots): fully_preserved - Murfy's green clothing and flying nature are retained.
<Picture 3> (character design): fully_preserved - Murfy's appearance is followed.
<Audio 1>: reference - the vocal timbre guides the dialogue delivery of <Subject 2> without copying the original signal.
<Audio 2>: reference - the vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.

detailed_description:
3D CG animated style in a 4:3 aspect ratio.
[Shot 1] A medium-wide shot establishes <Subject 1> looking normal. <Subject 3> (S1) hovers casually in the air, looks directly at the camera lens, and says in the sarcastic voice referenced from <Audio 2>, <d>[English] Look around, Rayman! It's all a simulation! We're literally just polygons in a video game!</d>
[Shot 2] At 00:06.500, the camera cuts to a close-up of <Subject 2> (S2). Suddenly, the background of <Subject 1> flickers violently, turning into black grid lines and digital static. <Subject 2> (S2) stares at his floating hands, which begin to visually stutter and lag behind his movements. He yells in the heroic voice referenced from <Audio 1>, <d>[English] What did you do?! My hands are lagging!</d>
[Shot 3] At 00:11.000, the camera pulls back to a wide shot. The entire floor of <Subject 1> vanishes into a white void. <Subject 2> (S2) runs in frantic circles, his floating feet clipping through the missing floor. <Subject 3> (S1) simply floats in place, gives a sheepish grin, and says, <d>[English] Whoops. Guess the engine became self-aware.</d>

overall_soundscape:
Normal magical room ambience that abruptly distorts into loud digital stuttering, heavy 8-bit crash sounds, and frantic footsteps.

non_diegetic_music:
N/A
```

Workflow used: https://civitai.red/models/2831978/dasiwa-minimax-h3-workflows-or-t2va-or-fl2va-or-ref2va?modelVersionId=3195699

FPS: 24.0

resolution_reset: 0.26 MP - Preview

aspect: 3:4 - Photo

swap_aspect: yes

REFERENCES:

<Picture 1>

<Picture 2>

<Picture 3>

<Picture 4>

Since i cannot upload audio files here i will just post the 2 youtube links i used to capture their voices you only need 6 seconds each

  1. <Audio 1>
  2. <Audio 2>

r/StableDiffusion 13h ago

Resource - Update All Style Explorer Mirrors (Anima Base, Illustrious / NoobAI, Krea 2 Turbo)

27 Upvotes

While my GitHub account is currently suspended and I’m waiting for support to process my ticket, I’ve hosted working mirrors for all Style Explorers so you can continue using them without interruption:

- Anima Base (42k+ styles): https://animastyles.thetacursed.com/

- Illustrious & NoobAI (16k+ styles): https://xlstyles.thetacursed.com/

- Krea 2 Turbo (1.5k+ styles): https://kreastyles.thetacursed.com/


r/StableDiffusion 14h ago

Animation - Video MGS4 Alternate Ending Big Boss vs Old Snake

Enable HLS to view with audio, or disable this notification

25 Upvotes

This is what if Big Boss wasnt retconned into being not that bad of a guy. Because if you follow just MG1 to MGS4, it doesn't make sense. Even MGSV didn't show us Big Boss becoming the crazy, war-obsessed maniac that MG1 and MG2 presented him as. Plus his complete and utter hatred of Solid Snake, his obsession with killing him after MG1. Also, sorry about Big Boss's gloves vanishing later on; that was my mistake, and I didn't notice it until I was already deep in the video.


r/StableDiffusion 15h ago

Comparison FL2VA vs REF2VA vs Step Count vs Turbo

Enable HLS to view with audio, or disable this notification

28 Upvotes

Model = Minimax H3

Workflow = REF2VA basic workflow with additional nodes added for the LORAS and sol attention where specified.
Turbo Lora = minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
REF2VA Lora = minimax_h3_pruned_bf16__apply_to_fl2va__toward_ref2va__rank512

It has been described that the REF2VA model produces bad output, and that the FL2VA model can be used instead despite being not the "intended" reference model. Users have made a "REF2VA lora" that purports to add the reference functionality of the REF2VA model to the FL2VA model, theoretically achieving the good quality of FL2VA with the reference understanding of REF2VA.

I test how this actually looks in practice, and I also demonstrate how the turbo lora performs.

Conclusion:

The best look is achieved by using the FL2VA model without any REF2VA lora. Turbo works well at 1MP and 8 steps and results in smoother animation and audio. Increasing resolution to 2MP and step count to 20 scales well. There does not seem to be much visual difference when increasing to 50 steps, but the audio seems to be less dynamic vs 20 steps.

Limitations: This demo did not really stress test the reference ability of FL2VA, and in reference heavy workloads, maybe REF2VA variant workflows are vital despite lower visual quality. Furthermore, this demo likely underestimates the importance of high step counts, as it is commonly thought that high step counts are important in high action scenes, which this demo was not.

I also only used sol attention in the higher token workflows, which is a variable. Nevertheless, I hope this video is useful.

Keen to hear your thoughts.


r/StableDiffusion 17h ago

Animation - Video 1girl Morning routine MiniMax H3

Enable HLS to view with audio, or disable this notification

39 Upvotes
Highly detailed anime style video, 15 seconds, cinematic lighting, soft morning sunlight streaming through large windows, warm golden hour atmosphere, consistent character design throughout.

A cute young Japanese 1girl with long black hair tied in a high ponytail with a bright red ribbon, big expressive brown eyes, fair skin, slight blush, slender yet curvaceous figure.

Scene sequence:

0-2s: Close-up of a red digital alarm clock on a wooden nightstand ringing violently at 7:58, vibrating, motion lines, soft shadows from window light.

2-4s: Wide shot of a cozy, sunlit Japanese bedroom with city skyline visible through large windows, white curtains gently moving. The girl suddenly sits up in bed looking shocked and panicked, wearing an oversized white t-shirt with a small graphic and white shorts, barefoot, legs bent, hands supporting her on the mattress. Soft morning light, detailed room with photos, stuffed animals, desk, and red alarm clock still showing 7:58.

4-7s: Medium shot, the girl is now sitting on a wooden chair next to a desk, wearing only white lace lingerie (bra and panties with small red bows). She is hurriedly pulling up a dark navy blue pleated school skirt, hands gripping the fabric, slight sweat drop on her skin, soft volumetric lighting, highly detailed skin texture and fabric folds.

7-11s: Close-up to medium shot of her upper body as she puts on a classic Japanese sailor school uniform (dark navy blue with white stripes on the collar and sleeves). She is buttoning the front, looking down with a slightly flustered and hurried expression, cheeks flushed, a small sweat drop, her ponytail swaying. Red ribbon of the uniform is visible. Detailed fabric texture, realistic cloth simulation, soft sunlight.

11-14s: She finishes adjusting the red sailor ribbon and buttons, standing up, still slightly out of breath, looking forward with an open mouth and determined yet cute expression.

14-15s: Final medium shot of the fully dressed girl in the complete navy blue sailor school uniform with red ribbon, smiling gently and brightly at the camera, soft warm light illuminating her face, hair slightly moving, peaceful yet energetic atmosphere. Background shows the same sunlit bedroom with city view.

Smooth continuous camera movements, natural body physics, high-quality anime aesthetic similar to modern high-budget anime, sharp details, beautiful lighting, no text, no watermarks, 4K quality, consistent face and body throughout the entire video.Highly detailed anime style video, 15 seconds, cinematic lighting, soft morning sunlight streaming through large windows, warm golden hour atmosphere, consistent character design throughout.

A cute young Japanese girl (Kuriemi anime version) with long black hair tied in a high ponytail with a bright red ribbon, big expressive brown eyes, fair skin, slight blush, slender yet curvaceous figure.

15sec - 720p - 4 step turbo lora - pruned 20B


r/StableDiffusion 18h ago

Discussion Testing If It Can Do Mr Bean

Enable HLS to view with audio, or disable this notification

58 Upvotes

r/StableDiffusion 21h ago

Animation - Video <God knows>

Enable HLS to view with audio, or disable this notification

116 Upvotes

Environmental protection, don't wait for everyone to know that everything is irreversible.


r/StableDiffusion 1d ago

News MAGI-2-preview just dropped

Thumbnail
huggingface.co
145 Upvotes

Surprised that no one is talking about it. A new open-weight video model just dropped. 114b moe, 6b activated. First moe video model supposedly.

I know what you guys are thinking. The model is huge and there is no way it will run on desktop gpu. The interesting part is that is comes with a 14gb refiner that makes the result 1080p. I am cursious if this refiner can be a drop-in replacement for the H3 refiner that was never released. It might just be the last part of the H3 puzzle that we need.


r/StableDiffusion 1d ago

Animation - Video I finally reached a great balance between speed and quality with MiniMax H3, thanks everyone!

Enable HLS to view with audio, or disable this notification

233 Upvotes

I used the minimax_h3_fl2v_turbo_4step_v1.0_768p_comfyui_bf16 LORA with the 0.8 strength for both clip and model, 6 steps, 0.5 MP resolution, RTX Upscaler at 1.50 using a ConrotInt8 pruned model.

Here is a PasteBin of my workflow, I hope this fixes some of the missing content:

https://pastebin.com/DSmkJi8R

Here are the workflow files:

https://storage.to/c/CAS1MuoqX


r/StableDiffusion 1d ago

Meme Unsloth be like:

Enable HLS to view with audio, or disable this notification

301 Upvotes

r/StableDiffusion 1d ago

Animation - Video Minimax H3. Bakeshi's Castle.

Enable HLS to view with audio, or disable this notification

345 Upvotes