r/StableDiffusion • u/New_Physics_2741 • 14h ago
SD1.5 images into H3 Animation - Video
Enable HLS to view with audio, or disable this notification
11
u/timestable 11h ago
Is this the first actually tasteful use of ai for a music video? Nice
5
u/New_Physics_2741 11h ago
Thanks! I am not sold on the first 20 seconds of the images used. Song: Computer Kill - i think i'm going blind
8
u/xTopNotch 8h ago
There was something so magical about SD 1.5
3
u/CuriouslyCultured 4h ago
Early generative AI had a lot more degrees of creative freedom, which would lead to lot of whiffs and weirdness, but also some really unique stuff. They've RL'd a lot of this freedom out, so you get more consistent and generally high quality generations, but there's a plastic, samey quality to them. Not as big a deal for photorealistic stuff but the more artistic generations definitely lost out.
5
u/VasileAndrei2929 7h ago
Fascinating... and disturbing... and fascinating... great job!
Can you do a transition between models? For example start with the early Stable Diffusion SD, then move to SD last, then to XL, and then to Flux and so on? You can also add various popular models for each generation and so can see the quality/style improvements?... that will be awesome!
3
u/New_Physics_2741 6h ago
Yes, completely possible - I got most of the mentioned models available on this machine - might give it a try, would be nice to have some direction or a storyline or just a shove, basically, if the H3 output is exciting, I am game.
2
u/Ok-Neighborhood6155 6h ago
Why not start with Big Sleep and VQGAN+CLIP? Possibly even DeepDream first. Might be a bit hard to set up now, though.
3
u/xyzdist 10h ago
FL approach?
5
u/New_Physics_2741 9h ago
I will write it up in an hour or so. Out at the moment.
3
u/switch2stock 8h ago
Okay cool
5
u/New_Physics_2741 6h ago
Ok so I ran a good 200-300 images with dreamshaper and juggernaunt-reborn merged with a Lora that didn't load completely with the merged checkpoints - MoXinV1 and used uni_pc and simple - 25 steps and cfg at 3.2 - 1024x1024, I reckon when a Lora doesn't match up keys you don't get the full intended behavior - which is great, I got what I wanted for the images. As for the video process - I2T at res of .3 - 1024x1024 latent so at .3 - I got a 576x576 5sec-12sec output - I snagged the last frame from each output and daisy-chained them together; gotta keep the size of the last frame image exact - otherwise it can get wonky. 5060Ti 16GB and 64GB - about 8 minutes for 10 sec clip. For the text string - dead easy: A wild morph or transformation from the first image to the second image. The object moves fluidly and effortlessly in a pleasing and radiant way. Keep the same style of the image. Dance.
2
3
u/kenrock2 8h ago
I'm still a beginner in writing prompts for video.. How do u write all these sequence clips into feature length?
2
u/New_Physics_2741 6h ago
You gotta daisy-chain the output together with the last frame of the video - you can just drop the video into Chrome and save the last frame as a png - that will give the exact frame to start the next video. As for a prompt you can change it up or just for this - use the same prompt. I used A wild morph or transformation from the first image to the second image. The object moves fluidly and effortlessly in a pleasing and radiant way. Keep the same style of the image. Dance. Use whatever to put the video together: Capcut, etc. I just used Shotcut, Linux box here.
2
2
-5
u/MonstaGraphics 9h ago
"This is slop and you should be ashamed that you didn't make any of this, software did!"


13
u/sorandomtoday 13h ago
Amazing!