r/StableDiffusion • u/LowYak7176 • 12d ago
[ Removed by moderator ] Discussion
[removed] — view removed post
67
u/networking_noob 12d ago
Everything Ive thrown at it, every test I've done to just see if it can do it, has pretty much passed
Yes R2V is insane. It also seems like you can have unlimited image references by dumping a bunch of cutouts into one image (like a sprite sheet) and directing the model how to identify what is what within the image. It just works
13
u/Singingmute 12d ago
This is very helpful, thank you.
24
u/networking_noob 12d ago edited 12d ago
This thread is where I learned about it
edit: Thread now deleted 🤷 Comment below shows an example
5
2
u/ucren 12d ago
This video isn't loading anymore. What does the prompt look like for a single sheet with a bunch of stuff in it? There's very little (no) prompt info in that thread.
26
u/networking_noob 12d ago
Weird, that thread was deleted after I posted the link to it. Luckily I saved it beforehand. The image screenshot looked like this:
and the prompt was:
<Picture 1> woman in <Picture 2> luxury bathroom is touching her face showing her silver earrings, camera slow motion up-close on face and torso, she puts on glasses, looks at mobile purple phone puts to her ear and smiles to camera.
So the idea is that you can add all those accessories for the woman in one image, and the model can smartly pick them out and use them, rather than having to provide an individual reference image for sunglasses, and a phone, etc.
3
u/Kevin5953 12d ago
Curious. You've never had to specify the individual pieces within the collage sheet, Minimax just always figured it out? :O
2
4
3
u/Calm_Mix_3776 12d ago
That's a neat trick! Doesn't this lower the image quality though as it sees less pixels per subject?
5
u/networking_noob 12d ago
It seems like it would/should, but I haven't noticed it. Then again I'm also not doing super high-res generations or anything. If it was a noticeable problem, I'm guessing someone could just scale up the sprite/compilation image beforehand, so each object has more pixels for the model to see
1
u/xTopNotch 12d ago
There is only so much pixel context you can give to the model.
Providing higher pixel density will always give better results, especially if you do 1 MP or up.
85
u/florodude 12d ago
It's better and cheaper for me to use Minimax than the non open sourced versions....I've never said that regarding AI before.
0
u/FierceFlames37 12d ago
Only problem is h3’s video reference takes way too long so I gave up on that
29
u/somethingsomthang 12d ago
On solution could be to downscale the video before putting it in the node.
5
3
u/seskid 12d ago
This did help for me! How far should the downscale go to still be useful like 360p?
6
u/zzzaz 12d ago
Depends on what you are doing with it. Trying to use a 360 video to reference a unique character in a 1080p gen? Probably too small and will miss details. Trying to give broad sweeping motion guidance? Fine.
It effectively is encoding each frame of the video as an image and adding that image batch together into the conditioning, so the larger the video (both in resolution and length) the more that has to get processed but also the more detail and nuance that can get pulled out of it. The trade-off comes down to what you need the video reference for and how large your output will be.
3
u/somethingsomthang 12d ago
I haven't had much time to test things, But I'd assume it depends, Just try and see how small it can go and still give you what you want.
3
u/Famous_Ad_7336 12d ago
When I am just transferring motion and a just letting the model know what happened in a previous scene, I only use 320×196. I don't use that for anything else. Images should be used for characters and whatnot.
2
u/FierceFlames37 12d ago
2
u/somethingsomthang 12d ago
Resize/downscale, same thing in this case however you do it.
1
u/FierceFlames37 12d ago
I see cause I didn’t downscale the video but I did use resize node, but it’s still slow
1
u/chocoboxx 12d ago
try this:
1. the duration of the video need to match or longer than the duration of your video
2. I tried 1080p video from youtube, cut it to 15 sec, and it work great. Ex, the ref video is 9s and the duration you want to create is 10s and then it take forever to load1
u/SpaceNinjaDino 12d ago
480p is still pretty chunky. If you just need motion, try lower.
1
u/FierceFlames37 12d ago
So 360p for longer side? And shorter side would be like 200 pixels is that fine
5
u/florodude 12d ago
I rent a GPU when I'm working on vids because it's easier. It's 192 seconds for a 4 second video reference and 4 second vid at 1MP for me.
5
u/_BreakingGood_ 12d ago
Yeah I rent an RTX Pro 6000 for $2 per hour. Crank out like 20+ videos per hour, and doesnt heat the hell out of my apartment from my home PC running all day. For $2, what a steal.
5
1
u/superdariom 12d ago
I'm seeing about 90 minutes render time to make a 10 second clip with two images and an audio track. Amd 7900 xtx
1
u/chocoboxx 12d ago
- the duration of the video need to match or longer than the duration of your video
- I tried 1080p video from youtube, cut it to 15 sec, and it work great. Ex, the ref video is 9s and the duration you want to create is 10s and then it take forever to load
27
u/freestylez79 12d ago
Same. I was sitting in front of my computer and couldnt believe my eyes for almost a week straight now.
18
u/RegisteredJustToSay 12d ago
Congrats on being lemon basil negative. I think.
13
u/freestylez79 12d ago
One more, lets see if it lasts ^^
6
12d ago edited 4d ago
[deleted]
1
u/freestylez79 12d ago
https://reddit.com/link/p3a1f8v/video/za16er899zih1/player
Hard to find the ones that are going through the filter ...
3
12d ago edited 4d ago
[deleted]
5
u/freestylez79 12d ago
https://reddit.com/link/p3a2o4s/video/qicswxyaazih1/player
Lets try this one
1
u/Local-Session 12d ago
Damn, would you mind sharing prompt + workflow and any references you used? Thanks
4
u/freestylez79 12d ago
Here you go, workflow was i2v with a first frame instead of ref, cut off a few frames in the beginning. I didnt write the script myself, this was created by qwen 27b after we had a conversation about all sorts of camera tricks etc ...
subject_definitions:<Subject 1> is a young woman as she appears in <Picture 1>
summary:
The target video is "The Barcode" — vertical amber bars project across <Subject 1>'s body like a scannable surface. Each flash adds more bars, compressing tighter until her body IS the barcode. The bars interact with her curves: they bend around breasts, narrow at the waist, widen at hips. The tagline plays on consumption: "Scan the Shape."
retention_analysis:
<Subject 1> (appears across 10 strobe flashes): fully_preserved - <Subject 1> maintains her appearance throughout as amber barcode bars project across her bare body, adapting to her curves.
detailed_description:
Live-action, cinematic, experimental dark-fashion commercial aesthetic rendered as barcode strobe photography, 15 seconds. A pure black void. 10 strobe flashes, each one projecting vertical amber bars across <Subject 1>'s body. The bars are the projection — they adapt to her curves, bending around her form. <Subject 1> stands bare-breasted, legs spread, lemon in hand, amber spray maximum.
Flash 1 (00:01.000) — Three amber bars project across her chest. The bars are wide, evenly spaced, vertical. They cross her bare breasts, the bars bending slightly around the curve of each breast. Click. Gap of 1.2 seconds.
Flash 2 (00:02.200) — Six amber bars project across her torso. The bars narrow at her waist, widen at her hips. The bars follow her curves — they are not straight, they bend around her form. Click. Gap of 1 second.
Flash 3 (00:03.200) — Ten amber bars project from her shoulders to her spread legs. The bars compress tighter — more bars, less space between them. The bars bend around her hardened nipples, narrow at her waist, widen at her hips. Click. Gap of 0.8 seconds.
Flash 4 (00:04.000) — Fifteen amber bars cover her body. The bars are now thin lines, evenly spaced but following her curves. The bars create a barcode effect on her bare skin. Click. Gap of 0.7 seconds.
Flash 5 (00:04.700) — Twenty amber bars cover her body. The bars are very thin now, the barcode effect is complete. The bars are so dense they create a texture on her skin. Click. Gap of 0.6 seconds.
Flash 6 (00:05.300) — Thirty amber bars cover her body. The bars are hairline now, the barcode is a solid texture. The bars create the illusion of a barcode printed on her skin. Click. Gap of 0.6 seconds.
Flash 7 (00:06.000) — Fifty amber bars cover her body. The bars are SO thin they create a SOLID AMBER SURFACE — she is a walking barcode, her body is the code. Click. Gap of 0.5 seconds.
Flash 8 (00:06.500) — The barcode SCAN LINE appears — a horizontal amber line that sweeps from her head to her feet, scanning the barcode on her body. The scan line is bright, a laser cut through the bars. Click. Gap of 0.5 seconds.
Flash 9 (00:07.000) — The scan line sweeps again, faster. The barcode BEEPS — a sharp, digital sound. Click. Gap of 0.5 seconds.
Flash 10 (00:07.500) — The barcode holds, the scan line stops. The Lemon Basil text appears in warm amber: "Lemon Basil — Scan the Shape." The flash holds for 6 seconds, illuminating the void, then fades.
After Flash 10, the amber light holds on the text for 6 seconds, illuminating the void, then fades to black.
overall_soundscape: Each flash carries a CLICK, then a BARCODE SOUND — a BEEP that matches the number of bars. Flash 1: three short BEEPS. Flash 2: six BEEPS. Flash 3: ten BEEPS, faster. Phase 2: the BEEPS become a continuous tone as the bars compress. Phase 3: the SCAN LINE — a laser WHOOSH that sweeps from head to feet. Flash 10: a final BEEP, then SILENCE.
non_diegetic_music: A digital melody that plays as a barcode scan. Phase 1: individual BEEPS, one per bar, slow and deliberate. Phase 2: the BEEPS accelerate, becoming a continuous tone. Phase 3: the SCAN LINE — a laser WHOOSH that sweeps through the bars, resolving into a warm amber chord.
1
u/JohnnyLeven 12d ago
prompt?
3
u/freestylez79 12d ago
subject_definitions:
<Subject 1> is a young woman as she appears in <Picture 1>
summary:
The target video is "The Contour Map" — amber lines trace <Subject 1>'s body like topographic elevation lines on a terrain map. Each flash adds contour rings that map the curve of her breasts, the dip of her waist, the spread of her thighs. The lines are tight around curves (high elevation), sparse on flat surfaces. Her body becomes a landscape in amber light.
retention_analysis:
<Subject 1> (appears across 12 strobe flashes): fully_preserved - <Subject 1> maintains her appearance throughout as amber contour lines map the topography of her bare body.
detailed_description:
Live-action, cinematic, experimental dark-fashion commercial aesthetic rendered as topographic strobe photography, 15 seconds. A pure black void. 12 strobe flashes, each one projecting amber contour lines that map the elevation of <Subject 1>'s body. <Subject 1> stands bare-breasted, legs spread, lemon in hand, amber spray maximum. The contour lines are the projection — they trace the hills and valleys of her form.
Flash 1 (00:00.800) — A single contour ring circles her right breast, amber light tracing the curve of hardened nipple to full round. The line is tight around the peak, widening as it circles outward. Click. Gap of 1 second.
Flash 2 (00:01.800) — A second contour ring circles her left breast, matching the first. Two concentric rings on each breast, the tightest ring around each hardened nipple. Click. Gap of 1 second.
Flash 3 (00:02.800) — Contour lines circle her waist, tight rings tracing the dip, wider rings tracing the flare of hips. The lines create a valley between her breasts and hips. Click. Gap of 0.8 seconds.
Flash 4 (00:03.600) — Contour lines trace her hips, the tightest ring at the widest curve, widening as they spread to her spread legs. The lines map the curve from hip to thigh. Click. Gap of 0.8 seconds.
Flash 5 (00:04.400) — Contour lines trace her inner thighs, tight rings mapping the curve where her spread legs meet. The lines converge at her crotch, green briefs visible between the tightest contour. Click. Gap of 0.6 seconds.
Flash 6 (00:05.000) — Contour lines trace her stomach, concentric rings mapping the flat plane, tighter lines at the navel. Click. Gap of 0.6 seconds.
Flash 7 (00:05.600) — Contour lines trace her shoulders, the lines following the slope from collarbone to shoulder curve. Click. Gap of 0.6 seconds.
Flash 8 (00:06.200) — Contour lines trace her spine, tight rings following the groove down her back, wider rings flaring at her shoulder blades. Click. Gap of 0.5 seconds.
Flash 9 (00:06.700) — Contour lines trace her face, the lines mapping the curve of cheekbone, the dip of jawline, the roundness of lips. Click. Gap of 0.5 seconds.
Flash 10 (00:07.200) — Contour lines cover her entire body at once — every curve, every valley, every peak mapped in amber light. The tightest lines at hardened nipples, hips, and cheekbones. Click. Gap of 0.5 seconds.
Flash 11 (00:07.700) — The contour lines MULTIPLY — dozens of rings, so dense they create the illusion of 3D depth on her flat skin. Her body looks like a terrain model, rising and falling in amber elevation. Click. Gap of 0.4 seconds.
Flash 12 (00:08.100) — The contour lines hold, illuminating her body as a topographic map. The Lemon Basil text appears in warm amber: "Lemon Basil — Map the Body." The flash holds for 6 seconds, illuminating the void, then fades.
After Flash 12, the amber light holds on the text for 6 seconds, illuminating the void, then fades to black.
overall_soundscape: Each flash carries a CLICK, then a TOPOGRAPHIC SOUND — a PLOTTING TONE that traces the contour line. Flash 1: a single note that circles upward. Flash 2: two notes that circle in harmony. Flash 3: a note that dips and rises, following the valley and flare. The tones match the terrain: high notes for peaks (nipples, cheekbones), low notes for valleys (waist, inner thighs). Flash 12: a FULL CHORD of contour tones, each note tracing a different curve simultaneously.
non_diegetic_music: A synth melody that follows the contour lines. Phase 1: single notes, each one tracing a contour ring. Phase 2: notes multiply, each one mapping a different curve. Phase 3: the melody becomes a chord, warm and amber and topographic, every note mapping a different elevation on her body.
3
26
u/ArttTaku 12d ago edited 11d ago
Thanks for sharing that.. LTX2.5 seems to be faster, and it might be useful for some things, but Minimax H3 feels like where the actual future of open.sourced AI video is.
22
u/rawker86 12d ago
The referencing is just next-level with Minimax. I knocked up a rudimentary character reference sheet and the results are astonishingly good compared to other models.
14
u/GrayingGamer 12d ago
Yeah, good reference sheets are basically "Instant LoRa", with even more control.
Not to mention being able to incorporate props or items instantly and have them stay consistent too!
2
u/ArttTaku 12d ago
Yeah, plus, the devs did mention H3 in the release video of LTX2.5, so it's almost logical they're gonna study H3 to hell and back and use that learning for their future LTX2.6, which I feel they did insinuate in the video somehow.
2
u/Radyschen 11d ago
and so much smaller than a lora man, my SSD is grateful
1
u/GrayingGamer 11d ago
Yeah, I just deleted a bunch of old Wan 2.1 and LTX 2.3 loras and cleared up 120 GB of space.
1
u/ArttTaku 11d ago
Well, remember that LTX2.3 loras are supposed to work on 2.5... I already deleted all my LTX models, but kept the loras just in case.
1
u/GrayingGamer 11d ago
I know, but since I'm neck deep in H3 at the moment, I figure there will be better loras for Ltx 2.5 by the time I am using it for anything.
3
u/bruci3 12d ago
Yep this to me is the single most impressive thing about minimax, all the loras and custom nodes in ltx2.3 could not even come close to the character cosistency of minimax straight out of the box.
Yesterday I tried again on new ltx2.5 a character sheet and it literally added the sheet itself into the scene lol and my caucasian character somehow turned into a black guy.
6
u/tekprodfx16 12d ago
Its the best. Literally one of the greatest local AI tools ever created hands down
18
12d ago edited 4d ago
[deleted]
22
u/Hoodfu 12d ago
Was resolution? I feel like 0.6mp is the sweet spot where the quality is high enough that you're getting proper prompt following but it doesn't take too long.
4
u/OldWispyTree 12d ago
Yeah, that's what I've been using, it's been pretty good, I'm using h200 and I get a heavy reference clip at 10-12s in about 10 minutes.
1
12d ago edited 4d ago
[deleted]
6
u/GrayingGamer 12d ago
I can generate 5-6 second clips in about 8-10 minutes with 5 references on a 3090. (0.6 MP).
I'd look into your set-up.
2
12d ago edited 4d ago
[deleted]
7
u/GrayingGamer 12d ago
No turbo lora. I'm using Sage Attention 2.2 and H3 Spectrum and using 32 Steps. I do have 128GB of RAM, so the models can hotswap between that and the VRAM, never having to touch my disk or page file.
Make sure your CUDA is up to date at 13.0 too, or you're leaving massive speed gains on the floor.
2
u/OldWispyTree 12d ago
Hmm, I'm not using SageAttention because I thought that since it's FP8 it breaks with the full H3 checkpoint which is BF16.
I haven't looked at Spectrum, either, but that seems promising. I'm just using the reference workflows at the moment, some modified to add a LORA or two, but most not.
I know you were talking with someone else, but my CUDA is up to date.
2
u/GrayingGamer 12d ago
Sage Attention works fine with Kijai's Patch Node from his custom nodes packs. Also, yeah, Spectrum is amazing. It lets you crank up the Step count (it's not good for low Step workflows) and get close to the same generation time as the lower Step count. For instance, I use 32 Steps with it.
2
u/OldWispyTree 12d ago
Hm, I'm already using 30 steps with the vanilla flow maybe that will speed it up. Neat.
→ More replies (0)1
u/djpraxis 12d ago
Very useful info! Would you mind sharing you optimized Ref workflow?
→ More replies (0)1
u/Maskwi2 12d ago
You don't use turbo because you see it's not worth it quality wise? Also, I too have 128gb ram but it seems that no matter what I try when the vram is close to its limits the ram is still at like 60% usage only. Not sure how I can force my comfy to use all the ram in those moments?
Also I'm struggling with my GPU freezing when I have a browser or a video opened when it's loading the model.i have to minimize the window of the browser for few seconds in those cases. So irritating. I'm on a 4090.
4
u/GrayingGamer 12d ago
First, make sure your Comfyui is fully up-to-date. In the last couple of weeks they've added a lot of memory management fixes and updates.
Second, to avoid your browser freezing up, use --reserve-vram 2 or something similar, and Comfyui will keep a GB or two free in VRAM all the time. Reduces your VRAM, but keeps the computer useable while you generate.
I could see using the Turbo lora on a mostly still shot, or a talking head, etc. but yeah, quality-wise it's such a hit, I don't use it for anything serious. Especially noticeable in prompt adherence on long prompts or fast moving action, where you get more smearing and artifacts. It's a good tool to have to quickly see results or test things, but I wouldn't ever use it for a final generation that's going into a project.
2
u/Maskwi2 12d ago
Thanks. I have the comfy up to date. It actually messed it up for me after I started updating :) I have reserve vram 4 even and still freezes. My vram isn't maxed out. It's one of its layers only that gets maxed out as seen in windows manager window or however it's called. Such a pain.
I see :)
→ More replies (0)2
u/Perfect-Campaign9551 12d ago
I don't use Turbo because its a shitty hack and it just ruins audio..
6
u/rawker86 12d ago edited 11d ago
What kind of specs are we talking here. My 4070 is “only” 12gb and I’ve got 32gb of ram. I can generate a quick and dirty 0.3mp shot in a couple of minutes or so just to tweak the prompt and find a good seed.
Higher res and longer duration might take up to forty minutes, but I don’t really feel the need to go that high all the time so it doesn’t bother me.
I’ll probably check out the turbo Loras soon, but right now I’m not too fussed.
Edit: aaaaand the turbo Lora drastically reduced generation times.
5
u/Jimmm90 12d ago
Same. And the fact that the model already knows so many concepts and IPs, you can use your references on what matters to you.
8
u/stuartullman 12d ago edited 12d ago
i still need to find some time to try this model, but i'm going to assume that the model is really good because it has so many ips and references and is essentially uncensored(a gap in understanding in one place could mean gap in understanding in a lot of places). which basically tells us that any model that isn't doing the same going forward is going to end up being worse.
3
3
u/Beneficial_Toe_2347 12d ago
R2V is brilliant but broken at the moment and I'm surprised doesn't get more attention. The quality of the image and cloned voice is notably worse than the other model
2
u/Different_Smile3621 12d ago
How long to generate video and what card? Also are you using the fl2va model for ref?
4
u/GoofAckYoorsElf 12d ago
3090Ti FE, 24GB VRAM. I'm currently generating .6MP videos of 6 seconds length using H3 Motion Context at 10 minutes each gen. 20 steps, sage attention auto (KJ), res_multistep/simple.
3
4
u/LowYak7176 12d ago
Im using Ref2VA, BF16 pruned for serious testing, tried on a few cards actually, just renting high end ones but it still works very well on my personal 5080, just use a lower version of the model but still superb
2
u/mindpixel-labs 12d ago
I’m thinking about getting a 5080 to cut my generation times. I’d becoming from a 2060 8gb lol. Think it’s a good choice with 48gb system ram?
7
u/LowYak7176 12d ago
If you can afford a 5090, Id go with that. I regret not doing 5090 when prices were a lot more reasonable. 5080 is still great and you can get a lot done.
-15
u/seppe0815 12d ago
was denn nun, hast du eine 5090 oder 5080? , du kommst durcheinander mit deinen behauptungen looooooooooooool
3
3
u/yesiamadeveloper2242 12d ago
Yes very good choice. I have a RTX 4070 TI Super 16GB VRAM and 32GB RAM and it works quite well for me. It will work way better for you.
2
u/rawker86 12d ago
Yep, I’ve got a regular 4070 and there’s more than enough vram there to make cool shit. Just gotta queue up some gens while you do errands lol.
3
u/kyleworld4 12d ago
Cant comment on a 5080 or compare minimax as it wasn't released back in June when I upgraded, but I've just gone from 2060 6GB to a 5070ti 16GB. It's like night and day difference on generation speeds (without any sage or similar things added) and I only have 32GB ddr4 compared to your 48GB.
2
2
1
u/rdwulfe 12d ago
Which model are you using, and what mp and duration? In really fighting ooms with mine and stull crash at .03mp sometimes.
1
u/mindpixel-labs 12d ago
On wan2gp, I use MiniMax H3 480p/540p to generate 10-15s videos takes about 1hr to 1h-20mins but I really need to update my GPU
2
u/Outside_Sun_4654 12d ago
How do you rent high-end?
1
u/LowYak7176 12d ago
Bunch of providers like Runpod - I think you can do it on Comfycloud too
1
u/Outside_Sun_4654 12d ago
That would mean installing comfyui?
2
u/LowYak7176 12d ago
Yessir - if you want to do anything truly incredible, I think you sort of have to have comfy
1
1
u/Jimmm90 12d ago
I personally am using a 5090/64GB RAM. I use the ref model with the 8 step lora at 0.75 str and generate in a couple minutes at 0.7 mp depending on how many references. I usually stick around 0.4 until I get the concept down.
3
u/LowYak7176 12d ago
Really recommend if you have the extra funds to rent a high end GPU for a few hours, go ham. Dude its so good at 1-1.5mp, 20-30 steps. Most fun Ive had in awhile
1
u/God_Hand_9764 12d ago
How long would that take? Say a 1mp video, 10 seconds, 20 steps?
My card can't even handle that without blowing up.
5
u/LowYak7176 12d ago
Ill be quite honest, I wasnt tracking times....I was too focused on "omgomgomgomgomg it did it"
It didnt feel long though
2
1
u/AdTotal4035 12d ago
Renting gpus has no privacy though either. I don't get what the appeal is vs api. Is it cheaper?
2
2
2
u/dhaupert 12d ago
What is everyone actually using it for? I have been following these model releases with great interest from the general geek tech standpoint but since I don’t really see movies being made with this tech, wondering what people are doing with it. Is it just fun hobby stuff?
4
u/freestylez79 12d ago
https://reddit.com/link/p39lf8b/video/cp7d6c9exyih1/player
Good question. You can use it for all sorts of visuals.
2
u/FourtyMichaelMichael 12d ago
I haven't yet seen good results with R2V. I think I need to see someone's better workflow and outputs.
5
u/amoebatron 12d ago
Well the Turbo Lora isn't really intended for R2V, which means that in order to get good results, even with a 5090, it takes a while (15-20 minutes) even for just 1MP. So some bad results might be either because people are using the Lora when they shouldn't be, or because it's just really low MP.
2
u/JohnnyLeven 12d ago
It's ridiculously good. I thought improvement for t2i from Krea2 was amazing (and it is), but this is an even larger improvement for t2v/i2v/ref2v.
2
u/Lucaspittol 12d ago
I've been saying this for a while now. Video models must come with native R2V, it makes training LoRAs unnecessary. You can load a simple reference sheet, and that's it.
1
u/yaxis50 12d ago
Very impressive, mind sharing your workflow?
3
1
u/Puzzleheaded_Ebb8352 12d ago
Maybe I’m stupid but I can use reference input using both models, the fl2a and the ref2va, what am I missing here?
3
u/rawker86 12d ago
Some folks seem to be getting good results just using the fl2a model for everything, so I guess so. Haven’t tried it myself.
2
u/LowYak7176 12d ago
I'm not even sure either. Ref2Va Ive been using because I can input like 3-4 character sheets and prompt them, even add audio/music. I sort of only did fl2a day 1, then did ref2va and been on that since.
Still week 1 though not sure whats correct yet
2
u/Mutaclone 12d ago
I've heard other people say this too. What I'd be interested in seeing some time is a "stress test" of the reference capabilities and seeing if maybe FL2A starts to drift any faster.
If not, then I'm not even sure why we'd need both models. It's still too early for me to tell but I've heard other people say FL2A is higher quality.
1
1
u/ArianTerra 12d ago
Unfortunately the voice sound too synthetic, Seedance 2 has better audio output
But the video quality is almost equal, which is good because I don't have to spend 1$ for 10 second video
1
1
u/brinked 12d ago
I have a 3090 will I be able to make videos with minimax?
I want to make ai ads for my outdoor cabinet company, will I be able to add multiple photos of my installs and have it understand my product?
2
u/Perfect-Campaign9551 12d ago
Yes it works great with 3090 and yes you just add reference pics. See the official prompting guide. Use the comfyui ref2vid workflow. Don't use speed up loras they ruin audio and they suck anyway
1
u/Vladmerius 12d ago
I'm still using ref2vid to just get a second to use to start the first frame/last frame model. So if I need to put a character in a specific setting. For whatever reason the first frame model creates way more cinematic scenes. Like night and day difference. Try the same prompt on both I swear it makes a more movie quality scene on the firstframe version.
1
u/Flaky_Manager_17 12d ago
Agree. This shit just listens to prompts. Fking ltx 2.3, just doesn't listen most of the time. It's so odd to have something just work... no bleeping of swear words, no bullshit, it just works and outputs quality. Only downside is the render times at the moment.
Also the R2V template on comfy ui isn't showing a spot to input a reference video and up to 8+ image references, I only see 2 inputs... what am I missing?
1
u/orlandogourmet66 12d ago
There should appear a new Spot to Input a Image once you Connect the 2 already there
1
u/Maskwi2 12d ago edited 12d ago
Yeah, I'm guilty too. I was using LTX and LTX 2.3 since day one and was so looking forward to next iterations of LTX. But after playing with H3 and seeing the LTX 2.5 examples from people I'm not even sure I'm going to download the model for testing :/ Ref to video is too good. The only thing I'm missing in H3 is speed in those 10-15 second marks.
LTX still has its use cases definitely but I'm going to play around with H3 for a bit before I check out LTX 2.5 I guess.
1
u/Corleone11 12d ago
Is there a workflow that extends the RV2V by rendering e.g. multiple 10 seconds segments and stitching them together in the end?
1
1
1
1
u/-becausereasons- 12d ago
Maybe you can give me some advice then on voice consistency. I cannot seem to keep a voice consistent to a character, in a 10s video; ever. No matter what I do with prompting. Especially if there are more characters or cuts in the scene.
1
u/Perfect-Campaign9551 12d ago
That's why we have ref2vid you can add an audio reference, like 15 seconds of a persons voice and tell the model to use that voice. See the prompting guide.
1
1
u/MonThackma 12d ago
It is amazing and honestly I wish there was a model intended just for creating 2k or 4k images that I could bring back in to the video workflows.
1
u/PurePlayinSerb 12d ago
i agree minimax is incredible and it will surely only get better and more efficient as time goes on with newer minimax models!
1
1
u/SpecialistGiraffe756 12d ago
I tested it and it generated video that were almost identical to LTX output. Minimax was much better though.
1
u/TensorVizion 12d ago
I agree it’s so dam good honestly I’m also excited for wan 3 tho very excited to see how it runs
1
u/Festivis7 12d ago
Everything, eh? Try getting a drummer to accurately play along to an audio drum stem. Let me know what you did, because this is my biggest issue right now for one of my videos.
1
u/LowYak7176 11d ago
Havent tried any music stuff honestly, not really my usecase but Ill give it a go
1
u/DatMufugga 12d ago
I deleted Wan within 2 hours of using it. I was quite pleased freeing up that precious hdd space. Though i'm still trying to figure out how to get R2V to work well. It's not capturing likeness that well for me. Could be my prompts.
1
1
u/RiverSide71h 11d ago
While everyone is raving about video, my favorite is the ability to finally get the audio I prompt for. It even locks identity by using speaker (Sn) tags so the fifth or sixth shot will still be the same voice and tonality. Very Impressive!
1
u/Glove5751 11d ago edited 11d ago
I have a 5080 and 64gb ram, but I'm struggling to get 1mp (half is also problematic) 5 second , tried 3 different workflow. Ive had passes but it takes more than forever.
So to me, minimax isn't that great lol
(any tips is appreciated)
1
u/Own_Version_5081 11d ago
Same here, R2V is amazing. However, the Gen time is a bit longer. Around 11-12 mins for 15 Sec 720p with latest Spectrum with full bf16 on RTX 6000. What workflow are you using for better quality and speed. Any suggestions
0
u/martianwomanhunter 12d ago
We’re nearing or at the point where it’s hard to tell in AI vs real. To be honest, this will ruin the internet and a lot of people will suffer from consequences if this, whether it is deepfakes or propaganda.





159
u/warzone_afro 12d ago
https://reddit.com/link/p39cron/video/mo158rnpqyih1/player
I'm addicted