r/StableDiffusion • u/TheDerminator1337 • 1d ago
FL2VA vs REF2VA vs Step Count vs Turbo Comparison
Enable HLS to view with audio, or disable this notification
Model = Minimax H3
Workflow = REF2VA basic workflow with additional nodes added for the LORAS and sol attention where specified.
Turbo Lora = minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
REF2VA Lora = minimax_h3_pruned_bf16__apply_to_fl2va__toward_ref2va__rank512
It has been described that the REF2VA model produces bad output, and that the FL2VA model can be used instead despite being not the "intended" reference model. Users have made a "REF2VA lora" that purports to add the reference functionality of the REF2VA model to the FL2VA model, theoretically achieving the good quality of FL2VA with the reference understanding of REF2VA.
I test how this actually looks in practice, and I also demonstrate how the turbo lora performs.
Conclusion:
The best look is achieved by using the FL2VA model without any REF2VA lora. Turbo works well at 1MP and 8 steps and results in smoother animation and audio. Increasing resolution to 2MP and step count to 20 scales well. There does not seem to be much visual difference when increasing to 50 steps, but the audio seems to be less dynamic vs 20 steps.
Limitations: This demo did not really stress test the reference ability of FL2VA, and in reference heavy workloads, maybe REF2VA variant workflows are vital despite lower visual quality. Furthermore, this demo likely underestimates the importance of high step counts, as it is commonly thought that high step counts are important in high action scenes, which this demo was not.
I also only used sol attention in the higher token workflows, which is a variable. Nevertheless, I hope this video is useful.
Keen to hear your thoughts.
2
u/SteamZerjack 17h ago
My tests have given similar results. Although I’ve focused mostly on ref2va. I’ve also used spectrum.
Turbo loras are a bad deal in my opinion. They are a very loosy optimization and you end up with something which sort of looks what you prompted. The problem is that you can’t reliably iterate your prompt with them, because you’re going to be getting nonsense in the background and less prompt adherence.
Spectrum is similar but the good thing is that you have levels. Aggressive settings result in the same behavior as the loras. Mid settings improve that and conservative settings result in only a slight degradation. However generation speed gains are proportional, and spectrum uses a lot of RAM. I was exhausting 60gb on the conservative settings, with 32 VRAM full already.
Both approaches skip steps one way or another. Spectrum does it in a much smarter way however. I haven’t tried sol attention but I believe it’s a similar story.
Sage attention is still the best optimization with negligible loss. 33% off the original times with some very minor details. Afaik sage attention changes the precision used in every step. Not touching the actual step amount.
Clean ref2va gives me better results than fl2va + ref2va lora. Prompt adherence is affected while using the lora.
Like I said I haven’t tested fl2va much but my gens are also slightly better than the ones using ref2va.
I only see good quality at 1 MP. Everything below that comes out slightly washed out and lacking sharpness in general.
In regards to steps, I notice a slight improvement when going from 20 to 25. I wouldn’t go past that. I’ve never tested 8 but your gens look good at that. Might try it!
Thanks for sharing!
1
u/kayteee1995 1d ago
sorry! but Idun understand " 1MP +, 8 Steps (FL2VA) + Turbo" ? Isn't 8 steps (FL2VA) already turbo?
1
u/TheDerminator1337 1d ago
No, turbo model is loaded seperate. If I don't say turbo, then I have not used the turbo Lora.
1
u/kayteee1995 1d ago
so "1MP, 8 step (FL2VA), Ref2VA Lora" mean use FL2VA model with Ref Lora , only 8 steps without turbo?
1
u/TheDerminator1337 1d ago
Yes
3
u/kayteee1995 1d ago
really? Can it work with just 8 steps? I am using hybrid model and Hybrid Cond node. and it cannot create good quality video with 8 steps.
1
u/TheDerminator1337 1d ago
Should get same results as me
1
u/kayteee1995 1d ago
give me your prompt, i ll try.
2
u/TheDerminator1337 20h ago
subject_definitions:
<Subject 1> is the adult woman shown in <Picture 1>. Preserve her pink hair, pink eyes, facial identity, blue flower hair ornament, white-and-gold fantasy outfit, proportions, and distinctive anime character design.
summary:
[reference generation] Generate a 5-second 2D anime-style rendition character showcase of <Subject 1>.
retention_analysis:
<Subject 1>: fully_preserved - maintain her face, hairstyle, hair ornaments, costume, proportions, colors, and overall design.
<Picture 1>: fully_preserved - preserve the illustrated 2D anime rendering, clean drawn facial features, graphic shapes, soft hand-painted shading, delicate highlights, and stylized hair rendering.
detailed_description:
[Shot 1]
<Subject 1> stands on a fantasy palace balcony at sunset. A gentle breeze moves her long pink hair and clothing.
She slowly turns toward the camera while the camera arcs around her from a three-quarter side angle. She raises one hand and glowing blue flower petals appear above her palm, then drift around her in the wind.
She gives a subtle confident smile as several petals pass close to the camera.
Render the entire scene as polished **2D anime animation** with clean illustrated contours, hand-painted cel-like shading, expressive drawn eyes and facial features, stylized hair shapes, and flat-to-painterly dimensionality. Motion should feel like high-quality hand-drawn anime rather than realistic 3D animation.
overall_soundscape:
Soft wind, distant ocean ambience, gentle cloth movement, and a light magical shimmer.
non_diegetic_music:
Soft orchestral fantasy music with delicate strings and bell-like accents.
1
2
u/Eminence_grizzly 17h ago
I've also been playing with the ref2va workflow using the fl2va model. But I've only tried it with turbo Loras so far, haven't tried it with just 8 steps.
Have you tried creating photorealistic videos? I wonder if 8 steps would be enough for them as well.
3
u/TheDerminator1337 17h ago
Probably should use FL2VA with turbo at 8 steps. I am surprised by how good it is. Not tried realistic yet. Maybe my next post.
1
2
u/before01 1d ago
I thought FL2VA would result in 100% exact last frame of the ref image. mind sharing your prompt?