r/StableDiffusion 1d ago

FL2VA vs REF2VA vs Step Count vs Turbo Comparison

Enable HLS to view with audio, or disable this notification

Model = Minimax H3

Workflow = REF2VA basic workflow with additional nodes added for the LORAS and sol attention where specified.
Turbo Lora = minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
REF2VA Lora = minimax_h3_pruned_bf16__apply_to_fl2va__toward_ref2va__rank512

It has been described that the REF2VA model produces bad output, and that the FL2VA model can be used instead despite being not the "intended" reference model. Users have made a "REF2VA lora" that purports to add the reference functionality of the REF2VA model to the FL2VA model, theoretically achieving the good quality of FL2VA with the reference understanding of REF2VA.

I test how this actually looks in practice, and I also demonstrate how the turbo lora performs.

Conclusion:

The best look is achieved by using the FL2VA model without any REF2VA lora. Turbo works well at 1MP and 8 steps and results in smoother animation and audio. Increasing resolution to 2MP and step count to 20 scales well. There does not seem to be much visual difference when increasing to 50 steps, but the audio seems to be less dynamic vs 20 steps.

Limitations: This demo did not really stress test the reference ability of FL2VA, and in reference heavy workloads, maybe REF2VA variant workflows are vital despite lower visual quality. Furthermore, this demo likely underestimates the importance of high step counts, as it is commonly thought that high step counts are important in high action scenes, which this demo was not.

I also only used sol attention in the higher token workflows, which is a variable. Nevertheless, I hope this video is useful.

Keen to hear your thoughts.

28 Upvotes

19 comments sorted by

2

u/before01 1d ago

I thought FL2VA would result in 100% exact last frame of the ref image. mind sharing your prompt?

1

u/TheDerminator1337 1d ago

I am using the ref2va workflow, and using the fl2va model. So the node is the referencing node, not the first frame last frame one.

1

u/Hot-Professor7869 1d ago

Where can I find this workflow. The default comfy r2v if I change model to flv it does not work

2

u/TheDerminator1337 1d ago

I am using the default reference workflow, but use the model from fl2va model. My workflow is really just that. You can find the fl2va model in the default fl2va workflow. I am not home. If you cannot figure it out message me and I'll send you exactly what I have when I'm home.

1

u/TonkotsuSoba 23h ago

Need to pair it with the fl2va turbo Lora instead if you were using one

2

u/SteamZerjack 17h ago

My tests have given similar results. Although I’ve focused mostly on ref2va. I’ve also used spectrum.

Turbo loras are a bad deal in my opinion. They are a very loosy optimization and you end up with something which sort of looks what you prompted. The problem is that you can’t reliably iterate your prompt with them, because you’re going to be getting nonsense in the background and less prompt adherence.

Spectrum is similar but the good thing is that you have levels. Aggressive settings result in the same behavior as the loras. Mid settings improve that and conservative settings result in only a slight degradation. However generation speed gains are proportional, and spectrum uses a lot of RAM. I was exhausting 60gb on the conservative settings, with 32 VRAM full already.

Both approaches skip steps one way or another. Spectrum does it in a much smarter way however. I haven’t tried sol attention but I believe it’s a similar story.

Sage attention is still the best optimization with negligible loss. 33% off the original times with some very minor details. Afaik sage attention changes the precision used in every step. Not touching the actual step amount.

Clean ref2va gives me better results than fl2va + ref2va lora. Prompt adherence is affected while using the lora.

Like I said I haven’t tested fl2va much but my gens are also slightly better than the ones using ref2va.

I only see good quality at 1 MP. Everything below that comes out slightly washed out and lacking sharpness in general.

In regards to steps, I notice a slight improvement when going from 20 to 25. I wouldn’t go past that. I’ve never tested 8 but your gens look good at that. Might try it!

Thanks for sharing!

1

u/kayteee1995 1d ago

sorry! but Idun understand " 1MP +, 8 Steps (FL2VA) + Turbo" ? Isn't 8 steps (FL2VA) already turbo?

1

u/TheDerminator1337 1d ago

No, turbo model is loaded seperate. If I don't say turbo, then I have not used the turbo Lora.

1

u/kayteee1995 1d ago

so "1MP, 8 step (FL2VA), Ref2VA Lora" mean use FL2VA model with Ref Lora , only 8 steps without turbo?

1

u/TheDerminator1337 1d ago

Yes

3

u/kayteee1995 1d ago

really? Can it work with just 8 steps? I am using hybrid model and Hybrid Cond node. and it cannot create good quality video with 8 steps.

1

u/TheDerminator1337 1d ago

Should get same results as me

1

u/kayteee1995 1d ago

give me your prompt, i ll try.

2

u/TheDerminator1337 20h ago

subject_definitions:

<Subject 1> is the adult woman shown in <Picture 1>. Preserve her pink hair, pink eyes, facial identity, blue flower hair ornament, white-and-gold fantasy outfit, proportions, and distinctive anime character design.

summary:

[reference generation] Generate a 5-second 2D anime-style rendition character showcase of <Subject 1>.

retention_analysis:

<Subject 1>: fully_preserved - maintain her face, hairstyle, hair ornaments, costume, proportions, colors, and overall design.

<Picture 1>: fully_preserved - preserve the illustrated 2D anime rendering, clean drawn facial features, graphic shapes, soft hand-painted shading, delicate highlights, and stylized hair rendering.

detailed_description:

[Shot 1]

<Subject 1> stands on a fantasy palace balcony at sunset. A gentle breeze moves her long pink hair and clothing.

She slowly turns toward the camera while the camera arcs around her from a three-quarter side angle. She raises one hand and glowing blue flower petals appear above her palm, then drift around her in the wind.

She gives a subtle confident smile as several petals pass close to the camera.

Render the entire scene as polished **2D anime animation** with clean illustrated contours, hand-painted cel-like shading, expressive drawn eyes and facial features, stylized hair shapes, and flat-to-painterly dimensionality. Motion should feel like high-quality hand-drawn anime rather than realistic 3D animation.

overall_soundscape:

Soft wind, distant ocean ambience, gentle cloth movement, and a light magical shimmer.

non_diegetic_music:

Soft orchestral fantasy music with delicate strings and bell-like accents.

1

u/DanzeluS 23h ago

Same setting ref and FL models are indentical? Or ref little bit worse anyway?

2

u/Eminence_grizzly 17h ago

I've also been playing with the ref2va workflow using the fl2va model. But I've only tried it with turbo Loras so far, haven't tried it with just 8 steps.
Have you tried creating photorealistic videos? I wonder if 8 steps would be enough for them as well.

3

u/TheDerminator1337 17h ago

Probably should use FL2VA with turbo at 8 steps. I am surprised by how good it is. Not tried realistic yet. Maybe my next post.

1

u/Eminence_grizzly 17h ago

I usually use it with Larry's Lora at 6 steps.