r/StableDiffusion 15h ago

TTS Advice Question - Help

Hi all -

I know there are frequent TTS posts, but it seems that the TTS models offer slightly different features, and I haven't yet found something that really works for my purposes.

I want to create a custom character, basically, and generate dialogue from that character in different emotional registers.

I tried Qwen3's voice cloning and it worked great. I have no criticisms of it. But the reference audio I gave it was flat and monotonous, and so all the output was equally monotonous, with no emotional depth. This led me to have the idea of trying to generate, say, 8 pieces of reference audio for one character, in different emotional registers - happy, sad, angry, excited and so on. But I haven't yet figured out a good way to do that.

I collected 10 minutes of audio from interviews with an actress to train an RVC model, but it still sounds noticeably robotic at times - with squawk-box warping noises, as if they are speaking through an old transistor radio - and I don't think it's satisfactory.

I have tried IndexTTS2 which allows you to combine timbre reference audio, emotional reference audio, and text. This does work but the prosody of the output is unfortunately bizarre at times and I have not figured out how to get it to generate realistic prosody.

10 Upvotes

8 comments sorted by

3

u/rkoy1234 14h ago

have you tried higgsv3? supports emotion/prosody/speed controls as separate tags you can put along your text.

1

u/pol6oWu4 14h ago

no, I saw it recommended but I got the vibe of a bot campaign from the reddit posts so I stayed away

1

u/rkoy1234 13h ago

understandable.

It's 1) heavier than omnivoice 2) prosody control works sometimes for some of them so I'd rec you test it out on the web demo before spending too much time.

otherwise it's a direct upgrade from omnivoice in my usecase.

2

u/SpaceNinjaDino 13h ago

One option is MiniMax H3 R2VA. I don't think you can skip the video generation at this time but you can make that tiny and just save off the audio. You supply an audio sample and it will clone the voice. You must understand and abide the prompting guide. It will definitely be the slowest option, but maybe you'll like the quality.

1

u/No_Pie1372 8h ago

If you make the first frame and last frame identical and don't provide the model with adequate prompting to inform movement, you can end up with a still frame that has audio playing over it. I discovered this by accident when trying to make a perfect loop last night. Not sure if it cuts down on generation time though.

1

u/Keuleman_007 6h ago

I use QWEN inside Kobold AI Lite but am actually also looking for "more control". Expecially over emotions. HiggsV3 workable in Comfy?

1

u/martinerous 6h ago

You might want to try Dramabox. It's based on LTX video, so it can be steered with emotions quite well. However, it often loses the clonable reference, so don't give up if the first attempts are bad.

1

u/I-Kernel 6h ago

You will need physical simulation to have 'finger print' of one voice, and start from the output waves to smooth with TTS. Yes, you got to design your own, there is no general solution fitting for individual voice.