Hey guys,

So TLDR is elevenlabs / play.ht is WAY too expensive for a realtime chat app, and we need an alternative. Guessing this is why character is rolling their own voice model, & obviously most apps can’t do that, so what are the alternatives here?

I’ve read zero shot prompting for TTS (inserting a sample at runtime) is part of the reason elevenlabs / play is so expensive, wheras finetuning on individual voices like character / OAI did and hosting those as their own model would be way faster and cheaper.

But couqi seems really slow from our finetune testing, even on an h100, and not only that but it’s not really… good. Does anyone know why, or there alternatives that chat apps are using? Is anyone working on better open source TTS? This seems totally overlooked compared to text where there’s so much competition right now, but is almost just as important. Shocked more people aren’t working on this! Thanks

  • UnoriginalScreenName@alien.top
    cake
    B
    link
    fedilink
    English
    arrow-up
    1
    ·
    1 year ago

    I just started down this rabbit hole and have a lot of questions. If you dont care about real time inference and just want a high quality voice clone, what’s the best option? I’m looking to do semi dynamic narration over video.