logoalt Hacker News

kamranjonlast Saturday at 10:07 PM1 replyview on HN

If you look at this chart here it seems the tiny model has a WER of ~12%… not sure about the micro model:

https://github.com/moonshine-ai/moonshine#when-should-you-ch...


Replies

yorwbalast Saturday at 10:35 PM

That's the error rate for STT, not TTS. TTS is generally easier than STT because you only need to produce one valid pronunciation and don't need to handle variation within and between individuals.