logoalt Hacker News

narrationbox • yesterday at 11:46 PM • 1 reply • view on HN

Plenty of models take in text + audio and spits out audio. It's the format of most newer generation accent conversion/voice cloning models.

What's your exact use case?


Replies

chr15m • today at 12:37 AM

Sound effects are completely different to voice, which those models are trained to output.