This is fun. The audio isn’t particularly believable. Not that it matters for this project.
I’ve used a few tts providers and although a technical marvel they are all still easily identifiable as ai. Humans are very good at spotting fakes.
This project might work better with mostly canned audio clips of real humans.
I disagree. I think many of the audio elements are highly believable for lay people.