Plenty of models take in text + audio and spits out audio. It's the format of most newer generation accent conversion/voice cloning models.
What's your exact use case?
Sound effects are completely different to voice, which those models are trained to output.
Sound effects are completely different to voice, which those models are trained to output.