logoalt Hacker News

xg15 • today at 12:00 AM • 3 replies • view on HN

Maybe a dumb idea, but how good are image models with manipulating spectrogram images? Then a workaround could be: convert input audio to spectrogram -> pass spectrogram + prompt to an image model -> convert modified spectrogram back to audio waveform.


Replies

Buttons840 • today at 12:09 AM

I wouldn't call it a dumb idea, but there's soooo much subtlety to sound that wont be visible in any reasonably sized image.

narrationbox • today at 1:26 AM

Could be interesting, though spectrogram voice gen is a step back in terms of advancement. Vocoder models that take in Mel Specs suffer from a lot of issues like hissing due to Griffin Lim issues. Would be curious to see if diffusion models can work on neural codecs directly.

I think Google had one called riffusion (the first version was designed for specs)