logoalt Hacker News

Inflect-Micro-v2: complete voice in 9.36M parameters

195 pointsby nateb2022today at 12:36 AM24 commentsview on HN

Comments

modinfotoday at 5:48 AM

This is amazing, the quality blow my mind for such small model! I just replaced my old onnx model with yours!

here my implementation with speech dispatcher and server: https://github.com/skorotkiewicz/inflect-speechd

thanks for shearing!

yjftsjthsd-htoday at 3:21 AM

Couple highlights:

> Complete local text-to-waveform speech synthesis under 10M parameters.

In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.

> English only, with one fixed male voice. This is not zero-shot voice cloning.

(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)

show 1 reply
K0balttoday at 6:43 PM

How heavy in inference on this? The model would easily fit on many microcontroller modules, I wonder if they could run it?

da-xtoday at 5:40 PM

I think we need more neurons in the human brain than parameters in this model for speech. I wonder what it says about the human brain vs LLM efficiency.

NetOpWibbytoday at 7:43 AM

The inflections are weird but this doesn't sound like a robot. Not bad!

billduebertoday at 2:06 PM

I keep seeing tts stories here. Is it just an interesting subset of the llm world, or is there a huge use case I’m somehow missing?

show 1 reply
tmalytoday at 2:01 AM

This is impressive. I wish there were a voice clone option.

show 1 reply
sudbtoday at 4:39 PM

this is extremely encouraging for individuals/small companies being able to train pareto-frontier TTS models (specifically compute required to run vs quality of model output)

StilesCrisistoday at 12:48 PM

I'd love to hear it but it seems your quota is exhausted.

show 1 reply
itaketoday at 6:29 AM

Amazing quality for small size, but definitely not that enjoyable to listen to.

IMHO, its at about the same quality level of historic TTS tools.

show 1 reply
phoenixrangertoday at 4:22 PM

amazing! was looking for something similar

jsomedontoday at 2:55 AM

amazing quality for such small size!

mcbetztoday at 8:55 AM

Alternative title: Text to speech in 9.36M, English only.

fintunertoday at 4:05 PM

[flagged]

ameliustoday at 5:49 PM

[dead]

zenith605today at 4:48 PM

[flagged]

afdsaifdoitoday at 3:37 PM

[dead]