logoalt Hacker News

ameliaquiningtoday at 5:19 AM0 repliesview on HN

STT models just turn audio into text, though; they don't interpret what the text means or figure out how to translate it into other representations like function calls. You would use a general-purpose model for that.