logoalt Hacker News

wolttamyesterday at 11:24 PM0 repliesview on HN

2 sparks currently run this model at 60 t/s single session, up to just over 100 t/s aggregate with concurrency of 4.

Going local has as opened up a world of use-cases I never would have entertained the idea of on metered/cloud usage. Privacy is a large part of it but, I also no longer think twice about whether to send a prompt or not based on the psychology of it costing money.

Cached input tokens on local inference are free, so I don’t care about running sessions up to 500k tokens and hundreds of turns (it’s rarely useful, but DSv4 remains surprisingly coherent up there)