logoalt Hacker News

Alifatisktoday at 3:40 PM2 repliesview on HN

> I'd gladly take A5B or A8B or even A10B as a sort of middle ground.

Whats up with focusing on the active param count? Do yall fiddle with the weights or something?


Replies

kennywinkertoday at 4:28 PM

Total param count decides how much vram you need to run it. Active param count decides how fast it runs. My 10 year old GPU can load quantized 35B or 27B, but it can’t process 27B parameters per token faster than 2-4tok/s, while it can do A3B at >40tok/s

show 1 reply
martinaldtoday at 3:44 PM

You can run these on CPUs at a somewhat reasonable speed.

show 1 reply