logoalt Hacker News

MaxikCZ • yesterday at 7:14 AM • 0 replies • view on HN

I totally get that, Its just, regardless of how bit-quantized it is, they are still pulling 40t/s from model sitting mostly in RAM instead of VRAM. If I understand correctly they split the network parts very deliberately between VRAM and RAM, and I wonder if your program, of which main feature is "get most of your hardware" is capable of similar feats, or if that performace is still locked for those willing to spend days experimenting manually.