While it's great to see tok/s go up as high as possible, I think it's important to consider the actual usability of these models when you quantize down to something like 2-bit. From what we've tested it seems like going below 4-bit quickly leads to serious issues with thinking, tool calls, and overall model coherence.
Our plan to enable running bigger models on less GPU memory in a way that'll remain productive is expert streaming. This will let you offload experts for MoE models to RAM or disk, and load them when needed. This can have some performance tradeoff, but is lossless.
I totally get that, Its just, regardless of how bit-quantized it is, they are still pulling 40t/s from model sitting mostly in RAM instead of VRAM. If I understand correctly they split the network parts very deliberately between VRAM and RAM, and I wonder if your program, of which main feature is "get most of your hardware" is capable of similar feats, or if that performace is still locked for those willing to spend days experimenting manually.