I understand that the interest in these extreme quants stems from wanting to maximize the capability the average user can get from a local LLM in this era of ludicrously expensive memory, but I wonder if this is maybe targeting the wrong axis?
We've been seeing various optimizations towards streaming, that have been much more impactful in the local AI space, e.g. MoE models where the less busy experts are offloaded to slow RAM or even pruned entirely, engram tables that can be read from NVMe instead of sitting around in RAM etc.
Maybe the trick with these extreme quants would be to increase the total parameter count while quanting individual weights, such that maybe the active parameter count comes down, or streaming weights from RAM or disk becomes more efficient, or cache behavior improves? Say, replacing a single 4bit/weight matmul with 3 1bit/weight operations that produce a much closer result than a single 1bit/weight matmul would.
I don't disagree, but that isn't really a quant anymore. That's just training a new model imo.