I understand that the interest in these extreme quants stems from wanting to maximize the capability the average user can get from a local LLM in this era of ludicrously expensive memory, but I wonder if this is maybe targeting the wrong axis?
We've been seeing various optimizations towards streaming, that have been much more impactful in the local AI space, e.g. MoE models where the less busy experts are offloaded to slow RAM or even pruned entirely, engram tables that can be read from NVMe instead of sitting around in RAM etc.
Maybe the trick with these extreme quants would be to increase the total parameter count while quanting individual weights, such that maybe the active parameter count comes down, or streaming weights from RAM or disk becomes more efficient, or cache behavior improves? Say, replacing a single 4bit/weight matmul with 3 1bit/weight operations that produce a much closer result than a single 1bit/weight matmul would.
I understand the obsession with low bit quantization, but it is empirically quite evident that it is not possible to compress models to less than 4 bit per weight without severe loss of capabilities².
It may be nice as an experiment, but it is obviously a very inefficient route for model training: spending all the flops on a saturated model only to prune its capabilties.
²As to why, I have seen few explanations. But the empirical evidence is there.
Can this produce a useful model? So far 1 bit quants have been less useful than smaller models that use the same memory
Their paper shows this comes with huge quality loss, but that doesn't make it a negative result by any means
Thought this was going to be on the original Little Bit paper, always nice to find out about a surprise sequel!
Has anybody tested this? Are there any available computed models to test?
Has anyone tried this on apple silicon M1-5? Any benchmarks/comps?
Perf goes from 80% to 47% on Wikitext-2. Also no comparisons to FP4 solutions that are able to maintain or exceed perf on the same dataset 80% perf with a 4.25-4.5 big budget.
I think more meaningful thing here would be a hybrid solution that went down to sub-bit representations when the informational representation does not need it (for example later layers) that still maintains task performance