logoalt Hacker News

yaloktoday at 12:16 AM0 repliesview on HN

sounds like a perfect fit for ASIC-optimized models (where matrix ops could be supported directly in BITCOS format, potentially) & achieving record power efficiency for on-device inference.

And it looks like per [0], a model needs only ~30% more weights to be at comparable quality, if quantization-aware training is done...

0. https://arxiv.org/pdf/2402.17764 - The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits