logoalt Hacker News

MaxikCZ • yesterday at 2:25 PM • 0 replies • view on HN

New models are trained with 8/4bit quantization in mind. Going from "native" 8 to 4 isnt as big of a step as going from 8 to 4 if native is full bf16.