logoalt Hacker News

IsTomyesterday at 11:43 AM2 repliesview on HN

How is 30B smaller than 27B?


Replies

LeBityesterday at 11:48 AM

It uses fractal compression

lostmsuyesterday at 12:27 PM

They say it is trained with quantization awareness, so it should only be 15GB or so. Qwen was only trained in FP8 with QAT.

UPD, NVM, got misled by comments here. It is actually almost 60 GB so much larger

show 2 replies