logoalt Hacker News

cpburns2009today at 2:49 AM0 repliesview on HN

Not full precision. I've only benchmarked 27B across Q3-6 quants using lm-eval. I lack the hardware to bench 27B at BF16 but I might be able to do Q8_0. I haven't gotten around to doing 35B. I really should upload my collection of results to Github or somewhere.

Here's a summary of what I have for 27B. I used unsloth's UD-Q{3-6}_K_XL quants across 11 evals. The values are pretty linear between Q3 and Q6.

    Qwen3.6-27B     Q3    Q6
    ARC-Challenge   97.0  97.0
    BIG-Bench Hard  57.9  59.3
    GPQA Diamond    77.8  83.3
    GSM8K           92.4  92.6
    Hendrycks Math  35.5  38.9
    HumanEval       80.5  85.4
    HumanEval+      75.0  79.3
    IFEval          87.3  88.0
    MBPP            75.2  77.2
    MBPP+           88.4  88.9
    MMLU-Pro        83.1  83.5