logoalt Hacker News

pu_peyesterday at 11:33 AM7 repliesview on HN

Based on the benchmarks, it seems that Muse Glimmer barely edges out against Qwen3.6 27B, except for tool-calling skills (MCP, etc.). I wouldn't be surprised if they released it now because they are afraid they wouldn't beat Qwen3.8 27B.


Replies

dofmyesterday at 7:21 PM

I am glad they released it because I think we need a competitive culture of open weights that isn't just geopolitics.

But I have to say, I quite like the way Muse Glimmer thinks and talks. It's a cocky bastard in tone, but it's quite good, and its thinking traces are relatively terse.

show 1 reply
mycallyesterday at 11:49 AM

Do AI companies make release plans based on upcoming other models like this? I would think all the processes that go into the repository and weight infrastructure pre-training, checkpointing, knowledge distillation, model compression, post training pipeline, ecosystem integrations, inference API, benchmarking, human eval/safety/alignment, docs, etc... all that dictates the release schedule.

show 5 replies
walrus01yesterday at 11:50 PM

It will also be very interesting to see some direct head to head benchmarks between qwen 3.6 27B (let's say all at Q8 XK quantization, using the GGUF that unsloth publishes as a baseline) vs 3.8 27B. Particularly in tool use, terminal use.

The whole class of what can reasonably fit in a single GPU is an interesting category of LLM, and based on the results I've seen from 3.6 35B A3B and 27B versus what existed a year prior, it seems there's a lot of room for advancement.

onlyrealcuzzoyesterday at 7:52 PM

I would hope that Qwen 3.8 is better. It's been 4 months, and we've seen almost no progress in this space.

As people have called out, Glimmer appears to be a trade-off rather than a clear winner.

And from what I've been reading, no one is expecting Qwen 3.8's model in this space to be a clear winner, but just slightly and marginally better.

That's a little concerning as DeepSeek v4 Flash proved at it larger sizes there's a ton of room left to compress knowledge.

If we don't see something that's substantially better in the ~30B param space soon - it would appear we might've saturated that size with knowledge.

show 3 replies
stevenhubertronyesterday at 5:29 PM

For so many non-coding workflows, tool calling is more important.

nojstoday at 12:28 AM

Qwen3.6 27B has really punched above its weight for a long time. It’s shockingly good for its size. Very excited to see what 3.8 can do.

kolbeyesterday at 2:34 PM

Qwen3.6 27B is the go-to medium sized model for coding, so beating it is not a small achievement

show 1 reply