logoalt Hacker News

Jgoauhyesterday at 8:08 AM2 repliesview on HN

Seems impressive, i believe better architectures are really the path forward, i don't think you need more than 100B params taking this model and what GPT OSS 120B can acchieve


Replies

CuriouslyCyesterday at 12:36 PM

We definitely need more parameters, low param models are hallucination machines, though low actives is probably fine assuming the routing is good.

NitpickLawyeryesterday at 8:24 AM

New arch seems cool, and it's amazing that we have these published in the open.

That being said, qwen models are extremely overfit. They can do some things well, but they are very limited in generalisation, compared to closed models. I don't know if it's simply scale, or training recipes, or regimes. But if you test it ood the models utterly fail to deliver, where the closed models still provide value.

show 1 reply