Does anyone know how token- and reasoning efficient it is? The charts don't show how many tokens were used in any benchmark.
> In this case, Qwen3.8-Max was asked to create the oh-my-cli project from scratch and, over a 10+ day long-horizon autonomous coding run, build a self-evolving harness.
They don't explain how successful that went but it's a bit hilarious seen that an Anthropic dev explained that it's been 15 days Claude was hard at work --with nothing to show yet-- trying to rewrite itself in another language.
"You rewrite Claude Code, we rewrite oh-my-pi."
"You're nowhere after 15 days, we do it in 10."
Sure, it's apples to oranges and all that. But part of me thinks they know fully well what they did there.
Now, if only we could afford a setup decent enough to run 2/3 instances at the same time...
[flagged]
[flagged]
[dead]
[dead]
ah so they distilled fable and sol, eh?
“self-evolves through feedback loops”
Does this mean they distilled Claude? Sounds like what Claude Code will often do.
Are these latest Qwen models still open weights or has Qwen moved away from that?
Tokenpocalypse canceled