logoalt Hacker News

Qwen 3.8 Omni Flash

201 pointsby jjcmyesterday at 11:05 PM73 commentsview on HN

Comments

mavamaartentoday at 8:08 AM

I'm wondering, is there a tool or something out there that helps me pick a model, in the vast sea of models out there these days? Every time I need a model for something I see the list on openrouter and I'm completely overwhelmed.

I'd love to be able to explain my use case, my cost preferences and have a tool select a few good models to try.

E.g. I wrote a tool that cleans out my email spam box. It classifies emails that are already flagged as spam, and if it's very obviously spam it removes it permanently (keeps a copy on disk though). And after x emails, it goes through the list of deleted spam mails and suggests email rules. What model would be best suited? I'd love to be able to explain this use case and get this info served to me. The list of models and the information about what they're good at is just too splintered and spread out. I landed on google/gemma-4-31b for now, because it's cheap and good enough and also supports Dutch and French a bit. But I can't realistically try them all.

show 5 replies
_ache_today at 2:57 AM

If the performances are comparable, and there is no evidence it's not.

in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47

That is a massive cost reduction.

Refs: https://www.alibabacloud.com/help/en/model-studio/model-pric... https://runware.ai/gemini-omni

syntaxingtoday at 1:59 AM

> audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash

Wow crazy if true. I think Gemini's audio capability and multi language was the "selling point" for a lot of people. Other capability also matches or exceeds 3.8 Flash.

They also made a new harness but github link seems to 404.

show 1 reply
conceptiontoday at 1:41 AM

3.8 Max is the most “grounded” model I think - talks generally normal, doesn’t go crazy and start doing things (I see you Gemini), has good design choices and isn’t overly nitpicky. But god it’s slow. And only available from Alibaba. Their token plan is stingy too. If I had to pick the “old reliable boring” LLM, a modern Claude 4.5 if you will, Qwen is my choice. Hopefully they don’t RL it to oblivion.

show 2 replies
podocarptoday at 6:56 AM

Please what is flash pro ultra and all these, can they just use semver or something

show 1 reply
tolugeniustoday at 1:33 AM

Curious if or when we'll see the Qwen4 series, one thing I love with Qwen is it comes a much larger range of sizes so I can experiment which extremely small llms.

show 1 reply
lxetoday at 3:37 AM

Looks like the harness repo is already removed?

andy_ppptoday at 5:41 AM

Why can the Chinese build models and Europe cannot? The algorithms behind this stuff are not that complicated, are they? Is it the cost of energy? The illegality and difficulty of obtaining all the data in the world? Lack of capital to start moonshot labs? Lack of optimism?

The Chinese just seem to have an ability to get it done without anywhere near the GPUs of the US and Europe can buy these GPUs.

I think relying on the US and China for AI is probably not ideal? For example I think Qwen have not released the Omni models as open weights in the past, it’d be good to know if they’re doing this here?

show 1 reply