I am very struck by the way open weights LLMs seem to reflect a culture.
I don't really enjoy the way Qwen writes prose, and I find its thinking a bit exhausting, though it clearly writes very good code.
I like the neutral, clear way the Gemma models write, which I sometimes use to get myself a "getting started" document on something I want to understand; it also summarises well. It is neutral, sensible, un-showy. It writes in a way that is fairly close to what I would use for documentation. The 12B and 26B models are also very good for talking about art and photography. Analysing my own photographic work has helped me more than I expected it to.
This model, honestly, has made me smile. It also feels like it is more creative at a given temperature than Gemma. I am trying to motivate myself to do something quite open-ended so I asked it about what other people's considerations might be in my situation, and at the risk of anthropomorphising, the things it has come up with feel like the work of a more curious mind, somehow. More eclectic. I have enjoyed testing it and I really want to test it more, which might help me get over a motivation hump there, too.
(I am also exploring its hard-wired policies by asking it to analyse some studio art nude work I have done; it definitely thinks out loud about its policies in a way I have not seen Gemma do.)
I ran a 9B over my like 100k photo library — it was very good at it. And extracting any text. All local.
I think we’re going to see a lot more “product“ focus in the future with deliberate attention paid to these kind of properties. Historically though there are some obvious differences, the focus has been on benchmark maximizing. As that saturates, I expect more interesting choices about writing and thinking style designed to be differentiators instead of a side effect. Kudos to the PM here for taking it in a different directions, there’s obviously been thought put into it.