Missing multimodal again?
It is so valuable in practise to be able to have the models see screenshots - I guess if they aren't in the benchmarks then nobody will focus on it. But it completely nixes these for some of my main use cases.
I can’t come up with a use case where I couldn’t extract the image details using another, multimodal model and pass it into the GLM’s context with as many details as I need.
I would assume that GLM 6 will be multimodal, but 5.x will be text-only.
Probably not what you're after, but I've considered having a separate small mm-model act as a seeing-eye dog for the bigger more capable one.