People have had surprising success adding vision to open-weight LLMs that ship without it, like DSV4 Flash [1] or GLM-5.2 [2]. Given this model is already vision-trained I expect that approach will work well here.
[1] https://old.reddit.com/r/LocalLLaMA/comments/1vl6ior/i_gave_...