Interesting. Wasn't Deepseek's founder saying that they had explicitly decided not to focus on multimodal models at all and were going text-only because they believed it was enough to achieve AGI?
Worth noting that deepseek has had a separate vision-capable model for some time, which also powers their chat interface's vision mode
I think you're thinking of Dario saying this about image generation.
It was explicitly said that they are pursuing multimodal support. A quote from the meeting transcript: https://github.com/demo-zexuan/liang-wenfeng-investor-meetin...
Earlier, the following was said, which might match more what you had in mind. It is difficult to tell who said what, since the speaker ids are missing.