why isn't this being compared to siglip2 (also from google)? because that one isn't fully multimodal? or because it's a different org/team?
well EG2 doesn't have an encoder so can't use it for OCR, for one
well EG2 doesn't have an encoder so can't use it for OCR, for one