I think their move makes a lot of sense. Frontier models can probably not advance much further with the datasets we have currently available. Most of the text that goes in is online chatter, images and some scientific texts. That's great for chatbots, knowledge retrieval and programming. But with that database genuine discovery is hard to do. I think for the next step in intelligence the models need to have access to much more data: Data from physics, chemistry, biology experiments - and so on. And they need the data in much higher fidelity than you can currently access. If we just feed AI with all the knowledge we've acquired it's much harder for it to become smarter than us - it basically needs it's own eyes, ears, nose and so on.
> Frontier models can probably not advance much further with the datasets we have currently available
I’ve been reading this sentiment on HN since GPT4o, yet models got better and better
> Data from physics, chemistry, biology experiments
There are PetaBytes of important scientific data locked in archival file formats. The first step is to make this efficiently readable.
https://www.earthmover.io/blog/virtual-zarr
https://news.ycombinator.com/item?id=46659254