logoalt Hacker News

traestoday at 10:23 AM1 replyview on HN

> one of the obvious applications of novel mathematical results is in building stronger AI models.

This gets repeated a lot and seems to be one of the primary stated goals of making AI solve math problems, but I still have no idea by what mechanism this is even supposed to happen. I guess they could make some minor improvements to matrix multiplication algorithms or whatever but I don't see what groundbreaking theorem could possibly significantly improve LLMs.


Replies

svaratoday at 11:34 AM

It's the kind of thing where it's sort of expected that you wouldn't know, right?

I think we don't really understand why deep learning works as well as it does, the thinking around that is, as far as I can tell, mostly a collection of empirical observations.

A fundamental theory of learning that can be used to predict optimal network architectures might enable smaller models that consume less energy.