Mathematics is like this. You read the first symbol in the paper, it is a wiggly triangle, what does that mean? Well you will find out that symbol means the Constant or Operator or Set or Function belonging to So-and-So with the unfortunate name. Well now you know it is called Grossediche’s Member, what does that mean? You will find out it is defined in these dozen lines in Grossediche’s seminal paper, you will need to read the entire paper to make sense of these dozen lines, you will need to read everything he published in this particular decade to make sense of the paper. Each of the dozen lines is jam packed with other symbols, for each of those you will have to repeat this entire process, with another stack of papers, from another unfortunately named mathematician. Now you have a firm grasp on Grossediche’s Member, you return to the original paper. You read the second symbol in the paper, it is a half-melted letter t, what does that mean? Well, …
Behind each symbol is a whole paper, behind each paper is a whole life’s work, and so on. With this in mind, it is perhaps not so surprising that language models operating on embeddings are extraordinarily well-suited to this particular task.
LLMs don't understand things as human mathematicians do, even though they are very good at finding analogies and similarities. Their advantage is a larger search space (experience) and search speed, not better understanding.