I think(?) you’ve already probably done a good job of explaining this criticism for semi-informed people. But can you dumb it down even more for those of us who are almost entirely out-of-the-loop?
> Training on the eval puzzles is cheating / “training on test”
> No this is false. “Training on test” specifically means training on the labels of test data. The labels were not trained on.
> Also, ARC is a metalearning benchmark, so you’re supposed to learn from the eval puzzles.
> Jargon: ARC has a set of train puzzles and a set of eval puzzles. Each puzzle has example pairs and test pairs. A pair consists of an input grid + output grid.
> The ARC, the label is only the test pair’s output grid in an eval puzzle.
> These labels were not trained on. They are hidden. You can delete it beforehand if you wish
I think what I gather here is that the test comes with one batch of training problems, which everyone agrees you can train on. But maybe the eval problems also come with input/output examples (to help define the problem) and training on those is controversial? I can’t see why it would be controversial but is that the criticism?
What you gather is correct, assuming by "the test" you mean the ARC benchmark in general. It was controversial because people are used to LLMs which are frozen at train time, where the eval problems are usually not trained on for various reasons like fragility (basically porridgeraisin's ans which is great)
Here's another explanation. Take the train dataset and test dataset of a benchmark
Train: {x_i -> f(x_i)}, Test: {x_j -> f(x_j)}
As long as f(x_j) in the test set is hidden, there is no "training on test". In a normal benchmark, each x_i is a single datapoint. But in metalearning benchmarks like ARC, x_i is the puzzle itself that has a train set and the test questions within it, hence the confusion and controversy
I won't weigh in on whether it's "cheating" but it is definitely benchmaxxing
Basically, you have a bunch of Q,A pairs in the training dataset. Here, it was trained to next-word predict the question itself, as well as next-word predict the answer given the question as prompt. This is bog-standard, no one's complaining.
In the test dataset's Q,A pairs, it was only trained to next-word predict the question itself, and it was not given the answer at all.
It was then evaluated by seeing if it is able to output A_test given the Q_test as prompt.
What would be cheating is training it to produce A_test (given Q_test as prompt) as well, since then you can always make a model that scores 100% by just memorising Q_test, A_test pairs.
The complaints online mostly stem from not reading that properly and assuming they trained on Q_test,A_test instead of just Q_test. This is further because these days large LLMs are inadvertently trained on many benchmark solutions even unintentionally due to the massive scale of data and the infeasibility of auditing it all. But none of that is the case here.
The reason you want to train on Q_test is because in these AR transformer models, they learn useful composable encodings of Q by simply learning to next-word predict Q. So you enable the model to learn composable encodings of the test questions, so that it can hopefully "connect it" to an earlier train problem it had seen, and adapt the solution it had seen for that, much like humans do in school exams.
Without this step, you are making it difficult for the model to "connect" the test question to a train question it had seen earlier, and then it still has to adapt the solution. This way, you precompute that "this test question is like this train question" and then during the exam you only have to do the adapting the solution part after a simpler "retrieval" process.
You can just think of next-word training Q_test as a "retrieval" process.
This practice often used in continual learning or "test time training" is not yet useful in general real world ML tasks due to the differences in memory and compute requirements, and more so the general fragility of training large neural networks in a streaming realtime way (as opposed to large data, batched), versus inferencing from a static neural network. It is due to that fragility that I believe (correct me if I am wrong) this guy had to train on a batch of Q_tests. If you enforced that you will not provide Q_test_2 before they answer Q_test_1, the performance will drop.
While the increased compute and memory is difficult to solve inherently, there are various efforts being made to fix the fragility, especially in reinforcement learning where this is called "streaming RL", there is revival of interest as seen in RLC 2026.
[Note] Arc-AGI-1 doesn't have any actual english words or such, but it's simpler to pretend it was a basic Q&A benchmark to explain the above
The point of ARC is essentially an "IQ Test" for AI systems. It is meant to cover abstract reasoning capabilities of generally-intelligent systems like LLMs. What the author did here was build a system that only solves ARC problems.
The other tension is the fact that this score is on the public eval set. In machine learning, you typically have 3 datasets: training, evaluation, and test. The training set is the dataset that's used to update the weights according to your loss function, you are "encoding" the patterns from the training set directly into your model. The eval set is what you use to track performance while training, it is NOT used to update model weights, but shows how well the model generalizes. The test set is a private holdout set that is only used when you're "done" developing your model. The difference between test and eval is information leakage: you can use performance against the eval set to modify your hyperparameters and model architecture to get better eval scores. So while the eval set doesn't directly update the weights, it can indirectly cause "overfitting" by tailoring your model to do well on the eval set. What you really want to see is the private test set performance, not the eval set. For all we know, this model could be ridiculously overfit on the eval set and perform poorly on the private test set.