> it would suggest that the model didn't become significantly more intelligent overall, but significantly better at solving those specific problems
We would actually need a test that shows the ability of a model to export its skills to more problems ("interdisciplinarity" etc.).