I find it unlikely that there's any direct training data for this. I'm talking about direct training data for using brush strokes like humans to construct an image.
Correct me if I'm wrong but this shows that painting capability is emergent, arising from unrelated training data thus very very compelling evidence that LLMs are actually intelligent.
One other commenter said that this capability was initially revealed by Anthropic employees, raising the possibility that they were RLd on it.
I'm also quite surprised to see that LLMs can do this. I guess it is possible that "make an image in MS paint" is a type of RL environment used for image understanding. This is one of those areas where people inside the labs have a very different view into how much models are generalizing.
It could have been fed videos of people painting.
Connecting it's capabilities to simulating the recreation of such paintings... that's I suppose interesting. But adding brush strokes to produce an image is probably well covered by video footage in its training data. e.g. all the Bob Ross videos in existence.