Exactly. It's got nothing to do with sensory input and everything to do with reasoning.
If someone asked me how many f's are in a word I hadn't seen before verbally, then a reasoned response would be that I don't know, but I estimate based on the syllables...or ask them to spell it out.
These are all the sorts of questions where general problem solving works, even if the conclusion is "I don't have enough data to speculate".
So that these models fall apart on it so readily means we're either grossly handicapping then with the requirement to "be helpful" or they just fail to recognize the problem and are just stochastically spitting out a high probability token sequence for the input.