Seems fairly trivially fixable to me, e.g. by allowing the LLM to call a tool to spell out a word.
... assuming you build the tool and then think that it's worth polluting context with making that tool available, and then that the LLM decides to actually use the tool. Tool parameter space and tool selection still remains a complicated topic.
Yes, they can already do this by writing code, and you can train them to know how/when to do this. Fundamentally though, it’s still a “what color is the air” type of question after tokenization.