Humor requires a sudden orthogonal leap from context. That's what a punchline is.
I think you're saying that you can eventually train models to arrive at that destination by training on existing jokes, effectively encoding these leaps as probabilities.
In that case, the model isn't actually making an intuitive/comedic leap; they're just following new probability chains in attempting to approximate examples they've seen in training.
I'm suggesting that something architecturally different is necessary to create a model which can make intuitive/comedic leaps.
Try to get a frontier model to write a clever, funny joke which hasn't been seen before. Or, try to get it to make an intuitive leap that leads to a novel discovery.
You can use them to guide your own efforts along these lines, as a sounding board. But with current architecture I just don't think either is possible for an LLM to do on its own.
> Try to get a frontier model to write a clever, funny joke which hasn't been seen before.
This is what I refer to when it comes to larger models like Fable 5 being funnier. They are more capable of doing that. They can deliver that "sudden orthogonal leap from context" of yours more reliably.
It's not a "fundamental inability" and never was. If you crank the scale up and a capability appears, "current architecture" was never the problem.