I would be interested to have a third X-axis with dollars.
Time is sometimes more about inference infrastructure (especially with systolic chips) than model quality (and providers tweak knobs to support higher batches at the expense of latency).
Tokens are not always fully equivalent between models.