Are there any projects that track how much usage of each model translates to how much percentage drop in weekly/5hr windows?
Not that I know of. AA's token use metrics (mentioned in this article) are indicative, however. They say explicitly here that the Grok models are notably token efficient. This is my experience.
Usage? Not exactly. But I tried to make something that can estimate dollars per tokens in actual usage while taking into account multiple factors.
https://harness.eveid.com/lazy-harness-cost-simulation