5 years of (200/month) tokens at that price, meanwhile an rtx 5090 pc is about half that… hmm
but i wonder how much these token costs are sustainable or not, it may be in the long term cheaper to have your own hardware if token costs go up (and hopefully hardware gets cheaper again)
The token price isn't the only reason to run a model locally though. You can do additional training to specialize or remove censorship that may be a no-no per TOS with cloud GPUs.
$2200 for a 64GB VRAM machine if you are willing to do a bit of work.
Those tokens aren’t guaranteed (esp. with RE and security tasks - rooted my own TV last week, Claude crapped out on “cyber safety” grounds; but also no guarantees about the model served - providers can pull a switcheroo on weights or quantization at any moment, and new options may not work for you), and you’re throwing money at entities that aren’t aligned with your interests instead of entities who are interested in actually empowering you.