> I dunno why Apple fans
You mean someone who uses a Mac? I find the term "Apple fan" is used as a way to attack the person.
> that can run full agentic loops, you may as well just use cloud inference for the price.
You can run full agentic loops fine with quantised models. In fact that is likely what you are doing with a 32GB PC Graphics card. Or what model are you using?
You use the right model for the right job. For example granite4.2 is optimised for agentic work and only needs 5GB of memory.
Gemma4 MLX runs fine with the larger model needing 19GB.
Prior to that I had OpenClaw (in Parallels VM) create an application with local models that worked the exact same way as created by Claude. It was more an experiment in how OpenClaw works, hence the VM.
> And 57 gb is split across 3 cards quite easily,
I assume you are talking about a good graphics card. A good 32GB will run you $2K a card, so that $6K to beat out a laptop of similar price, and where the difference doesn't matter.
[edit]
Anyway my main point is Macs work fine for local models. I have a 6 year old machine that proves that.