logoalt Hacker News

taylorhou • today at 12:38 AM • 1 reply • view on HN

Running an inference network across 11 Macs (Teale.com), so I pointed my orchestrator at the repo.

The autotuner is the real thing - kernel-level search, config budget split by measured time share, winners cached per device/toolchain. Rotating resident weights across layers to dodge the hot-cache trap is a nice touch.

One question on model ranking: the fit scores look like predictions built from cost constants measured on a single M4 Max, not per-device measurements. How do you rank models across genuinely mixed hardware? Does the estimator improve from actual runs over time? That gap between predicted and measured fit is what eats mixed machine fleets alive.

Shameless plug since magnitude's goal is highly relevant to what i'm working on: teale.com - distributed inference across fleets of macs. If you're running local models on more than one box, check it out with your agent!


Replies

anerli • today at 8:19 AM

Hi, the model ranking scores are based on whatever machine Magnitude is actually running on. So it will account for your specific memory capacity and performance characteristics to recommend appropriate models. The estimator is purely to help filter and recommend a model. The tuning is the only part that actually effects real performance, and is done automatically whenever you download a new model.