logoalt Hacker News

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

191 points • by anerli • yesterday at 5:37 PM • 96 comments • view on HN

Hey HN, Anders and Tom here. We're building Magnitude, an inference engine for agents that optimizes itself to run as fast as possible on your hardware. It works on Mac, Linux, and Windows on any hardware and is up to 2x faster than llama.cpp.

We're both software engineers and previously built an open source browser agent to 4k+ GH stars and 100k+ downloads. We increasingly wanted to run it on local models, but found that no inference engine worked for our use case.

Inference engines today all make a performance tradeoff. They are either:

- Built for batched inference on datacenter hardware at the cost of single-session performance (vLLM, SGLang) - Designed for broad compatibility instead of optimizing for specific hardware (llama.cpp, Ollama) - Specialized for specific hardware or models but lacking engine completeness (oMLX, ds4)

Plus none of them are designed for running agents locally. Sessions are long, several often run at once, and you still want to use your computer for other things.

Magnitude is built for maximum performance on your hardware and running local agents:

- On-device compilation and tuning: Kernels are written with flexible parameters that are tuned on your actual device before the model runs. This gives you broad hardware compatibility with the same performance ceiling as hardware-specific kernels.

- Focus on best architectures: We write our tunable, highly efficient kernels for the most popular open-weights families. This allows us to achieve and surpass the performance of hardware or model specialized engines, without forcing ourselves to over-generalize at the cost of performance.

- Dynamic memory allocation: Magnitude reserves only enough memory up front to hold model weights. As your agent sessions grow, the memory heap dynamically increases, and frees itself when agents stop. Your hardware can still be used for other stuff while agents run.

- Hybrid paged attention: We borrow the best ideas from engines like SGLang to allow concurrent sessions to share prefix caches, but optimize placement for memory-adjacency so single-session performance doesn't suffer.

Magnitude is fully open source (Apache 2.0). We built it in Rust, including a custom GPU kernel runtime and autotuner. We take inspiration from the best innovations in inference from academics (e.g. FlashAttention, FlashInfer, TurboQuant) as well as other engines (e.g. SGLang radix attention) to reach the performance ceiling.

Benchmarked against llama.cpp with Qwen 3.6 35B A3B (4 bit), 64k context, no speculative decoding:

Metal (Mac M4 Pro 48 GB) - 92% faster decode (30 tok/s → 57 tok/s) - 9% faster prefill (466 tok/s → 507 tok/s) - 28% less per-agent memory usage

CUDA (DGX Spark) - 19% faster decode (49 tok/s → 58 tok/s) - 23% faster prefill (2,033 tok/s → 2,507 tok/s) - 27% less per-agent memory usage

Magnitude ships as a desktop app that you can easily connect with whatever agents you already use (Pi, OpenCode, Hermes, Codex, and more). It automatically runs models on demand when these agents actually need them, and shuts them down after inactivity. Here's what it looks like: https://www.youtube.com/watch?v=0qE8BWEZu7o

We're excited to push Magnitude further to let you run bigger models on the same hardware while continuing to improve performance. Our plans include:

- Expert streaming: store experts on RAM or disk and load them just-in-time. This lets you run models bigger than what otherwise would fit on your GPU.

- Kernel compiler: our current kernels tune a few parameters to fit your hardware. We can take this further with a fully custom compiler that automatically chooses how to fuse kernels and which implementations to use, to make it fit to your hardware even better.

- Multi-device utilization: Make the best possible use of all hardware on a system (CPU, GPUs, RAM, disk) by detecting these and automatically solving for the best model layout.

We'd love for more people to try it out and give us feedback. Feel free to comment here, we'll be around all day!


Comments

lxe • yesterday at 6:12 PM

On my local inference box I have a perpetual codex thread open in my llama.cpp checkout that I periodically ask to take a look at currently pending llama.cpp PRs, do some research on latest MTP, Dflash and other prediction or attention optimizations, do research on the latest model quants and finetunes, take a look at localLlama Reddit threads and just do essentially a sweep of the frontier.

Then it rebuilds latest llama.cpp, grabs the PRs it finds relevant to test against, and then it performs a benchmark and finalizes the upgrade and verifies what model, variant, or even a separate finetune that we should be running.

Occasionally, it performs its own optimizations and commits, which then gets superseded by pull requests and merged code that essentially validates the model's own optimization directionality.

➕ show 1 reply
kmike84 • yesterday at 6:07 PM

This seems to be a good idea. However, beating llama.cpp on speed is a low bar :)

I found it to be a good baseline, but at least on Mac there was always something way faster, and/or with better memory requirements - like you said, ds4, omlx, mtplx, etc. It seems if you use local LLMs for real, there is very little reason not to use one of the more optimized engines.

3 main failure modes I observed in the engines:

* Not using best available spec decoding

* Using too much VRAM for KV cache (e.g. KV cache used to take almost nothing in ds4, but huge amount of VRAM on unsloth/llama.cpp for deepseek models)

* Degraded performance at large context sizes - benchmarks at 4K or 32K are awesome, but at realistic 100-200K it's slower than some stupid baseline

➕ show 2 replies
openamer • today at 7:06 PM

Congrats on the launch. We went down a similar path with OpenAmer (open source) and the lesson that stuck: for self-optimizing agents, the bottleneck is not making the agent smarter, it is verifying that each claimed outcome actually happened. We ended up routing every result through a heartbeat subsystem that re-checks it before it enters shared memory - without that, the agent happily builds on its own hallucinated successes. Curious whether Magnitude's self-optimization loop includes a verification stage or relies on the benchmark score alone.

kmike84 • yesterday at 10:18 PM

How accurate are speed estimates in the UI? I'm asking because for Qwen 3.8 (Q8) the speed numbers cited in the UI look quite poor:

  Estimated speed on your machine
  Context tokens Tokens / sec
  25 000 17
  50 000 16
  75 000 16
  262 144 12

262K number is ok, but for lower context sizes (<128K) it's about 2x slower than the numbers I'm getting from real mtplx sessions for qwen3.8 q8 (mac m5 max).

Is it a lack of optimizations, or incorrect numbers, or a benchmark artifact (e.g. something which is harder for spec decoding than usual agentic sessions)?

➕ show 2 replies
mncharity • yesterday at 6:56 PM

Fwiw, top of my own pain-point list (I suppose given the first item, that's a pun) includes:

External/policy-based throttling for temperature control. Unthrottled, my laptop bottom goes skin-burn hot. But fixed compute caps can have non-linearly dreadful performance impacts in particular cases. Plan is a runtime knob, to replace manual limits-kludgery.

I'll use models which barely fit in VRAM+RAM, and are order-1 tok/s slow. So tool call step overhead can be painful - a world where `ls` costs tens of seconds. Plan is blending harness plugins with inference loop, for "no, don't stop - I already have the call result for you - just keep going" (and also some logit games).

➕ show 2 replies
bythreads • today at 6:49 AM

Ok so i took the time to benchmark this on the following on my m5 max 128gb:

Qwen3-4B-Instruct-2507-4bit Qwen3.5-35B-A3B-4bit Qwen3.5-9B-MLX-4bit Qwen3-Reranker-0.6B-4bit Qwen3-Coder-30B-A3B-Instruct-4bit qwen2.5:0.5b

and the results are what i kinda expected to begin with, this adds next to nothing? - also the repo was pivoted from a playwright sub assembly to this not long ago - so my conclusion - THIS MIGHT be worth some watching if you have a model where no-one!, has optimized it at all - and where it does not use anything native to your platform.

results (averages)

VIA rapid-mlx :8902 (MLX) Decode: 175 tok/s TTFT: 64 ms prefill (~760 tok cold): 594 ms

Magnitude 0.2.1 (GGUF/llama.cpp+Rust) Decode: 161 tok/s TTFT: 111 ms prefill (~760 tok cold): 669 ms

➕ show 1 reply
sebastienburel • today at 5:42 AM

On a Mac the baseline I'd want is MLX, not llama.cpp. llama.cpp isn't the fast path on Apple Silicon for most models people run locally, so a speedup over llama.cpp could still be slower than mlx_lm. Do you have that number?

Second, more important for agents: decode speed is rarely what hurts. It's resending the same system prompt plus tool schemas every turn. Does self-optimizing cover prefix cache reuse across requests, or is it kernel and layout tuning only?

And is the endpoint OpenAI-compatible? My runtime already talks to llama.cpp and LM Studio through that wrapper, so drop-in is the difference between trying it tonight and not.

➕ show 2 replies
herf • yesterday at 6:27 PM

I have two NVIDIA GPUs (16GB+16GB) here, and it detects them each twice (says I have 4 GPUs). But then, it says most models are too big (anything >8GB?) and seems to run only on one GPU (5070ti).

Unfortunately even with my 5070ti, llama.cpp seems to be about 20-30% faster at decode, running as:

set CUDA_VISIBLE_DEVICES=0 build\bin\Release\llama-server -hf google/gemma-4-12B-it-qat-q4_0-gguf -ngl 99 --no-mmproj-offload -mg 0 -c 262144 -fa on --host 0.0.0.0

➕ show 2 replies
msdz • yesterday at 6:15 PM

Congratulations on the launch, it looks like an impressive product and tool!

Q: From my (very, very limited!) understanding, I’m under the impression that part of the “inference engine inertia” is that model- or at least architecture-specific code is required for most, if not each new open-weight model coming out.

Assuming I got that right, do you plan on supporting everything vLLM/llama.cpp can do, such that Magnitude becomes a drop-in replacement for as many (economically/pareto-viable) models as possible, or do you want to focus on the best possible support for only a select few models/classes of models?

➕ show 1 reply
happybox2016 • yesterday at 11:20 PM

2x llama.cpp" on what, an M3 Max? llama.cpp's metal kernels already saturate memory bandwidth. Real agent bottleneck isn't single-stream tok/s — it's KV cache for 5+ concurrent 128k contexts on 24GB VRAM. Who's actually running multi-agent locally? A) Single session only B) 2-3 agents C) 5+ agents D) Gave up,

➕ show 2 replies
nateb2022 • yesterday at 5:51 PM

Any source on the benchmarks/methodology besides the image? There's a ton of variance possible in llama.cpp's performance depending on how it was configured. I'd also like to see benchmarks against MLX.

➕ show 2 replies
larodi • today at 6:52 AM

Everyone focusing on the specs side, but how is this enterprise going to make money, given it is a YC cohort company?

c7b • yesterday at 7:35 PM

Cool idea! Do you happen to have benchmarks for Strix Halo (AMD Ryzen AI Max+ 395)? I take it that Qwen3.8-Flash-Next is not supported?

And a more general question: does your engine detect and optimize for custom setups like multiple (possibly different) GPUs, eGPUs,...? Because if all you have is a stock major system like a Mac or DGX Spark, that's all you're going to care about, and there are a lot of highly optimized single-hardware engines out there that will be hard to beat in the long run. Something that automatically adapts to custom systems that don't have their own subreddits could really fill a gap.

➕ show 1 reply
thoughtpeddler • today at 5:00 AM

For those running local models on macOS with Apple Silicon, how does Magnitude differ from what Apple's own first-party Core AI now does during its "specialization" procedure, wherein it performs some kind of "model optimization and conversion" (vis-a-vis AOT compilation) into a Core AI "compiled model" file that is optimized for Apple Silicon (leveraging custom Metal 4 kernels, or so Apple says), per WWDC labs from this year that discuss this? [0]

--

[0] https://developer.apple.com/videos/play/wwdc2026/326/

➕ show 1 reply
singh_abinashi • yesterday at 11:03 PM

Curious how the evals for this work on real agent workloads versus synthetic benchmarks. In my experience, agent cost and latency profiles change a lot once there's a tool-use loop involved, because the token distribution gets much burstier than a single prompt. Did you evaluate on multi-step tool-calling traces, or mostly single-turn?

aitoolcrux • today at 6:14 AM

Self-optimizing inference is one of those problems where marginal gains compound—small improvements in batching, KV-cache reuse, and routing across thousands of requests add up fast.

The hard part isn't the optimization itself—it's measuring whether a change actually helps across the long tail of request patterns. Most inference benchmarks show huge gains on popular workloads, but production traffic has a fat tail where naive optimizations hurt latency.

Interested in how you handle regression detection when the optimizer changes between requests.

teabee89 • yesterday at 6:29 PM

How does this compare to ZML's llmd https://zml.ai/llmd/ ?

➕ show 1 reply
lin7c • today at 3:23 AM

One thing I'd want to see in the evals is per-turn latency across a full agent trajectory, not just end-to-end time. In my experience the workload flips mid-run: early turns are prefill-heavy (big system prompt, tool schemas), late turns are short decodes against a huge KV cache, so a config that's optimal for turn one can be badly wrong by turn thirty. The self-tuning idea is interesting, but I'm curious whether the tuning happens per-request or per-trajectory. With prefix-cached tool schemas the win should compound; without it you're re-solving the same optimization problem every call.

mrtsepelev • today at 10:08 AM

Congrats on launch! Tried it on the gemma-4-26b-qat-4bit model. Was indeed faster on token generation then on oMLX (82.8 tok/s vs 76.5 tok/s), but the prefill time was ~2.6x slower (709 tok/s vs 1843 tok/s). Don’t use any acceleration on the oMLX. Macbook M5 Pro, 48 gb

thoughtpeddler • today at 4:48 AM

This just makes me think about optimizing model weights to run as 'close to the metal as possible' (i.e. within or 'just above' a UEFI boot environment, like the NightRun project), so that there isn't any OS-level overhead either. If we're optimizing, let's optimize! Curious though, maybe the OS doesn't impose much of a burden here? Open to hearing what other tinkerers think...

➕ show 1 reply
hypercube33 • yesterday at 6:35 PM

From your description looks like this isn't for AMD or Strix Halo at all? Also one of the things I'm not sure of but definitely plays a huge factor is the variant of the model you download - how does this help select the fastest version for your specific hardware / context size?

➕ show 1 reply
cedricd • yesterday at 7:27 PM

Looks interesting! Is there any way to skip or speed up the 'Assessing Models' step? I'm unable to download anything because it's been taking forever. I'm sure you could apply some quick heuristics or do a lookup or something to filter models. Or trust the user a bit more -- I already know which models fit on my machine. As it stands I'm stuck at that step and can't use the app.

Maybe have it run silently in the background and assess on demand when a user selects / attempts to download a model. It's not quite clear why all need to be assessed before I can download the first model to try.

➕ show 1 reply
steinvakt2 • today at 1:36 PM

Any way to use this for speech-to-text? To gain faster whisper inference for instance, without losing accuracy?

➕ show 1 reply
theParadox42 • today at 7:17 AM

I remember when this company was just doing agentic playwright style browser interactions

MaxikCZ • yesterday at 7:57 PM

if fully custom compiler would find best settings for given setup, upload the setup to mothership and allow new peers to download it as good starting point.

Can it do all the shenanigans that allows to run qwen flash on 12GB vram over 40 toks like people seems to be getting in this thread?: https://www.reddit.com/r/LocalLLaMA/comments/1wp7zyb/qwen38f...

➕ show 1 reply
malshe • today at 2:52 PM

Can this be used for fine-tuning models? I have a M4 Pro Mac mini with 64 GB RAM.

➕ show 1 reply
digitaltrees • today at 12:25 AM

Do you support splitting models across devices so larger models can run on clusters?

I am building propelcompute.com an open router for private hardware and experimented with exo labs to run large models on for Mac studios and plan on doing the same with nvidia and amd. Id love to integrate your inference engine into the system but built gpu is critical.

➕ show 1 reply
kenzic • yesterday at 5:55 PM

How long does tuning take (on an M3 MacBook Pro for example)?

➕ show 1 reply
karlkloss • today at 11:39 AM

It couldn't detect my hardware at all, Win11 Intel Core Ultra 7 notebook with Nvidia RTX PRO 1000.

➕ show 1 reply
amirhesham • yesterday at 5:51 PM

Oh this is so cool. Curious about the business model, too.

➕ show 1 reply
chzblck • today at 3:13 AM

Sounds interesting would love to test it out but here's what I got when I first tried to get some models downloaded.

On a 64gb Ram and 5080 machine the biggest model suggested was Qwen 9B

can hit 90+ tps on the MoE 35b but mag thinks it wont fit.

➕ show 1 reply
NKosmatos • today at 9:35 AM

Nice one! Tried it and unfortunatelly there are no small models that can fit my 16GB RAM or GTX1650 4GB GPU. Yeah, I know that this configuration is not meant to be used for AI/LLM work, but it would be good to provide support for some smaller models so that us plebeians can also play a bit with what you techbros are used to ;-) There are many small/very small models out there and I'm sure you could add a couple just for playing around and experimenting.

paulgerhardt • yesterday at 7:00 PM

Trying to run this but keep hitting bugs. Can you open up issues reporting on your repo?

➕ show 1 reply
sgtwompwomp • yesterday at 5:52 PM

This is dope, is this kind of like Wafer.ai but for local models? As in a coding agent optimizes the kernels so the local model runs continuously better? Cause that is compelling if so. If it’s more simple that’s cool too

➕ show 1 reply
loclol101 • today at 4:55 AM

Will it support multi-agent setup across heterogeneous devices (macbook pro, RTX5090, mac mini, etc)?

➕ show 1 reply
p-e-w • yesterday at 5:44 PM

What is the business model?

➕ show 1 reply
taylorhou • today at 12:38 AM

Running an inference network across 11 Macs (Teale.com), so I pointed my orchestrator at the repo.

The autotuner is the real thing - kernel-level search, config budget split by measured time share, winners cached per device/toolchain. Rotating resident weights across layers to dodge the hot-cache trap is a nice touch.

One question on model ranking: the fit scores look like predictions built from cost constants measured on a single M4 Max, not per-device measurements. How do you rank models across genuinely mixed hardware? Does the estimator improve from actual runs over time? That gap between predicted and measured fit is what eats mixed machine fleets alive.

Shameless plug since magnitude's goal is highly relevant to what i'm working on: teale.com - distributed inference across fleets of macs. If you're running local models on more than one box, check it out with your agent!

➕ show 1 reply
andrethegiant • yesterday at 11:19 PM

Congrats on the launch!

Zetaphor • yesterday at 11:39 PM

Please consider adding support for Qwen 3.8 Flash Next

➕ show 1 reply
yolandac • yesterday at 5:59 PM

does it allow us to run larger models that weren't possible before?

➕ show 1 reply
nullbio • today at 10:20 AM

Why is everyone talking about Mac like it's the only hardware people use?

➕ show 2 replies
eventuallyworth • today at 10:12 AM

[dead]

alescalaios • today at 9:09 AM

[dead]

tinykit • today at 5:34 AM

[dead]

ContinuityLab • today at 12:39 AM

[dead]

ls-a • yesterday at 9:17 PM

[dead]