If anyone's looking for something a bit more minimal, I can't recommend hax [0] enough. No MCP, agent just gets a shell tool, simple config, whole thing's in C.
It says "local models" - do you know if it makes any special effort to fit work within a small context window?
That looks interesting, built it (took seconds) and it is much smaller than Pi in terms of what it needs to work (when i installed Pi in a fresh Debian container it downloaded ~500MB of stuff, which isn't exactly what i had in mind when i read it is minimalistic :-P but it is a container so i didn't care much).
I'm just running it now in its own source with Qwen 3.8 27B and llama-server and asked it to analyze the code itself. I'm mainly curious to see how it handles context compaction during tasks (what Pi does is almost seamless and AFAICT it isn't anything particularly fancy so i'd expect Hax to do something similar) and i guess asking it to analyze a whole C codebase would help trigger that with a 131,072 context. Unfortunately it seems to be missing some "context usage" indicator while it does stuff (it shows context usage in the prompt but not while working), but i guess if it does manage to analyze the C code properly, i can ask it to add that :-P and see how it fares (from my use of Pi i'm positive Qwen 3.8 27B can do all that stuff, so it'd mainly be up to the harness).
EDIT: also i wonder if it works nicely if it is possible to convert Pi transcripts to Hax - i have a few "in progress" and i'd like to continue where i left from, though while both seem to use JSONL for the transcripts i'm not sure if they're compatible
EDIT2: hrm, it tried to use more than available tokens during a compaction and stopped there expecting me to increase the limit (i can, but what if i couldn't?) and restart the llama-server. Pi sometimes does hit it but it manages to recover by itself without requiring any input by me (or to increase llama-server's limit).