logoalt Hacker News

lxe • yesterday at 6:12 PM • 1 reply • view on HN

On my local inference box I have a perpetual codex thread open in my llama.cpp checkout that I periodically ask to take a look at currently pending llama.cpp PRs, do some research on latest MTP, Dflash and other prediction or attention optimizations, do research on the latest model quants and finetunes, take a look at localLlama Reddit threads and just do essentially a sweep of the frontier.

Then it rebuilds latest llama.cpp, grabs the PRs it finds relevant to test against, and then it performs a benchmark and finalizes the upgrade and verifies what model, variant, or even a separate finetune that we should be running.

Occasionally, it performs its own optimizations and commits, which then gets superseded by pull requests and merged code that essentially validates the model's own optimization directionality.


Replies

anerli • yesterday at 7:57 PM

Yeah we heavily leverage coding agents for optimizing our kernels. Since it's highly verifiable and takes time to measure we often leave multiple running and improving performance on different model architectures.

Definitely still helps to reference relevant academic work as well, or even just encouraging the agent to make bigger structural leaps, otherwise it will often get stuck working on low impact micro-optimizations.

➕ show 2 replies