logoalt Hacker News

syntaxingyesterday at 2:53 PM2 repliesview on HN

Really looking forward to this, 27B is a struggle with a strix halo and Laguna 2.1 can do stupid things for tooling calls.


Replies

cpburns2009yesterday at 2:58 PM

Yeah 27B is way too slow for the Strix Halo. Laguna was better but still slow when I tried it. Qwen3.6 35B is still the best today.

show 4 replies
corysamayesterday at 5:00 PM

So, I know https://cactuscompute.com/needle is designed only to enable tool calling on tiny devices. But, I wonder if anyone has used it as a CPU-side mediator between a tool and a GPU-side local LLM making semi-natural-language tool requests...