I've been experimenting with it on a M4 Pro 24G for the last few hours and it's been very promising using 32k context. getting around 30-40 tps
With Qwen3.8 27B I could not get anywhere near 32k context window, that made it very unusable for agentic coding, although it was very smart.