Astra-Ultra? Even the largest open model to date (Kimi K3) is nowhere close to Astra level, and it will be quite slow even on the highest-spec M5 Ultra, with achievable speeds of about 0.5 tok/s at most due to having to stream weights from SSD (~13 GB/s on the highest storage capacity M5 Max machines so far). This is OK for doing simple Q&A in the background but it's far from a genuine coding experience. You'd have to test batching of multiple thinking streams in order to try and raise overall tok/s via layer-wise reuse of the streamed weights (and this is where the "Ultra" part sort of becomes relevant; Kimi series models have good support for agent swarms) but this would decrease single-session performance even further. It would only be usable for background jobs, though the hardware would then have a chance of paying for itself if it was fully used on a 24/7 basis.
> You'd have to test batching of multiple thinking streams in order to try and raise overall tok/s via layer-wise reuse of the streamed weights
isn’t this very straightforward to do..? I thought batching for Qwen models is already proven out.
> but this would decrease single-session performance even further
Well let’s take Qwen 3.8 27B. Throughput for M3 at 8 agents is 4x compared to single agent. [1]
It’s really not clear to me that 8 concurrent agents at half speed will be worse task completion latency than 1 agent.
And that’s M3 studio benchmarks, not even M5 ultra, and without the many software improvements we will see
If you haven’t tried Qwen 3.8 27B xhigh on a task you might not get the hype. Idk.
If you’ve tried doing this and don’t like it sure, and be specific about what isn’t effective, but let’s not speculate.
[1]: https://omlx.ai/benchmarks/performance/69kzkrv8?utm_source=c...