logoalt Hacker News

quantoyesterday at 11:42 AM2 repliesview on HN

Is your Qwen3.6 locally run on your laptop? What kind of tokens/s are you getting from your laptop GPU?


Replies

nozzlegearyesterday at 9:29 PM

Qwen is running on my Mac Studio, an M1 Ultra 64gb. My harness (oh-my-pi) on my laptop is configured to use the models hosted on my local network, since it's just a MacBook Air 16gb and probably incapable of running anything useful itself.

I get about 45-55 tokens per second using Qwen with this setup. I could probably squeeze out more if I messed around with the settings, but I'm mostly using oMLX's defaults for the model.

saagarjhayesterday at 11:46 AM

I have Qwen3.6 35B-A3B on my laptop and it does 60 tokens/s

show 1 reply