It's running a very small, non-reasoning model at the moment. But more generally, almost all LLMs argue on the hardware/model they are/are on.
What would tokens/sec performance look like for a reasoning model? An order of magnitude slower?
Which model? Or how many active parameters?
What would tokens/sec performance look like for a reasoning model? An order of magnitude slower?