It also seems to have far more tokens per second than needed for general "close the blinds" "tool_call(blinds, CLOSED)".
I do wonder if more smartness could be had by using sparser experts.... And possibly even having some kind of expert switching penalty to try to reduce the amount of data read from read only flash memory by encouraging subsequent tokens to use already loaded experts.