It's definitely cool that you can get any reasoning whatsoever out of such a small model. That said, its reasoning is "interesting":
Query: "Make the living room dark" Agent: "User wants lights on in living room. 'dark' implies dim. Room 'living room', action 'on'." (And on every test I did, it just completely ignored the "brightness" parameter)
It also appears to have no concept of what a door or light actually is, whenever the query diverges from "Lock door X" or "Turn on light X", it tries to shoehorn whatever additional context is given into the device name:
Query: "Lock out the vacuum salesman at the front door" Agent tries to lock "front door vacuum salesman"
"The way you talk really makes me appreciate silence" is classified as "positive" with 82% confidence.
that model is 14MB large what do you expect. but I agree it's funny regardless
Ok, this is genuinely funny, we will fix these as we iterate, thanks lol.