I also wonder how small a LLM trained on catching only subject (e.g. living room) and action (light on) from text input could be compared to needle - the json wrapping could be done afterwards using templates.
The problem is the target device, an LLM can't run on an average TV well.
The problem is the target device, an LLM can't run on an average TV well.