It seems like the more honest comparison would be to OCR the screen and send that as input to the LLM?