That's perfectly aligned with my point, thanks for the opportunity to expand.
The hacking agents being tested have goals beforehand, from the frontier lab or from a superior agent, that they execute immediately.
But the perceived experience most people have is a chatbot, which is the encyclopedia form.
Ah, are you saying that because most people don’t interact with agents, they aren’t aware that LLMs can have initiative and pursue goals?
I think the line is blurring though, mainstream chat interfaces are adding more and more “agentic” features.
ChatGPT will happily execute code in a sandbox, search the web and design downloadable PDFs purely through the standard OpenAI chat interface. They can also send you emails or do tasks on a repeated schedule.