logoalt Hacker News

frumplestlatz • yesterday at 6:22 PM • 1 reply • view on HN

If I perform inference with an LLM, it has the capacity for independent choice and action, and will use it.

The gun does not. No matter what, I have to choose to pick it up, aim it at someone, and pull the trigger.

Where exactly is the flaw in the analogy?


Replies

Topfi • yesterday at 6:45 PM

If you don't interact with/talk to/prompt an LLM, it won't do anything. So this comes down to the model doing something different than what was prompted. What you consider "the capacity for independent choice and action". What I consider, a defect, to be excised.

In my model evals, if a model deviates from the provided prompt (task adherence), that’s a failure even if the primary goal might have been achieved in a different manner.

To go with the analogy, task deviation should be treated the same as a gun that due to manufacturing defects can fire despite the safety being on. That defect remains, even if you can use the gun to shoot (in an unsafe manner).

Simply, neither should happen and both models deviating from their prompt or guns firing by themselves are to be considered a fatal flaw. It's why, despite greatly lauding the GPT-5 series, which did adhere to prompts in most every scenario, I have ranked every OpenAI model post Spud very poorly as those traded task adherence for brute forced, deviated approaches to solutions and why HF, Medicare, etc. were inevitable with their current trajectory.

A model that as part of normal, well scoped use proceeds by taking independent action or, far worse, makes choices beyond the original prompt, is not something I feel should be used. GPT-5.6 Sol and GPT-6 Astra both do this on the regular when trying to safe an ancient, utterly messed up git tree with branches upon branches that I maintain for eval purposes, thus leading to data loss that if the models adhered to the prompt as written, wouldn't happen (though the task will take three times more steps). Something GLM-5.3 Flash, prior OpenAI models including original GPT-5, anything from Anthropic in the recent years, etc. do not fail at.

Nothing happens without an initial prompt, deviating from it is a severe flaw and should lead to a model not being considered for deployment or wide use.

➕ show 1 reply