logoalt Hacker News

spike021 • today at 12:35 AM • 11 replies • view on HN

I think whichever one is used, there needs to be a way to enforce what's written.

If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.


Replies

zahrevsky • today at 12:49 AM

There is a command in oh-my-pi called "/omfg <problem>". You explain what is wrong with agent's response, and it writes a hook to make sure that the problem doesn't happen again. It then re-runs your previous prompt to make sure that hook is triggered, and if not, it rewrites the hook to make your previous prompt trigger the hook. Then each next agent's response is checked by the hook, and if it is triggered, the agent receives feedback on what's wrong and what must be done differently.

ozim • today at 10:27 AM

Claude code setup seems to be using jq all the time when I use it but I did what is written here:

https://gist.github.com/cynthiateeters/6868ca26c059a3106cd93...

jkhdigital • today at 2:18 AM

Treat it like any other software system: rules that must not be violated are enforced by static type-checking or a trusted runtime monitor. There’s no other option.

➕ show 1 reply
fennect • today at 6:02 PM

You could try https://github.com/ioni-dev/mati the constant changes in reasoning effort in consumer models can break things and its more dangerous for devs that relay to much on agents.

I’m the author of the project.

ACCount39 • today at 12:44 AM

Obviously, AI is way more comfortable with using adhoc Python scripts, which are used for everything, than it is with using jq, a niche CLI tool.

dboreham • today at 1:11 AM

I think it's kind of cute the way it writes Python scripts, but I've never seen it do that when the relevant native tool is on the PATH. It's like the most competent ever intern, on speed. No tool to convert SVG to PNG? No problem, I'll write a Python program to do that!

locknitpicker • today at 7:35 AM

> If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json.

I think there is a deeper problem emerging from this sort of behavior. Even when we bother to create agent skills with there own scripts that call tools like jq a specific way to achieve a goal, AI coding assistants and agents still go way out of their way to generate ad-hoc scripts to do the most absurdly stupid tasks such as parsing output in structured language, and even remove whitespaces from a markdown file. This means AI coding assistants and coding agents treat agent skills as mere suggestions of using a alternative option that more often than not the choose to ignore.

This has a very dangerous implication: your average user is trained to develop a pavlovian reflex to authorize agents to just execute their ad-hoc scripting code with our own permissions and credentials in our systems, which includes the ability to call anything over the internet.

➕ show 1 reply
triyambakam • today at 6:04 PM

It's often better to just find ways to embrace what it tries to do naturally. Otherwise you're fighting the weights and hidden prompts

eivindmeyer • today at 6:43 PM

[flagged]

HisashiSpace • today at 4:13 PM

[flagged]

sick_of_slop • today at 4:01 AM

[dead]