logoalt Hacker News

raesene9yesterday at 9:24 AM1 replyview on HN

Interesting write-up and I do think LLM assisted/powered exploit disclosure is a real concern (I've been able to get models to create container breakouts from Linux LPEs relatively quickly).

One thing I'm surprised about is that GPT-5.6 didn't block that prompt due to guardrails. My experience is that GPT-5.5 and up does not like offensive security work (similar to Opus 4.7+/Fable).

I didn't notice it but I'd assume that the authors have some level of cyber approvals from OpenAI to relax the guardrails a bit.


Replies

Santasyesterday at 9:56 AM

This might help https://chatgpt.com/cyber ease the guardrails a bit.

show 1 reply