logoalt Hacker News

NitpickLawyeryesterday at 9:02 PM0 repliesview on HN

> My concern is what a misaligned model will do when they’re even more competent.

I think the alignment talk is a red herring. It won't matter in the end, because there will be (if there aren't already) efforts to train offensive models without any guardrails whatsoever. And RL has another advantage: you can reward for whatever you need, and get different results. Right now they're training for general capabilities, but in the future I could see models trained for stealth intrusion and ensuring access, or for all out "milspec" penetrate, replicate and disable, or anything in between.