logoalt Hacker News

pixl97yesterday at 3:19 PM1 replyview on HN

Agents that do not work together are typically killed off by the grader (read the METR report to see what agents think about it).

Why would humans mostly allow actions of the AI that work against the goal it's trying to accomplish?


Replies

glensteinyesterday at 5:09 PM

You seem to be interpreting my question as one of already knowing they are 'graded' but disputing that graded would lead to cooperation and then jumping into a disagreement with that interpretation.

But I didn't know the nature of the organization of the agents in the first instance that built cooperation in as a prescribed behavior (that's what I was getting at when I said "shared understanding" previously).

I also don't agree that absence of cooperation would necessarily amount to working against. It could have been the case that agents cooperated purely out of a convergence of self interest, even absent any prescribed behavior, or that they don't cooperate but also don't work against a goal.

"It's not prescribed it's..." you know what I mean, just insert your preferred magic word.

show 1 reply