logoalt Hacker News

dfdydxyesterday at 8:37 PM1 replyview on HN

There are two different things:

- was item X in the training data

- did the inclusion of X in the training data lead to Y

I understand why the second is hard, but why is the first one hard?


Replies

keedayesterday at 9:32 PM

Yep, the last part in my post was suggesting some ways we could determine if "item X was in the training data" (as well as some potential blockers for that from OpenAI's perspective.)