logoalt Hacker News

OscarCunninghamyesterday at 6:56 AM5 repliesview on HN

They're trained on human data. I would expect them to emulate human biases as closely as possible.


Replies

avianlyricyesterday at 10:13 AM

This is where harness, and the fact that a machine can be endlessly prompted to try again comes in.

Even if an LLM starts by pursuing things that follow human bias, continuous failures and re-prompting to try something different will eventually force it to consider things outside of what ever biases it has.

You can do the same thing to a human. But most people would consider it unethical to lock someone in a box and force them to keep trying to solve the same problem over and over again until they figure it out.

2b3a51yesterday at 9:47 AM

Your comment stopped me in my tracks a little bit.

Is a 'bias' in a piece of writing generally a property of word to word choice and sentence to sentence construction or is it something more nebulous? Especially in terms of the appreciation of mathematics and someone's hesitance about publishing a mathematical argument they think is ugly or brute forced in some way.

show 1 reply
SiempreViernesyesterday at 7:58 AM

Is it? I'd expect most of the training set to be synthetic data extrapolated from a small set of human authored texts.

show 1 reply
epolanskiyesterday at 11:50 AM

It's more complex than that, especially as post training is often goal based.

show 1 reply
catigulayesterday at 2:53 PM

No, they are not. Specifically, this was a method based specifically on learning from scratch, like most modern AI models.

Why do you think it's called Alpha ZERO?

show 2 replies