logoalt Hacker News

applfanboysbgontoday at 3:53 AM0 repliesview on HN

> It has the ability to rapidly experiment and potentially succeed at tasks through trial-and-error, but not without constantly corralling it in the correct direction

It kind of could, given the above. Part of an LLM's advantage is that, much like a calculator or a Chess engine, it can iterate over a finite problem space far, far faster than a human can. That much is expected of a useful computing tool.

It is worth noting that we know literally nothing about how the counterexample was achieved. Technically speaking, the person who tweeted it could have solved it themselves with zero LLM assistance and then attributed it to Fable to boost their IPO and ensuing payday. I'm not saying that's actually what happened, but it's hard to draw conclusions without any transparency about the degree of human involvement.

My daily experience certainly does not reflect that of prompting a superhuman intelligence when it routinely flubs commands and destructively drops the PATH of its vm, or bypasses an instruction about passing tests by burning millions of tokens constructing a completely new test suite that rubberstamps its own work when it can't pass the real tests.