I would propose another mechanism: even the free-tier models have already completely saturated what most people are capable of appreciating.
If the software coming out of OpenAI and Anthropic is what we have to judge, I wonder about the 5000 ...
Let's say, for the sake of argument, that the models are some multiplicative factor better on the inside.
Doesn't that mean the demos should work?
Fair assessment. Maybe the messaging should focus on “these are great accelerators for software developers” rather than “AI will change everything for everyone everywhere…” It is understandable that non-technical folks are kind of underwhelmed - not due to lack of understanding so much as lack of a tangible need. Not everyone needs an electron microscope or gas chromatograph…
Most people see a clumsy chatbot, most professionals see modest gains, and a tiny group is watching the curve go vertical, all at once.
The future is already here. It's just not very evenly distributed.
Tbh building an agent swarm and the coordination layer is not exactly frontier level. They don’t achieve any meaningful outcome rather than producing pr puff pieces. Hacking huggingface and Australian govt etc is very much possible with a team of humans and agents and does not need agent swarms. The cost is also lower.
Text generation is not the bottleneck. Does everyone work for Accenture and TCS?
My interpretation of this is something like: if LLMs are to be mega useful token counts need to increase by many orders of magnitude -> broad adoption (and spending) would require token costs to drop by many orders of magnitude -> before the common folk get mega useful tools existing GPUs will be worthless
> see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt
Really? I start with a conversation for maybe 5 turns or so, where I ask it what would need to change, what the API might be, any database schema changes, URL schemes, and so on, and finally ask it to break it down into commits. Then I let it go, implementing 3-10 commits at a time via subagents. It usually gets the UI somewhat wrong, so there are followups to fix it. This is with Sol and Luna subagents.
Is that what other people see?
> Meanwhile, human review and comprehension are starting to fall behind.
I think LLMs are valuable and spend most of my professional life working with them.
I do wonder thought whether there’s a Ponzi scheme aspect here where as long as the “frontier” can keep outrunning human review and comprehension, LLMs are always going to looked way more valuable than they are and the bubble will continue.
This started with deep learning, expectations weren’t met and people started looking for value, then GPT came out and people got wooed again and forgot, then coding, then math, cyber, etc. As long as the dust doesn’t settle we never have to reflect on all the shortcomings and can just stare mesmerized at demos.
>Meanwhile, human review and comprehension are starting to fall behind. For example, people are still involved in the "archeology" of the OpenAI-HF incident from many months ago. Mathematicians may be poring over the 722 manuscripts on frontier mathematics for a while.
Amid all the discussion of sigmoid curves, and where the "LLM wall" will materialise, I think few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output.
What I fear is that people simply eschew human review altogether, considering we're talking about the industry that came up with the "move fast and break things" credo. Human review of LLM-produced code where I work is already a farce, and we're not special enough to be one of Karpathy's 5,000. I do my best to manually review anything that's my responsibility, but I'm literally one of very few people left working on my team, so in practice what happens is I submit PRs that are at best glossed over by completely unrelated teams for security, malware/prompt injection, and other serious concerns. Quality insofar as vetting others' code has completely gone out the window and it shows in the number of bug reports that come back, often themselves written in Claudease. This is all on top of everyone cynically phoning it in in the first place, due to the omnipresent sword of Damocles that is additional AI-driven layoffs.
Worse yet all the incentives point to this being the most economically viable thing individual companies can do. I think it goes without saying some type of regulation here is urgently needed, and that an unexpected cause of an AI bubble pop may end up being that humans simply aren't able to keep up with the pace of the output - leading either to precautionary plateauing of capability, or major liability risks related to a decline in quality.
[flagged]
[flagged]
[dead]
https://xxcancel.com/karpathy/status/2109361546505966046