logoalt Hacker News

redox99today at 5:27 PM0 repliesview on HN

Other than attention optimizations and other minor changes, the top Chinese models (which are way better than gemini) have basically the same architecture as GPT2. Of course RL is key for agentic workloads, but I'd say it's correct that progress has been mostly scaling models,adding more data and cleaning it better.