As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale.
I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilliant but didn't quite make the cut" ideas.
They must be trying some truely bonkers stuff to be able to land this much architecture novelty in their full releases.
It was my favourite part of the original R1 paper - they had a section on other reasoning approaches that they had tried, which people had speculated o1 used, (like MCTS and Process Reward Models).
This is adapted from Microsoft research's YOCO. It was known for a while(2024!).
Yes, credit to Deepseek for actually scaling it up and releasing a frontier flash LLM.
Edit: the rest of this thread has become a US China infowar theory culture war. I am not of either of these countries and the above comment isnt meant to implicitly support either "side".
quant HFT is pretty decent mental exercise and it has given them “deep” brain muscles. that’s my take.
Its CEO allegedly holds a 84% stake and he's the same guy who founded the hedge fund that funds it.
Deep pockets + simple control = perfect culture to just hire talent and let them go wild without worrying about financial viability, as long as the king CEO is fine with it that is