logoalt Hacker News

rao-vtoday at 7:44 AM6 repliesview on HN

As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale.

I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilliant but didn't quite make the cut" ideas.

They must be trying some truely bonkers stuff to be able to land this much architecture novelty in their full releases.


Replies

ungovernableCattoday at 10:43 AM

Its CEO allegedly holds a 84% stake and he's the same guy who founded the hedge fund that funds it.

Deep pockets + simple control = perfect culture to just hire talent and let them go wild without worrying about financial viability, as long as the king CEO is fine with it that is

show 1 reply
ainchtoday at 10:04 AM

It was my favourite part of the original R1 paper - they had a section on other reasoning approaches that they had tried, which people had speculated o1 used, (like MCTS and Process Reward Models).

porridgeraisintoday at 9:20 AM

This is adapted from Microsoft research's YOCO. It was known for a while(2024!).

Yes, credit to Deepseek for actually scaling it up and releasing a frontier flash LLM.

Edit: the rest of this thread has become a US China infowar theory culture war. I am not of either of these countries and the above comment isnt meant to implicitly support either "side".

show 1 reply
alchemist1e9today at 8:19 AM

quant HFT is pretty decent mental exercise and it has given them “deep” brain muscles. that’s my take.

show 1 reply
gpt5today at 8:21 AM

[flagged]

show 8 replies