logoalt Hacker News

yogthosyesterday at 3:36 PM1 replyview on HN

The real question is whether it's easier to improve the software side instead. There are likely a lot more optimizations possible in terms of model architecture, and if there is a compute bottleneck, then it's going to put a lot of pressure on Chinese labs to address the problem using more efficient designs.


Replies

tedd4uyesterday at 5:51 PM

And one cool trick is that if one develops step function more efficient training, one can still “open source” the model without revealing the training techniques.

show 1 reply