logoalt Hacker News

minimaltomtoday at 4:10 PM1 replyview on HN

Architecture thread! Afaict they continue to use gated attention + delta net, which was also adopted+adapted by K3, but im surprised theres no improvements to the residual stream (deepseek are using manifold hyper-connections, kimi have attention residuals) ?

Perf improvements seem to all come from training?


Replies

anana_today at 4:19 PM

As was the case with GLM 5.3, it seems that there is still much juice to be squeezed from post-training