logoalt Hacker News

Kimi K3 Architecture Overview and Notes

213 pointsby ModelForgetoday at 3:48 PM23 commentsview on HN

Comments

thatsgcaseytoday at 7:48 PM

Sabastian Raschka is one of the great LLM researchers/authors. I highly recommend his substack

show 2 replies
Ilaurenstoday at 8:49 PM

"Interestingly, Kimi K3 got rid of all RoPE layers and uses NoPE (No Positional Embeddings) everywhere instead."

It just baffles me that this even works at all. Doesn't it just become a token soup? Is attention that precise that a second token can tell its the second token just because it learns to accumulate something in the embedding space without any sort of inductive bias?

show 2 replies
souravsspacetoday at 8:00 PM

i like your detailed breakdown. Thanks. <3

alealvarezargtoday at 7:28 PM

Great breakdown. After using Kimi extensively, it's fascinating to see how architectural choices like KDA and NoPE translate into such strong real-world performance. Really impressive engineering.

myronkeir1968today at 8:56 PM

[dead]

myronkeir1968today at 8:55 PM

[flagged]

pullruntoday at 8:55 PM

[dead]

gokohltoday at 4:56 PM

Interesting that they went NoPE everywhere — everyone else hedges with RoPE in the local layers. Feels like the linear-attention stuff (Kimi Delta) is quietly doing the positional work so they can get away with it. Curious to see if it holds up at frontier scale.

show 4 replies