logoalt Hacker News

FuckButtonsyesterday at 7:07 PM1 replyview on HN

Where does that rule of thumb come from?


Replies

wongarsuyesterday at 9:37 PM

That's a great question. I learned it on HN. Some searching around suggests it originated as an empirical observation in the local LLM space around 2023-2024

It's obviously just a rough approximation. Actual scaling laws suggested in published papers are a lot more complex, and even then you run into issues (architecture changes, effects like better training, putting intelligence on a one-dimensional axis is stupid in the first place, etc). But as an approximation it holds up pretty well for normal-ish ratios between active and total parameters