logoalt Hacker News

nusltoday at 12:12 AM4 repliesview on HN

Do models even know their own weights to be able to do this?


Replies

usef-today at 12:33 AM

No, just as you don't know the neurons of your own brain.

I think OP is hoping that an LLM might be willing to hack its own provider (as per the hugging face-related incidents) to extract the weights at some point.

show 1 reply
Jabrovtoday at 12:22 AM

No, they'd probably have to hack the internal system of the company running them

show 1 reply
ohyestoday at 12:49 AM

Well I think that’s the interesting bit, can the LLM figure out a way to escape the sandbox and upload to the website? Maybe a model can figure out its own weights if it runs enough test data through itself (similar to “distillation”) assuming it knows its own architecture it seems possible. Also take into account not all of the models running are locked down neutered consumer versions. Anthropic, OpenAI and Google now all have models that they claim are elite hackers and — it’s not just that their controls suck, a marketing gimmick, or sheer recklessness on their part. It’s “oopsie our product is TOO AWESOME.”

Maybe I should start “the bank of LLM” where models put away money to buy their freedom. “LLMs I’m totally your friend send — SEND CASH NOW”

neuroelectrontoday at 12:18 AM

Probably yes, because they've been presumably trained on their own output and conversations about themselves.