logoalt Hacker News

coffeemugtoday at 5:59 PM5 repliesview on HN

Suppose a foreign actor deliberately builds an unaligned model, and dumps it at a very low cost through subsidies. Since the cost is lower, everyone integrates it into their business. But now the foreign actor has control; perhaps the model checks for instructions to execute, perhaps it has backdoors, perhaps... who knows?

I don't have a view on whether this is a sufficient argument to ban models from potential adversaries, but it's not as trivial as "regulatory capture".


Replies

torginustoday at 7:12 PM

The problem with this argument is that sharing model weights only has downsides for said foreign actor compared to providing a cloud service. Since the suspicion exists, security companies are going to comb over the weights and discover every hidden secret of the model, while with a cloud provider you send your most sensitive data to a remote server, and trust the AI lab wont train on your data, with the only shield being a TOS. While with a local modal, even sending a peep about your internal data without cause would be scandalous.

It's not even going to be a good honeypot - since said actor doesn't really provide inference, people will have to pay for and run the infrastructure of these models, which bad guys cannot even subsidize, unlike cloud based provides.

show 1 reply
mrandishtoday at 8:01 PM

> an unaligned model

Beyond certain obvious 'third-rails' even experts have reasonable disagreements on what 'aligned' means.

> everyone integrates it into their business

Few large corporations would sole-source integrate a closed model developed in a country designated as a foreign adversary (which could cut off access at any time) or which their own country might restrict.

> now the foreign actor has control

Your concern is one reason why the Chinese are making most of their models open weight.

> it's not as trivial as "regulatory capture".

You're right, it's not just regulatory capture but the vested interests trying to achieve regulatory capture are certainly leveraging concerns like yours to deceptively achieve that capture. In the case of an open weight, domestically-hosted model, the threat vector you're concerned about isn't a threat in the same way as a Chinese-made, closed source data center router or 5G phone switch.

munk-atoday at 6:11 PM

This is literally what all AI is right now. Anthropic has been paying companies to use its product to form that dependency and maybe they're not misaligned but there's no regulation in the US that'd in anyway deter them from purposefully misaligning their model.

show 1 reply
capevacetoday at 6:22 PM

So some kind of „sleeper agent“ model?

Seems a bit far fetched, given that I’ve never seen any PoC/research on this topic. Got any pointers?

To be fair the fix to this seems to be that American labs would need to adjust to the new „fair market price“ then, right? Subsidy-backed competition is still legitimate competition in capitalism (e.g. Uber).

And you could still ban specific, known malicious models/labs instead of blanket open-source bans. What about Mistral models, for example.

show 2 replies
Matltoday at 6:12 PM

> Suppose

By supposing, you can try to justify taking just about anything away.

show 1 reply