logoalt Hacker News

Tepixtoday at 10:27 AM8 repliesview on HN

Keep a close eye on abliterated and "heretic" open weight models. They will be outlawed first.


Replies

roenxitoday at 10:56 AM

It is not feasible. They never made much of an inroad against torrents and that is a much easier target than abliterated models. As the linked website shows; the process to abliterate a model can be as simple as

pip install -U heretic-llm && heretic Qwen/Qwen3.5-4B

let alone people just putting the weights up in a torrent. All assuming that someone even tried to ban abliterated models.

show 1 reply
totetsutoday at 9:29 PM

At the moment is there any legislature that is seriously pressing regulation to as you say outlaw .. or ban outright AI models that do not have guard rails built in? i know there is a lot of moves about this for things used in critical infrastructure.. but i thought no one is really saying ban these things completely.

Ajedi32today at 2:40 PM

They picked a good name for fighting that. The optics of trying to outlaw heresy probably aren't great. ;)

show 1 reply
apitoday at 12:10 PM

This is the test. If the speech that's easiest to dislike is legal, then we all have free speech.

IMO math is free speech, and outlawing math is censorship.

petratoday at 11:56 AM

Like they've outlawed drugs? Illegal weapons? Hacking?

thih9today at 10:42 AM

I'm not sure what is your point. It reads as defeatism to me but I'm not sure.

Could you elaborate? Do you find it good or bad? What actions can be taken?

show 1 reply
ben_wtoday at 11:00 AM

Good.

If you think closed source software/binaries only is bad, wait until you see how awful the state of the art is with a clear-as-mud bucket of matrix weights.

We know it's possible to train an LLM to secretly respond to certain trigger phrases, and last I checked these could only be detected with the assistance of whoever chose those phrases.

The trigger condition for such backdoors is not something anyone can do a systematic brute-force check for, for the same reason we had to invent LLMs in order to do natural language processing: combinatorial explosion.

Passing around open weight models from known sources is already asking you to trust those sources; because of how difficult this is to do correctly even without deliberately inserting such things, we still don't know if China has already put such trigger conditions into their models despite headlines such as these: https://venturebeat.com/security/deepseek-injects-50-more-se...

Regardless of if it was deliberate or not, we don't know if we caught all of these misbehaviours. We don't know how to.

And note, I'm not saying "and therefore you should trust the Big Name Models". If open weight models score 2/100 in this context, closed ones score 1/100.

show 1 reply
luxpirtoday at 11:36 AM

Agree. I took a look at these last few months, did a write-up: https://languageops.com/blog/ai-safety-pdoom-local-vs-fronti... and I don't know if I agree or not on outlawing completely, but I think an age restriction *at least* like for alcohol, firearms and driving would be not unwise.

show 2 replies