logoalt Hacker News

patrick0dyesterday at 12:09 PM1 replyview on HN

Inspired by parameter golf and speedrun approaches I make the case for picking loss functions like a wallclock for LoRA on AI safety targets. The result when I tried it was a functional distillation of an Sparse AutoEncoder into a 5.3MB probe. I have a technical writeup below about it if anyone is interested.

https://www.lesswrong.com/posts/PagGF8roBJmjLunsX/competitiv...


Replies

Vineeth147today at 12:12 AM

[dead]