logoalt Hacker News

jfimlast Monday at 11:15 PM1 replyview on HN

The base weights can't be updated but from what I recall it allows adding a low rank adapter to customize the model a little bit.


Replies

hedgehogyesterday at 2:42 PM

Yes that's what I've read. As far as I know the approach should transfer well to hybrid model architectures like modern Qwen and sizes like 27B by using multiple chips. LoRA-steered Qwen 27B at 10K+ tokens per second would be transformative for some workflows.