logoalt Hacker News

om8yesterday at 3:01 PM0 repliesview on HN

This project needs webgpu -- I did it on cpu about a year ago.

My demo uses 2 bit quantization to run llama3 models on any device with enough ram.

https://galqiwi.github.io/aqlm-rs/