This project needs webgpu -- I did it on cpu about a year ago.
My demo uses 2 bit quantization to run llama3 models on any device with enough ram.
https://galqiwi.github.io/aqlm-rs/