logoalt Hacker News

cpfohltoday at 12:55 PM1 replyview on HN

I’m still slightly confused on what this adds.

Let’s say I wanted to run a full size open weight model. I have a 128GB m3 max laptop.

Does this basically load layers in and out on demand? So I still have to download the full model to disk, but the RAM requirements go way down? The readme calls out that one still needs to connect HuggingFace, which leads me to believe that maybe you don’t even need to download the full model?


Replies

dofmtoday at 1:11 PM

If you point it at a huggingface model identifier, it will download it, I assume. No way around that.

It reads like it is keeping only the core and the active layer loaded at any one point, and streams layers from disk; there are several other solutions like this and if my understanding is right, this is probably better than an mmap implementation or just streaming experts in.

show 2 replies