logoalt Hacker News

Show HN: Shoehorn – Quantize any model down to run on your machine

67 pointsby rhgraysoniilast Tuesday at 2:29 PM17 commentsview on HN

Working on Mac, Linux, and Windows now. I include a simple GUI to find new models and get things built and set up. It is working quite well across a few models for me. The GitHub README and DESIGN.md files go into detail of the how/why and it's working remarkably well so far. https://github.com/notactuallytreyanastasio/shoehorn


Comments

puttycattoday at 9:14 PM

This is really impressive. Can you say a bit about the underlying process? I'm guessing this is post-training qantization? Isn't PTQ also resource-intensive? (Ie might not work on any machine)

sscarduziotoday at 8:27 PM

The project name is perfect!

hmokiguesslast Tuesday at 3:49 PM

Reminds me of https://github.com/AlexsJones/llmfit

show 1 reply
jedbrooketoday at 7:02 PM

I gotta laugh at some of the models it suggests, for example:

> AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF

you’re telling me you managed to fit Fable 5 into just 4B?

show 1 reply
akshay_akulalast Wednesday at 2:05 PM

This is interesting. I wonder how it could work with something like https://github.com/JustVugg/colibri.

mbuchel-hnlast Tuesday at 3:19 PM

does this work similar to airllm? i am wondering how it would handle something like quantizing kimi k3 on a budget of 8 gbs, or is that something you are not attempting to solve yet?

show 1 reply
jaylanelast Tuesday at 6:32 PM

tried it out but based on the model sizing result i got i got an insufficient memory error when the server started running

show 1 reply
kelvo_ranyesterday at 1:41 AM

[dead]