logoalt Hacker News

delicious_appleyesterday at 5:25 PM1 replyview on HN

I am running it on a single RTX 3090 (24GB VRAM).

Some folks on Reddit are having the same experience: https://www.reddit.com/r/LocalLLaMA/comments/1vkm42m/muse_gl...

It uses an order of magnitude less VRAM at longer contexts which is a huge advantage over Qwen 3.6 27B


Replies

r0b05today at 1:56 AM

Seems like that's the tradeoff with this model. Close to 27b intelligence while using less vram.