alt
Hacker News
slim
•
today at 12:58 AM
•
0 replies
•
view on HN
llama.cpp can run MoE with some layers in vram and some layers in ram