I can see that and I don't know your setup, but there are people pushing >70t/s with MT...

pferdone • today at 11:31 AM • 2 replies • view on HN

I can see that and I don't know your setup, but there are people pushing >70t/s with MTP on a single 3090, with big contexts still >50t/s. 64k is not a lot for agentic coding, and IIRC 128k with turboquant and the likes should be possible for you. r/LocalLLM/ and r/LocalLLaMA/ are worth a visit IMO.

EDIT: just found this recipe repo, may wanna give it a go: https://github.com/noonghunna/club-3090

EDIT-2: this can also shave off a lot of context need for tool calling -> https://github.com/rtk-ai/rtk

Replies

gchamonlive • today at 1:50 PM

I managed to execute with vllm successfully, but it breaks opencode on simple "what's this repo?" task. On oh-my-pi it wont event execute because omp sends multiple system prompts. I'll try with llama.cpp later and see if it works more reliably.

gchamonlive • today at 12:04 PM

will give more info in the post

EDIT: thanks for the links!

alt Hacker News

Replies