Running Q3 on my AMD RX 9070XT. 32k context and 32/TPS. Apart from the context window preventing it from doing any large tasks, this thing is seriously powerful. I could probably push it to 64k context. Local open models are the future, and I am definitely getting a more powerful card. Very fun!
Could you comment more on how you set this up? I have a mostly idle 9070XT I use for gaming, and I was considering using it with the newer local open models. Many thanks.
What are you offloading to ram (or even CPU)? I’m using a 9080 (not XT) and having trouble with context/token rates
Running Q3 on 5060ti with 64k context. It runs great
9070XT operator here: I'm using llama.cpp with the same model and quant and I'm getting 87,000 for my context limit. I tried the Unsloth models but they lowered it to around 30-40K so I went back to upstream.
I'm on Linux and using some sort of unholy mess of ROCM libraries that I don't understand.