> I was trying to use rocm with llama.cpp
completely offtopic but is rolling with rocm worth it? I spend a fair bit monthly on rental gpus for projects and going to upgrade at home instead, AMD has some solid winners here pricewise but get conflicting reports about using it for ML in 2026.
once upon a time it seemed unthinkable to use anything but nvidia but seems to have come a long way since I last looked, probably would be just pytorch and gemma 31B
I get the feeling the situation is only going to improve longer term so might be a good time to just do it
Support has gotten much better in just the last couple months. I just got a 9070 XT and can't count the number of times I've installed a package and the changelog made me think how much it would have sucked to be doing this a year ago.
> completely offtopic but is rolling with rocm worth it?
It's so fucking easy.
From AMD: https://lemonade-server.ai/
Then you can easily throw a openweb-ui container in front, and then connect to the openweb-ui via your mobile app of choice (if you want chat, otherwise you just point your harness of choice at the lemonade server api endpoint).
I have so much to share on this topic. Will keep it short.
ROCm promises a 30-50% prompt processing speedup. This is REALLY important for my workflow so I've been trying to get this shit to work for months. But no release before v10 worked well enough with any engine for it to matter.
The llama.cpp release binaries for ROCm (10) FINALLY work on gfx1501 and its relatives (with the correct shell variables), but the prompt processing boost doesn't materialize and the token generation speed decreases.
There continues to be a chronic problem across all engines with the ROCm integration for UMA devices. The good news is that some improvements have been made to that end for Vulkan, so more recent llama.cpp Vulkan binaries are now faster.