I have so much to share on this topic. Will keep it short.
ROCm promises a 30-50% prompt processing speedup. This is REALLY important for my workflow so I've been trying to get this shit to work for months. But no release before v10 worked well enough with any engine for it to matter.
The llama.cpp release binaries for ROCm (10) FINALLY work on gfx1501 and its relatives (with the correct shell variables), but the prompt processing boost doesn't materialize and the token generation speed decreases.
There continues to be a chronic problem across all engines with the ROCm integration for UMA devices. The good news is that some improvements have been made to that end for Vulkan, so more recent llama.cpp Vulkan binaries are now faster.