logoalt Hacker News

npodbielskitoday at 8:35 PM1 replyview on HN

And it fails on rocm of course. This engine is such a hassle on AMD.


Replies

trouve_searchtoday at 10:08 PM

That's AMD's fault.

RDNA4 is pretty similar to CDNA4, yet over a year after the release of "pro AI" cards like the r9700, they had basic kernels lacking in vllm (like w4a16 int4 kernels) while they were implemented in the datacenter CDNA4 cards.

AMD hardware runs well on llama.cpp because basically anything runs on llama.cpp, especially with vulkan. It's not high praise of AMD's software team to say llama.cpp runs well on their hardware