You can follow the whole comparison without leaving this page, but the numbers themselves live at Phoronix. Nearly a year ago, Phoronix found a number of models where llama.cpp's Vulkan back-end beat the ROCm of that time on RDNA4. It has now rerun the comparison on a Radeon AI PRO R9700 and a Ryzen AI Max+ 395 (Strix Halo) system, which means the old answer may no longer hold.
The Phoronix ROCm vs. Vulkan Lemonade review has the actual results, and this article does not repeat them. Instead, it covers what separates the two back-ends in practice, so you can choose a default and confirm it on your own machine.
What the Phoronix test setup tells you
Phoronix used Lemonade 2026.39.1, which bundles llama.cpp b10825, and chose it for easy reproduction through its built-in benchmarking. The two test systems were a System76 Thelio Major with a Radeon AI PRO R9700 and a Framework Desktop with a Ryzen AI Max+ 395. Phoronix notes that the hardware is completely different and the software stacks are slightly different, so the two systems give distinct looks at Vulkan vs. ROCm.
The earlier RADV Vulkan vs. ROCm result shows that the gap moves with software versions. Any result, including a new one, has a shelf life.
Setup, hardware support and maintenance
The Vulkan back-end asks less of your system. It runs through the graphics driver stack you probably already have, typically Mesa's RADV on a desktop Linux install. You need a working Vulkan driver and a llama.cpp build with Vulkan enabled. No separate compute stack is required.
The ROCm back-end (llama.cpp's HIP path) adds a layer. You need a ROCm release that supports your GPU, matching user-space libraries, correct permissions for the compute device nodes, and a llama.cpp build compiled for your GPU architecture. When any of those drift out of alignment, the usual failure is a GPU that goes undetected or a silent fall back to CPU.
That is general background, not a claim about any specific ROCm release. The documentation for your ROCm version defines which GPUs it officially supports, so check it before you commit. Some GPUs can work outside the official list, but then you own the troubleshooting.
Maintenance differs too. With Vulkan, driver updates arrive with your distribution's Mesa. With ROCm, you track a separate release train. Lemonade bundles llama.cpp, which can spare you from compiling it yourself. Check its documentation for how it handles each back-end on your system. This pairs well with ai website builder vs ai app builder: how to choose explained.
Prompt processing vs token generation
llama.cpp reports two numbers, and they stress the GPU differently. Prompt processing (often shown as pp in llama-bench) chews through your input in large batches and leans on raw compute and matrix-math kernels. Token generation (tg) produces one token at a time and is usually limited by how fast the GPU can read the model weights from memory.
Because the bottlenecks differ, a back-end can lead in one and trail in the other. That is an inference from how the workloads behave, not a result from the Phoronix review. In practice, your workload decides which number to weight.
If you paste long documents, run retrieval pipelines or feed big codebases into context, prompt processing dominates your wait time. If you chat with short prompts and long answers, you feel generation speed. Measure both, and record memory use and stability alongside them. A back-end that is slightly faster but crashes on long contexts is not faster in any way you care about.
Discrete GPU vs unified memory
A discrete card like the Radeon AI PRO R9700 has its own VRAM. The model and its KV cache either fit there or they do not, and spilling into system RAM usually hurts. On such a card, ask how much context you can fit and how each back-end allocates memory for the same model.
Strix Halo works differently. The GPU shares system memory with the CPU, so the question shifts from "does it fit in VRAM" to "how much system RAM can the GPU address, and how does each driver path handle large allocations." GPU-addressable memory depends on your kernel, firmware settings and driver versions. The Phoronix review does not cover this, so check current guidance for your distribution.
Large models make the choice matter more here. A unified-memory machine may run models that would never fit on a typical discrete card. Test your biggest model on both back-ends before trusting results from a small one. Behavior at the memory limit can differ even when small-model speeds look similar. This pairs well with our guide to raspberry pi 5 alternatives after the $77.50 price hike.
How to run a fair A/B test
A fair test changes exactly one thing: the back-end. Everything else stays fixed. Lemonade's built-in benchmarking, which Phoronix used, is the easiest route. Check Lemonade's own documentation for the exact invocation in your version, since flags change between releases.
If you prefer llama.cpp directly, build two binaries from the same commit, one with Vulkan and one with HIP, then run llama-bench on each. The commands below are illustrative and were not run. Confirm flags with llama-bench --help on your build.
## Illustrative only: same model, same settings, different binary
./build-vulkan/bin/llama-bench -m model-Q4_K_M.gguf -p 512,4096 -n 128 -ngl 99 -r 5
./build-hip/bin/llama-bench -m model-Q4_K_M.gguf -p 512,4096 -n 128 -ngl 99 -r 5
Use this checklist so the comparison holds up:
- Pin versions. Record the llama.cpp commit or build number, the Lemonade version, the kernel, the Mesa version and the ROCm version.
- Match the model. Use the identical GGUF file and quantization on both runs, not two downloads of nominally the same model.
- Match the settings. Keep GPU layers, context size, batch sizes and flash-attention settings the same. Defaults can differ between back-ends, so set them explicitly.
- Warm up. Discard the first run, which includes kernel compilation and cache effects.
- Repeat. Run at least five repetitions and look at the spread, not only the average.
- Test both prompt lengths. Include a short and a long prompt, plus a long context if that is how you work.
- Record the extras. Note peak memory use, any errors or fallbacks, and whether outputs look sane.
- Control the machine. Close other GPU workloads and keep power and cooling conditions consistent.
Also read: pgo + lto explained: do compiler tweaks really help? — background
Troubleshooting and fallbacks
If ROCm does not detect your GPU, check the basics first. Confirm your GPU appears in rocminfo, that your user belongs to the groups that own the compute device nodes, and that your ROCm release lists your GPU as supported. If llama.cpp starts but runs at CPU speed, the HIP back-end likely failed to initialize, so read the startup log for the device list. When in doubt, switch to Vulkan and keep working while you debug.
If Vulkan picks the wrong device, such as an integrated GPU on a system with a discrete card, list devices with vulkaninfo --summary and read the device names llama.cpp prints at startup. llama.cpp offers device-selection options, and its Vulkan back-end has historically honored an environment variable for choosing devices. Check the documentation for your build, since the exact name can change.
If results look wrong on either back-end, such as garbled output or crashes at long context, change one variable at a time. Try a different quantization, a smaller context or an adjacent llama.cpp build before blaming the hardware.
Which back-end to start with
The table gives starting points only. Results vary by model, quantization and version, so treat it as a first guess and verify it with the test above.
| Your situation | Suggested starting back-end | Why | | --- | --- | --- | | Easiest setup on any Radeon | Vulkan | Uses the existing graphics driver stack, with no separate compute install | | Discrete GPU officially supported by your ROCm release | Test both | Support exists, so measure instead of assuming | | GPU outside ROCm's supported list | Vulkan | Avoids unsupported-GPU troubleshooting | | Long prompts, RAG, big contexts | Test both, weight prompt processing | Prompt speed can differ from generation speed | | Short prompts, long chat replies | Test both, weight token generation | Generation is mostly memory-bound | | Strix Halo with very large models | Test both with your largest model | Unified-memory allocation behavior can differ |
If you want a working setup today and your GPU is unsupported, start with Vulkan. If you run a ROCm-supported discrete card or a Strix Halo machine and want every last bit of speed, install both and let your own A/B test decide. Then compare your results with the R9700 and Strix Halo numbers in the Phoronix Lemonade ROCm vs. Vulkan review, and rerun your test whenever you update llama.cpp, Mesa or ROCm.



