A Zen 6 "IBS Memory Profiler" headline sounds like a niche CPU footnote, but it points at a hard kernel problem. On a machine with fast and slow memory, Linux has to guess which pages deserve the fast spot. At LPC 2026 in Prague, an AMD engineer pitched new Zen 6 hardware as a way to guess better. The questions below run in order, with a clear line between what is confirmed and what is not.
What is the news, and what is actually confirmed?
Phoronix reports that an AMD engineer talked up a new IBS Memory Profiler feature found on Zen 6 at LPC 2026 in Prague. The talk was framed around hot page detection and promotion on Linux (Phoronix). According to that report, the feature also appears in a public AMD technical document.
That is all that is established. Register layouts, sampling modes, overhead, kernel patch status and any performance figures from the talk remain unconfirmed, so this explainer states none of them. Read what follows as background for the talk and slides, not a summary of them.
What is hot page detection, and why does tiered memory need it?
A "hot" page is one the workload touches often. A "cold" page sits idle. On a single-tier machine the distinction matters little. On a tiered system it decides performance: local DRAM is faster than remote-socket DRAM, and both are usually faster than CXL-attached memory.
Linux presents slower memory as separate NUMA nodes, so the kernel can place pages on any of them. Promotion moves a hot page from a slow node to a fast one. Demotion moves cold pages the other way, which frees fast memory instead of forcing the system to reclaim or swap. Both depend on one input: an accurate answer to "which pages are hot right now?"
That answer is hard to get. A large server holds an enormous number of 4 KiB pages, and the CPU does not report each load to the kernel. Every detection method therefore trades cost against fidelity.
How does NUMA balancing find hot pages, and where does it fall short?
Automatic NUMA balancing works by sabotage. The kernel periodically marks slices of a process's address space as inaccessible. The next touch triggers a minor "hinting" fault, and the fault handler learns which CPU and node accessed the page. Over time the kernel can migrate pages toward the tasks using them, or tasks toward their memory.
The tunable is described in the sysctl documentation (kernel sysctl docs). A memory-tiering mode also exists for promoting pages from slow nodes.
The limits follow from the mechanism. Faults cost CPU time and can stall the application at an unlucky moment. The scan sweeps address ranges, so recency is coarse. A page touched once after a scan looks much like one touched a million times, because the fault fires only on the first access. The kernel also rate-limits migration to avoid thrashing, which slows its reaction to phase changes.
How does DAMON differ?
DAMON (Data Access MONitor) samples the hardware "accessed" bit in page tables instead of forcing faults. It groups memory into regions and adapts their size, so monitoring overhead stays bounded however much RAM the system has. The kernel documentation describes its configuration and interfaces (DAMON admin guide). DAMON-based schemes can then act on what the monitor observes.
The trade-off is resolution. An accessed bit is binary per sampling interval, so DAMON sees "touched or not", not how many times or from which CPU. Region-based aggregation can also blur a hot page sitting next to cold neighbors. The migration actions available depend on your kernel version, so read the documentation that matches the kernel you run. For more on this, see read about how to check your amd p-state driver mode on linux.
What is AMD IBS, and what might a memory profiler add?
Instruction-Based Sampling (IBS) is an AMD hardware facility that tags instructions as they flow through the pipeline and records details about the tagged ones. The kernel exposes it through perf, and on supported systems perf list shows IBS-related events. Unlike page-table tricks, the hardware reports real executed operations. For loads and stores, that can include the address touched.
That is why the idea suits tiering. A sample says "this CPU just accessed this address" with no fault and no scan. Sampling also scales with the sample rate, not with memory size. The costs are statistical noise at low sample rates, interrupt overhead at high ones, and vendor specificity: code written for IBS does not help Intel or Arm hosts.
What the new Zen 6 profiler changes is unconfirmed. One inference, not a finding: a dedicated memory profiler could lower sampling overhead or return richer memory data than today's IBS op sampling. Only the AMD document and the talk can settle whether it does, and whether Linux uses it for promotion.
How do the three approaches compare?
The table is a general characterization, not benchmark data. It rates hardware sampling on its general design, not on any Zen 6 result.
| Method | Overhead profile | Accuracy | Maturity | |---|---|---|---| | NUMA balancing hinting faults | Per-fault CPU cost; can interrupt applications | Coarse recency; first touch only per scan | In mainline; tiering mode is a newer option | | DAMON (accessed-bit sampling) | Bounded by design, tunable | Region-level, binary per interval | In mainline; available actions vary by kernel version | | Hardware sampling (IBS) | Interrupt cost scales with sample rate | Per-access addresses, but statistical | General IBS perf support exists; Zen 6 profiler support unconfirmed |
What should you check before buying or configuring around this?
Do not buy hardware for this feature yet. The checklist below separates facts from open questions. See full coverage of rocm or vulkan for llama.cpp on amd gpus? how to decide for additional background.
| Item | Status | Where to verify | |---|---|---| | AMD engineer discussed IBS Memory Profiler on Zen 6 at LPC 2026 | Reported | Phoronix | | Feature is described in a public AMD technical document | Reported by Phoronix; primary document not yet checked | AMD's developer documentation | | Which Zen 6 products include it | Unconfirmed | AMD product documentation | | Mainline kernel support for using it in promotion | Unconfirmed | linux-mm mailing list, kernel release notes | | Performance gains over NUMA balancing or DAMON | Unconfirmed | LPC 2026 slides and recording | | Behavior inside VMs and cloud instances | Unconfirmed | Hypervisor documentation |
Beyond the feature itself, confirm three things:
- Your CXL platform exposes memory with correct performance attributes.
- Your firmware is current.
- Your distribution kernel is recent enough for any tiering code you plan to use.
Wait for a released kernel and test on your own workload before making a production decision.
How can you inspect your current tiering setup?
These commands are read-only. The author has not executed them, and output varies by distribution and kernel, so verify the results on your own system.
Also read: related topic: best webdav mount tool for windows: pick and test one
numactl --hardware
lscpu | grep -i numa
ls /sys/devices/system/node
cat /proc/sys/kernel/numa_balancing
cat /sys/kernel/mm/numa/demotion_enabled
ls /sys/devices/virtual/memory_tiering
grep -i ibs /proc/cpuinfo | head -n 3
perf list 2>/dev/null | grep -i ibs
The first three show node count, distances and CPU-less memory nodes, which usually indicate CXL or similar capacity tiers. The sysctl and demotion files show whether balancing and demotion are on. The memory_tiering directory may not exist on older kernels.
The last two show whether the CPU advertises IBS and whether perf exposes it. A missing line there says something only about that machine and kernel, not about Zen 6. Background on node performance attributes is in the NUMA performance documentation.
Does a hardware profiler mean CXL memory will behave like DRAM?
No, though headlines like this one invite that reading. Better hot page detection improves the decision about what to place where. It does not change the latency or bandwidth of the slow tier.
A workload with a small, stable hot set can benefit a great deal from good promotion. One that touches a huge working set uniformly has no hot set to promote, and no profiler fixes that. Migration also costs bandwidth and CPU time, so faster detection can hurt if the policy migrates too eagerly. Expect the hardware to feed the policy, and expect the policy to need tuning.



