On a NUMA machine the cost of a memory access depends on which node holds the page and how busy the interconnect is. Carrefour measures the traffic and moves or replicates pages to balance it, a holistic view rather than the usual locality-only one; the CACM article surveys what still goes wrong in modern kernels.

When the CPU and the GPU share one physical memory, the memory-management methods designed for discrete cards can cost more than they save. I measured them on integrated systems, including a negative result on accelerating garbage collection on integrated GPUs, and lately unified-memory performance across edge, consumer and datacenter platforms.

Off-the-shelf processing-in-memory hardware moves computation next to the DRAM. Our case study asks which workloads actually gain, and what the system software needs to make it usable.

A framework that adds publicly observable dynamic data sources to the kernel's entropy without degrading what it already has.