Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 7 min read

Page Tables Are Not Free: The 1/512 Rule That Breaks Under Shared Mappings

A 4 KiB data page costs 8 bytes to map. That is a ratio of 1 to 512. When one address space maps a page once, the page table overhead is roughly one 512th of the memory it describes, which is why most developers never th

A 4 KiB data page costs 8 bytes to map. That is a ratio of 1 to 512. When one address space maps a page once, the page table overhead is roughly one 512th of the memory it describes, which is why most developers never think about it. The ratio is not a law. It is an assumption, and the assumption is that each page is mapped exactly once. The moment many address spaces map the same physical page, the ratio stops holding and page tables can consume as much memory as the data they describe.

Linux kernel developers have hit this wall repeatedly over two decades. The pattern is consistent, the workloads are not exotic, and the numbers are large enough to take down production servers. This article walks through that history in release order, because each case sharpens the same point: page table memory is a real, measurable cost, and it grows with the number of address spaces, not just the amount of data.

The Arithmetic Behind the Ratio

Start with the units. A standard 4 KiB page needs one page table entry to be mapped. On a 64-bit x86 kernel that PTE is 8 bytes. So a single mapping costs 8 bytes of page table memory for 4 KiB of data, which is 1/512 of the mapped size.

If N address spaces map the same page, each one needs its own PTE for it. The kernel does not share PTEs across processes. Every task has a pointer to its own mm_struct, and each mm_struct owns its own page table tree. So the cost scales with N even though the data page is shared once.

Run the arithmetic for the boundary case. If 512 processes map the same page, they need 512 PTEs. At 8 bytes each that is 4 KiB, which is exactly the size of a data page. The page tables now cost as much as the data they map.

That is the whole story in one number:

data page size:    4096 bytes
PTE size:          8 bytes
mappings to parity:  4096 / 8  = 512

Converted to code, the rule is a one-line function of how many address spaces map the same page:

pte_bytes = 8          # a PTE on 64-bit x86
page_bytes = 4096      # a 4 KiB data page
map_count = 512

# cost of mapping one page in N address spaces, in pages
pte_pages = (map_count * pte_bytes) / page_bytes   # 1.0 when N = 512

When you see a server with hundreds of processes attached to a large shared region, the page table overhead is no longer a rounding error.

The Assumption Sits Inside the Kernel

The 1/512 figure is not an arbitrary target. It comes from the geometry of the page table itself. A multi-level page table maps a virtual address through several levels, and each level holds pointers to the next, so the total overhead is a small fraction of the addressable range. Linus Torvalds made exactly this point in an early kernel mailing-list discussion about hashed page tables, arguing that a tree structure keeps neighboring entries adjacent so a single cache-line fill covers several TLB entries at once. Memory size, in that framing, was a secondary concern.

His own master's thesis acknowledged the memory side. Virtual memory mappings must be memory-efficient so the mapping information does not consume physical memory that could otherwise serve file-system caching or user programs. The 1/512 ratio is the kernel's way of buying structure with a modest fixed tax. The tax is low when the working set is mapped once. It stops being low when the same working set is mapped many times.

Andrea Arcangeli, 2002: A 64 GiB Machine Out of Memory

The first real production case is instructive because the machine had plenty of free RAM. Users with 64 GiB x86 machines kept running out of memory even though a large amount of memory was idle. Their workloads ran hundreds of processes, each mapping the same 1 GiB of shared memory.

The data pages were shared, but every process needed its own page tables to map them. On that architecture the page tables lived in a region with a hard ceiling, so the overhead collided with a physical limit. The fix was to move page tables out of that constrained region. It was an edge case created by the combination of a 32-bit constraint and a feature that widened physical addressing beyond what the platform normally supports.

The lesson survives. The workload was not malicious. It was a normal shared-memory deployment, and the page tables quietly doubled the memory footprint of the shared region.

Khalid Aziz, 2022: An Oracle Server and 878 GB of Page Tables

Twenty years later the same shape reappeared at a much larger scale. An Oracle database server with 512 GB of RAM was hitting the out-of-memory killer. When 1500 or more clients attached to a 300 GB shared global area, each client mapped the shared region with its own page tables.

With 8-byte PTEs, two thousand processes mapping the same 4 KiB page need 16 KiB of PTEs for that single page. In the worst case, where every process maps the whole shared area, the PTEs alone would consume roughly 878 GB. The proposed fix was to let processes share page tables, but the mechanism that would share PTEs across processes has not been merged into mainline.

Again the cause was not a leak. The data pages were shared correctly. The cost was purely the per-address-space page table replication.

Qi Zheng, 2021: MADV_DONTNEED Leaves Empty Page Tables Behind

The third case does not involve shared mappings at all. A single process reported 590 GiB of resident set size and 110 GiB of page tables, when mapping that much memory should only need about 1.2 GiB of PTEs. That is roughly a hundred times the expected overhead.

The workload used allocators that return memory to the kernel with madvise(MADV_DONTNEED) instead of munmap(). The distinction matters. Unmapping memory eventually frees the page tables. Advising the kernel that a range is not needed frees the data pages and clears the PTEs, but keeps the page table structures allocated. Empty page tables piled up.

A patch to reclaim those page tables took several rewrites and was eventually merged years later. The case is a reminder that a page table is a resource you have to release deliberately, and that the API you call determines whether it is released.

Huge Pages Are the Lever Database Operators Pull

Page table memory is proportional to the number of mappings, so the direct fix is to map more per entry. That is exactly what huge pages do. A 2 MiB huge page still needs a single PTE, so the overhead per byte drops by a factor of 512.

Database operators reached for this lever when page tables showed up in production incidents. One reproduction with PostgreSQL was forced to kill backends on a machine with 192 GiB of RAM and a large shared buffer. Page tables grew from tens of MiB to over 25 GiB as backends touched more of the cache, memory got pressured, and the machine started swapping. Switching to huge pages kept page tables around 61 MiB for the same workload.

A recent measurement with ClickHouse made the growth explicit. On a machine with 128 GiB and a 32 GiB shared buffer, each backend scanning a cached table added 31 MiB of page tables. At 200 connections, page tables took more than 6 GiB. With huge pages, the same 200 connections cost about 111 MiB.

Huge pages are not free. Reserving them needs unfragmented memory, and under load the same request can routinely fail. Transparent huge pages can stall on allocation, and kernel developers have debated whether defragmentation should be the default for years. There is no parameter-free win here, only a trade.

NUMA Machines Reintroduce the Same Argument

The memory cost of page tables has a latency twin on multi-node machines. When a thread runs on a node that does not hold the page tables, a TLB miss has to walk page tables in remote memory.

The research consensus is that you pay one way or the other. One approach replicates the entire page table tree on every node, which costs page-table memory times the number of nodes, and every change has to update every copy. Another approach copies a PTE to a node only when one of its threads faults on that node, and each page table page keeps a list of the nodes that hold a copy. Work on remote page tables showed they can slow an application as much as remote data, and the same effect was reproduced on an 8-socket machine with 8 TB of RAM.

So the design space is a balance between two costs: keep one copy and pay the latency of reading remote tables on every miss, or keep a copy per node and pay the cost of keeping them in sync, plus the memory multiplication.

What to Check in Your Own Workload

The cases above share a single, checkable symptom. Ask whether many processes, threads, or containers map the same large region. If the answer is yes, page table memory is worth measuring rather than assuming.

Linux exposes the metric directly. A single process reports its page table size, and the kernel tells you whether transparent huge pages are available:

# page table memory for one process (kB)
grep VmPTE /proc/<pid>/status

# is transparent huge page support enabled, and to what degree?
cat /sys/kernel/mm/transparent_hugepage/enabled

The VmPTE line is the per-process number to watch. On a shared-memory workload it grows with the number of address spaces, not with the amount of data, which is exactly the failure mode described above.

There is no single answer here, only a lever. This article presents a handful of real production incidents and the trade-offs involved, rather than a universal fix, because the right choice depends on the workload and the hardware. What is universal is the measurement.

The ratio of 1 to 512 is a convenient assumption. Under shared mappings it is a floor, not a cap.

Originally published on Dispatch.

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.