Dev.to Security 🔐 Cybersecurity 👁 0 📖 12 min read

What Is a Shadow Stack? How Can CPUs Stop Return-Address Hijacking?

Every time a function returns, the processor has to trust a piece of data sitting in writable memory. That data is the return address. It tells the processor where execution should continue after the function finishes.

Every time a function returns, the processor has to trust a piece of data sitting in writable memory.

That data is the return address. It tells the processor where execution should continue after the function finishes. On conventional architectures like x86-64, CALL pushes this address onto the normal stack. RET reads it back and jumps there. The whole mechanism depends on that value being intact.

If an attacker can overwrite it, the processor does not notice. It reads the corrupted address and jumps to it, exactly as designed.

The question shadow stacks answer is this: if the attacker can corrupt the normal stack, how can the system maintain a trusted record of where a function is actually supposed to return?

What a Normal Function Call Actually Does

When a function is called on x86-64, CALL does two things: it pushes the address of the instruction following the call onto the stack and transfers execution to the called function. The stack pointer RSP decrements to make room for this return address.

Inside the called function, the stack frame grows downward as local variables and other state are allocated. The return address sits above all of that, placed there during the call.

When the function executes RET, the processor reads the value at the current top of the stack, loads it into the instruction pointer RIP, and increments RSP. Execution continues at whatever address was retrieved.

That is the architectural contract. CALL stores where execution should resume. RET retrieves it.

The problem is where that value lives. The return address is stack memory. Stack memory is writable. And writable memory can be corrupted.

When the Normal Stack Becomes the Attack Surface

Consider a function with local buffers, variables, and saved state. All of that lives on the stack, and so does the return address. The stack frame might look roughly like:

Higher addresses
-------------------
|  Return Address  |  ← what RET will use
-------------------
|  Saved RBP       |
-------------------
|  Local Variable  |
-------------------
|  Buffer          |  ← if this overflows...
-------------------
Lower addresses

A memory-corruption bug, a buffer overflow or a write beyond a buffer's bounds, can overwrite adjacent memory. If the write extends far enough up the stack, it reaches the return address. When the function returns, RET reads the corrupted value and jumps there.

This is return-address hijacking. The program continues executing, but at a destination the attacker chose rather than the one the compiler intended.

This is also the conceptual foundation of return-oriented programming. Instead of injecting new code, an attacker redirects returns through sequences of existing instructions in executable memory, chaining them together by controlling what appears on the stack.

Why NX Alone Is Not Enough

NX (no-execute) and DEP (data execution prevention) mark data regions non-executable. The processor refuses to execute instructions from pages flagged that way. Placing shellcode on the stack and jumping to it fails because the stack is non-executable.

But return-address hijacking does not require executable data pages. The attacker redirects execution into already-executable code. The program's own code, shared libraries, the C runtime, all of this is legitimate and executable. ROP exploits this by chaining small sequences from within that legitimate executable memory.

NX answers the question "is this page executable?" It does not ask "should execution actually be here right now, given how this program is supposed to run?"

That is the question a shadow stack addresses.

The Shadow Stack Idea

The core insight is straightforward: keep a second record of expected return addresses, stored somewhere the attacker cannot easily corrupt.

When a function is called, the return address goes to the normal stack as usual. It also goes to a separate structure, the shadow stack. This second copy is protected by mechanisms that ordinary stack writes cannot reach.

When the function returns, the system checks whether the address on the normal stack matches what the shadow stack says it should be. If they agree, the return proceeds. If they disagree, something has tampered with the normal stack, and the return is blocked.

Function call:
    Normal stack ← return address
    Shadow stack ← protected copy of return address

Function return:
    Normal stack → provides return address
    Shadow stack → provides expected return address
    Compare:
        Match → return proceeds
        Mismatch → control-protection fault

The security property depends entirely on the protection of the shadow stack. Simply copying return addresses to another memory region accomplishes nothing if the attacker can modify that region too. The shadow stack has to be protected by a mechanism that ordinary memory writes cannot bypass.

Why the Security Property Holds

An attacker with a memory-corruption primitive can generally write to any writable memory the process can reach. The normal stack is writable. That is the problem.

The shadow stack needs to be in a different category. Not just a different location, but enforced differently. On hardware-assisted implementations, the shadow stack is backed by memory that the processor treats as read-only for ordinary stores and only updates through specific call/return mechanisms. A normal MOV targeting shadow-stack memory faults. Only CALL, RET, and the specific instructions that manage shadow-stack entries can modify it.

This is not the same as putting the shadow stack at a secret address and hoping the attacker does not find it. Address secrecy is fragile. Enforced protection at the hardware level is not: the protection holds regardless of whether the attacker knows where the shadow stack is, because the enforcement mechanism rejects the write before it can happen.

The shadow stack is not just another copy of the return address. It is a protected source of truth maintained by the processor itself as part of the call/return mechanism.

Software Shadow Stacks

Hardware assistance is not the only option. A shadow stack can be implemented in software by maintaining protected metadata representing expected return addresses. Compiler instrumentation can insert code at each function entry and exit to update and check this metadata.

The challenge is making the software-maintained shadow stack actually resistant to corruption. Options include placing it in memory with page protections that make accidental overflows unlikely to reach it, or using a segmentation register to reference it through a protected pointer. Neither is as strong as hardware enforcement.

Software approaches also have to handle everything the hardware would normally manage automatically: context switches, signal delivery, exceptions, threads, setjmp/longjmp, dynamic code, and more. Each of these represents a case where the shadow-stack state has to be correctly maintained across a non-standard control-flow transition. Getting any of these wrong can cause false positives, missed detections, or crashes.

The overhead is also real. Checking a shadow stack at every function return costs something, and doing it in software costs more than doing it in hardware.

Hardware-Assisted Shadow Stacks

The processor already knows exactly when a call happens and when a return happens. It processes CALL and RET directly. Adding shadow-stack behavior to those instructions is a natural fit.

With hardware support, the processor maintains shadow-stack state as part of the call/return mechanism. No instrumentation required. The processor automatically records expected return addresses during calls and validates them during returns. The shadow stack lives in protected memory that ordinary stores cannot modify.

This removes most of the complexity that makes software shadow stacks difficult. Hardware support removes much of the per-call and per-return bookkeeping from software, but the operating system and runtime still have to manage shadow-stack state correctly across context switches, exceptions, and other non-standard control-flow transitions. The protected memory enforcement is a hardware property, not a software policy.

Intel CET

Intel's Control-flow Enforcement Technology provides hardware shadow stacks as one of its two major components.

When CET shadow stacks are enabled, CALL records the return address in both the normal stack and a hardware-managed shadow stack. RET compares the address from the normal stack against the shadow stack. A mismatch raises a control-protection exception (#CP).

The shadow stack memory is backed by pages with a specific page-table attribute. Ordinary stores to those pages fault. Only the processor's call/return instructions and a small set of shadow-stack management instructions can modify shadow-stack entries legitimately.

CET also includes Indirect Branch Tracking (IBT), which addresses forward-edge control flow. IBT requires that valid targets of indirect calls and jumps be preceded by an ENDBR instruction. The two mechanisms are distinct and address different parts of the control-flow problem.

Shadow Stack: backward-edge protection. Returns must match the shadow-stack record.

IBT: forward-edge constraint. Indirect branches must land on marked valid targets.

Deploying CET requires operating system support, which Windows has had for some time and Linux has been adding, plus toolchain support to ensure that code works correctly with both shadow stacks and IBT.

ARM and Pointer Authentication

ARM has its own approach to return-address protection in Pointer Authentication (PAC), but it is worth being precise: PAC is not a shadow stack.

Pointer Authentication uses cryptographic instructions to sign and authenticate pointers, including return addresses. A function's prologue can sign the return address using a key stored in a system register, and the epilogue can authenticate it before using it. If the return address has been tampered with, authentication fails and the resulting invalid pointer cannot be safely used as the intended return target.

This protects return-address integrity through authentication rather than through a separate protected copy. Both mechanisms protect backward-edge control flow, but they do so through entirely different means. PAC authenticates a pointer in place. A shadow stack maintains a separate protected record to compare against.

ARM has also introduced its own shadow-call stack mechanism in certain contexts, which is closer to the traditional shadow-stack model, but the exact mechanisms differ by architecture and configuration.

Shadow Stacks vs Stack Canaries

Stack canaries and shadow stacks both relate to stack corruption detection, but they are not the same.

A canary places a secret value at a known location in the stack frame, between local data and the return address. Before returning, the runtime checks that the canary value has not changed. If something overwrote memory continuously from the local buffer toward the return address, it had to pass through the canary. A changed canary means something corrupt happened.

Canaries ask: was this stack region modified?

Shadow stacks ask: does the return address match the protected record established at call time?

A canary detects corruption by observing that an adjacent value has changed. A shadow stack maintains an independent authoritative copy of what the return address should be. These are different security properties. An attacker who can perform a targeted write past the canary without disturbing it, or who finds a way to write directly to the return address from a non-contiguous location, defeats the canary but not the shadow stack. The shadow stack comparison fails regardless of how the normal stack was modified.

Shadow Stacks vs CFI

Shadow stacks and CFI are complementary, not synonymous.

CFI establishes a policy describing valid control-flow destinations and checks indirect transfers against that policy. It answers "is this destination allowed?" The validity of a destination is determined by static analysis of what the program was supposed to do.

A shadow stack answers a different question: "is this return address the one that was established by the corresponding call?" It does not evaluate whether the destination is in some permitted set. It checks whether the value matches the protected record from when the function was actually called.

CFI primarily addresses forward-edge control flow: indirect calls, function-pointer calls, virtual dispatch. Shadow stacks primarily protect backward-edge control flow: returns. They cover different parts of the control-flow surface, and modern systems that want comprehensive protection deploy both.

What Happens When the Stacks Disagree

The mechanism in detail:

  1. A function call records the legitimate return address in both the normal stack and the protected shadow stack.
  2. Some memory-corruption bug overwrites the return address on the normal stack.
  3. The function executes to completion and reaches RET.
  4. The processor reads the normal stack return address. It has been corrupted.
  5. The processor compares it against the shadow-stack entry. They disagree.
  6. The processor raises a control-protection fault.
  7. The operating system handles the fault, typically by terminating the process. The corrupted return is never used. The function does not return to the attacker's chosen destination. The fault fires before execution can be redirected.

What Shadow Stacks Do Not Protect

Shadow stacks protect the integrity of return addresses. That is their scope.

They do not protect against:

Arbitrary data corruption. If a bug corrupts data the program uses but does not affect return addresses, the shadow stack has nothing to check.

Function-pointer attacks. A corrupted function pointer that is invoked through an indirect call is a forward-edge problem. Shadow stacks do not cover forward-edge transfers.

Every form of control-flow hijacking. An attacker who can corrupt a vtable pointer, a callback stored in a data structure, or any other indirect-call target has an avenue that shadow stacks do not directly address.

Logic bugs. If the program's control flow is architecturally correct but semantically wrong, shadow stacks cannot help.

This is why CFI matters alongside shadow stacks. Shadow stacks close the return-address gap. CFI constrains where indirect calls can go. Together they cover more of the control-flow surface than either does alone.

The Practical Reality

Deploying shadow stacks across real systems involves more than flipping a hardware feature on.

Legacy binaries compiled without shadow-stack support may not work correctly with hardware enforcement. Mechanisms exist to mark shadow-stack compatible code, but older software may need updates or compatibility modes.

Unusual control flow matters. setjmp and longjmp bypass the normal call/return mechanism and require explicit shadow-stack management. Signal handlers delivered while a function is executing need the shadow stack state handled correctly. Threads each need their own shadow stack. JIT-compiled code may generate RET instructions in ways that need special handling.

Operating system, compiler, and runtime support are all required for correct deployment. Getting one layer wrong can cause crashes on legitimate programs or create gaps in the protection.

Performance overhead exists but is generally modest for hardware-assisted implementations, because the shadow stack operations happen in hardware as part of CALL and RET.

The Bigger Picture

Return addresses are not ordinary data. They are pieces of control-flow authority. When the processor reads one and jumps to it, it is deciding what the program does next. Placing that authority in ordinary writable memory and trusting it implicitly is a structural vulnerability.

A shadow stack changes that model. It maintains trusted control-flow state through a mechanism that ordinary writes cannot corrupt, and checks the return path against that state before committing to it.

The deeper lesson is about what modern systems are increasingly protecting. ASLR makes addresses unpredictable. NX stops code injection. Canaries detect some corruption. CFI constrains where execution can go. Shadow stacks protect the integrity of how execution gets there.

These are all different answers to variations of the same question: what does the program's execution path actually allow?

As software has gotten better at protecting data and preventing obvious bugs, the remaining attack surface has increasingly concentrated in control-flow metadata. The return address is a small piece of information, but it has outsized influence over what the program does. Systems that treat it as trusted architectural state rather than ordinary writable data are harder to subvert precisely because they stop treating an execution primitive like regular memory.

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.