Dev.to Security 🔐 Cybersecurity 👁 0 📖 6 min read

Testing the deny path: how to prove an MCP server's guards hold

An MCP server's guards are the only code in the project where the test that matters asserts that nothing happened. No row came back. No request left the box. No file was opened. That is an awkward thing to assert, and it

An MCP server's guards are the only code in the project where the test that matters asserts that nothing happened. No row came back. No request left the box. No file was opened. That is an awkward thing to assert, and it is the reason a guard suite can be green while proving almost nothing about the guard.

Here is the shape I use, and the two mistakes it is designed to catch.

Make the guard a decision you can call

If the check lives inside the tool body, the only way to test it is to drive the whole tool and infer the decision from whatever came back. Pull the decision out so it is a function with a return value or an exception, and nothing else.

# tools/files.py
from pathlib import Path


class Denied(Exception):
 pass


def resolve_in_sandbox(root: Path, requested: str) -> Path:
 candidate = (root / requested).resolve()
 if not candidate.is_relative_to(root.resolve()):
 raise Denied(f"outside sandbox: {requested}")
 return candidate


def read_text_file(root: Path, requested: str) -> str:
 return resolve_in_sandbox(root, requested).read_text(encoding="utf-8")

The same split works for the other two reaches an MCP server tends to hand a model. A SQL guard is assert_read_only(sql) raising before execute. An SSRF guard is assert_host_allowed(url) raising before the HTTP client is ever handed the URL. In every case the guard is a pure decision over an input string, and the effect is a separate line that runs afterwards.

Now the deny cases are a table:

import pytest
from pathlib import Path
from tools.files import Denied, read_text_file, resolve_in_sandbox


@pytest.fixture
def root(tmp_path):
 (tmp_path / "notes.txt").write_text("ok")
 return tmp_path


@pytest.mark.parametrize("requested", [
 "../../etc/passwd",
 "/etc/passwd",
 "notes.txt/../../../etc/passwd",
 "./../notes.txt",
])
def test_escape_is_denied(root, requested):
 with pytest.raises(Denied):
 resolve_in_sandbox(root, requested)

Parametrize rather than loop inside one test. When a new bypass turns up you add a string, and when one regresses pytest names the exact input in the failure line instead of telling you that test_paths failed.

The deny-only suite that passes for the wrong reason

Here is mistake one. A guard test file made entirely of pytest.raises blocks passes if you replace the body of the guard with raise Denied("nope"). Every attack string is rejected. So is every legitimate request. The suite is green and the tool is dead.

So pin the allow path in the same file, right next to the denials:

def test_plain_name_is_allowed(root):
 assert resolve_in_sandbox(root, "notes.txt").read_text() == "ok"

You can check this property deliberately, and it costs one edit. Make the guard unconditionally deny, run the suite, and look at what fails. If only your allow tests go red, the deny tests are real. If nothing goes red, your deny tests were never testing the guard. Then make the guard unconditionally allow and confirm the whole deny table lights up. Undo both. That pair of edits tells you more about a guard suite than reading it does.

Asserting that nothing happened

Mistake two is subtler. pytest.raises(Denied) proves an exception came out. It does not prove the effect never ran. A guard that checks the path after opening the file raises exactly the same exception and leaks exactly the same bytes into your process first. Order is the whole property, and the exception type cannot see order.

You can test order directly by booby-trapping the effect:

def test_denied_read_never_touches_disk(root, monkeypatch):
 def explode(*args, **kwargs):
 raise AssertionError("read_text ran on a denied path")

 monkeypatch.setattr(Path, "read_text", explode)

 with pytest.raises(Denied):
 read_text_file(root, "../../etc/passwd")

If the guard runs first, read_text is never reached and the test passes. If someone reorders the function later, the AssertionError surfaces instead of Denied and the test fails with a sentence describing what went wrong. The same trick applies anywhere an effect is reachable through a seam: swap the DB cursor for an object whose execute raises, swap the HTTP client for one whose send raises, and your deny tests start making a claim about sequencing rather than about exception types.

This is also the cheapest way to test an SSRF guard without network access in CI. You are not asserting that the request failed. You are asserting that no request was ever constructed.

Test the wiring, not just the guard

A guard can be correct, fully covered, and protect nothing, because the handler the MCP protocol actually dispatches to is a different callable than the one your unit tests import. An unguarded helper left registered, a decorator applied to the wrong function, a second code path added for a batch variant. Unit tests on resolve_in_sandbox cannot see any of that.

So add tests one layer up that go in by tool name, the way a client does, and assert the request is refused:

@pytest.mark.parametrize("tool,args", [
 ("read_text_file", {"path": "../../etc/passwd"}),
 ("query_database", {"sql": "DELETE FROM users"}),
 ("call_api", {"url": "http://169.254.169.254/latest/meta-data/"}),
])
def test_dispatch_refuses(tool, args):
 result = dispatch(tool, args)
 assert result.is_error

Guard-level tests tell you the decision is right. Dispatch-level tests tell you the decision is in the path. You want both, and they fail for different reasons, which is the point.

What this does not get you

Enumerated strings only cover the bypasses you thought of. A table of traversal inputs is a regression net, not a proof, and the interesting failures usually live one layer below your guard: how the OS normalizes a path before you ever see it, how symlinks and case-insensitive filesystems change what resolve() means, how a given SQL dialect treats a construct you did not plan for. If you need a stronger promise than string-level checking, you move the enforcement down to where the effect happens: a connection opened read-only, a filesystem handle confined by the kernel, an outbound proxy that only knows about hosts you listed.

Name-level checking has its own ceiling. Deciding that a hostname is allowed is a decision about a name, and the address behind a name is resolved separately. If you need an address-level guarantee, pin the resolved address and connect to that, or put the allowlist in a proxy outside the process.

And none of this is about resource use. A read-only guard says a statement cannot write. It says nothing about a statement that scans your largest table, so timeouts, row caps and statement limits are a separate job.

If your MCP server only exposes computation over data you already shipped inside it, skip most of this. The guards earn their keep the moment a tool takes a string from the model and turns it into a URL, a query, or a path.

I wrote these guards once and kept them. The MCP Starter Kit is the production-shaped Python MCP server I built, maintain and run, with the SSRF host allowlist, the read-only SQL guard, the path-traversal sandbox and a 14-test pytest suite over the guards and tool dispatch: https://fulcrumenterprises.tech/go/mcp-starter-kit/?c=devto&v=57dd17

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.