Dev.to AI 🤖 Ai 👁 0 📖 5 min read

A Schema Is Not an Enforcement Boundary

Your agent has a YAML contract. The schema validates. A reviewer approves it. The file says refunds require human approval. Now ask a different question: what stops the refund function from running without that approval

Your agent has a YAML contract. The schema validates. A reviewer approves it. The file says refunds require human approval.

Now ask a different question: what stops the refund function from running without that approval?

If the answer is "the contract says so", you have a declaration. You have not yet shown a control.

In The Delegation Boundary, I argued that the boundary belongs in the workflow, not in the prompt. This is the engineering follow-up: how to tell whether a declared boundary reaches execution.

1. Separate four claims

These are different claims with different evidence:

  • Declared: a contract describes what should happen.
  • Validated: a checker accepts the contract's structure and required fields.
  • Enforced: the execution path checks the relevant rule before the side effect and stops on refusal.
  • Observed: a test or runtime record shows what happened on a particular path.

A valid schema does not establish that execution consults it. A decision log does not establish that every path passed through the gate. A signed contract does not establish that its rules cover the risk you care about.

Each is useful. None substitutes for the next.

2. Use our own registry as the example

Hlinor Agent Registry makes this distinction explicit in its README section, "What is enforced at runtime".

The compiler reads the sources named in the manifest. PolicyChecker evaluates compiled action allowlists, blocklists, resource patterns, and supported typed policies. A named policy with no compiled policy behind it remains declarative context.

That last sentence matters. Putting no-customer-pii-in-logs in a policy list does not, by itself, inspect or redact a log message.

The repository also contains contracts that support review or standalone validation without participating in a PolicyChecker decision. Capabilities can be compiled as inventory without becoming an authorization gate. Declared budgets and rate limits are not automatically enforced by the checker.

The right question is not "does the repository have a schema for this?" It is "which code reads this field, on which path, and what happens when the rule fails?"

3. Test the side-effect boundary

The registry's starter example has a strict agent and a typed policy requiring an approval for refunds. You can generate and compile it locally:

pip install hlinor-registry
hlinor-registry init
hlinor-registry compile --manifest registry.yaml --output bundle.json

The following adapts the README example. It uses a list instead of a payment API, so no money moves:

from hlinor_registry import GovernanceDeniedError
from hlinor_registry.integrations.decorators import governed

executed = []

@governed(
    agent_id="my-agent",
    action="refund_payment",
    bundle_path="bundle.json",
    resource=lambda call: f"order/{call.kwargs['order_id']}",
    signals=lambda call: call.kwargs.get("approval", {}),
)
def refund(*, order_id: str, amount: int, approval=None):
    executed.append((order_id, amount))

try:
    refund(order_id="1234", amount=5000)
except GovernanceDeniedError as denial:
    print(denial.decision.reason_code)

assert executed == []

I ran this against repository revision a9af5e5. It printed POLICY_SIGNAL_MISSING, and the assertion passed: the wrapped function body did not run.

That is evidence for a narrow claim. This invocation went through the wrapper, which refused it before execution. It does not prove that an application has no unwrapped refund path, that its credentials cannot be used elsewhere, or that the process itself cannot be changed.

For your own workflow, replace the list with a fake client and assert that the downstream call count stays at zero on refusal. Checking only the returned decision misses the thing that matters: whether the effect happened.

4. An approval-shaped object is not an authenticated approval

The direct PolicyChecker signals API accepts claims supplied by its caller. It can check the required role, request binding, and freshness. It cannot establish that a human actually granted the approval.

I tested this distinction too. A locally constructed signal naming refund_payment:order/1234 allowed the wrapped call. Reusing it for order/9999 was refused. Changing the amount while keeping order/1234 was accepted by this starter path.

There is no mystery here: the signal names the action and resource, not the complete refund arguments. A role string is not an authenticated person. An order identifier is not an approved amount.

For higher assurance, the registry has a separate experimental BoundTool path with normalized argument validation and optional detached signed approvals, trusted keys, and replay guards. Those controls must be wired into the invocation. Their existence elsewhere in the repository changes nothing about the simpler wrapper above.

Even that binding path is not independent deployment attestation or proof that a callable's hidden side effects match its declaration. The README and security policy describe these limits. Read both before making a security claim.

5. Review one workflow, end to end

For a refund, deployment, or customer-facing send, write down:

  1. The actual effect: which function and credential can change the outside world?
  2. The gate: where does refusal stop that function, and can another path bypass it?
  3. The binding: does the checked resource come from the arguments the function actually uses? Does approval cover the amount, recipient, or other consequential parameters?
  4. The trust source: who can supply approval signals, replace the bundle, change trusted keys, or lower the minimum accepted revision?
  5. The failure behavior: what happens on missing approval, malformed input, a broken bundle, or unavailable evidence?
  6. The evidence: which negative tests show that the downstream call did not happen?

For the registry CLI, 0 means allowed, 1 means denied, and 2 means no decision was reached. Both refusal and evaluation failure should stop the effect in a fail-closed workflow, but keep them separate in reporting. A broken bundle is an operational fault, not proof that a policy successfully denied an action.

The boundary is where the effect stops

Keep the schema. Keep the signed bundle. Keep the reviewable contract. They make intent visible and changes inspectable.

Then follow that intent into the exact execution path and test refusal there.

A useful engineering claim sounds like this: "For this workflow, with these inputs and this bundle, the gate refused the request and the downstream function did not run. These other paths remain outside that evidence."

That is less impressive than "our agents are governed". It is also something another engineer can check.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.