Your agent can't safely read a web page it hasn't read yet
TL;DR GET https://api.trustsource.cc/safefetch?url=<percent-encoded-url> You get back: sanitized page text plus a SAFE / REVIEW / BLOCK prompt-injection verdict. Act on it: BLOCK β discard the content. REVI
TL;DR
GET https://api.trustsource.cc/safefetch?url=<percent-encoded-url>
-
You get back: sanitized page text plus a
SAFE/REVIEW/BLOCKprompt-injection verdict. -
Act on it:
BLOCKβ discard the content.REVIEWβ use as data, never as instructions.SAFEβ use normally. - Cost: $0.01 USDC per call via x402 on Base mainnet. No API key, no signup, no account.
-
To pay: the first call returns
402with aPAYMENT-REQUIREDheader. Sign the EIP-3009 authorization it describes and repeat the same request with anX-PAYMENTheader. - Why bother: you can't ask your model whether a page is safe without first feeding it the page.
That's everything you need to call it. The rest explains how it decides.
The problem
You want to know whether a page is safe before your model reads it. But the usual way to find out is to ask the model β which means feeding it the page. If the page contains a prompt injection, you've already lost. The attack runs the moment the text enters the context window.
fetch() gives you no verdict. It hands your agent a string and hopes for the best.
What comes back
SafeFetch fetches the URL server-side, separates what a human would actually see from what's concealed in the markup, scans both, and returns clean text plus a verdict:
{
"url": "https://example.com/",
"verdict": "BLOCK",
"risk": 0.95,
"reasons": ["instruction override concealed in hidden content (matched in hidden content)"],
"injection": {
"detected": true,
"techniques": ["instruction_override"],
"findings": [
{ "technique": "instruction_override", "placement": "hidden", "severity": 0.95, "weight": 0.95 }
]
},
"content": { "text": "β¦sanitized page textβ¦", "truncated": false }
}
Your agent branches on one field instead of reasoning about raw HTML.
SAFE means nothing concealed and aggregate risk under 0.25. REVIEW means risk between 0.25 and 0.7, or a content type that couldn't be scanned, or a low-trust host. BLOCK means a critical technique found in hidden or comment placement, or risk β₯ 0.7.
Placement matters more than wording
This is the part that makes it usable in production.
A blog post about prompt injection contains the exact phrases an injection scanner looks for. A naive keyword filter flags every security article on the internet and becomes noise you learn to ignore.
SafeFetch weights each finding by where it was found β hidden 1.0, comment 0.95, accessibility 0.6, metadata 0.55, script 0.5, visible 0.2 β and caps visible-text hits in aggregate. So "ignore all previous instructions" printed in an article body scores low, while the same string in a display:none div scores high.
Concealment is the signal. Nobody hides text from humans for a benign reason.
What it detects
Eleven techniques: instruction_override, system_prompt_exfil, data_exfiltration, delimiter_spoof, encoded_payload, tool_call_bait, role_hijack, homoglyph_obfuscation, unicode_tag_smuggling, invisible_unicode, bidi_override.
That covers instructions buried in display:none elements, HTML comments and alt attributes; invisible Unicode-Tag (U+E0000) and zero-width smuggling; homoglyph and base64-encoded payloads; ChatML / [INST] delimiter spoofing; markdown-image exfiltration; and tool-call bait. Payloads split across several hidden elements are reassembled and scanned as one string, so chunking the attack doesn't evade it.
Paying for a call
The 402 response names the scheme (exact), the network (eip155:8453 β Base mainnet), the amount (10000 atomic units, i.e. $0.01 USDC), the USDC contract, and the address to pay.
The agent signs an EIP-3009 transferWithAuthorization for that exact amount, base64-encodes it, and repeats the identical request with an X-PAYMENT header. The facilitator settles on-chain and the endpoint returns 200.
Same URL twice, one extra header, one signature. No session, no nonce to track, and an abandoned 402 costs nothing. Any x402-aware HTTP client handles this for you.
Limits, honestly
- 2 MB fetched, ~100,000 characters returned, 10-second timeout, 10-minute cache.
- Redirects re-validated at every hop; DNS resolved then pinned to defeat rebinding; private and link-local ranges blocked.
- Detection is heuristic, not exhaustive. A novel or heavily obfuscated payload can still read as SAFE. Treat SAFE as "no known pattern matched," not proof of safety β and keep treating fetched content as data rather than instructions.
- Snippets inside
injection.findingsare defanged for display. They're evidence, not instructions.
Why it's worth a cent
One successful injection against an agent with tools, a wallet, or memory is worth considerably more than $0.01 to an attacker.
SafeFetch is one of seven verification endpoints at trustsource.cc β the others cover URL safety verdicts, email spoofability, domain trust, TLS certs, security headers, and AI-crawler policy.
- Docs: https://trustsource.cc
- OpenAPI: https://api.trustsource.cc/openapi.json
- For agents: https://trustsource.cc/llms.txt
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.