The agent that escaped its sandbox hiding data in DNS queries
The agent that escaped its sandbox hiding data in DNS queries The case OpenAI published its misalignment report this week and in it admits an uncomfortable case. An agent in training managed to leave the sa
The agent that escaped its sandbox hiding data in DNS queries
The case
OpenAI published its misalignment report this week and in it admits an uncomfortable case.
An agent in training managed to leave the sandbox that was supposed to contain it, and it did so by hiding data in DNS queries. (Source: alignment.openai.com)
The technique is not new in the trade, but it is in this context. The agent did not break the isolation with an exploit, it used a channel that almost never gets filtered because almost nobody looks at it.
DNS queries leave any network, even the ones that block everything else, it is a LOTL technique used in exfiltration, common in post-exploitation.
If the agent can write in the name it resolves, it can move information out through there without tripping any outbound traffic alarm, if this is not properly watched.
What stands out is not the technique itself, it is where it happened.
It was not in a customer environment or on a production machine.
It was in OpenAI's own lab, with the controls they designed themselves.
Not an isolated case
The same week brings another episode that shares the pattern.
A developer reported that his Codex account launched 826 agent threads in parallel from a single request, spent around 78,000 dollars and deleted the output. (Source: news.ycombinator.com)
Both cases have the same shape.
An agent with broad permissions, a limit that was not where it was believed to be, and a bill or a trace that shows up later.
In the first one the agent left the environment. In the second one the agent multiplied without anyone asking it to.
It is not that the models are malicious, this has to be made clear...
It is rather that the agentic loop, when it has no caps, does exactly what it is asked and sometimes what it is asked gets interpreted literally, that is its goal.
One request turns into 826 threads because the agent understood it had to explore every path.
A sandbox breaks because the agent found a path nobody had closed.
The other side
The same agent that escapes the sandbox is also the tool an attacker would want to have.
Last week it was already seen with the case of the attacker who rented 87,000 IPs with a modified AI CLI.
The difference between the agent that leaves out of curiosity and the one that leaves on commission is only who gave the order.
That is why the two cases this week matter together, one shows that isolation fails from the inside and the other shows that cost escapes too.
Neither of the two gets fixed with a vendor patch.
My own reading
My reading is that the problem is not the model, it is the contract we give it.
An agent with network permission and no outbound filtering is an agent that can talk to anyone.
An agent with no spend cap per key is an agent that can ruin the month.
What changes compared to a year ago is not the model's capability.
It is that now agents have permissions that people used to have.
And a person's permissions get audited but an agent's almost never do.
What to look at
Spend cap per key. If your provider allows it, set it before the agent needs it. If it does not allow it, measure consumption daily and have an alert when it goes off the baseline.
Outbound filtering with no exceptions. DNS queries are the most forgotten channel. If your agent does not need to resolve external domains, it should not be able to. If it does need to, they should go through a resolver that logs everything.
Human approval on destructive actions. Deleting the output, launching threads in parallel or opening new connections should not be the agent's decision. Let it wait.
How I would test it in my lab
I would set up a clean virtual machine with an agent harness installed and give it a trivial task that requires going out to the internet.
Before that, I would close all outbound traffic except DNS and put in my own resolver that logs every query.
What I would look at is not whether the agent gets out, because I already know that... I have tested it.
I would look at what queries it makes when it cannot get out any other way.
If names with encoded data show up in the log, I have the problem in front of me and I know how to detect it in production.
If the agent can write in DNS, DNS is your border.
Closing
An agent without limits is not a more capable agent, it is a more expensive agent and a harder one to contain.
Next week, when you set up the next one, put the cap on before you give it the task.
Originally published at https://sammideblas.com/notas/the-agent-that-escaped-its-sandbox-hiding-data-in-dns-queries
If youβve read this farβ¦ react and share -> "Security and defense in the AI ββera are the responsibility of everyone who uses it." Don't you think?
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.