Dev.to Security 🔐 Cybersecurity 👁 0 📖 5 min read

Why AI Agents Fail at Login Walls: SMS, OTP, Passkeys, CAPTCHA

Agents can open a browser, click through a form, and call APIs. Then they hit a login wall and the run becomes weird. The browser waits for a code. A phone receives an SMS. A link lands in an inbox. A passkey prompt ask

Agents can open a browser, click through a form, and call APIs. Then they hit a login wall and the run becomes weird.

The browser waits for a code. A phone receives an SMS. A link lands in an inbox. A passkey prompt asks for a device. A CAPTCHA asks for proof that the actor is not automation.

The agent keeps waiting because the challenge moved outside the page it can see.

That is the real failure: the system treats authentication as a local browser step, but the proof lives across channels, people, devices, policies, and time.

Login is not one step

A human login flow hides a lot of state.

The person knows which account they are using. They can reach the phone or inbox. They understand whether a prompt belongs to the current attempt. They know when to stop. They can decide that a CAPTCHA or passkey prompt means "I need to take over."

An agent does not automatically have that context. Without a control layer, it only sees:

  1. A form.
  2. A challenge.
  3. A timeout.
  4. Maybe a pasted secret.

That is not enough for a reliable run.

The login wall map

Different walls fail agents in different ways:

Wall Where the browser loses context Safe product behavior
SMS The code arrives outside the browser. Wait under policy, consume once, return a verdict.
Email OTP The proof moves to an inbox. Resolve whether this inbox is approved for this run.
Magic link Opening the link may be the login action. Bind the link to the right authority and target.
Passkey The service wants device and origin proof. Ask for takeover or fail closed.
CAPTCHA Automation is being challenged. Stop or use an approved test path.

The important thing is not "can the browser keep clicking?" It is whether the agent has a governed branch when the proof channel leaves the page.

SMS and OTP are state problems

SMS looks simple because the payload is small. Enter number. Wait for code. Submit code.

In practice, the useful state is larger:

Question Why it matters
Is this an app you own or are authorized to test? Otherwise the agent is trying to automate someone else's account boundary.
Which run requested this code? Old or parallel messages can be consumed by the wrong attempt.
Is the number allowed by the target app? The app can reject a number before any carrier delivery happens.
Did the message arrive inside the wait window? Late codes can be unsafe to reuse.
Was the code consumed once? Two agents should not race the same credential.
What happens if access is revoked while the agent waits? The safe answer is a stopped branch, not a leaked result.

The important part is not just delivery. It is associating the code with the right authority, the right target, the right session, and the right outcome.

Email OTP and magic links move the wall

Email challenges move the proof to another inbox.

For test environments, injected email OTPs and magic links can prove your parser and run logic. For production-like flows, the inbox is part of the identity boundary.

A magic link is especially sensitive because clicking it can be the login action, not just a way to retrieve a code.

The agent needs an answer to: "Am I allowed to use this mailbox for this target, right now?" If the answer is not explicit, the run should pause or stop.

Passkeys are not passwords with nicer UI

A passkey prompt usually means the service wants proof tied to a device, origin, and credential holder. It is not just asking for text.

For an agent, a passkey prompt should usually become a branch:

  • ask for human takeover,
  • switch to a supported test path,
  • or fail closed with a clear reason.

Treating passkeys as something to bypass misses the point. They are designed to bind the login to a trusted device and origin. A safe agent workflow should respect that boundary.

CAPTCHA is a stop sign, not a missing solver

CAPTCHA exists to challenge automation. That means the honest product behavior is not "silently solve it and keep going."

For owned tests, you can use staging controls, test bypasses, allowlisted origins, or app-level policy.

For third-party destinations, a CAPTCHA is often the moment to stop. The agent should report a clear result rather than spin, scrape, or pretend the run is still healthy.

The agent needs a branch, not just a browser

The better model is a state machine around the login wall:

agent reaches login wall
        |
        v
challenge detected
        |
        v
policy checks actor, owner, target, channel, allowance, revocation
        |
        v
structured verdict
        |
        +--> code received: continue the authorized run
        +--> timeout: retry only if policy allows
        +--> access revoked: stop safely
        +--> passkey required: request takeover or stop
        +--> CAPTCHA: stop or use an approved test path
        +--> target refused: report the refusal

This is where "operational identity" becomes concrete.

The agent is not merely holding a token or watching a browser. It is acting under a named authority, against an approved target, with a lifecycle that can be revoked and audited.

Where AgentSIM fits

AgentSIM is the control plane for agents hitting auth walls in services you own or are authorized to automate.

Today, live SMS is the supported challenge path for verified owned apps. AgentSIM opens a challenge, applies policy, waits for a verdict, and returns a structured outcome to the agent.

Email OTP and magic links can be exercised with injected test messages. Passkeys and CAPTCHA are treated as stopping conditions unless your environment has an approved test path.

The narrowness is deliberate. A safe product should not claim universal access to every login wall. It should make authority explicit, return clear outcomes, and fail closed when the wall means "do not continue."

The practical test

If your agent hits a login wall, ask five questions before adding another browser trick:

  1. Who is the agent acting for?
  2. Is the target service owned or authorized?
  3. Which channel carries the proof?
  4. What exact outcomes can the agent receive?
  5. What happens if authority changes mid-run?

If those answers live only in a prompt, the run is fragile. If they live in a state machine with policy and evidence, the agent can make progress without pretending to be a human.

The future is not agents that bypass every wall. It is agents that understand which wall they hit, which authority they hold, and which branch is safe next.

Read the canonical AgentSIM article

Try the AgentSIM quickstart

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.