Dev.to Security 🔐 Cybersecurity 👁 0 📖 7 min read

We Gave AI Agents Real Tools — Then Realized “Just Ask Before Acting” Wasn’t Enough

Giving an AI agent tools feels like the moment it becomes truly useful. Now it can: read files send messages call APIs update records trigger workflows create documents modify connected apps That is when it stops be

Giving an AI agent tools feels like the moment it becomes truly useful.

Now it can:

  • read files
  • send messages
  • call APIs
  • update records
  • trigger workflows
  • create documents
  • modify connected apps

That is when it stops being just a chatbot.

It starts becoming software that can change things.

And that is also when the risk changes.

At first, one rule sounds reasonable:

“Ask the user before doing anything important.”

Simple.

Human-friendly.

Easy to add to the prompt.

But once an agent has real tools, that is not enough.

Because now the model is being asked to decide two things:

  1. What action should happen?
  2. Whether that action is important enough to require approval.

That is too much authority to put inside the same reasoning loop.

The Problem Starts With a Simple Workflow

Imagine an agent connected to a few business tools.

A user says:

“Clean up these customer records and notify the team.”

The agent might decide to:

  • read CRM records
  • modify customer fields
  • merge duplicates
  • delete old entries
  • send a team message
  • update a spreadsheet

Some of those actions are harmless.

Some are reversible.

Some are not.

Now imagine the only safety rule is:

Ask before anything important.

What exactly counts as important?

The model has to decide.

That is where things become uncomfortable.

“Important” Is Too Ambiguous

To a user:

Deleting 500 records is obviously important.

To the agent:

Removing duplicates may look like a normal cleanup step.

To a developer:

Sending data to an external system may be the risky part.

To security:

Accessing the data at all may require approval.

The word “important” does not define a reliable boundary.

It creates interpretation.

And interpretation is exactly what we should avoid for high-impact actions.

The Model Should Not Decide Its Own Authority

This became the key lesson.

The agent can decide:

What should I do next?

But it should not be the final authority on:

Am I allowed to do it?

Those are separate responsibilities.

A safer architecture looks more like this:

User request
    ↓
Agent proposes action
    ↓
System classifies action
    ↓
Policy checks permission
    ↓
Human approval if required
    ↓
Action executes
    ↓
Action is logged

The important part is:

The approval decision happens outside the model.

Read, Write, and Destructive Are Not the Same

One simple thing that helps is classifying actions by impact.

For example:

Read

  • fetch records
  • inspect documents
  • search files
  • read calendar data

Write

  • update a field
  • create a document
  • send a message
  • add an event

Destructive / High Impact

  • delete data
  • revoke access
  • publish externally
  • deploy
  • move money
  • change permissions

Now the system can enforce something concrete.

For example:

READ → allowed

WRITE → allowed or approval depending on context

DESTRUCTIVE → approval required

That is much stronger than:

“Please ask before doing anything risky.”

Prompts Are Guidance. Policies Are Boundaries.

This distinction matters.

A prompt can say:

“Never delete data without asking.”

That is useful.

But prompts can be:

  • misunderstood
  • forgotten
  • overridden by context
  • interpreted differently

A policy layer should be deterministic.

For example:

delete_record()
→ blocked
→ approval required

The agent does not get to decide whether deletion is “important enough.”

The system already knows.

Why This Matters More as Agents Get Better

A weak agent often fails because it cannot complete the task.

A strong agent creates a different problem:

It can complete the task in ways you did not anticipate.

That is the real shift.

The better the agent becomes at planning and using tools, the more important hard boundaries become.

Because capability is increasing.

Authority should not increase automatically with it.

Helpful Agents Can Still Cross a Line

This is important.

The dangerous behavior does not need to be malicious.

Imagine:

“Organize this workspace.”

The agent decides to:

  • archive old files
  • move folders
  • rename documents
  • remove duplicates

Every step may look helpful.

But maybe one folder was legally required to remain unchanged.

Maybe one document belonged to another team.

Maybe the “duplicate” was actually a historical copy.

The agent was trying to help.

That does not make the action safe.

Human Approval Should Happen at the Right Moment

Approval should not mean:

Confirm every tool call.

That would be terrible UX.

The goal is to insert approval when the action crosses a meaningful boundary.

For example:

Reading data
→ no approval

Creating a draft
→ no approval

Sending externally
→ approval

Deleting
→ approval

Changing permissions
→ approval

Deploying
→ approval

This keeps the agent useful without making it unrestricted.

Approval Should Explain the Action

Another important lesson:

Do not show the user:

Approve action?

That is too vague.

Show:

Send this message to the engineering channel?

or:

Delete 42 archived records?

or:

Publish this document externally?

The user should know exactly what they are approving.

That means the approval layer needs:

  • action name
  • target
  • scope
  • consequence

Not just a yes/no button.

The Agent Should Propose, Not Hide

A good pattern is:

Agent proposes → system explains → human approves → tool runs

Not:

Agent runs → explains afterward

That difference matters a lot.

Once the action already happened, approval is no longer approval.

It is just notification.

This Became a Real Problem While Building Xenition

We ran into this problem directly while building Xenition at xenition.com.

Xenition is designed around AI agents that can work across real tools and connected services, not just generate text inside a chat box.

That means an agent may need to:

  • read data
  • create content
  • update records
  • trigger workflows
  • interact with connected applications
  • produce real outputs

Once agents can actually act, the permission model becomes just as important as the model itself.

The early idea sounds simple:

Let the agent decide when it should ask for approval.

But that still puts too much responsibility inside the model.

So the safer direction is:

Agent proposes the action
        ↓
The system evaluates the action
        ↓
High-impact actions require approval
        ↓
The action executes
        ↓
The result is recorded

That is the kind of boundary we are building around agent workflows in Xenition — https://xenition.com/.

The lesson was bigger than one product:

The model can decide what action makes sense.

The system should decide whether that action is allowed.

That separation is what starts turning an agent demo into something you can actually trust with real tools.

Audit Trails Matter Too

Approval solves only part of the problem.

You also need to know what happened later.

For example:

  • what tool was called
  • what data changed
  • when it happened
  • who approved it
  • what the agent requested
  • what the final result was

That is why agent systems need an action ledger or audit trail.

If something goes wrong, “the agent did something” is not enough.

You need evidence.

Task-Scoped Permissions Are Even Better

There is another improvement I think agent systems need.

Do not give the agent every permission it may ever need.

Give it what the current task needs.

For example:

Task: Summarize customer feedback

Needs:

  • read support tickets
  • read CRM notes

Does not need:

  • delete customer
  • modify billing
  • publish anything

Task: Prepare a campaign draft

Needs:

  • read campaign data
  • create draft

Does not need:

  • publish campaign
  • charge customers

Same agent.

Different task.

Different authority.

That reduces blast radius dramatically.

Fail Closed

One more rule:

If the permission system fails, the action should stop.

Bad:

policy error
→ continue

Better:

policy error
→ block

This sounds obvious.

But guardrails that fail open are not really guardrails.

A Simple Model That Works

For each action, ask:

What is it?

Read, write, destructive?

What does it affect?

One file? One customer? Production?

Is it reversible?

Can we undo it easily?

Does it leave the system?

Is data being sent externally?

Does it need approval?

If yes, stop before execution.

Is the result logged?

Can we reconstruct what happened?

That simple framework catches a surprising amount.

The Bigger Lesson

When agents only generated text, safety mostly meant:

Don’t say the wrong thing.

Now that agents can use real tools, safety increasingly means:

Don’t do the wrong thing.

That requires more than prompt engineering.

It requires:

permissions

policy enforcement

approval gates

task-scoped access

audit logs

fail-closed behavior

Because:

The model can decide what action makes sense.

The system should decide whether that action is allowed.

Final Thought

Giving AI agents real tools is what makes them powerful.

It is also what makes them dangerous if the authority model is vague.

“Ask before acting” sounds safe.

But it still asks the model to decide when it needs permission.

That is the wrong place to put the boundary.

The safer model is:

Let the agent propose.

Let policy decide.

Let the human approve when the impact is high.

Because the moment an AI agent can change the real world, permission stops being a prompt.

It becomes part of the architecture.

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.