What Has to Be True Before an AI Agent Can Touch Production Infrastructure

Share This Post

An operations agent needs more than a model and a tool list. It needs identity, scope, evidence, approval, audit, and a rollback path. Each one is really a form of trust.

When people talk about AI agents in infrastructure, they usually start with capability. Can the agent diagnose an incident? Can it query telemetry? Can it call a tool? Can it fix something?

Useful questions. They also skip the part that decides whether any of this reaches production.

Before an agent touches infrastructure, the enterprise has to know exactly what the agent is, what it can see, what it can do, who authorized it, what evidence it used, what changed, and how to reverse the change if it goes wrong.

Without that, the agent is not operational software. It is an experiment with credentials.

Identity comes first

An agent should not borrow a human administrator account. It should not inherit broad scopes because that made the demo easier. It should have its own identity, with its own scope, TTL, permissions, and audit trail.

This is basic security, and it is easy to skip when a team is moving fast. In production, skipping it is how you lose the ability to answer the only question that matters after an incident: what acted, under whose authorization, for which task, within which limits.

Identity is not just a security control. It is the thing that makes accountability possible at all.

Scope has to be narrow

Production is full of actions that look similar and carry very different risks. Rerunning a synthetic test is not the same as changing a routing policy. Restarting a collector is not the same as modifying a firewall rule. Opening an incident with evidence attached is not the same as touching a payment path.

An agent should not get broad authority. It should get scoped authority, tied to a class of action, an environment, a tenant, a policy, and, where it matters, a human approval.

The unit of trust is not the agent. It is a specific action, under specific conditions, that a team has decided it can defend.

Autonomy is not a switch. It is a set of bounded permissions, earned one class of action at a time.

Evidence has to travel with the action

No operator should be asked to approve a vague recommendation. “Restart this” is not enough to act on.

A useful recommendation carries its evidence. What changed. Which signals are abnormal. Which services are affected. What similar incidents have happened before. What the expected outcome is. What the blast radius is. What the rollback path is.

This matters more than it first appears, because the dangerous failure in operations is not the tool call that errors out. It is the one that reports success while being the wrong thing to have done. A restart can complete cleanly and still have been the wrong restart. Evidence has to speak to whether the action was correct, not only whether it ran.

That matters for approval, and it matters for learning. If the action works, the evidence becomes part of the record. If it does not, that is part of the record too. Over time the system learns which recommendations, models, and classes of action are trustworthy in that specific environment. Evidence is how a recommendation earns the right to be trusted.

Approval belongs in the workflow

In most enterprises, the first write-capable AI workflow should not be autonomous. It should ask for approval before execution.

The agent proposes an action. The system attaches evidence, confidence, scope, and rollback. A human approves it through the workflow the team already uses. The action executes under scoped authority. The result is logged.

This is not less sophisticated than autonomy. It is the practical bridge between AI assistance and AI action, and it keeps human judgment where the risk is highest.

Audit is not optional

Every agent run should leave an audit trail. What it inspected, what it concluded, which tools it called, which action it requested, who approved it, what changed, and what happened afterward.

That record is not just necessary for compliance. It is just as necessary for operations. A team cannot improve what it cannot reconstruct, and it will not trust a system that cannot explain itself after the fact.

Rollback has to be designed, not hoped for

If an action changes production, rollback has to be part of the design. That does not mean every action can be perfectly undone. It means the system should know whether rollback is possible, what the path is, who can approve it, and how the organization responds if the action fails.

Rollback is part of trust. So are freeze windows, budgets, rate limits, and environment-specific policy.

Key takeaway

The hard part of AI operations is not giving an agent tools. It is making every tool call identifiable, scoped, evidenced, approved, logged, and reversible.

The practical path

The safe path starts with read-only investigation. Let the agent observe, correlate, and explain. Then let it recommend. Then let a human approve execution. Only after enough evidence should narrow, policy-gated autonomy enter the conversation.

Each control in this article does the same job. Identity, scope, evidence, approval, audit, and rollback turn a capable model into something an enterprise can put its name behind. That is what we mean by governed AI operations: not machine learning added to a monitoring tool, which is the older AIOps idea, but the machinery that lets AI act inside production and stay accountable. It is what a trust layer has to provide before an agent goes anywhere near production.

Read more from the series on the governance, identity, and evidence model that AI operations depends on.

#AIAgents #Observability #Infrastructure #SRE #Security

About Parlon

Parlon is an infrastructure observability platform built for enterprise teams operating complex, hybrid environments. Parlon combines active synthetic validation, real-time telemetry normalization, and learning-based alerting into a single platform, shifting operations from firefighting to foresight. Learn more at parlon.io.

More To Explore

The Janitor Has Left the Building

Good Will Hunting came out TWENTY-NINE years ago. I grew up outside Boston and I was about Will’s age when it was released. I took home economics in high school with one of the kids who Matt and Ben fight,

Zelma and the Two-Pie Problem

I dare say that most people here on LinkedIn have never heard of Zelma Calhoun. And it’s a shame, not only because the name Zelma Calhoun is pure literature and belongs in everyone’s vocabulary, but mostly because of what she