An operations agent needs more than a model and a tool list. It needs identity, scope, evidence, approval, audit, and a rollback path. Each one is really a form of trust.
When people talk about AI agents in infrastructure, they usually start with capability. Can the agent diagnose an incident? Can it query telemetry? Can it call a tool? Can it fix something?
Useful questions. They also skip the part that decides whether any of this reaches production.
Before an agent touches infrastructure, the enterprise has to know exactly what the agent is, what it can see, what it can do, who authorized it, what evidence it used, what changed, and how to reverse the change if it goes wrong.
Without that, the agent is not operational software. It is an experiment with credentials.
Identity comes first
An agent should not borrow a human administrator account. It should not inherit broad scopes because that made the demo easier. It should have its own identity, with its own scope, TTL, permissions, and audit trail.
This is basic security, and it is easy to skip when a team is moving fast. In production, skipping it is how you lose the ability to answer the only question that matters after an incident: what acted, under whose authorization, for which task, within which limits.
Identity is not just a security control. It is the thing that makes accountability possible at all.
Scope has to be narrow
Production is full of actions that look similar and carry very different risks. Rerunning a synthetic test is not the same as changing a routing policy. Restarting a collector is not the same as modifying a firewall rule. Opening an incident with evidence attached is not the same as touching a payment path.
An agent should not get broad authority. It should get scoped authority, tied to a class of action, an environment, a tenant, a policy, and, where it matters, a human approval.
The unit of trust is not the agent. It is a specific action, under specific conditions, that a team has decided it can defend.
Autonomy is not a switch. It is a set of bounded permissions, earned one class of action at a time.
Evidence has to travel with the action
No operator should be asked to approve a vague recommendation. “Restart this” is not enough to act on.
A useful recommendation carries its evidence. What changed. Which signals are abnormal. Which services are affected. What similar incidents have happened before. What the expected outcome is. What the blast radius is. What the rollback path is.
This matters more than it first appears, because the dangerous failure in operations is not the tool call that errors out. It is the one that reports success while being the wrong thing to have done. A restart can complete cleanly and still have been the wrong restart. Evidence has to speak to whether the action was correct, not only whether it ran.
That matters for approval, and it matters for learning. If the action works, the evidence becomes part of the record. If it does not, that is part of the record too. Over time the system learns which recommendations, models, and classes of action are trustworthy in that specific environment. Evidence is how a recommendation earns the right to be trusted.
Approval belongs in the workflow
In most enterprises, the first write-capable AI workflow should not be autonomous. It should ask for approval before execution.
The agent proposes an action. The system attaches evidence, confidence, scope, and rollback. A human approves it through the workflow the team already uses. The action executes under scoped authority. The result is logged.
This is not less sophisticated than autonomy. It is the practical bridge between AI assistance and AI action, and it keeps human judgment where the risk is highest.
Audit is not optional
Every agent run should leave an audit trail. What it inspected, what it concluded, which tools it called, which action it requested, who approved it, what changed, and what happened afterward.
That record is not just necessary for compliance. It is just as necessary for operations. A team cannot improve what it cannot reconstruct, and it will not trust a system that cannot explain itself after the fact.
Rollback has to be designed, not hoped for
If an action changes production, rollback has to be part of the design. That does not mean every action can be perfectly undone. It means the system should know whether rollback is possible, what the path is, who can approve it, and how the organization responds if the action fails.
Rollback is part of trust. So are freeze windows, budgets, rate limits, and environment-specific policy.
Key takeaway
The hard part of AI operations is not giving an agent tools. It is making every tool call identifiable, scoped, evidenced, approved, logged, and reversible.
The practical path
The safe path starts with read-only investigation. Let the agent observe, correlate, and explain. Then let it recommend. Then let a human approve execution. Only after enough evidence should narrow, policy-gated autonomy enter the conversation.
Each control in this article does the same job. Identity, scope, evidence, approval, audit, and rollback turn a capable model into something an enterprise can put its name behind. That is what we mean by governed AI operations: not machine learning added to a monitoring tool, which is the older AIOps idea, but the machinery that lets AI act inside production and stay accountable. It is what a trust layer has to provide before an agent goes anywhere near production.
Read more from the series on the governance, identity, and evidence model that AI operations depends on.
#AIAgents #Observability #Infrastructure #SRE #Security
About Parlon
Parlon is an infrastructure observability platform built for enterprise teams operating complex, hybrid environments. Parlon combines active synthetic validation, real-time telemetry normalization, and learning-based alerting into a single platform, shifting operations from firefighting to foresight. Learn more at parlon.io.