Intelligence Is a Feature. Reliability Is a Discipline.

Share This Post

AI platforms will become essential production infrastructure. That makes observability, resilience, and accountability more important, not less.

Intelligence Is a Feature. Reliability Is a Discipline.

Last week, Anthropic’s platform saw elevated errors across multiple models for a stretch of the afternoon. By its own status page, the impact window ran about four hours, with most of the errors concentrated in the first hour. Predictably, the internet had feelings, some about the future of AI, a few about the future of civilization.

Let me lower the temperature. This is not an argument against Claude, against Anthropic, or against AI. We use Claude across our own business, for prototyping, research, content, GTM agents, and analysis, and it earns its place. Downtime is not a sin. Every serious platform has outages, and the people running these ones care about reliability as much as anyone in our industry.

The interesting question is not whether AI platforms go down. It is what the rest of us have quietly built on top of them while assuming they wouldn’t.

Downtime is normal. The expectations around it aren't.

Spend a few minutes on the public status pages and the picture is consistent. Anthropic’s own page shows yesterday’s outage was not the first disruption this month, or even this week.

Each publish their own histories, and each one records incidents. That is not a knock on any of them. It is what running production systems at scale looks like. Every platform I have depended on, including the ones my own teams built, has a status page with yellow days on it.

It helps to be precise about what a reliability target promises, because “three nines” gets used as if it meant “always on.” Three nines is 99.9% availability, which still allows for roughly 8.8 hours of downtime a year, or about 43 minutes a month. It is a strong target. It is not zero. And providers report against such targets in good faith over different windows, services, regions, and definitions, so their published percentages are not directly comparable and should not be read as a leaderboard.

The durable point holds regardless of the figures, which are accurate only as of the moment you read them. Incidents are a normal property of every provider, so upstream availability is an architectural dependency, not a background assumption. The math does not change just because the workload got smarter.

The dependency beneath the product

This is where it gets commercially interesting. An entire generation of companies is being built on a small number of model providers, and their product, customer experience, margins, valuation, and investor returns can rest in part on decisions they do not make and cannot see coming.

Availability is the visible one, because it trends on the day it fails. The quieter dependencies matter more. Token pricing moves. Rate limits and capacity tighten when you least want them to. Models get updated and behavior shifts in ways your evals didn’t catch, or get deprecated and the version your product was tuned around is scheduled for retirement. Safety and policy boundaries change. Regional availability and data handling terms decide which customers you can serve, and which your security team will let you serve at all. And when something breaks, your ability to diagnose it is bounded by how much visibility the provider gives you.

None of these are reasons to avoid building on AI. They are reasons to know what you are standing on. A startup whose reliability, cost, and roadmap are partly inherited from an upstream vendor has a business model with a coupling it may not have priced. That is a strategy question before it is an engineering one.

A model is not a service architecture

Access to a capable model is not the same as a dependable production system, and the gap between the two is where operational maturity lives.

A model gives you intelligence on demand. A service architecture gives you the boring, load-bearing parts around it: defined support, contractual commitments, security controls, auditability, predictable behavior, escalation paths, and a human who is accountable when the thing misbehaves. The demo shows you the first. Production runs on the second. As we wrote in AI Operations Has a Trust Problem, Not an Intelligence Problem, the assistant is not the architecture. The frontier is not whether a model can answer a question. It is whether an enterprise can govern the system around the model well enough to put its name behind the result.

Why observability matters more, not less

There is a tidy narrative that AI will absorb observability too: why watch your systems when a model can tell you what is wrong. It has the direction backwards. Every layer of automation and abstraction adds things that can fail silently and widens the gap between a failure and the symptom you notice. When your product depends on an external inference endpoint, a regional route, a rate limit, and a model version you do not control, you have more dependencies to understand, not fewer. Yesterday’s outage is a small example. Teams that could watch their inference path degrade knew before their users did. The rest found out from a support queue.

That is the case we make in AI Observability: AI infrastructure has no external observer by default, because the model cannot see its own outage.

Parlon’s role here is narrow, and worth stating plainly so it is not mistaken for something larger. We do not make any AI provider infallible; no outside party can. What we provide is an independent view of the infrastructure, networks, endpoints, and external services that AI-enabled workflows depend on, so that when something slows down an operator can tell an application failure from network degradation, a provider incident from a routing problem, and a dependency issue from ordinary noise. Knowing which of those you are looking at is most of the work at 2am.

We do not make any AI provider infallible; no outside party can. What we provide is an independent view of the infrastructure, networks, endpoints, and external services that AI-enabled workflows depend on.

A prototype is not an operating model

You can now vibe-code a plausible operations tool in an afternoon, and that is a real advance. For experiments, internal workflows, and one-off investigations, fast beats formal, and the teams working this way are often the sharpest in the building.

It stops being enough the moment something has to not fail. A tool built over a weekend does not arrive with an SLA, a support contract, a security review, an audit trail, or anyone on the hook when it returns the wrong answer mid-incident. As we argued in What Has to Be True Before an AI Agent Can Touch Production Infrastructure, the hard part of AI operations is not giving an agent tools. It is making every action identifiable, scoped, evidenced, approved, logged, and reversible. A capable model without those controls is not operational software. It is an experiment with credentials.

That is the old, unglamorous line between a compelling demonstration and a dependable service. Buyers who run mission-critical systems learned it long ago, which is why, as I wrote in Enterprises Still Want AI. They Want It on Their Terms, they are past the demo and asking the practical questions: where does the data go, what does this cost at scale, who approves an action, and who picks up the phone when it breaks.

AI and SaaS will converge, not compete

The framing of AI versus SaaS is a false choice, and I suspect it will read as quaint in a couple of years.

AI is becoming the reasoning and action layer. It will supplement observability, accelerate investigation, sharpen correlation, and power recommendations and remediation. We are building toward that deliberately.

Production SaaS remains the durable operational layer underneath: data collection, monitoring, historical context, permissions, governance, workflow, support, and accountability. Model capability is advancing faster than the operational, governance, and accountability systems built around it, and that gap is where the risk concentrates. The need for a dependable substrate grows as the models improve, because you are handing the intelligence higher-stakes work. The pattern that wins is the two layers together, the smart one doing more and the accountable one making it safe to.

The point

Adopt AI, enthusiastically. Put it into research, prototyping, analysis, and the operational workflows where it earns its keep. Then engineer for the world as it is, not the world in the keynote.

Every dependency can fail, and the intelligent ones are no exception. The mature move is not to distrust AI. It is to build so that when a provider has a bad afternoon, and periodically one will, you find out from your own instruments, not from your customers, and you already know what to do next.

Intelligence is a feature. Reliability is a discipline. The companies that treat them as different problems, and solve both, are the ones that will still be standing after the next status page turns yellow.

About Parlon

Parlon is an infrastructure observability platform built for enterprise teams operating complex, hybrid environments. Parlon combines active synthetic validation, real-time telemetry normalization, and learning-based alerting into a single platform, shifting operations from firefighting to foresight. Learn more at parlon.io.

More To Explore

The Janitor Has Left the Building

Good Will Hunting came out TWENTY-NINE years ago. I grew up outside Boston and I was about Will’s age when it was released. I took home economics in high school with one of the kids who Matt and Ben fight,

Zelma and the Two-Pie Problem

I dare say that most people here on LinkedIn have never heard of Zelma Calhoun. And it’s a shame, not only because the name Zelma Calhoun is pure literature and belongs in everyone’s vocabulary, but mostly because of what she