For years, the observability industry has focused on one primary objective: collecting more telemetry. Organizations invested in metrics, logs, traces, events, and increasingly sophisticated monitoring platforms to improve visibility into their systems. That investment has paid off. Modern engineering teams can now observe virtually every component of a distributed application, from Kubernetes clusters and cloud infrastructure to APIs, databases, and individual microservices.
Yet despite this unprecedented visibility, one fundamental challenge remains. During a production incident, engineering teams often spend far more time trying to understand what is happening than they do actually fixing the problem.
The issue is not a lack of information. In many organizations, there is almost too much of it. A single customer-facing issue can generate dozens of alerts across applications, infrastructure, cloud services, databases, networking components, and security systems. Each alert may be accurate. Each may point to a legitimate symptom. But responders are still left with the difficult task of determining which alerts are connected, which are simply secondary effects, and where the investigation should actually begin.
This is one of the defining operational challenges facing modern engineering organizations. The next evolution of observability is no longer about collecting more telemetry. It is about transforming that telemetry into operational understanding.
Operational understanding is the ability to move beyond isolated signals and see the complete operational picture. It allows engineering teams to quickly recognize how seemingly independent events relate to one another, understand the scope of an incident, and begin resolving the underlying problem with confidence. As organizations continue adopting AI-assisted operations and, ultimately, Autonomous Operations, this understanding becomes increasingly important. AI systems, like engineers, can only make good decisions when they begin with a clear and coherent picture of what is actually happening.
That is why New Relic believes operational understanding represents the next critical layer in the evolution of intelligent observability.
The Hidden Cost of Fragmented Operations
Modern software systems were never designed to fail in simple ways. Applications today are composed of hundreds of interconnected services running across multiple cloud providers, Kubernetes clusters, containers, APIs, databases, and third-party services. While this architecture enables organizations to innovate faster and scale more effectively, it also creates an entirely new level of operational complexity.
Consider what happens when a single customer-facing service begins experiencing elevated latency. Within seconds, monitoring systems may detect increased response times, infrastructure utilization, failed transactions, application exceptions, Kubernetes resource constraints, database performance degradation, and network anomalies. Each monitoring rule behaves exactly as intended by generating an alert. From the perspective of the monitoring platform, everything is working correctly.
From the perspective of the engineer receiving those alerts, however, the situation is very different.
Instead of immediately understanding the operational problem, the responder is presented with a growing collection of independent notifications that must be evaluated, compared, and organized before meaningful investigation can even begin. Different teams may receive different alerts describing different symptoms of the same underlying issue. Multiple engineers often begin separate investigations, each working from only a portion of the available information. Valuable time is spent assembling the story rather than solving the problem.
This hidden operational tax is rarely measured directly, but it affects nearly every engineering organization. It increases mean time to resolution, duplicates investigative effort, slows collaboration, and consumes engineering resources that could otherwise be spent improving reliability or delivering new customer capabilities.
Many organizations understandably refer to this as alert fatigue. In reality, the challenge runs much deeper.
The problem is not simply that there are too many alerts. The problem is that there is insufficient operational understanding connecting those alerts into a meaningful picture.
Better Signals Are Only Part of the Answer
Over the past several years, the observability industry has made significant progress toward improving the quality of individual alerts. Dynamic thresholds, anomaly detection, machine learning, and AI-assisted recommendations have helped organizations reduce unnecessary notifications while creating more meaningful operational signals.
This represents an important step forward. Better alerts allow engineering teams to focus on events that truly require attention instead of constantly filtering through false positives and low-value notifications.
However, improving individual signals solves only part of the larger operational challenge.
Even when every alert is accurate, responders must still determine how those alerts relate to one another. Does a database alert explain the application slowdown, or is it simply another symptom? Are infrastructure alerts describing the same customer impact already identified elsewhere? Which signals represent the underlying problem, and which simply reflect downstream consequences?
These questions continue to require significant human reasoning because most monitoring systems present alerts individually rather than collectively.
This distinction is increasingly important within New Relic’s Intelligent Observability strategy.
Smart Alerts focuses on improving the quality of operational signals. By helping organizations build better alert conditions and intelligently identify monitoring gaps, Smart Alerts reduces unnecessary noise while increasing confidence that important operational events are detected quickly and accurately.
That stronger signal foundation is essential. But trusted signals alone do not automatically create trusted understanding.
Engineering teams still need a way to assemble those signals into a coherent operational picture before investigation begins.
Building the Foundation for Autonomous Operations
This is where the next evolution of autonomous operations begins.
Rather than asking responders to manually determine which alerts belong together, the observability platform itself should help assemble the operational picture. Related alert conditions should no longer appear as isolated events requiring independent investigation. Instead, they should be intelligently correlated into actionable issues that represent the broader operational situation confronting the engineering team.
Operational context is valuable for today’s engineering teams, but its importance becomes even greater as organizations adopt AI throughout their operational workflows.
Artificial intelligence excels at analyzing information, identifying patterns, and accelerating decision-making. However, AI systems face the same challenge as human responders when presented with fragmented operational data. If the operational picture must first be reconstructed from dozens of independent alerts, neither humans nor AI agents can reason efficiently about the problem.
Autonomous Operations therefore requires more than intelligent agents. It requires progressively higher-quality operational intelligence at every stage of the workflow.
New Relic views this progression as a continuous evolution rather than a collection of independent products.
Telemetry captures what is happening across applications, infrastructure, services, and cloud environments. Smart Alerts transforms that telemetry into trusted operational signals that deserve attention. Compound Alerts builds upon those trusted signals to create operational understanding by correlating related alert conditions into actionable issues. Ground Truth enriches those issues with trusted operational context drawn from across the technology environment. Finally, Autopilot uses that understanding and context to assist engineers with investigation today while laying the groundwork for increasingly autonomous operational workflows in the future.
Each layer improves the quality of operational intelligence before passing it to the next.
Seen individually, each capability delivers meaningful customer value. Seen together, they represent a new approach to intelligent observability—one that focuses not simply on collecting more operational data, but on continuously improving the quality of operational understanding.
The Next Chapter of Intelligent Observability
Observability has already transformed the way organizations understand their systems. The superhuman era is where the next opportunity is to transform how they understand operational problems.
As environments continue growing in complexity and AI becomes a foundational component of engineering operations, the organizations that succeed will not necessarily be those collecting the most telemetry. They will be the ones that can transform that telemetry into understanding, context, and ultimately action.
That is why New Relic sees Compound Alerts as more than another alerting capability. It represents an important step in the industry’s evolution from signal collection to operational understanding. By helping engineering teams begin every investigation with a clearer picture of the problem, Compound Alerts reduces operational friction today while establishing the foundation for the Autonomous Operations systems of tomorrow.
The future of observability will not be defined by the volume of data we collect. It will be defined by how effectively we transform that data into understanding that people, and increasingly AI, can trust, reason over, and confidently act upon. That is the future New Relic is building toward, one intelligent layer at a time.
本ブログに掲載されている見解は著者に所属するものであり、必ずしも New Relic 株式会社の公式見解であるわけではありません。また、本ブログには、外部サイトにアクセスするリンクが含まれる場合があります。それらリンク先の内容について、New Relic がいかなる保証も提供することはありません。