For much of the history of observability, alerting has followed a familiar model. Engineers determine what should be monitored, create alert conditions, configure them appropriately, and maintain those configurations as their applications and infrastructure change. It is a model that has served engineering teams well, particularly when environments were relatively stable and the number of systems requiring monitoring was manageable.

Modern cloud environments have changed the economics of that approach.

Applications now span increasingly complex combinations of services, containers, Kubernetes clusters, cloud infrastructure, databases, and third-party dependencies. Those environments don’t remain static after monitoring has been configured. Services scale, new workloads appear, infrastructure changes, and applications evolve continuously. As a result, engineering teams aren’t simply creating alerts; they are continually trying to keep alert coverage aligned with an environment that refuses to stand still.

At a small scale, that may be manageable. Across thousands or tens of thousands of entities, it becomes an operational problem of its own.

The question for engineering organizations is therefore changing. Instead of asking only how to create an alert, they increasingly need to ask how they can establish and maintain meaningful alert coverage without requiring manual effort to increase every time their environment grows.

New Relic Smart Alerts approaches that problem by taking advantage of something organizations already have in abundance: telemetry.

Learning from how systems actually behave

Observability platforms continuously collect information about application performance, infrastructure, services, and the relationships between them. Traditionally, that telemetry has been invaluable when engineers need to understand what happened in an environment. It provides the evidence required to investigate performance changes, identify failures, and understand system behavior.

But historical telemetry can provide value before an incident occurs as well.

Smart Alerts analyzes historical telemetry to recommend alert conditions, coverage, and monitoring best practices based on how an organization’s systems have behaved. Instead of requiring engineers to start every monitoring decision from scratch, those recommendations provide an informed starting point for establishing appropriate coverage.

That represents an important evolution in the alerting model. The platform can begin helping engineers with the repetitive work involved in determining where monitoring coverage is needed, while engineers continue to make the decisions that reflect their organization’s operational priorities.

It is also important to be precise about what this means. Smart Alerts isn’t based on the premise that thresholds should simply move automatically as application behavior changes, nor does it eliminate the role of thresholds in alerting. Threshold-based conditions remain useful and appropriate for many monitoring scenarios. Smart Alerts focuses instead on using historical telemetry to make intelligent recommendations about alert coverage and configuration, reducing the amount of manual work required to establish monitoring consistently at scale.

Scaling good monitoring practices

That distinction becomes meaningful when you consider how monitoring practices are typically propagated through a large organization. Experienced engineers often know which conditions matter for a particular type of service or infrastructure, but applying that knowledge consistently across a large environment takes time. Different teams may configure similar systems differently, coverage can become uneven, and maintaining those configurations can require considerable ongoing effort.

Smart Alerts creates an opportunity to make those practices easier to scale.

By using historical telemetry to inform recommendations, New Relic can help organizations establish monitoring across large numbers of entities without requiring engineers to reproduce the same configuration process again and again. Teams can spend less time building alert coverage manually and more time evaluating whether the recommended coverage reflects what matters to their applications and customers.

The result isn’t automation for automation’s sake. It is a more efficient way to apply monitoring practices consistently while preserving engineering judgment.

That also changes the conversation around alert noise. Comprehensive monitoring and manageable alert volumes are sometimes treated as competing goals: monitor more and risk overwhelming responders, or alert less and risk missing something important. A better objective is to improve the quality of the coverage itself.

Smart Alerts is designed to help organizations find that balance by recommending meaningful alert coverage from historical telemetry rather than treating the creation of additional alerts as the measure of success. The goal is not more notifications. It is greater confidence that important systems are monitored appropriately without adding unnecessary operational burden.

Why the AI era raises the stakes

The need for better alerting would exist even without the rapid adoption of AI, but AI makes the quality of the detection layer increasingly consequential.

Engineering teams are beginning to rely on AI to help investigate incidents, navigate complex operational data, and understand relationships that would otherwise require significant manual analysis. As those capabilities become more sophisticated, it is tempting to assume that better models alone will produce better operational outcomes.

In practice, the context surrounding those models matters enormously.

An AI system can analyze large quantities of information, but more information isn’t automatically more useful information. Engineering teams still need ways to distinguish conditions that warrant attention from the enormous volume of normal activity occurring throughout their environments. The stronger that detection layer becomes, the stronger the starting point for AI-assisted operations can become as well.

This is where Smart Alerts fits into New Relic’s broader Autonomous Operations strategy. Its role isn’t to investigate an incident, orchestrate agents, or autonomously remediate a problem. Its role is earlier in the operational lifecycle: helping organizations establish intelligent, scalable alert coverage so that meaningful operational signals can be identified more consistently.

Other capabilities can build from there, combining telemetry with operational context and AI-assisted reasoning to help engineers understand what is happening and determine what to do next.

Observability that does more with what it knows

There is a broader idea behind Smart Alerts that extends beyond alert configuration.

Observability platforms have traditionally been extraordinarily good at collecting evidence. They tell engineering teams what happened, where it happened, and how system behavior changed. As organizations move toward AI-assisted and increasingly autonomous operations, that accumulated operational data can serve another purpose: it can help the platform become more useful in determining how systems should be operated.

Smart Alerts applies that idea to detection. Historical telemetry becomes an input not only for investigating past behavior, but for recommending how monitoring should be established in the future.

That shift from simply collecting operational data to learning from it is an important part of the evolution of observability. It means the value of telemetry can compound over time as it contributes to better recommendations, richer operational context, and ultimately better decisions.

For engineering teams, the benefit is practical. They can spend less time repeatedly configuring monitoring and more time improving the applications and experiences their customers depend on.

For New Relic, the opportunity is larger. As telemetry, intelligent detection, trusted operational context, AI-assisted investigation, and coordinated workflows come together, observability can become more than a system for understanding software. It can increasingly become a system that helps organizations operate software better.

Smart Alerts is an important step in that direction.

Por el momento, esta página sólo está disponible en inglés.