For years, reliability came down to a simple question: Is the system up and running? 

But modern systems have made reliability more complex than uptime alone. As applications become more distributed and interconnected, a system can be technically available and still deliver a poor experience if it is slow, degraded, or failing somewhere in the customer journey. That, in turn, has shifted the reliability conversation from uptime alone to customer experience, business impact, and a system's ability to detect and respond to failure.

For New Relic, this evolving definition of reliability hits close to home. Our customers rely on New Relic to help them build and operate reliable systems. Airlines, retailers, streaming platforms, and businesses across every industry use our platform to monitor and protect critical digital experiences. When an airline is processing thousands of bookings, when a retailer is mid-way through a flash sale, when a streaming platform is handling a live event and traffic has spiked to 100x, those teams are watching New Relic dashboards in real time. That means they aren't just relying on New Relic to tell them when something has failed. They're relying on us to help them understand what's happening, identify problems early, and respond before those problems become customer-impacting incidents.

That's what raises the stakes for us. When our customers depend on New Relic to monitor and protect critical digital experiences, our reliability becomes part of theirs. That’s why we continuously obsess over improving the reliability bar. Over the last few years, New Relic has improved its reliability by over 50%. To improve our own uptime, we focused on a few key areas that saw a strong return on investment (ROI). It required treating reliability as something we engineer deliberately: measuring continuously, designing  for resilience from the start, and automating the response so that speed of recovery doesn't depend on who happens to be on call. 

Let's talk about the practices that helped us engineer a more reliable platform.

Track what matters

Reliability starts with knowing what to measure. That means tracking not just when things break, but how quickly you catch it and how many customers feel the impact. At New Relic, we treat incident metadata with the same rigor as application performance telemetry, putting every potential degradation in service in NRDB

Across every incident, we focus on four numbers that tell us how well we're actually performing:

  • Time-to-detect (TTD): How long from the moment something goes wrong to the moment we know about it.
  • Time-to-mitigate (TTM): How long from detection to customer impact mitigation.
  • Percent of customers impacted: The impact radius of any given incident, which tells us whether we're dealing with a widespread outage or an isolated failure.
  • Change failure rate: How often changes to our systems result in an incident or customer impact.

Every incident is tagged with a rich set of attributes such as which products and features were affected, which team owned the system that failed, and what triggered the event, whether that was a code change, a configuration update, a hardware failure, or a third-party dependency. All of this lives in NRDB as custom events. Because every service is instrumented with New Relic, we can correlate incident history with service levels, change markers, cloud costs, and other telemetry we collect, then query that data to identify reliability trends. Over time, those patterns can reveal recurring failure modes, reliability hotspots, or categories of change that consistently introduce risk. The data alone may not fix anything. But it makes the patterns visible, and visible patterns are what you act on.

Design for resilience from the start

One of the most common mistakes companies make with reliability is treating it as something to retrofit. You build the system, ship the product, and then try to make it stable after the fact. That approach rarely works and at scale, it becomes nearly impossible, because the surface area of potential failure grows faster than any team can patch it reactively.

At New Relic, resilience is part of the design process from the start. Major architectural changes go through a formal review that includes an explicit audit for failure modes. Can this service handle a sudden surge in backpressure? What happens when a downstream dependency slows down or goes dark? Where are the hot spots, and how do we partition around them?

We run a cell architecture specifically to contain the blast radius of any failure. If a problem surfaces in one cell, traffic can be re-routed to another cell Our routing layer uses an active-active design so that no single system is a critical dependency. These decisions are deliberate design choices made to avoid a single point of failure. The philosophy is simple…anticipate foreseeable failures and build the mitigations in before they happen. Assume something will go wrong, and make sure the system is ready for it when it does.        

But designing for resilience also means accounting for degradation in external services we depend on. These systems can be difficult to control, so the goal is to use telemetry to identify problems early and design the system to contain their impact. Here are some examples of how we apply that principle to dependencies and infrastructure we can't fully control.                      

  1. Designing around external dependencies

    Some external dependencies such as content delivery networks (CDNs) are critical infrastructure for customer-facing systems, including major telemetry and observability providers like New Relic. CDN outages are outside our direct control, and our answer to that is not to accept the risk, but to architect around it.

    We work with multiple CDN providers simultaneously through an active-active multi-CDN architecture and continuously collect telemetry from all the providers.  When that telemetry indicates degradation with one provider, our system automatically reroutes traffic to the healthy provider, helping maintain ingest continuity. By using provider health and performance telemetry to make failover decisions automatically, we can achieve higher reliability than either provider could deliver independently. We’ve noticed a 60% drop in CDN-related customer impact since implementing this architecture.

    The same thinking applies to any external dependency. Instead of trying to control each and every part of your stack, expect the failure and architect around it. And customers can apply this same principle to their own telemetry, using it not only to understand what is happening but also to make informed decisions and trigger automated responses. Turn your telemetry into action and you can contain failures faster and protect the customer experience.

  2. Protecting against cloud infrastructure degradation

    Infrastructure problems don't always start with a complete failure. Cloud infrastructure can degrade in different ways and often shows warning signals before a service fails. One example is a node that begins exhibiting TCP errors as its underlying hardware degrades.

    New Relic monitors TCP metrics across nodes every few seconds as one way to detect these signals. When a node crosses a degradation threshold, an internal automation service moves workloads off that node before it turns into an outage. By catching degradation at the signal level rather than waiting for an incident, this removes an entire class of failures from the table. The underlying hardware gets replaced without any customer impact or anyone getting paged. It's the same principle as the CDN redundancy work, use telemetry data to identify when things start to go wrong and automate the response. The more failure modes you can handle automatically, the less your reliability depends on how quickly someone can respond.

Systematic alerting to reduce time to detect (TTD)

Being responsible for the observability of thousands of companies across virtually every industry gives New Relic a unique vantage point on what makes detection slow. The same two patterns come up again and again, regardless of company size or engineering maturity. 

  • Alerting gaps

    New services spin up faster than alert coverage is added, and no single team has full visibility into what is and isn't covered. Those gaps become the places incidents hide until customers find them. New Relic addresses this with Smart Alerts, which continuously scans for services lacking adequate coverage and surfaces blind spots before they become incidents. Alerting best practices are also encoded as Terraform modules layered on top of the New Relic Terraform provider, so new services inherit a solid alerting baseline by default rather than starting from zero.

  • Missing test coverage on complex pipelines

    Individual services can look perfectly healthy while a customer flow is broken somewhere in a multi-team pipeline. This is one of the harder detection problems to solve since unit or service-level tests only validate what each team can see but no single team owns the entire pipeline.

    One way to catch these failures is end-to-end synthetic monitoring. At New Relic, we use what we call agent-to-query checks. Each check sends test data into New Relic the same way a customer's agent would, then runs a query to confirm that data shows up in the results. That way, each check runs through the full multi-team pipeline, covering every hand-off between teams, so it catches failures at the integration points that individual service tests never reach. If the data doesn't come back as expected, the check fails and the alert fires. Finding which stage broke, and which team owns it, is then a job for the telemetry each service already sends to New Relic.

By focusing on alerting coverage and end to end testing, we have improved our own TTD by 40% over the past year.    

Reducing time to mitigate (TTM) with Change Tracking and the SRE Agent

Faster detection only goes so far if mitigation still depends on someone waking up, reading a dashboard, figuring out what changed, and deciding what to do. At scale, with more services, more teams, and more incidents all competing for the same on-call engineers, that bottleneck becomes a serious reliability problem. The goal is to remove humans from the loop on everything that can be automated, so people can focus on the problems that genuinely require judgment. New Relic tackles the mitigation bottleneck through two capabilities:

  • Change tracking: The first thing any engineer asks when something breaks is what changed. Most production incidents trace back to a recent deployment, configuration update, or infrastructure change. New Relic Change Tracking records those events as markers in NRDB alongside all other telemetry. When an incident fires, the relevant changes show up in the same timeline, without anyone having to cross-reference deployment logs, release notes, or Slack threads.
  • Autopilot: New Relic Autopilot uses AI to help engineers investigate incidents faster. It triages incoming alerts, gathers context such as related entities and recent deployments, and digs through traces, logs, and metrics to identify the likely root cause. It then recommends next steps, such as rolling back a deployment or adjusting resource allocations.

Together, they shorten the time from detection to resolution while reducing how often engineers are pulled into incidents that can be handled automatically.

What this means if you're a New Relic customer

Every practice described in this blog is one we apply to the platform you trust your systems with. The same architectural decisions, the same detection systems, the same automation, they govern our systems the same way we recommend they govern yours.

That’s what makes this more than best practice advice. We depend on New Relic for our own observability, which means consistent uptime is as mission critical for us as it is for you. By following the same practices, you can deploy with confidence knowing our system and yours will stay up and running. 

Pour le moment, cette page n'est disponible qu'en anglais.