Verisk minimizes alert fatigue with New Relic MCP and AWS DevOps Agent

As a leading strategic data analytics and technology provider for the global insurance industry, Verisk manages business-critical applications where uptime and regulatory compliance are paramount. However, maintaining consistent alert configurations and keeping pace with feature innovations threatened to swamp engineering teams in operational noise.

By pairing the AWS DevOps Agent with the New Relic Intelligent Observability platform, Verisk established an automated, intelligent watchdog. Supported by New Relic’s robust telemetry data platform as the core foundation, Verisk successfully slashed alert noise, reduced mean time to resolution (MTTR), and discovered considerable cost optimizations.

The challenge: Overcoming the noise of scaling infrastructure

As Verisk grew its footprint, high-frequency alerts and flapping errors were firing continuously, creating alert fatigue and the risk of masking critical system issues. To scale effectively without adding dedicated headcount to manage monitoring infrastructure, Verisk needed a fast, automated way to bridge the visibility gap and isolate root causes.

The solution: A synergistic partnership powered by New Relic telemetry

Elliot Markowitz, Vice President of DevOps and QA at Verisk, and his team addressed these hurdles by deploying the AWS DevOps Agent as an orchestrator, seamlessly integrated with New Relic’s Model Context Protocol (MCP) server. While the AWS DevOps Agent operates as the front-line troubleshooting arm, it relies extensively on New Relic’s comprehensive telemetry data and underlying monitoring framework to function. When an application metric breaches a production threshold, the New Relic engine registers the anomaly and triggers a PagerDuty incident. PagerDuty then does two things at once: it alerts the on-call engineer and hands execution to the AWS DevOps Agent. While the engineer is being paged, the agent uses the New Relic MCP server to query the central telemetry database, identifying the precise time of failure to determine why alert conditions were met.

The integration establishes a common semantic language between the two systems. Now, when a New Relic alert triggers an incident in PagerDuty, the AWS DevOps Agent automatically initiates root-cause research before an engineer ever steps in. The findings are sent back to PagerDuty and, at times, Microsoft Teams, pushing updates to engineers’ mobile devices. To reduce background noise, Verisk also configured the system to generate weekly artifact reports identifying high-frequency alert noise. DevOps and QA managers review these reports weekly to mute redundant notifications, combine related conditions, and optimize thresholds. Markowitz emphasized how critical New Relic's context and data mapping are to this automated workflow.

"We continue to rely on New Relic for our metrics and alerts, so having the DevOps agent speak the same language was incredibly important. When a New Relic alert fires, the DevOps agent immediately begins troubleshooting, often before an engineer is engaged. The resulting insights help us resolve issues faster and with less effort. That seamless integration between the two platforms has streamlined our incident investigation process and delivered meaningful value to our workflow."

Impact and results

Dramatic reduction in alert noise and incident fatigue

By implementing weekly noise-filtering artifact reports, Verisk successfully optimized its alert architecture. Total weekly alert volume dropped by approximately 59% from its historical peak.

Accelerated incident resolution (MTTR)

By moving from manual troubleshooting to automated root-cause investigations, Verisk heavily optimized its response velocity. Excluding outlier incidents requiring significant manual engineering changes, Verisk’s mean time to resolution (MTTR) fell sharply by nearly 50%.

Unprecedented alert correlation rates

The telemetry data processing engine excels at grouping linked alerts stemming from a single failure event. Over a 12-week period, the system investigated all baseline incidents and successfully correlated an additional 45% of duplicate incidents back to their original root causes, saving many engineering hours that would otherwise be spent piecing events together manually.

Data ingestion and contract cost optimization

Verisk also used its New Relic integration to audit its own data pipelines and maximize ingestion efficiency. The system flagged hidden redundancies, such as AWS accounts concurrently running both polling and metric streams, excessive successful trace captures, unnecessary health check traces, and other optimizations. Addressing these insights allowed Verisk to immediately reduce its weekly telemetry data ingestion costs by 17%.

Looking ahead

With even more stable infrastructure monitoring and a highly productive engineering team, Verisk is preparing for Phase 2 of its integration. Moving beyond troubleshooting, Markowitz and his team plan to use the New Relic integration as an ecosystem orchestrator connecting directly with CI/CD pipelines and internal APIs. For enterprise leaders looking to navigate the superhuman era of software scale, Verisk’s success shows how combining an advanced automation agent with New Relic’s intelligent telemetry layer is an effective strategy for modern operational excellence.

Starten Sie noch heute kostenlos

Derzeit ist diese Seite nur auf Englisch verfügbar.