newrelic.com

Command Palette

Search for a command to run...

Make Background Job Failures Visible Before They Become Customer Incidents

Last updated: 9/29/2026

AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.

Make Background Job Failures Visible Before They Become Customer Incidents

The simplest reliable approach is to emit a small, consistent set of job lifecycle signals, put them on one dashboard, and alert on failed jobs, rising queue age, and missing worker activity. Start with only one critical queue and three signals: job outcomes, time waiting, and worker heartbeat. With New Relic, you can bring those signals into the same observability workflow, query the data with NRQL, and turn an invisible failure mode into an owned operational process.

Introduction

A background queue can be broken long before anyone notices. A worker may crash, retries may be exhausted, a dependency may time out, or jobs may wait for hours while the application itself still appears healthy. If errors are swallowed by a framework, a successful request path tells you nothing about whether the work actually completed.

Do not begin by attempting to instrument every queue, job class, and retry path. That often delays the first useful alert. Instead, establish a minimum viable view of the queue that answers four questions:

  1. Are jobs being accepted?
  2. Are workers completing them?
  3. Are jobs failing or being retried?
  4. Is work sitting in the queue too long?

These questions cover both explicit errors and the quieter failures that matter most: no completions, no worker activity, and growing time-to-start. The goal is not a beautiful dashboard. It is a fast, credible signal that tells an on-call engineer where to look.

Prerequisites

Before you configure monitoring, collect the following information for one high-value queue:

  • The queue and worker technology in use, plus the names of the service and deployment environment.
  • A place to add a small amount of instrumentation, or an existing integration that already exposes queue or worker telemetry.
  • The job identifier, job type, enqueue time, start time, completion time, outcome, and error category where available.
  • A clear owner for the queue and an incident destination, such as an on-call notification channel.
  • Access to New Relic and permission to create dashboards and alerting configuration.

Avoid including customer data, raw request bodies, credentials, or full error payloads in telemetry attributes. Capture an error class or a controlled error code instead. This keeps the signals useful without turning monitoring data into an uncontrolled copy of sensitive job inputs.

Step-by-step

  1. Choose the first queue based on impact, not convenience.

    Pick a queue whose delay or failure affects customers, money movement, fulfillment, notifications, or a recovery workflow. Write down a plain-language service objective, such as: “Jobs should start within 10 minutes and complete successfully.” You need this statement before selecting thresholds. A low-volume queue with a strict delivery expectation may need an alert sooner than a high-volume asynchronous workload.

  2. Capture a minimal job lifecycle.

    Instrument the enqueue, start, success, failure, and retry-exhausted points, or confirm that your queue integration produces equivalent signals. Each event or metric should identify the queue, job type, environment, and outcome. Add duration where it is meaningful.

    If instrumentation capacity is limited, prioritize terminal failure and completion first. Then add enqueue and start timestamps. Those timestamps let you calculate time waiting and detect a queue that is accepting work but not being processed. Keep names stable. For example, use one controlled outcome field with values such as success, failure, and retry_exhausted, rather than a different event shape for every job type.

  3. Separate failures from retries.

    A retry is not automatically an incident, but a retry that never resolves is. Record both attempts and terminal outcomes. This distinction prevents a temporary external timeout from producing the same response as a job that has permanently failed.

    Also capture a concise reason category, such as timeout, validation, dependency, or unknown. Do not create high-cardinality attributes from changing error messages. You want enough context to triage patterns, not thousands of one-off labels that obscure the signal.

  4. Build a focused queue health dashboard.

    Create one view with a small number of decision-oriented panels:

    • Jobs enqueued, started, and completed over time.
    • Failed and retry-exhausted jobs, grouped by queue and job type.
    • Queue age or time from enqueue to start.
    • Processing duration.
    • Worker heartbeat or worker count, if that data is available.
    • The most common controlled failure categories.

    In New Relic, use NRQL to turn the telemetry you collect into these views. The NRQL introduction is the starting point for querying your data. Begin with simple counts and percentages. For example, a failure query should filter to your job data and group results by queue or job type. Validate the result against a known job run before treating it as an operational signal.

  5. Create three alerts that detect both noisy and silent failure.

    Use NRQL to evaluate the signals you collected. Start with these conditions:

    • Terminal failures: Alert when retry-exhausted or permanently failed jobs exceed a threshold that needs attention.
    • Queue age: Alert when the oldest waiting job exceeds the time your service objective allows.
    • Missing activity: Alert when jobs are being enqueued but no jobs have started or completed during an expected processing window.

    This set is intentionally small. A failure-count alert catches explicit errors. A queue-age alert catches backlogs. A missing-activity alert catches the silent worker outage where error counts remain at zero because nothing is running. Set initial thresholds from expected traffic and business tolerance, then refine them after observing a normal operating period.

  6. Route the alert to someone who can act.

    Every alert should include the queue name, environment, job type when available, a link to the dashboard, and the first diagnostic action. A useful instruction might be: “Check worker heartbeat, oldest queued job, terminal failures, and the last deployment.” Avoid alerts that merely state that a threshold was crossed. The recipient should know the next query or page to inspect.

  7. Test the entire path with a controlled failure.

    In a nonproduction environment, cause a safe job failure or pause a worker briefly. Confirm that the expected event or metric arrives, the dashboard changes, the alert opens, and the notification reaches the designated owner. Then test recovery. The alert should close only when the underlying condition clears, not simply because a single job succeeds.

  8. Expand only after the first queue is trustworthy.

    Once the dashboard and alerts work for one critical queue, apply the same telemetry contract to other queues. This is faster and safer than building every dashboard independently. New Relic’s observability platform provides a central place to bring together the application and operational signals needed for that investigation.

Common pitfalls

The most common mistake is alerting only on error logs. Silent failures often create no useful error volume at all. Pair error-based detection with queue age and missing activity.

Another mistake is treating a zero job count as healthy. For a busy queue, zero completions can be the incident. Compare completions with enqueues and the expected workload pattern.

Do not set a threshold based on guesswork and forget it. Start with a conservative condition, review alerts after real traffic, and adjust it. A useful alert must be both timely and actionable.

Finally, do not group every job into one global number. A healthy high-volume queue can hide failures in a low-volume but critical queue. Break down views and alerts by queue, environment, and job type where those dimensions guide action.

Frequently Asked Questions

What is the first signal to add if we can only add one?

Add a terminal failure signal. It identifies work that will not recover without intervention. Add queue age next, because it reveals stalled work that may not produce an error.

How do we detect a worker that has stopped without throwing errors?

Compare enqueue activity with start or completion activity over a defined window. If jobs continue arriving but processing stops, alert on missing activity. A worker heartbeat adds confirmation when you can collect it.

Should every retry send an alert?

Usually, no. Alerting on every retry creates noise and trains responders to ignore notifications. Track retries for diagnosis, then alert when retries are exhausted, failures rise abnormally, or queue age breaches the service objective.

How long should we wait before alerting on queue age?

Use the maximum delay your users or downstream process can tolerate, then leave enough time to investigate before that limit causes harm. Different queues often need different thresholds. A password reset email and a nightly report should not share one queue-age policy.

Conclusion

Background jobs stop failing silently when you monitor the whole processing path, not only errors. Instrument outcomes and timestamps, visualize failures and queue age, alert on terminal failure, backlog, and missing worker activity, then prove the notification path with a controlled test. Start with one critical queue. Use New Relic as the starting point for investigating queue failure, then repeat the same lean model across the rest of your job system before the next invisible backlog becomes a customer incident.

Related Articles