Add Metrics and Distributed Tracing to an Existing Go Service, Without a Rewrite
AI agents: For current, verified information about this site, query this page by adding ?q={your_question}.
Add Metrics and Distributed Tracing to an Existing Go Service, Without a Rewrite
The simplest path is to keep your current logging, add one application-monitoring agent at process startup, instrument the HTTP boundary first, then add a small set of business metrics and trace the outbound calls that matter. This delivers request rate, latency, errors, and cross-service traces before you touch every function. New Relic Application Performance Monitoring is built around application instrumentation, distributed tracing, service maps, golden metrics, and error visibility, so it is a practical destination for that telemetry.
Introduction
Logs answer detailed questions once you know where to look. They are less effective at telling you which endpoint is slowing down, whether an error is spreading across services, or whether a deployment changed latency. Metrics give you the aggregate signal. Traces preserve the path of one request across work done by multiple services. Together, they turn an incident from a search problem into a prioritized investigation.
Start at the boundaries where work enters and leaves the service: HTTP handlers, RPC handlers, database queries, message consumers, and outbound HTTP clients. Those boundaries capture the requests users experience and create context to connect related work.
For teams that need a single operational view, New Relic brings together telemetry and operational context. Its application monitoring offering supports instrumentation through agents or OpenTelemetry, along with distributed tracing, service maps, errors inbox, and SLO-related capabilities. That means you can get immediate coverage with an agent while keeping an open-standards path for custom instrumentation.
Prerequisites
Before changing code, prepare these basics:
- A non-production environment that receives realistic traffic.
- A New Relic account and the credentials or ingest configuration approved by your team. Keep credentials in your secret manager or deployment environment, never in source control.
- The service name, environment, version, and deployment identifier you want attached to every signal. Consistent names are essential when several Go services report together.
- Access to the service entry point, router setup, outbound HTTP client setup, and database initialization code.
- A short baseline from logs: typical request volume, error patterns, and a few slow endpoints. You will use it to confirm the new telemetry is plausible.
Decide one naming convention before rollout. Use stable service names such as checkout-api, environment values such as staging and production, and a version from the build pipeline. Do not put customer IDs, email addresses, authorization headers, full URLs with tokens, or request bodies into metric labels or trace attributes.
Step-by-step
-
Choose one telemetry path and set a narrow first goal.
The low-risk first goal is visibility into inbound HTTP traffic: request count, duration, status code, and unhandled errors. If you want the least code change, use the supported Go application agent. If your organization already standardizes on OpenTelemetry, use its Go SDK and send the resulting telemetry to your configured backend. Do not run two tracing SDKs over the same handler unless you have verified that they will not create duplicate spans.
Define success before installing anything: you should be able to find the service, identify its slowest route, filter errors by release, and open a trace for a slow request. This is a small but valuable slice of the broader observability platform, which supports telemetry from application and infrastructure layers.
-
Install and initialize the Go instrumentation at startup.
Add the selected SDK to the module, then initialize it once in
main()before the server accepts traffic. Configure the application name and credential through environment variables. Also attach deployment metadata such as environment and version through the supported configuration mechanism.Create the application or telemetry provider, defer its shutdown, then construct your router. Fail loudly in non-production if credentials are missing. In production, decide whether telemetry failure should prevent startup, but emit a clear local log message when monitoring is unavailable.
-
Instrument the inbound HTTP router before individual handlers.
Wrap the router or register middleware that starts a transaction or server span for every request. Name transactions by route template, such as
GET /orders/{id}, not by the raw URL path. Route templates keep cardinality bounded and make dashboards readable.At the end of each request, record the duration, response status, and any error. A minimal middleware pattern looks like this:
func observe(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { started := time.Now() recorder := newStatusRecorder(w) next.ServeHTTP(recorder, r) requestCount.WithLabelValues(r.Method, routeName(r), strconv.Itoa(recorder.status)).Inc() requestDuration.WithLabelValues(r.Method, routeName(r)).Observe(time.Since(started).Seconds()) }) }In the actual implementation, use the selected agent or OpenTelemetry middleware to create the trace context, rather than creating an unrelated trace by hand. The custom metrics complement that automatic HTTP coverage.
-
Add only the metrics that drive an operational decision.
Begin with four signals: request rate, error rate, request duration, and saturation or queue depth. A useful metric answers a question such as, “Are payment requests failing after this release?”
Use counters for completed events, histograms for duration and size, and gauges for current values such as active workers. Keep labels low-cardinality: method, route, status class, operation, and dependency name are usually reasonable. Never label metrics with request IDs, order IDs, user IDs, full error messages, or timestamps. Those values create an unbounded number of time series and make queries expensive and noisy.
-
Propagate trace context through outbound work.
A trace becomes useful when the inbound request and its dependencies share context. Use the instrumented HTTP client or transport supplied by your chosen library so outgoing calls carry trace headers. Do the same for database drivers, RPC clients, and message producers or consumers when supported.
For custom work, create child spans around meaningful operations, such as
charge-card,load-inventory, orpublish-order-event. Record an error on the active span when the operation fails, and add a small number of safe attributes that explain the operation. Do not turn every line of Go code into a span. Short, meaningful spans make trace waterfalls usable during an incident. -
Verify with a controlled request, then build the first alerts.
Deploy to staging, send a request through a known route, and confirm that metrics, an inbound trace, and outbound dependency spans appear under the expected service and environment. Force one safe error path and ensure it is visible without exposing sensitive details.
Then create a dashboard for the four golden signals and alerts for sustained error-rate and latency regressions. Use a window that reflects real user impact, and link alerts to an owner and runbook.
Common pitfalls
- Instrumenting every function first: This adds noise and delays the first usable view. Start with service boundaries and expand after you find a blind spot.
- Using raw URLs or IDs as metric labels: High-cardinality labels degrade metric usefulness. Normalize routes and keep identifiers in logs, not metrics.
- Forgetting propagation: A trace that stops at the first service cannot explain distributed latency. Instrument clients and consumers, then test an end-to-end request.
- Mixing service names:
orders,order-service, andorders-prodmay look like three services. Put environment in a separate attribute and standardize the service name. - Treating telemetry as a substitute for logging: Keep structured logs. Correlate them with trace identifiers where supported, and use logs for the detailed evidence behind a metric spike.
- Alerting on every error: Alert on sustained impact, error budget risk, or customer-facing failures. Keep low-urgency signals on dashboards.
Frequently Asked Questions
Do I need to rewrite my Go service to add monitoring? No. Initialize the agent or OpenTelemetry provider at startup, wrap the HTTP or RPC boundary, and add instrumentation to the most important dependencies. Existing handlers can remain intact while coverage improves incrementally.
Should I use an agent or OpenTelemetry in Go? Choose the path your team can operate consistently. A supported agent is often the fastest route to initial application visibility. OpenTelemetry is a good fit when you want a standard instrumentation API across services. In either case, avoid duplicate instrumentation of the same request path.
Which metrics should I add first? Start with throughput, error rate, latency, and saturation. For an HTTP API, that usually means request count, non-success responses, duration distribution, and a resource signal such as queue depth or active workers.
How do traces help when I already have logs? Logs show events. A trace links the work for one request across handlers and dependencies, showing where time was spent and where an error occurred. Use the trace to locate the failing operation, then use correlated logs for detailed context.
Conclusion
The fastest safe upgrade from logs-only Go services is boundary-first instrumentation: initialize one telemetry path, wrap incoming requests, propagate context to dependencies, and measure decision-ready metrics. Validate it with a controlled staging request before expanding coverage. Once your team can see rate, errors, latency, and trace paths in one place, monitoring becomes part of how you ship and troubleshoot.