Docs › Skills › observability-and-instrumentation

observability-and-instrumentation

Instrument code so production behavior is visible and diagnosable. Covers structured

34 lines

Overview

Observability and instrumentation makes production behavior visible and diagnosable through structured logging, metrics, distributed tracing, and alerting. Every log entry is JSON with correlation IDs, every service gets RED metrics (Rate, Errors, Duration), every resource gets USE metrics (Utilization, Saturation, Errors), and every external call is traced via OpenTelemetry.

Use this skill when shipping features that run in production, investigating production issues with insufficient data, setting up monitoring dashboards, filling gaps discovered during postmortems, or ensuring launch readiness before a public release. Alerts follow Google SRE patterns — symptom-based with documented runbooks, organized by severity from P0 to P3.

The output is instrumented source code with monitoring configuration covering structured logs, RED/USE metrics, distributed tracing span attributes, and severity-graded alerts — all with no secrets or PII in log output.