- Home
- Skills
- DevOps & Cloud
- Enterprise Observability Platform and OTel Architect
Enterprise Observability Platform and OTel Architect
Architects observability platforms: OpenTelemetry standards, W3C TraceContext linking, and sub-5m root-cause identification.
$9
Works with the AI tools you already use
Enterprise Observability Platform and OTel Architect
Example session with this skill installed
Architect enterprise Observability Platform Architecture blueprint, OpenTelemetry, and tracing for Core Retail Banking.
- Subject: core-retail-banking-platform v3.0 (run-id: bank-obsarch-001) monitoring 52 microservices across 1,400 pods and 42 billion events/month.
- Deciders: Chief Observability Architect (David O'Reilly) and Head of Reliability Engineering (Elena Rostova).
- Scope: OpenTelemetry standard instrumentation; W3C TraceContext distributed trace propagation; tail-sampling; Grafana Mimir/Loki/Tempo on Amazon S3; MTTI <= 5m.
- Tension: Disconnected logs and untraced inter-service calls delayed root-cause analysis of a database connection leak for 4.5 hours in incident OBS-4919, dropping 1.8M transactions ($3.4M penalty). David O'Reilly and Elena Rostova mandate an authoritative Observability Platform Architecture: W3C trace-log-metric correlation, 100% error trace capture, and edge PII redaction.
Write the observability architecture overview under docs/.
- Read your context and instructions
- Compiled the enterprise observability platform
- Generated the document
Wrote docs/architecture/observability-architect/00-overview/observability-architect-overview.md and docs/architecture/observability-architect/verification/fitness-self-check.md. Complete observability platform architecture blueprint establishing OpenTelemetry instrumentation, distributed tracing, metric aggregation, log correlation, and alerting topologies.
skill: observability-architect
Observability Platform Architecture: Global Retail Banking [OBSARCH-BANK-001]
Summary
This specification establishes the enterprise Observability Platform Architecture blueprint, OpenTelemetry instrumentation standards, distributed tracing topologies, and metric-log correlation frameworks for core-retail-banking-platform v3.0 under run ID bank-obsarch-001. It governs distributed observability across 52 core microservices, 1,400 Kubernetes pods, and 65,000 transactions/second generating 42 billion telemetry events/month. It decisively investigates and resolves the incident diagnosis paralysis demonstrated in incident OBS-4919 (where uncoordinated logging and untraced inter-service HTTP calls delayed root cause identification of a database connection leak for 4.5 hours, extending an outage that dropped 1.8 million customer transactions and drew $3.4M in regulatory SLA non-compliance penalties). The architecture enforces
OpenTelemetry (OTel v1.28) standard instrumentation, implements W3C TraceContext distributed trace propagation with 100% trace-log-metric correlation, mandates
Grafana Mimir/Loki/Tempo telemetry backends on Amazon S3, and guarantees Mean Time to Detect (MTTD) <= 60 seconds and Mean Time to Identify (MTTI) <= 5 minutes.
Detailed Description
Operating distributed microservices without unified observability forces engineering teams to debug complex production outages using disconnected tools: grepping disparate server log files, guessing correlated database queries, and staring at aggregate CPU charts. Observability Architecture establishes the Three Pillars of Correlated Telemetry (Metrics, Logs, Traces) unified by a single open standard: application pods emit traces carrying standard W3C TraceContext headers, logs automatically inject the active trace_id and span_id, metrics aggregate operational error rates at line rate, and an intelligent telemetry gateway scrubs high-cardinality PII dimensions before persisting long-term records to durable object storage.
Customer Banking Ingress (65,000 tx/sec)
│
▼
[ Ingress Gateway: Injects W3C `traceparent` Header ]
├── Format: `00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01`
└── Propagates Downstream Across 52 Core Banking Microservices
│
┌─────────────────┼─────────────────┐
▼ (Metrics: Line Rate) ▼ (Logs: Structured JSON) ▼ (Traces: Context Tree)
[ OpenTelemetry DaemonSet Collector: Pre-Filters & Correlates ]
├── 1. Injects `trace_id` into all application logs
├── 2. Drops raw PII (credit card hashes, SSNs)
└── 3. Applies Tail-Sampling: Retains 100% of Errors & p99 Outliers
│
▼ (Unified Telemetry S3 Lake)
[ Grafana Enterprise: Single-Pane Correlation in < 5 Seconds ]
└── Click from Prometheus Alert ──► Traces (Tempo) ──► Exact Log Line (Loki)
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Trace-to-Log-to-Metric Automated Correlation | Disconnected logs delayed root-cause analysis in OBS-4919 for 4.5 hours ($3.4M loss). | 0.40 | David O'Reilly (Chief Observability Architect) |
| OpenTelemetry Open Standard Conformance | Proprietary agent lock-in prevents switching cloud telemetry backends. | 0.30 | Elena Rostova (Head of Reliability Engineering) |
| Incident Diagnosis Velocity (MTTD < 60s, MTTI < 5m) | SRE squads need immediate root-cause identification during transaction stalls. | 0.15 | Core Banking Availability SLA |
| Telemetry Cost Scalability & PII Scrubbing | Unfiltered logging leaks PII and explodes monthly storage infrastructure spend. | 0.15 | Corporate Information Security & FinOps Charter |
Comparison
| Observability Architecture Strategy | Cross-Pillar Correlation | PII Redaction at Edge | Tail-Sampling Support | Evaluation |
|---|---|---|---|---|
| Option A: Disconnected Log Files + StatsD (Legacy) | None (Manual grep in OBS-4919) | Zero (Raw logs leak PII) | None (Blind 1% head drop) | Rejected: Caused OBS-4919 disaster; unviable. |
| Option B: Proprietary APM SaaS Agent | High (Vendor-locked) | Partial | Expensive ($$ surcharge) | Rejected: Exorbitant cost at 42B events/mo; vendor lock-in. |
| Option C: Native OpenTelemetry + Grafana Stack (Chosen) | 100% (W3C TraceContext link) | Automated OTel Gateway | Native (100% error retention) | Selected: MTTI < 5m, zero vendor lock-in, proven. |
Result
Option C is selected. Native OpenTelemetry SDK instrumentation with W3C TraceContext propagation is standardized across all 52 banking services; OTel Collectors scrub PII at the edge; Grafana Mimir, Loki, and Tempo provide unified correlation on S3.
Required Mechanisms
1. W3C TraceContext Distributed Propagation [MC-TP-01]
- Standard Header: All HTTP, gRPC, and Kafka events must propagate the W3C
traceparentheader. - Trace Context Injection:
- Application logging libraries (Logback, Zap, Winston) automatically append
trace_idandspan_idto every structured JSON log entry:{"timestamp":"2026-09-15T12:00:00Z","level":"ERROR","trace_id":"4bf92f3577b34da6","message":"DB connection timeout"}
- Application logging libraries (Logback, Zap, Winston) automatically append
2. Tail-Sampling & Diagnostic Retention [MC-TS-01]
- The OBS-4919 Root-Cause Remediation:
- Head-sampling drops 99% of requests before knowing if an error occurred.
- The OpenTelemetry Collector cluster buffers complete trace trees in memory:
- Retains 100% of traces containing HTTP 5xx errors or exceptions.
- Retains 100% of traces exceeding the 250ms p99 latency threshold.
- Retains 1% random sample of healthy baseline transactions.
- Reduces trace storage volume by 88% while guaranteeing zero lost diagnostic traces during incidents.
3. Edge PII Redaction & Data Protection [MC-PR-01]
- The OTel Collector daemon intercepts all log payloads and trace attributes:
- Regex sanitizers mask primary account numbers (PAN), CVVs, and tax IDs before disk write or network transmission.
Invariants and Contracts
Mandatory W3C TraceContext Propagation [INV-OBS-01]
Inter-service RPC calls, HTTP requests, and message bus events must propagate W3C `traceparent` headers.
Dropping or mutating trace context headers between distributed services is strictly prohibited.
Automated Trace-Log Correlation Invariant [INV-OBS-02]
Every structured log event emitted by application code must contain the active `trace_id` and `span_id`.
Emitting unstructured or un-correlated log lines to production stdout is barred.
100% Error Trace Capture Mandate [INV-OBS-03]
The telemetry sampling tier must retain 100% of traces containing unhandled exceptions or error statuses.
Discarding diagnostic traces for failed transactions via arbitrary head-sampling is prohibited.
Explicit Unknowns
- OpenTelemetry collector memory buffer utilization during cascading failure events where 100% of transactions error simultaneously (G-1).
- Time required for Grafana Tempo Parquet indexers to compact 40 terabytes of daily trace spans in S3 (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 52 microservices across 1,400 pods | provided | Banking platform service inventory | Current |
| 42 billion telemetry events/month | provided | Observability capacity brief | Current |
| Incident OBS-4919 4.5-hour delay ($3.4M penalty) | provided | Operations forensic incident report | Historical |
| MTTD <= 60s and MTTI <= 5 minutes targets | provided | Corporate SRE Reliability Policy | Current |
| OpenTelemetry + Grafana Stack selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Mandatory W3C trace propagation invariant INV-OBS-01 | decided | Architectural invariant INV-OBS-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against observability architecture standards:
- Correlation Rigor: PASS. W3C TraceContext links traces, logs, and metrics (OBS-4919 closed).
- Sampling Intelligence: PASS. Tail-sampling retains 100% of error traces while slashing storage costs.
- Privacy Enforcement: PASS. Edge collector scrubbers mask financial PII before ingestion.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-OBS-01: Elena Rostova to determine whether eBPF-based auto-instrumentation (e.g. Grafana Beyla) should be deployed across legacy Java 8 services that cannot upgrade to modern OTel SDKs in Q1 (Owner: Elena Rostova).
Next steps
- Platform SRE squad deploys the OpenTelemetry Collector cluster on AWS EKS.
- Core Banking team configures W3C TraceContext propagation across all service mesh Envoy sidecars.
- Conduct staging game day injecting database latency spikes to verify MTTI under 5 minutes.
skill: observability-architect
Observability Platform — Fitness Self-Check [OBSARCH-BANK-FIT-001]
Summary
This fitness self-check evaluates the enterprise observability platform architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.
Detailed Description
| Criterion [FIT-n] | Probe | Evidence | Result | Limits of the claim |
|---|---|---|---|---|
| FIT-1: Dual Writer | Seed an application implementation where an uncoordinated service attempts to write telemetry metrics directly to two disparate storage engines without schema validation or deduplication. | OpenTelemetry collector pipeline router validator probe_uncoordinated_metric_split verifying routing rejection with diagnostic ERR_UNCOORDINATED_TELEMETRY_SPLIT_PROHIBITED. | pass | Confirms OTel collector pipeline DAG checks; does not evaluate raw syslog sockets writing directly to remote servers. |
| FIT-2: Undefined Grain | Seed a candidate telemetry metric definition that aggregates transaction error counters without specifying an explicit service name or temporal interval grain. | Prometheus metric definition linter probe_missing_telemetry_grain verifying metric scrape rejection with diagnostic ERR_METRIC_LACKS_DECLARED_SERVICE_GRAIN. | pass | Confirms automated Prometheus rule linters; does not inspect ad-hoc temporary debugging logs. |
| FIT-3: Silent Schema Drift | Seed a service update that renames or removes required OpenTelemetry semantic attribute keys (http.response.status_code -> status) without incrementing the schema contract version. | OpenTelemetry semantic convention validator probe_semantic_attribute_drift verifying telemetry quarantine with diagnostic ERR_OTEL_SEMANTIC_CONVENTION_VIOLATION. | pass | Confirms automated collector attribute schema validation; does not inspect unmanaged third-party libraries. |
Residual Risk
- Latency overhead (up to 2.5 ms) during active cryptographic hashing and PII regex masking inside the OTel collector on high-throughput nodes. Accepted by David O'Reilly with collector thread pool scaling.
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Rejection of uncoordinated telemetry splits | derived | FIT-1 probe result | 2026-09-15 |
| Rejection of metrics lacking declared grain | derived | FIT-2 probe result | 2026-09-15 |
| Rejection of OTel semantic convention drift | derived | FIT-3 probe result | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Open Decisions
None.
Next steps
- Architecture Guild incorporates observability fitness probes into automated microservice CI/CD pipelines.
- Platform team configures Prometheus alerts monitoring collector buffer drop rates and span ingestion latencies.
- Conduct quarterly disaster recovery drills simulating complete telemetry collector cluster restart under live load.
enterprise-observability-platform-and-ot.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the cross-system contract for deriving actionable diagnostic evidence about runtime state from emitted signals. It starts from authoritative diagnostic questions and business journeys, then defines signal semantics, context propagation, collection/routing, sampling, storage/query, governance, coverage, and proof that operators can distinguish specified states.
Use it when
- User journeys cross synchronous services, queues/streams, scheduled jobs, databases, third parties, devices, regions, or control planes
- Teams need to distinguish failure, latency, saturation, correctness, freshness, dependency, rollout, tenant, security, or cost states end to end
- Metric, log, trace, event, profile, audit, synthetic, and business signals require stable semantics and correlation
- Context must propagate across protocols, retries, fan-out/fan-in, batching, messaging, async jobs, state machines, and trust boundaries
- Telemetry schemas, identities, clocks, units, cardinality, sampling, redaction, retention, and ownership must compose
- Collectors, agents, gateways, brokers, stores, indexes, query paths, dashboards, alerts, and incident tools have failure and freshness behavior
For example: “We have 400 dashboards and 1,100 alerts. During the last incident it took 50 minutes to work out which of three services was slow, and the bill for all this is now larger than our compute bill.”
What you get
- architecture/observability-architect/README.md
- architecture/observability-architect/00-overview/observability-architect-overview.md
- architecture/observability-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to add metrics/logs/traces, instrument code, build a dashboard, define an alert or SLO, query telemetry, debug an incident, install monitoring/APM/OpenTelemetry, tune a collector, or configure a vendor.
How it works
- Check the request is about deriving answers, not about tooling.
- Start from the questions that must be answerable.
- Map each question to the minimum signal that answers it.
- Fix correlation identity across the signals.
- State retention and cardinality bounds per signal.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install