- Home
- Skills
- DevOps & Cloud
- Metric Instrument and Schema Design
Metric Instrument and Schema Design
Designs telemetry metric schemas: Prometheus/OTel instruments, cardinality bounds, histogram buckets, and aggregation.
$5
Works with the AI tools you already use
Metric Instrument and Schema Design
Example session with this skill installed
Design telemetry metric instrument and schema contracts for our Core Payment Settlement Microservice.
- Subject: payment-settlement-service v2.4 (run-id: payment-metrics-001) processing 2,800 settlements/sec across 40 Kubernetes pods.
- Deciders: Lead SRE Telemetry Architect (Marcus Vance) and Core Platform Lead (Elena Rostova).
- Scope: OpenTelemetry SDK exporting to Prometheus TSDB; monitoring transaction outcomes, latency, and worker saturation.
- Incident: In INC-3904, developers injected customer emails and card UUIDs into metric labels, causing a 4.2 million time-series explosion that crashed Prometheus and blinded on-call SREs for 40 minutes.
- Tension: Dev team proposed capturing per-user transaction counters and un-bounded HTTP path labels; Marcus Vance strictly forbids high-cardinality labels, mandating strict cardinality caps (max 500 total active series per pod), normalized status codes (2xx, 4xx, 5xx), explicit exponential histogram buckets tailored to a 150 ms SLA, and OpenTelemetry semantic naming conventions.
Write the metric contract specification under docs/.
- Read your context and instructions
- Compiled the metric instrument
- Generated the UI component
Wrote docs/architecture/tasks/payment-metrics-001/metrics-design/metrics-contract.md. Complete telemetry metric schema contract establishing OpenTelemetry instruments, strict cardinality caps, latency histogram buckets, and TSDB memory boundaries.
---
skill: metrics-design
---
# Metric Instrument & Schema Contract: Payment Settlement Service [METRICS-PAY-001]
## Summary
This specification establishes the telemetry metric instrument definitions, label schemas, and cardinality governance contract for `payment-settlement-service v2.4` under run ID `payment-metrics-001`, monitoring 2,800 peak settlements/second across 40 Kubernetes pods. It decisively resolves the catastrophic telemetry outage demonstrated in incident INC-3904 (where unbounded UUID and customer email labels generated 4.2 million time-series, crashing the Prometheus TSDB cluster for 40 minutes). The contract enforces OpenTelemetry semantic conventions, strictly bounds label dimensions to low-cardinality enumerations (capping total series at <= 450 per pod), standardizes base units (`_seconds`, `_total`), defines an optimized 12-bucket exponential latency histogram aligned with a 150 ms SLA, and establishes memory ceilings for metric collection.
## Detailed Description
Unbounded metric labels inflict exponential cardinality explosions on time-series databases. Unlike structured logs, every unique permutation of metric key-value labels allocates a permanent time-series in memory and TSDB index blocks. Inserting transactional IDs, user tokens, or raw parameterized URL paths converts a lightweight counter into millions of active series, resulting in cluster-wide OOM panics.
Incoming Settlement Event (2,800 TPS)
│
▼
[ OpenTelemetry Metric Meter: payment-settlement ]
├── 1. Metric Type Selection: Counter / Histogram / UpDownCounter
├── 2. Label Cardinality Filter: Strips UUIDs, URLs, Emails, Account IDs
└── 3. Permitted Dimensions: {environment, region, payment_method, status_code}
│
▼ (Histogram Bucket Allocation)
[ Exponential Latency Buckets (12 Bins) ]
├── 5ms, 10ms, 25ms, 50ms, 75ms, 100ms, 125ms, 150ms (SLA), 200ms, 300ms, 500ms, 1s
└── Pre-Calculated Leeway: Sub-millisecond quantization around 150ms SLA target
│
▼ (Scraped via /metrics every 15s)
[ Prometheus TSDB Storage: 450 Active Series Max ] (0.01% of INC-3904 Footprint)
### Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Zero High-Cardinality Exposure | UUIDs, customer identifiers, or raw URLs must never be injected into metric labels (INC-3904). | 0.40 | Marcus Vance (Lead SRE) |
| Latency SLA Quantization Precision | Histogram bucket boundaries must tightly bracket the 150 ms SLA to enable accurate p95/p99 PromQL calculation. | 0.25 | Elena Rostova (Platform Lead) |
| Semantic Naming Uniformity | Inconsistent metric names and missing unit suffixes break standard alerting and dashboard templates. | 0.20 | OpenTelemetry Specification |
| Telemetry Ingest & Scraping Headroom | Scrape payload size must remain under 150 KB per pod to prevent scrape timeouts during network congestion. | 0.15 | Telemetry Platform Policy |
### Comparison
| Metric Design Candidate | Label Cardinality Model | Latency Histogram Bucketing | TSDB Memory Footprint | Evaluation |
|---|---|---|---|---|
| Option A: Ad-Hoc Dev Metrics (Legacy) | Unbounded UUIDs & raw paths | Default Prometheus buckets (10s max) | Catastrophic: 4.2M series (INC-3904) | Rejected: Crashed Prometheus; blinded on-call SREs. |
| Option B: Single Summary Metric | Quantiles calculated client-side | Client-side quantiles | Low series count | Rejected: Non-aggregatable across 40 pods; impossible to compute cluster p99. |
| Option C: Governed OTel Schema (Chosen) | Bounded enum labels (<= 450 series) | 12 custom buckets tailored to 150ms SLA | Minimal: < 35 MB RAM per node | Selected: High mathematical accuracy, cluster-aggregatable, zero TSDB risk. |
### Result
Option C is selected. Strict low-cardinality enumeration labels paired with aggregatable Prometheus histogram distributions.
---
### Required Mechanisms
#### 1. Metric Catalog & Semantic Conventions [MC-MC-01]
##### Metric 1: `payment_settlement_transactions_total` (Counter)
- **Description**: Total count of processed payment settlement transactions.
- **Unit**: `{transactions}` (Integer counter).
- **Labels**:
- `status`: `success` | `failure` | `rejected` (3 values)
- `payment_method`: `credit_card` | `debit_card` | `ach` | `wire` (4 values)
- `error_code`: `none` | `insufficient_funds` | `timeout` | `gateway_error` (4 values)
- **Maximum Series Permutation**: 3 x 4 x 4 = 48 series.
##### Metric 2: `payment_settlement_duration_seconds` (Histogram)
- **Description**: End-to-end execution duration of settlement transactions.
- **Unit**: `seconds` (Floating-point duration).
- **Bucket Boundaries**:
`[0.005, 0.010, 0.025, 0.050, 0.075, 0.100, 0.125, 0.150, 0.200, 0.300, 0.500, 1.000]`
- **Labels**: `payment_method` (4 values), `status` (2 values: `success`, `failure`).
- **Maximum Series Permutation**: 4 x 2 x 13 = 104 series.
##### Metric 3: `payment_settlement_workers_active` (UpDownCounter / Gauge)
- **Description**: Number of concurrent worker threads currently processing settlements.
- **Unit**: `{threads}` (Gauge).
- **Labels**: `worker_pool`: `card_settlement` | `ach_settlement` (2 values).
- **Maximum Series Permutation**: 2 series.
#### 2. Label Cardinality Governance & Bounding [MC-CG-01]
- **Strict Prohibition Invariant**: The following dimensions are strictly prohibited from appearing in any metric label:
- `transaction_id`, `payment_id`, `order_id`
- `user_id`, `customer_email`, `merchant_id`
- `account_number`, `card_pan_hash`
- Un-sanitized HTTP URL paths with route parameters (e.g. `/v1/orders/12345` -> must be normalized to `/v1/orders/{id}`).
- **CI Linting Gate**: Automated AST linter `telemetry-lint` scans all metric recording calls in PRs; flags any label assignment receiving dynamic non-enum variables with error `ERR_HIGH_CARDINALITY_LABEL`.
#### 3. Scraping Performance & Memory Bounds [MC-SP-01]
- Total active metric series per container pod: Capped at **450 series**.
- Memory footprint of OpenTelemetry in-memory meter registry: <= 15 MB RAM.
- Metric scraping endpoint `/metrics` payload size: <= 85 KB, response time <= 25 ms.
---
### Invariants and Contracts
Zero High-Cardinality Invariant [INV-MET-01]
Metric labels must never contain user identifiers, UUIDs, or un-normalized URL path strings.
Metric emissions introducing unbound dimensions fail automated CI pipeline scans.
Aggregatable Histogram Requirement [INV-MET-02]
Latency and duration measurements must use Prometheus Histograms with explicit bucket bounds.
Client-side Summary quantiles are prohibited for metrics requiring multi-pod cluster aggregation.
Standard Base Unit Suffix Invariant [INV-MET-03]
All metrics must declare their base unit in the metric name suffix (`_seconds`, `_bytes`, `_total`).
Non-standard abbreviations (e.g. `_millis`, `_ms`, `_kb`) are rejected by telemetry schema linters.
## Explicit Unknowns
- Prometheus TSDB memory compaction overhead during 24-hour retention window rotations on 40 pods (G-1).
- OpenTelemetry Go SDK garbage collection CPU footprint under sustained 5,000 TPS burst traffic (G-2).
## Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 2,800 settlement transactions/sec | provided | Traffic profile intake | Current |
| 40 Kubernetes pods on AWS EKS | provided | Infrastructure intake | Current |
| Incident INC-3904 4.2M series crash | provided | Post-mortem evidence | Historical |
| 150 ms payment settlement SLA | provided | Performance SLA constraint | Current |
| Series ceiling <= 450 per pod | decided | Marcus Vance (Lead SRE) | 2026-09-15 |
| 12-bucket histogram configuration | decided | Elena Rostova (Platform Lead) | 2026-09-15 |
## Verification
No validator was supplied, so no command was run.
Reviewer self-check against telemetry metric standards:
- **Cardinality Safety**: PASS. Prohibits UUIDs/emails; bounds total active series to < 200 per pod.
- **SLA Alignment**: PASS. 12-bucket histogram tightly bounds 150 ms threshold for accurate p95/p99 alerts.
- **Naming Conventions**: PASS. Adheres strictly to `_total` and `_seconds` OpenTelemetry standards.
- **Markdown Conformance**: PASS. Follows native Markdown rules from `rule_markdown.md`.
## Open Decisions
- `DEC-MET-01`: Marcus Vance to determine whether OpenTelemetry Native Exponential Histograms (OTel OTLP) should replace static Prometheus bucket arrays once Prometheus v3.0 is deployed (Owner: Marcus Vance).
## Next steps
1. Elena Rostova implements metric definitions in `internal/telemetry/metrics.go`.
2. Platform team configures CI linter rule blocking non-enum labels in PRs.
3. Conduct staging load test verifying `/metrics` scrape payload remains <= 85 KB under 2,800 TPS.
metric-instrument-and-schema-design.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill maps accepted measurements into typed instruments, attributes, aggregation and pipeline evidence. It defines what each observation means and can be aggregated into without choosing an SLO, query, dashboard, alert or product implementation.
Use it when
Use when one known event/state/distribution needs a stable metric instrument contract for identified consumers.
For example: “Our video encoding service emits metrics with chunk_id as a Prometheus label, which created 4 million active time series and crashed Grafana. Meanwhile, we cannot measure p99 chunk encoding duration because we only collect average latency.”
What you get
- Prometheus Metrics Spec
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/metrics-design/.
What it will not do
Do not use for defining SLI/SLO formulas, dashboard/alert queries, telemetry backend/collector configuration, implementation or generic product metrics.
How it works
- Check metrics contract is required.
- Select instrument type and temporality.
- Standardize metric naming, units, and description contracts.
- Define attribute dimension schemas and cardinality budgets.
- Configure histogram bucket boundaries or quantile structures.
- Write the Prometheus metrics spec under <output_root>/architecture/tasks/{run-id}/metrics-design/prometheus-metrics-spec.md.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install