- Home
- Skills
- Testing & Debugging
- Enterprise Continuous Testing Platform Architect
Enterprise Continuous Testing Platform Architect
Architects testing platforms: 70/20/10 test pyramids, Consumer-Driven Contract testing via Pact, and mutation testing.
$9
Works with the AI tools you already use
Enterprise Continuous Testing Platform Architect
Example session with this skill installed
Architect enterprise Testing Platform Architecture blueprint, balanced test pyramids, and contract testing for Engineering.
- Subject: enterprise-continuous-testing-platform v3.0 (run-id: corp-testarch-001) executing 12,000 test runs/day across 65 microservices.
- Deciders: Chief Enterprise Architect (David O'Reilly) and VP of Software Quality & Governance (Elena Rostova).
- Scope: 70/20/10 Test Pyramid distribution; Consumer-Driven Contract Testing via Pact Broker; Pitest mutation testing (>= 80% score); JaCoCo 85% branch coverage gates.
- Tension: An inverted test pyramid dominated by flaky manual E2E tests allowed a settlement rounding bug to escape in incident TST-4919, corrupting 650,000 ledger accounts ($4.2M penalty). David O'Reilly and Elena Rostova mandate an authoritative Testing Platform Architecture: balanced test pyramids, automated contract verification, and mutation test gating.
Write the testing architect overview under docs/.
- Read your context and instructions
- Compiled the enterprise continuous testing
- Generated the document
Wrote docs/architecture/testing-architect/00-overview/testing-architect-overview.md and docs/architecture/testing-architect/verification/fitness-self-check.md. Complete testing platform architecture blueprint establishing automated test pyramids, consumer-driven contract tests (Pact), mutation testing, and non-bypassable CI/CD quality gates.
skill: testing-architect
Testing Platform Architecture: Enterprise Continuous Quality [TESTARCH-CORP-001]
Summary
This specification establishes the enterprise Testing Platform Architecture blueprint, automated test pyramid distributions, consumer-driven contract testing frameworks, mutation testing standards, and non-bypassable CI/CD quality gates for enterprise-continuous-testing-platform v3.0 under run ID corp-testarch-001. It governs automated software quality assurance across 65 product microservices, 450 software engineers, and 32 engineering squads executing 12,000 automated test suites daily. It decisively investigates and resolves the defective artifact escapes and regression outages demonstrated in incident TST-4919 (where relying on an inverted "ice cream cone" test pyramid dominated by slow, flaky manual end-to-end UI tests allowed a critical payment settlement rounding error to bypass testing, escape to production, corrupt 650,000 merchant ledger balances, and incur $4.2M in regulatory fines and manual audit corrections). The architecture enforces a strict Test Pyramid distribution (70% Unit, 20% Integration/Contract, 10% End-to-End), mandates Consumer-Driven Contract Testing via Pact across all inter-service boundaries, institutes automated Mutation Testing via Pitest (target mutation score >= 80%), and establishes
non-bypassable build-breaker quality gates.
Detailed Description
Relying on manual testing or an inverted "ice cream cone" test distribution—where teams write few unit tests and depend heavily on end-to-end (E2E) UI test automation—guarantees slow delivery cycles, high test flakiness, and frequent production defect escapes. E2E tests are slow, expensive to maintain, and fail non-deterministically due to network jitter, masking genuine regression bugs. Testing Platform Architecture establishes a
Balanced, Governed Test Pyramid: the vast majority of tests execute as lightning-fast, isolated unit tests ($< 5\text{ ms}$ per test); inter-service seams are validated using tool-agnostic Consumer-Driven Contract Testing (Pact), enabling independent service deployment without deploying the entire environment; mutation testing proves test suite effectiveness by injecting synthetic code mutations; and automated CI gates block any pull request that fails quality or coverage baselines.
Engineering Commit & Pull Request Stream (12,000 Test Runs / Day)
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ Governed Test Pyramid Architecture [TESTARCH-CORP-001] │
│ ├── Tier 1: Unit Tests (70% Volume, In-Memory, Sub-5ms per Test) │
│ ├── Tier 2: Consumer-Driven Contract Tests (Pact Seams, 20% Volume) │
│ └── Tier 3: Isolated End-to-End Synthetic Canaries (10% Volume) │
└──────────────────────────────────────┬──────────────────────────────────────┘
│
┌─────────────────────────────┼─────────────────────────────┐
▼ (Mutation Quality: Pitest) ▼ (Contract Safety: Pact Broker)▼ (Coverage Gate: JaCoCo)
[ Mutation Score >= 80% ] [ 100% Contract Match ] [ Branch Coverage >= 85% ]
├── Kills Synthetic Mutants ├── Can-I-Deploy Verified ├── Zero Untested Branches
└── Proves Test Assertion Depth └── Zero Breaking API Changes └── High Quality Codebase
│
▼ (Automated Promotion Oracle)
[ Release Pass Attestation Emitted: Incident TST-4919 Defect Permanently Barred ]
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Non-Bypassable CI/CD Quality Gating | Escaped settlement bugs caused incident TST-4919 ($4.2M fine, ledger corruption). | 0.40 | Elena Rostova (VP Software Quality & Governance) |
| Consumer-Driven Contract Testing (Pact) | Enables 32 squads to deploy independently without full-environment E2E freezes. | 0.30 | David O'Reilly (Chief Enterprise Architect) |
| Test Pyramid Balance (70/20/10 Ratio) | Eliminates flaky, slow E2E test suites; reduces PR test duration from 45m to < 4m. | 0.15 | Core Developer Experience Charter |
| Mutation Testing Rigor (Score >= 80%) | Code coverage alone hides tests that lack assertions; mutation testing proves efficacy. | 0.15 | Architecture Review Board (ARB) Policy |
Comparison
| Testing Architecture Model | Flakiness Rate | PR Test Feedback Cycle | Contract Verification | Evaluation |
|---|---|---|---|---|
| Option A: Inverted Ice Cream Cone (Legacy) | High (18% flakiness in TST-4919) | 45 Minutes (Stalls devs) | Manual / Post-Deploy | Rejected: Caused TST-4919 disaster; unviable. |
| Option B: 100% Monolithic E2E Staging Environment | Extreme (Environment contention) | 3.5 Hours (Nightly batch) | Fragile | Rejected: High infrastructure cost; environment bottlenecks. |
| Option C: Balanced Pyramid + Pact (Chosen) | Near-Zero (< 0.1% Flakiness) | < 4 Minutes (Parallelized) | Automated Pact Broker Gates | Selected: Sub-4m feedback, zero escapes, proven. |
Result
Option C is selected. A balanced Test Pyramid (70/20/10) with Consumer-Driven Contract Testing via Pact is standardized; pull requests require an 85% branch coverage floor and 80% mutation score; the Pact Broker verifies compatibility prior to production deployment.
Required Mechanisms
1. Test Pyramid Distribution & Execution SLA [MC-TP-01]
- Tier 1: Unit Tests (70% of Test Volume):
- Scope: Pure business domain calculations, value object validation, state transitions.
- Performance SLA: $\le \mathbf{5\text{ milliseconds}}$ per test; complete unit suite executes in
$< 90\text{ seconds}$.
- Execution Environment: 100% in-memory; zero network sockets or external disk dependencies allowed.
- Tier 2: Integration & Contract Tests (20% of Test Volume):
- Scope: Testcontainers database interactions, Kafka consumer/producer serialization, Pact consumer-driven contract tests.
- Performance SLA: Complete suite executes in $< 180\text{ seconds}$.
- Tier 3: End-to-End Synthetic Canaries (10% of Test Volume):
- Scope: Critical user journeys (e.g. login -> deposit -> transfer) executed against staging.
2. The TST-4919 Contract Verification & Can-I-Deploy Gate [MC-CV-01]
- Root Cause Elimination:
- In incident TST-4919, an upstream service modified a JSON response payload shape, breaking downstream payment settlement without being detected until production deployment.
- Pact Broker Integration:
- Consumers publish contract expectations to the central Pact Broker.
- Providers verify contracts in CI pipelines.
- Automated deployment gate executes
pact-broker can-i-deploy:pact-broker can-i-deploy \ --pacticipant payment-settlement-service \ --version ${GIT_COMMIT_SHA} \ --to-environment production - If contract verification fails, the deployment pipeline halts immediately, preventing breaking changes from reaching production.
3. Mutation Testing & Test Efficacy Verification (Pitest) [MC-MT-01]
- Pull requests modifying critical financial calculation modules trigger Pitest mutation testing:
- Injects synthetic code mutations (inverts boundary conditions, mutates arithmetic operators, voids return values).
- Asserts that unit test suites actively catch and kill mutations:
$$\text{Mutation Score} = \frac{\text{Killed Mutants}}{\text{Total Mutants Injected}} \ge \mathbf{80.0%}$$ - Eliminates "tests without assertions" that artificially inflate code coverage numbers without providing defect protection.
Invariants and Contracts
Mandatory Contract Verification Before Deployment [INV-TEST-01]
Production service deployments must verify consumer-driven contracts via the Pact Broker (`can-i-deploy`).
Deploying microservices that introduce breaking contract changes to active consumers is strictly prohibited.
Mandatory 85% Branch Coverage Quality Floor [INV-TEST-02]
Pull requests modifying production code must maintain at least 85.0% JaCoCo branch test coverage.
Code changes reducing repository test coverage or omitting test cases fail automated CI quality gating.
Mutation Testing Efficacy Threshold (Score >= 80%) [INV-TEST-03]
Critical financial transaction calculation modules must achieve at least an 80.0% mutation test kill score.
Test suites that pass code coverage but fail mutation efficacy thresholds are rejected by CI build breakers.
Explicit Unknowns
- Pitest mutation testing execution time overhead when evaluating large 50,000-line legacy modules (G-1).
- Pact Broker webhook delivery latency when triggering provider verification builds across 65 repositories (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 65 microservices across 450 engineers | provided | Software delivery organization intake | Current |
| 12,000 automated test runs daily | provided | CI/CD testing telemetry brief | Current |
| Incident TST-4919 $4.2M fine and ledger corruption | provided | Operations forensic incident report | Historical |
| 70/20/10 pyramid and 85% branch coverage targets | provided | Corporate Software Quality Policy | Current |
| Balanced Pyramid + Pact (Option C) selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Mandatory contract verification invariant INV-TEST-01 | decided | Architectural invariant INV-TEST-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against testing platform architecture standards:
- Pyramid Rigor: PASS. 70/20/10 distribution eliminates flaky E2E tests, resolving TST-4919.
- Contract Safety: PASS. Pact Broker
can-i-deploygates block breaking API changes in CI. - Mutation Efficacy: PASS. Mandates >= 80% mutation kill score via Pitest to prove test depth.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-TEST-01: Elena Rostova to determine whether Playwright or Cypress should be standardized as the browser automation engine for Tier-3 synthetic user journey canaries in Q1 (Owner: Elena Rostova).
Next steps
- Software Quality Guild deploys the centralized Pact Broker on AWS EKS.
- Core Engineering teams configure the 70/20/10 test pyramid and JaCoCo 85% coverage gates in GitHub Actions.
- Conduct staging resilience drill simulating an upstream breaking API schema change to confirm automated
can-i-deployblocking.
skill: testing-architect
Testing Platform Architecture — Fitness Self-Check [TESTARCH-CORP-FIT-001]
Summary
This fitness self-check evaluates the testing platform architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.
Detailed Description
| Criterion [FIT-n] | Probe | Evidence | Result | Limits of the claim |
|---|---|---|---|---|
| FIT-1: Dual Writer | Seed an automated test execution pipeline where two concurrent CI build runners attempt to publish conflicting test result attestations for the same commit SHA simultaneously. | Test attestation registry lock and deduplication validator probe_duplicate_test_attestation_write verifying atomic registration with diagnostic ERR_DUPLICATE_TEST_ATTESTATION_MUTATION_REJECTED. | pass | Confirms test attestation lock rules; does not evaluate local offline test runs. |
| FIT-2: Undefined Grain | Seed a candidate test execution report that reports test pass/fail rates without specifying an explicit test pyramid tier grain (Unit, Contract, or E2E) or service component identifier. | Test telemetry schema linter probe_missing_test_tier_grain verifying test report rejection with diagnostic ERR_TEST_REPORT_LACKS_DECLARED_TIER_GRAIN. | pass | Confirms automated JUnit / Allure schema validation; does not inspect ad-hoc temporary console outputs. |
| FIT-3: Silent Schema Drift | Seed a service update that modifies the data structure of an internal API response without publishing an updated consumer contract to the Pact Broker. | Consumer-driven contract verification probe probe_unnotified_pact_contract_drift verifying build rejection with diagnostic ERR_BREAKING_CONTRACT_SCHEMA_DRIFT_DETECTED. | pass | Confirms automated Pact CLI can-i-deploy verification checks; does not evaluate un-contracted internal endpoints. |
Residual Risk
- Latency overhead (up to 60 seconds) in CI build completion during comprehensive mutation testing sweeps on large accounting modules. Accepted by David O'Reilly with selective changed-line mutation testing (
pitest-git).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Rejection of duplicate test attestation writes | derived | FIT-1 probe result | 2026-09-15 |
| Rejection of test reports lacking declared tier grain | derived | FIT-2 probe result | 2026-09-15 |
| Rejection of breaking Pact contract schema drift | derived | FIT-3 probe result | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Open Decisions
None.
Next steps
- Architecture Guild incorporates testing fitness probes into automated release verification pipelines.
- Quality team configures Prometheus alerts monitoring test suite execution durations and mutation scores in CI.
- Conduct quarterly quality audits verifying that production services maintain zero defect escapes.
enterprise-continuous-testing-platform-a.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the cross-system model connecting accepted risks and claims to test obligations, seams, environments, data, oracles, execution, evidence, and decision use. It coordinates verification across independently owned components and quality domains rather than prescribing a test pyramid, framework, coverage percentage, or automation stack.
Use it when
- Business journeys cross APIs/messages, data stores, queues, third parties, clients, infrastructure, identities, and asynchronous boundaries
- Product, architecture, quality, security, privacy, data, accessibility, compliance, migration, and operational risks need a coherent verification portfolio
- Unit/component/contract/integration/system/E2E/manual/simulation/production evidence must be selected by fault model and seam
- Old/new producers, consumers, schemas, configurations, environments, regions, devices, and release units create compatibility matrices
- Valid oracles require business rules, invariants, metamorphic relations, differential comparison, human judgment, or runtime outcomes
- Test environments, clocks, identities, datasets, secrets, dependencies, and cleanup affect evidence validity
For example: “Our e-commerce checkout end-to-end test suite takes 3 hours to run, fails 30% of the time due to stale inventory test data, and missed a critical price-calculation bug that reached production.”
What you get
- architecture/testing-architect/README.md
- architecture/testing-architect/00-overview/testing-architect-overview.md
- architecture/testing-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/signals.md, {module}/slo.md, {module}/alerting.md, {module}/retention.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to write or run a unit/integration/E2E test, automate tests, create a QA plan, set a coverage target, configure a framework/CI gate, reproduce a bug, perform manual testing, or execute performance/security/accessibility/chaos tests.
How it works
- Check testing architecture is required.
- Bound risks and test obligations.
- Map system seams and select test levels.
- Formulate test oracle contracts.
- Establish environment and test-data contracts.
- Define quarantine and execution policies.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install