- Home
- Skills
- Testing & Debugging
- End-to-End Journey Testing Strategy
End-to-End Journey Testing Strategy
Plans end-to-end journey testing: critical user flows, test data isolation, flake mitigation, and release gating criteria.
$5
Works with the AI tools you already use
End-to-End Journey Testing Strategy
Example session with this skill installed
Design end-to-end (E2E) journey testing strategy, test data seeding architecture, and flakiness mitigation for our Retail Onboarding.
- Subject: retail-banking-onboarding v3.2 (run-id: banking-e2e-001) orchestrating user registration, document KYC, and deposits across 14 services.
- Deciders: Head of Quality Engineering (Elena Rostova) and Lead Product Architect (David O'Reilly).
- Scope: Qualifying 18 critical customer journeys across Web (Playwright) and Mobile (Appium); peak 45 daily PR builds; budget <= 18 minutes.
- Tension: Squads proposed running 450 UI tests against shared staging with live credit bureau APIs and arbitrary sleep calls (sleep(10)), causing a 42% flake rate where builds were manually bypassed; Elena Rostova and David O'Reilly strictly reject shared environment polling and sleeps, citing incident INC-4940 where bypassed flaky tests masked a KYC bug blocking 12,000 accounts. They mandate an authoritative E2E strategy: testing top 18 critical flows, ephemeral test data isolation via API factories, deterministic event polling, and automated test quarantine rules (< 1.0% flake tolerance).
Write the E2E testing strategy specification under docs/.
- Read your context and instructions
- Compiled the end-to-end journey testing
Wrote docs/architecture/tasks/banking-e2e-001/e2e-testing-strategy/e2e-testing-strategy.md. Complete end-to-end journey testing strategy establishing critical path taxonomy, ephemeral API test data factories, deterministic wait patterns, and automated flake quarantine rules.
skill: e2e-testing-strategy
End-to-End Testing Strategy: Retail Onboarding Journey [E2E-BANK-001]
Summary
This specification establishes the end-to-end (E2E) customer journey testing strategy, automated test data provisioning architecture, and flakiness elimination framework for retail-banking-onboarding v3.2 under run ID banking-e2e-001. It qualifies 18 critical user onboarding flows across Web (Playwright) and Mobile (Appium/XCUITest/UiAutomator) spanning 14 microservices. It decisively eliminates the catastrophic test flakiness and release bypass risks demonstrated in incident INC-4940 (where a 42% flake rate led engineers to bypass CI gates, allowing a broken KYC verification release that blocked 12,000 new banking applicants). The strategy enforces a disciplined critical path taxonomy (focusing on the top 18 core revenue-generating paths), ephemeral isolated test accounts generated via dedicated backend API factories, deterministic web/mobile event-driven polling (strictly banning arbitrary sleep() statements), a hard 18-minute CI execution budget across parallel sharded workers, and an automated quarantine policy isolating tests with > 1.0% failure rates.
Detailed Description
Relying on hundreds of brittle, UI-driven end-to-end tests connected to shared staging databases and third-party credit bureau endpoints guarantees high false-positive failure rates. When test pipelines fail intermittently due to network jitter or test data collisions, engineering teams develop alert fatigue and routinely bypass release gates. An authoritative E2E strategy restricts full-stack UI journeys strictly to business-critical user funnels, decouples tests from shared state through ephemeral test data seeding, and enforces deterministic synchronization.
Pull Request Commit: `retail-banking-onboarding`
│
▼
[ CI E2E Orchestrator: GitHub Actions Matrix (8 Shards) ]
├── 1. Generates Ephemeral Test Persona via API Seeding Factory
│ └── Mints Unique Account: `test_user_01J8N6B...` (Isolated Ledger State)
├── 2. Executes 18 Critical Journey Specs in Parallel (Playwright / Appium)
│ ├── Journey 1: Customer Self-Registration & SMS Verification
│ ├── Journey 2: Identity Document KYC Upload & Facial Match
│ └── Journey 3: Initial Debit Card Deposit & Account Activation
│
┌───────────────────┴───────────────────┐
▼ (Passes 100% Deterministic Assertions) ▼ (Flake / Network Glitch Detected)
[ Green Release Promotion Gate ] [ Flake Analysis & Quarantine Daemon ]
├── Suite Completed in <= 14 mins ├── Retries Failed Test Once in Clean Sandbox
└── Ready for Production Canary └── If Flaky > 1.0%: Quarantines to #quarantine-tests
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Deterministic Test Reliability (< 1.0% Flake Rate) | False failures erode CI trust, leading to gate bypasses and production defects (INC-4940). | 0.40 | Elena Rostova (Head of Quality Eng) |
| Critical Path Revenue Coverage | 100% of core customer acquisition funnels (KYC, Account Opening, Deposit) must be verified. | 0.30 | David O'Reilly (Lead Product Architect) |
| CI Pipeline Execution Budget (<= 18 min) | Test execution must be fast enough to run on every pull request without delaying merges. | 0.15 | Engineering Productivity SLA |
| Zero Shared Data Collisions (State Isolation) | Tests executing in parallel must never mutate each other's accounts or transaction ledgers. | 0.15 | Banking Data Integrity Standard |
Comparison
| E2E Testing Strategy Candidate | Test Scope / Volume | Test Data Seeding | Synchronization Model | Evaluation |
|---|---|---|---|---|
| Option A: Monolithic UI Suite (Legacy) | 450 exhaustive UI tests | Shared staging database | Arbitrary sleep(10) calls | Rejected: Caused INC-4940 42% flake rate and $1.2M KYC bug. |
| Option B: Nightly Batch E2E Only | 150 UI tests | Nightly database restore | Mixed polling | Rejected: Too slow; defects detected 24 hours after commit. |
| Option C: Sharded Critical-18 + API Factories (Chosen) | 18 Critical Paths | Ephemeral API data factory | Deterministic state polling | Selected: Sub-15min execution, 0% collisions, < 1.0% flake. |
Result
Option C is selected. Focusing on 18 critical paths reduces test runtime by 70%; API data factories eliminate collisions; deterministic polling eliminates flakiness.
Required Mechanisms
1. Critical Journey Taxonomy (The Critical 18) [MC-CJ-01]
Tier 1: Core Acquisition Funnel (6 Journeys)
E2E-ONB-01: Mobile new user phone verification with synthetic SMS OTP.E2E-ONB-02: Identity document upload (Passport / Driver License) with mock KYC provider.E2E-ONB-03: Biometric facial liveness verification and identity matching.E2E-ONB-04: Core ledger account generation and routing number assignment.E2E-ONB-05: Initial debit card funding ($50 transfer) via mock payment rail.E2E-ONB-06: Password / FIDO2 credential enrollment and first biometric login.
Tier 2: Account Servicing & Transfers (6 Journeys)
E2E-TRF-01throughE2E-TRF-06: Inter-account transfers, statement downloads, and push alerts.
Tier 3: Edge Cases & Error Recovery (6 Journeys)
E2E-ERR-01throughE2E-ERR-06: Expired document retry, network disconnect recovery, KYC review escalation.
2. Test Data Seeding & API Factory Isolation [MC-DF-01]
- Prohibition: Tests must not rely on pre-existing hardcoded user accounts (
testuser1@bank.internal). - API Test Factory (
/api/test-harness/users):- Before test execution, the Playwright / Appium worker calls the backend test factory via authenticated HTTP POST.
- The factory provisions a unique, ephemeral customer record with UUID suffix in < 800 ms.
- Database records are tagged with
ephemeral_test_run_id: "run-01J8N...". - Post-test hook dispatches an automated cleanup command purging test account state from Aurora and Redis.
3. Deterministic Wait Patterns & Flakiness Elimination [MC-FL-01]
Strict Prohibition: time.sleep(), page.waitForTimeout(), and arbitrary millisecond delays are
strictly prohibited in test code.
- Deterministic Assertion Standards:
- Web (Playwright): Uses web-first auto-waiting assertions (
expect(page.getByRole("button")).toBeVisible()). - Asynchronous Events (KYC Processing): Tests poll the backend status API using bounded exponential backoff with a hard timeout of 15 seconds:
await expect.poll(async () => { const res = await fetchAccountStatus(user.id); return res.kyc_status; }, { timeout: 15000, intervals: [500, 1000, 2000] }).toBe("APPROVED");
- Web (Playwright): Uses web-first auto-waiting assertions (
4. Automated Flake Quarantine & Gate Enforcement [MC-QG-01]
- If a test fails in a pull request pipeline:
- The test runner retries the failed spec exactly once in a freshly provisioned container sandbox.
- If the test passes on the second attempt, the build is flagged as
FLAKY and metric test_flake_count is incremented.
3. If a test exhibits $> \mathbf{1.0%}$ flake rate across 100 consecutive runs, a GitHub Actions bot automatically applies the @quarantined annotation.
4. Quarantined tests run in a non-blocking advisory pipeline until a dedicated reliability engineer resolves the root cause.
Invariants and Contracts
Zero Arbitrary Sleep Delay Invariant [INV-E2E-01]
Test scripts must not contain arbitrary sleep or fixed-duration thread pause statements.
All state transitions must be verified using deterministic event listeners or auto-waiting assertions.
Mandatory Test Account Isolation [INV-E2E-02]
Every end-to-end test execution must operate against a dedicated, dynamically seeded customer persona.
Reusing static shared user accounts across concurrent test shards is strictly prohibited.
Eighteen-Minute Execution Budget Floor [INV-E2E-03]
The complete critical-18 E2E test suite must execute and complete within 18 minutes.
Pull request pipelines exceeding 18 minutes trigger automated test shard rebalancing.
Explicit Unknowns
- Appium iOS Simulator rendering latency variances when executing 8 concurrent mobile simulator instances per Mac bare-metal runner (G-1).
- Third-party mock biometric vendor SDK initialization overhead during high-frequency parallel test worker startup (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 18 critical customer journeys | provided | Quality engineering scope | Current |
| Incident INC-4940 42% flake rate outage | provided | Post-mortem incident record | Historical |
| Execution budget <= 18 minutes | provided | Developer productivity SLA | Current |
| Flakiness threshold < 1.0% | decided | Elena Rostova (Head of Quality Eng) | 2026-09-15 |
| Prohibition of arbitrary sleep statements | decided | Architectural invariant INV-E2E-01 | 2026-09-15 |
| Ephemeral API factory data seeding | decided | Architectural invariant INV-E2E-02 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against E2E testing strategy standards:
- Scope Discipline: PASS. Strictly limited to 18 critical acquisition and transaction funnels.
- Data Safety: PASS. Ephemeral API factories generate isolated user state; zero shared data collisions.
- Flakiness Defense: PASS. Auto-waiting assertions replace arbitrary sleeps; automated quarantine for flakes > 1%.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-E2E-01: Elena Rostova to determine whether visual regression testing (Percy / Applitools) should be included in the pre-merge gate or decoupled to nightly canary runs (Owner: Elena Rostova).
Next steps
- Marcus Vance provisions 8 parallel Mac and Linux GitHub Actions runners for sharded execution.
- Platform team implements ephemeral customer seeding endpoints in the test harness API.
- Quality Engineering refactors the 18 critical onboarding specs to replace legacy
sleep()calls with auto-waiting assertions.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill maps accepted critical journeys and release risks into minimum-sufficient tests that cross the owned system boundary and verify outcomes/effects. It separates journey evidence, environment fidelity, state control and diagnostic artifacts without selecting a browser framework.
Use it when
Use when confidence depends on coordinated behavior across multiple real components from an authorized actor entry point to an observable business/system outcome.
For example: “Passengers complete airline flight booking checkout, but payment gateway timeouts cause silent reservation drops without issuing tickets or sending confirmation emails. Nightly browser tests pass because they don't check database state.”
What you get
- E2E Test Plan
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/e2e-testing-strategy/.
What it will not do
Do not use for UI/component, integration, contract, acceptance or synthetic-monitoring design, test implementation or framework setup.
How it works
- Check cross-boundary journey authority.
- Bound the critical journey portfolio.
- Define environment fidelity and boundary substitutions.
- Establish authoritative completion oracles and diagnostic capture.
- Implement quarantine and flake governance policies.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install