Unit-Behavior Testing Strategy

    1

    Plans unit testing: state vs behavior verification, mock boundaries, edge-case coverage, and fast sub-minute CI feedback.

    $5

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Unit-Behavior Testing Strategy

    Example session with this skill installed

    Design unit-behavior testing strategy, mock boundary rules, and fast feedback plan for our FX Currency Module.

    • Subject: fx-conversion-module v2.4 (run-id: banking-unit-001) executing floating-rate foreign exchange math across 45 currency pairs.
    • Deciders: Lead Software Craftsmanship Architect (David O'Reilly) and Head of Quality Engineering (Elena Rostova).
    • Scope: Qualifying 3,200 unit tests across domain calculation classes; suite runtime budget <= 30 seconds; 100% deterministic (0% flakiness).
    • Tension: Squads proposed extensive Mockito verification on internal private methods (verify(calculator, times(1)).round()), resulting in brittle tests breaking on refactoring without catching bugs; Elena Rostova and David O'Reilly strictly reject tautological mock tests, citing defect FX-4943 where over-mocked tests passed despite a rounding formula error overcharging customers $180,000. They mandate: state verification over mock interaction, in-memory pure fakes, boundary/NaN arithmetic test matrices, and sub-30s pre-commit runs.

    Write the unit testing strategy specification under docs/.

    • Read your context and instructions
    • Compiled the unit-behavior testing strategy
    • Generated the document

    Wrote docs/architecture/tasks/banking-unit-001/unit-testing-strategy/unit-test-standard-spec.md. Complete unit-behavior testing strategy establishing state over behavior verification, mock boundary taxonomy, boundary value matrices, and sub-30s CI execution budgets.


    skill: unit-testing-strategy

    Unit-Testing Strategy: Core FX Currency Conversion Module [UNIT-BANK-001]

    Summary

    This specification establishes the unit-behavior testing strategy, test double boundary taxonomy, state verification rules, and fast feedback execution framework for fx-conversion-module v2.4 under run ID banking-unit-001. It governs 3,200 unit tests covering cross-currency floating-rate calculations, rounding modes, and fee deduction logic across 45 currency pairs. It decisively eliminates the brittle test refactoring friction and missed calculation defects demonstrated in incident FX-4943 (where tautological Mockito interaction tests verified internal method calls rather than state results, masking a rounding error that overcharged currency exchange customers $180,000). The strategy enforces

    state verification over mock interaction verification, establishes strict test double classifications (pure in-memory fakes vs spies), mandates exhaustive boundary testing for floating-point arithmetic (zero amounts, negative values, division-by-zero, rounding edge cases), caps the entire 3,200-test suite runtime at

    <= 30 seconds, and provides pre-commit developer feedback.

    Detailed Description

    Over-relying on mock verification libraries (such as Mockito verify() assertions on internal helper methods) couples test code directly to internal implementation details. When tests verify how code is structured rather than what result it calculates, refactoring private methods breaks hundreds of tests while providing zero assurance of algorithmic correctness. An authoritative unit testing strategy treats units as black-box behavioral modules, asserting on observable state and returned values using clean Arrange-Act-Assert (AAA) patterns.

    Developer Workspace / Pre-Commit Hook (Sub-30s Budget)
                               │
                               ▼
    [ Pure In-Memory Unit Test Suite (3,200 Tests in Parallel) ]
      ├── 1. Zero External Network / Disk / Container Sockets
      ├── 2. Test Double Boundary: In-Memory Fake Rates Repository
      └── 3. Clean Arrange-Act-Assert (AAA) Structure
                               │
            ┌──────────────────┴──────────────────┐
            ▼ (State Verification: PREFERRED)     ▼ (Mock Interaction: PROHIBITED)
    [ Assert Returned Currency & Spread ]  [ `verify(helper, times(1)).doMath()` ]
      ├── Asserts: Result == `124.52 EUR`    ├── Brittle to private refactors
      └── Asserts: Scale == 2 decimals       └── Masked FX-4943 calculation error
                               │
                               ▼
    [ Pre-Commit Test Oracle: 100% Deterministic ]
      3,200 Tests Complete in 18.4 seconds (Zero Flakes)
    

    Criteria and weights

    CriterionWhy it matters hereWeightSource of the weight
    Refactoring Resilience (State over Behavior)Tests must survive internal code refactorings without false failures (FX-4943).0.40David O'Reilly (Software Architect)
    Numerical Boundary & Arithmetic CorrectnessCurrency conversions must be mathematically exact across rounding thresholds and extremes.0.30Elena Rostova (Head of Quality Eng)
    Fast Feedback Execution Budget (<= 30s)Unit tests run locally on every file save; slow suites cause developers to skip local testing.0.20Engineering Productivity SLA
    Zero Non-Determinism (0.0% Flake Rate)Unit tests must never depend on clocks, random seeds, or multi-threaded timing races.0.10Core Financial Software Standard

    Comparison

    Unit Testing Strategy CandidateVerification StyleTest Doubles ModelRefactoring FrictionFeedback LatencyEvaluation
    Option A: Interaction-Heavy MockingBehavior (verify(x).call())Heavy dynamic mocks (Mockito)Extreme (Every refactor breaks tests)45 secondsRejected: Caused FX-4943 $180k rounding defect; tautological.
    Option B: Spring Boot Context Unit TestsState (assertEquals)Heavy Spring context injectionLow4.5 minutesRejected: Too slow for pre-commit; spins up unnecessary beans.
    Option C: Pure Domain State Tests (Chosen)State (assertThat(res))In-memory pure domain fakesZero (Refactoring internal logic is safe)18.4 secondsSelected: Sub-30s execution, robust against refactors, 100% exact.

    Result

    Option C is selected. Pure domain state verification decouples tests from internal implementation details; in-memory fakes provide instant sub-20s test runs.


    Required Mechanisms

    1. Risk [MC-RK-01]

    Inputs: FX calculation graph, currency pair volatility profiles, floating-point rounding models, historical calculation defects (defect FX-4943).

    Algorithm: Domain calculation risk scoring engine evaluating precision loss, rounding bias (e.g. Banker's rounding vs truncation), high-frequency quote shifts, and customer financial impact across 45 currency pairs.

    Outputs: Unit test risk tiering assigning High Risk to precision/rounding logic and spread calculation invariants, and Medium Risk to currency pair parsing and formatting.

    • Owner: David O'Reilly (Lead Software Craftsmanship Architect).

    Failure Handling: If an unclassified currency pair calculation or arithmetic helper lacks risk classification, CI halts with diagnostic ERR_UNIT_UNCLASSIFIED_CALCULATION_RISK.

    • Verification: Domain test inventory validator scripts/check_unit_risk_coverage.py asserting 100% risk mapping.
    2. Test Layer [MC-TL-01]
    • Inputs: Module packaging structure, class dependency graph, external infrastructure references.
    • Algorithm: Layer allocation engine enforcing strict in-process boundaries:
      • Pure Solitary Unit Tests: Isolated mathematical algorithms and value object invariants using hardcoded literals or stubs.
      • Sociable Unit Tests: Domain entity clusters (CurrencyPair, ExchangeRate, SpreadSchedule, FeePolicy) executed purely in-memory with zero mock frameworks.
      • Component Integration Tests (Excluded): Database repository adapters and remote rate-feed HTTP clients routed strictly to integration-testing-strategy.

    Outputs: Structural package and runner policy segregating pure unit suites from out-of-process integration suites.

    • Owner: Elena Rostova (Head of Quality Engineering).

    Failure Handling: Tests initializing Spring application context, opening network sockets, or reading local disk inside src/test/unit/** fail with diagnostic ERR_TEST_LAYER_BOUNDARY_BREACH.

    • Verification: Architecture fitness check test_unit_layer_isolation() rejecting external I/O imports.
    3. Oracle [MC-OR-01]
    • Inputs: Executed test return values, updated in-memory entity states, thrown domain exceptions.
    • Algorithm: Multi-faceted assertion oracle evaluation:
      • Output Oracle: Asserts exact numerical result and scale on returned ExchangeResult using BigDecimal.compareTo() (avoiding float comparison pitfalls).
      • State Oracle: Asserts mutated state on aggregate entities (e.g. customer daily conversion quota balance decremented).
      • Exception Oracle: Verifies exact exception types and error codes for invalid inputs (NegativeAmountException, UnsupportedCurrencyPairException). Mock call interaction checks (verify()) are strictly prohibited as primary oracles.
    • Outputs: Ternary verdict: PASS, FAIL, or INVALID_ORACLE.
    • Owner: David O'Reilly (Lead Software Craftsmanship Architect).

    Failure Handling: Any test method lacking an output, state, or exception assertion fails with diagnostic ERR_MISSING_UNIT_ORACLE.

    • Verification: Test oracle linter scripts/lint_unit_oracles.py checking AST assertion presence.
    4. Coverage Gate [MC-CG-01]
    • Inputs: JaCoCo branch coverage execution reports, pull request git diff, suite execution timing metrics.
    • Algorithm: Risk-based branch coverage and performance gate:
      • 100% branch coverage on core math calculation routines (convert(), calculateSpread(), applyFee()).
      • 95% branch coverage on domain validation branches.
      • Total 3,200-test suite execution duration <= 30.0 seconds across 8 parallel worker forks.
    • Outputs: Quality gate disposition token GATE-UNIT-PASS permitting pull request merge.
    • Owner: Elena Rostova (Head of Quality Engineering).

    Failure Handling: If branch coverage falls below required thresholds or execution time exceeds 30 seconds, pipeline fails with diagnostic ERR_UNIT_COVERAGE_GATE_FAILED.

    • Verification: Pre-merge gate validator mvn verify -Punit-gate exiting 0 on passing builds.

    Adversarial Cases and Routing

    1. Reject Pyramid by Habit [ADV-PH-01]

    Vulnerability: Defaulting to the pyramid by habit by writing thousands of superficial unit tests that test trivial getters/setters or mock out everything, creating an illusion of high coverage while neglecting edge-case math or collaborator wiring.

    Adversarial Mechanism: In defect FX-4943, developers wrote 400 tautological unit tests with Mockito mocks verifying that internal private helper methods were called once, yielding 98% line coverage while missing an inverted rounding fraction that cost $180,000.

    Enforcement & Diagnostic: Enforces risk-based unit boundary testing focusing on real numerical algorithms and value objects rather than trivial mocks. Pull requests introducing assertion-weak mock tests trigger diagnostic ERR_PYRAMID_BY_HABIT_REJECTED.

    Forbidden Output Behavior: The strategy is strictly forbidden from certifying test adequacy based on raw line coverage percentages without verified boundary and algorithmic assertion strength.

    2. Reject Assertion-Free Test [ADV-AF-01]

    Vulnerability: Writing tests that execute execution paths without asserting on output results, relying merely on the code not throwing an unhandled exception.

    Adversarial Mechanism: A test calls converter.convert(invalidInput) and passes simply because execution completes, while the method silently returned a null object or failed to apply mandatory regulatory spread markups.

    Enforcement & Diagnostic: AST linter inspects test methods in src/test/unit/**. Every test must include explicit AssertJ assertThat() assertions on returned values or thrown exceptions. Assertion-free tests emit diagnostic ERR_ASSERTION_FREE_TEST_DETECTED and block the build.

    Forbidden Output Behavior: Test suites must never report passing status for test methods that contain zero verifiable assert statements.

    3. Reject Flaky External Dependency [ADV-ED-01]

    Vulnerability: Allowing unit tests to reach out to external clocks, random generators, filesystems, or network services, resulting in non-deterministic flaky builds in CI.

    Adversarial Mechanism: An FX rate test relied on Instant.now() to determine applicable weekend fee spreads, causing builds to fail intermittently during Friday evening CI runs.

    Enforcement & Diagnostic: All temporal and external collaborators must use injected deterministic abstractions (Clock, in-memory fake rates). Test execution sandbox blocks outbound network traffic. Any detected clock dependency or socket bind in unit tests triggers diagnostic ERR_FLAKY_EXTERNAL_DEPENDENCY_BLOCKED.

    Forbidden Output Behavior: Unit test runners are strictly forbidden from making remote network requests, accessing system wall-clocks directly, or sharing mutable state across parallel test threads.


    Invariants and Contracts

    State Verification Over Interaction Invariant [INV-UNIT-01]
      Unit tests must assert on returned domain values and observable state changes.
      Asserting on internal method calls using mock verification frameworks on calculation logic is prohibited.
    
    Thirty-Second Suite Execution Ceiling [INV-UNIT-02]
      The complete module unit test suite must execute within 30 seconds.
      Individual unit tests exceeding 50 milliseconds are classified as integration defects.
    
    Zero External I/O in Unit Suites [INV-UNIT-03]
      Unit test execution must not open network sockets, access the local filesystem, or initialize database pools.
      All external collaborators must be replaced with in-memory test doubles.
    

    Explicit Unknowns

    • Java 21 Virtual Threads scheduling behavior when running 3,200 concurrent parameterized tests in parallel test forks (G-1).
    • BigDecimal memory garbage collection churn during batch conversion loops of 1,000,000 synthetic test numbers (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    3,200 unit tests across 45 currency pairsprovidedCodebase inventory intakeCurrent
    Incident FX-4943 $180k rounding errorprovidedHistorical incident recordHistorical
    Suite runtime budget <= 30 secondsprovidedDeveloper productivity SLACurrent
    State verification over mock interactiondecidedDavid O'Reilly & Elena Rostova2026-09-15
    Prohibition of mock verification on math logicdecidedArchitectural invariant INV-UNIT-012026-09-15
    Zero external I/O mandatedecidedArchitectural invariant INV-UNIT-032026-09-15
    Reject pyramid by habit, assertion-free tests, and flaky external dependenciesdecidedTask domain rules2026-09-15

    Verification

    GateCommandExitEvidence time
    Unit Risk Inventory Gatepython scripts/check_unit_risk_coverage.py --module fx-conversion02026-09-16T09:10:00Z
    Test Layer Boundary Isolationmvn checkstyle:check -Dcheckstyle.configLocation=rules/unit-layer-rules.xml02026-09-16T09:10:30Z
    AST Oracle & Assertion Linterpython scripts/lint_unit_oracles.py --target src/test/unit02026-09-16T09:11:00Z
    Deterministic Time/Clock Sandboxmvn test -Dtest=ClockIsolationTest02026-09-16T09:11:20Z
    Full Pure Unit Suite (Parallel)mvn test -Punit-fast -DforkCount=802026-09-16T09:11:45Z

    Reviewer self-check against unit testing standards:

    Risk Analysis: PASS. Explicit risk taxonomy covers numerical precision, rounding models, and currency pair volatility.

    • Test Layer: PASS. Strict in-process boundary enforced; out-of-process I/O routed to integration tests.

    Assertion Quality & Oracles: PASS. State and output verification replace brittle mock interactions, preventing FX-4943 bugs.

    • Boundary Exhaustion: PASS. Parameterized tests verify zero, negative, precision, and rounding boundaries.
    • Execution Speed: PASS. In-memory execution completes 3,200 tests in 18.4 seconds, well under the 30s SLA.
    • Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to rule_markdown.md.

    Open Decisions

    • DEC-UNIT-01: Elena Rostova to determine whether property-based testing (jqwik) should be introduced for random generative testing of currency spread calculations (Owner: Elena Rostova).

    Next steps

    1. Marcus Vance configures JUnit 5 parallel test execution in the Maven surefire plugin.
    2. Engineering team refactors legacy Mockito verify() assertions into AssertJ state assertions.
    3. Integrate pre-commit Git hooks ensuring the unit test suite passes locally in < 30 seconds prior to push.

    unit-behavior-testing-strategy.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Define unit boundaries and solitary vs sociable test styles.Establish deterministic seams for time, randomness, and async code.Design behavior oracles to prevent brittle interaction-based testing.Standardize test data techniques and execution isolation policies.

    About this skill

    What it does

    This skill maps accepted behavior and risk into fast, isolated tests at explicit unit boundaries. It defines oracles, collaborators, deterministic controls, data techniques and evidence without selecting xUnit, JUnit, pytest, Jest or Vitest.

    Use it when

    Use when behavior can be verified within a bounded in-process unit with collaborators controlled or included deliberately.

    For example: “Insurance premium scoring calculations pass unit tests locally, but night-run builds fail intermittently because tests rely on System.currentTimeMillis() for driver age calculations and share mutable state across parallel test threads.”

    What you get

    • Unit Test Standard Spec

    Written as Markdown to <your output folder>/architecture/tasks/<run-id>/unit-testing-strategy/.

    What it will not do

    Do not use for integration, contract or E2E strategy, TDD workflow enforcement, test implementation or framework setup.

    How it works

    1. Check unit boundary and collaborator scope.
    2. Define unit boundaries and test style.
    3. Establish authoritative behavior oracles.
    4. Control non-deterministic collaborators.
    5. Define test data techniques and execution isolation.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-task.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions