Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIGitHub CopilotGitHub CopilotGemini CLIGemini CLI+20 more

    Trading Strategy Overfitting Detector

    1

    Detect selection bias, parameter overfitting, and backtest fragility in trading strategies using a formal audit framework.

    $9.99

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    0 installsSecurity scanned
    Trading Strategy Overfitting Detector

    Trading Strategy Overfitting Detector

    Example session with this skill installed

    Audit this EMA crossover strategy. IS Sharpe is 2.1, OOS is 0.4. I tested 500 parameter combinations across 10 years of SPY data. The best setting was a 12/26 EMA. Here is the optimization table.

    • Read your context and instructions
    • Compiled the trading strategy overfitting
    • Generated the document

    The audit identifies High Overfitting Risk for the submitted EMA strategy. The extreme degradation between In-Sample (IS) and Out-of-Sample (OOS) metrics, combined with a high trial count, suggests the result is a selection artifact.

    ## Trading Overfitting Analysis Report

    CategoryScoreFinding
    Final Verdict34/100High Overfitting Risk
    Research Provenance5/15High trial count (500) without multiple-testing adjustment.
    Parameter Robustness4/15Sharp optimum at 12/26; performance drops 60% at 13/27.
    OOS Integrity8/15Significant IS/OOS divergence suggests OOS was not a true holdout.
    Rank Stability3/10Top 5% of IS candidates failed to maintain top-quartile rank OOS.
    Concentration9/10Profits distributed across multiple regimes; low tail dependence.

    Analysis Highlights

    Parameter Surface: The 12/26 setting sits on a "parameter cliff." Moving the fast EMA by +1 period results in a 42% reduction in Sharpe, indicating the strategy is tuned to specific noise rather than a broad signal.

    Winner's Gap: The selected strategy is a 3-sigma outlier compared to the median candidate in the 500-trial sweep. This is a classic indicator of Selection Bias.

    IS → OOS Degradation: Retention is only 19%. A collapse of this magnitude typically invalidates the underlying research hypothesis.

    Search Breadth: Testing 500 combinations on a single asset (SPY) significantly inflates the probability of finding a "lucky" backtest.

    Next steps

    1. Perform a parameter perturbation test to see if any stable plateaus exist in the 500-trial universe.
    2. Test the 12/26 logic on an uncorrelated asset class (e.g., Gold or Treasury Futures) to check for universal edge.
    3. Conduct a fresh walk-forward analysis with strictly partitioned training and testing windows.

    trading-strategy-overfitting-detector.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Quantify selection bias by analyzing total research search breadth.Identify parameter cliffs where minor changes break strategy performance.Validate the integrity of OOS holdout data and detect contamination.Score strategy robustness using a standardized 100-point rubric.

    About this skill

    The problem

    Trading strategies often perform well in backtests only to fail in live markets. This happens because researchers unknowingly select historical winners from noise through excessive parameter tuning and dataset reuse.

    What it does

    • Analyzes research provenance to quantify the actual search breadth behind a strategy.
    • Identifies parameter "cliffs" where small changes in settings lead to catastrophic performance collapse.
    • Audits OOS (Out-of-Sample) integrity to detect holdout contamination and dataset reuse.
    • Evaluates rank stability and profit concentration to determine if an edge is structural or accidental.
    • Provides a standardized Trading Overfitting Robustness Score based on evidence-backed criteria.

    Why this beats prompting it yourself

    General LLMs tend to be overly optimistic about backtest results and miss subtle selection biases. This skill uses a specialized diagnostic framework that forces a skeptical reconstruction of the research process, preventing the "hallucination of alpha" that occurs when simple prompts ignore multiple-testing risks.

    Use cases

    • Vetting a new quantitative strategy before committing capital.
    • Auditing a vendor's backtest claims for signs of hidden parameter optimization.
    • Identifying if a strategy's returns are dangerously concentrated in a few lucky trades.
    • Verifying the validity of walk-forward validation results.

    Known limitations

    Does not compute Deflated Sharpe Ratio (DSR) or Probability of Overfitting (PBO) natively, though it assesses their applicability. Requires historical performance data to be provided by the user.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 26 days ago

    • Passed all security checks, Safe to install

    Listed26 days ago

    What's inside

    Frequently Asked Questions