- Home
- Skills
- Trading & Investing
- Trading Strategy Overfitting Detector
Works with the AI tools you already use
Trading Strategy Overfitting Detector
Detect selection bias, parameter overfitting, and backtest fragility in trading strategies using a formal audit framework.
$9.99
Trading Strategy Overfitting Detector
Example session with this skill installed
Audit this EMA crossover strategy. IS Sharpe is 2.1, OOS is 0.4. I tested 500 parameter combinations across 10 years of SPY data. The best setting was a 12/26 EMA. Here is the optimization table.
- Read your context and instructions
- Compiled the trading strategy overfitting
- Generated the document
The audit identifies High Overfitting Risk for the submitted EMA strategy. The extreme degradation between In-Sample (IS) and Out-of-Sample (OOS) metrics, combined with a high trial count, suggests the result is a selection artifact.
## Trading Overfitting Analysis Report
| Category | Score | Finding |
|---|---|---|
| Final Verdict | 34/100 | High Overfitting Risk |
| Research Provenance | 5/15 | High trial count (500) without multiple-testing adjustment. |
| Parameter Robustness | 4/15 | Sharp optimum at 12/26; performance drops 60% at 13/27. |
| OOS Integrity | 8/15 | Significant IS/OOS divergence suggests OOS was not a true holdout. |
| Rank Stability | 3/10 | Top 5% of IS candidates failed to maintain top-quartile rank OOS. |
| Concentration | 9/10 | Profits distributed across multiple regimes; low tail dependence. |
Analysis Highlights
Parameter Surface: The 12/26 setting sits on a "parameter cliff." Moving the fast EMA by +1 period results in a 42% reduction in Sharpe, indicating the strategy is tuned to specific noise rather than a broad signal.
Winner's Gap: The selected strategy is a 3-sigma outlier compared to the median candidate in the 500-trial sweep. This is a classic indicator of Selection Bias.
IS → OOS Degradation: Retention is only 19%. A collapse of this magnitude typically invalidates the underlying research hypothesis.
Search Breadth: Testing 500 combinations on a single asset (SPY) significantly inflates the probability of finding a "lucky" backtest.
Next steps
- Perform a parameter perturbation test to see if any stable plateaus exist in the 500-trial universe.
- Test the 12/26 logic on an uncorrelated asset class (e.g., Gold or Treasury Futures) to check for universal edge.
- Conduct a fresh walk-forward analysis with strictly partitioned training and testing windows.
trading-strategy-overfitting-detector.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
The problem
Trading strategies often perform well in backtests only to fail in live markets. This happens because researchers unknowingly select historical winners from noise through excessive parameter tuning and dataset reuse.
What it does
- Analyzes research provenance to quantify the actual search breadth behind a strategy.
- Identifies parameter "cliffs" where small changes in settings lead to catastrophic performance collapse.
- Audits OOS (Out-of-Sample) integrity to detect holdout contamination and dataset reuse.
- Evaluates rank stability and profit concentration to determine if an edge is structural or accidental.
- Provides a standardized Trading Overfitting Robustness Score based on evidence-backed criteria.
Why this beats prompting it yourself
General LLMs tend to be overly optimistic about backtest results and miss subtle selection biases. This skill uses a specialized diagnostic framework that forces a skeptical reconstruction of the research process, preventing the "hallucination of alpha" that occurs when simple prompts ignore multiple-testing risks.
Use cases
- Vetting a new quantitative strategy before committing capital.
- Auditing a vendor's backtest claims for signs of hidden parameter optimization.
- Identifying if a strategy's returns are dangerously concentrated in a few lucky trades.
- Verifying the validity of walk-forward validation results.
Known limitations
Does not compute Deflated Sharpe Ratio (DSR) or Probability of Overfitting (PBO) natively, though it assesses their applicability. Requires historical performance data to be provided by the user.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 26 days ago
- Passed all security checks, Safe to install