More screenshots

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIGitHub CopilotGitHub CopilotGemini CLIGemini CLI+17 more

    experiment reproducibility auditor

    1

    Audit and verify the reproducibility of ML experiments by mapping code, data, and environment lineage.

    Secure checkout via Stripe

    0 installsSecurity scanned

    See it in action

    You say

    Audit this MLflow run (ID: 5f2b8) against Git commit a1b2c3d. Here are the DVC data hashes for the training set. Can we reproduce the accuracy result of 0.94?

    Your agent does

    ARTIFACT: EXPERIMENT-REPRO-AUDIT VERDICT: Incomplete lineage. Findings:

    • Code: [Verified] Matches commit a1b2c3d.
    • Data: [Verified] DVC hashes match.
    • Environment: [Assumed] Python 3.9 (missing manifest). Implementation: Re-export conda.yaml from run 5f2b8 to verify environment consistency.

    What you get

    Audit model lineage to ensure results match specific code and data versions.Identify missing environment manifests or uncommitted changes in ML runs.Trace DVC data hashes back to specific training experiment outcomes.Generate formal reproducibility reports for model handovers or audits.

    About this skill

    The problem

    Machine learning experiments often fail to replicate because of missing lineage between code, data, and environment. Teams struggle to verify if a specific model artifact actually matches a specific MLflow run or Git commit.

    What it does

    • Maps Git commits, MLflow metadata, and DVC hashes into a single evidence chain.
    • Identifies missing links in the reproduction pipeline before scoring results.
    • Flags statements as Verified, Inferred, or Assumed to prevent false confidence.
    • Generates a structured reproducibility audit with implementation steps for recovery.

    Frameworks & tools

    MLflow, DVC, Git, and various experiment manifests or logs.

    Why this beats prompting it yourself

    Standard LLM prompts often assume tool presence implies configuration maturity. This skill enforces a strict evidence contract that refuses to infer capability from tool names and explicitly tracks unverified assumptions through a specialized suite trace.

    Use cases

    • Auditing a model run to ensure it can be rebuilt from scratch.
    • Identifying why a local experiment doesn't match a remote baseline.
    • Verifying lineage for regulatory compliance or peer review.
    • Mapping disparate metadata from multiple ML tools into one source of truth.

    Known limitations

    Does not debug failed training runs or select orchestrators. Does not execute deployments or handle credentials.

    How to install

    Drop the file into your AI Agent. Works with Claude, Cursor, ChatGPT, and 20+ more.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean today

    • One-time purchase, yours forever

    Listedtoday

    Frequently Asked Questions