- Home
- Skills
- Agents & Orchestration
- Goal & Scoreboard — Stop Overnight Agent Runs Drifting
Works with the AI tools you already use
Goal & Scoreboard — Stop Overnight Agent Runs Drifting
No finish line and no self-check means the run is a lottery ticket. This fixes both in 60 seconds.
$12
Goal & Scoreboard — Stop Overnight Agent Runs Drifting
Example session with this skill installed
I want to run an autonomous agent to optimize my site's performance and accessibility tonight. How do I make sure it doesn't just waste tokens and break things?
- Read your context and instructions
- Compiled the goal & scoreboard
- Generated the document
I have analyzed the request. To ensure this autonomous run is reliable, we need to move from a vague direction to a concrete scoreboard. An unsupervised performance run without a finish line will likely result in regressions.
Goal & Scoreboard Definition
1. Finish Line (State)
Every page in the /routes directory achieves a mobile Lighthouse score of 90+ for Performance, 100 for Accessibility, and 100 for Best Practices.
2. Scoreboard (Signal)
The agent will execute npm run lighthouse-ci after every optimization attempt. A "Pass" is defined as the JSON output showing all scores meeting the thresholds; a "Fail" is any score below the target or a non-zero exit code from the CI tool.
3. Retry Cap
The agent may attempt up to 4 optimization cycles. If the scoreboard does not show a "Pass" after the 4th cycle, the agent must stop, revert to the last stable git commit, and generate a report detailing which specific metrics failed.
Next steps
- Verify Environment: Ensure
lighthouse-ciis installed and globally accessible in the agent's shell. - Baseline Check: Run one manual scoreboard check now to see the current scores before the agent starts.
- Initialize Run: Start the agent with these constraints hard-coded into its system prompt for the session.
goal-scoreboard-stop-overnight-agent-run.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
Long autonomous AI runs fail for one reason far more than any other: the model can't tell whether it's done or whether it's wrong. It generates, assumes success, and errors pile up quietly.
This skill fires before any long, vague, or unsupervised handoff — goal modes, overnight agent jobs, scheduled tasks, client deliverables — and forces three lines into writing:
- The finish line, stated as a STATE rather than a direction
- A scoreboard the model can read by itself — a test that exits 0/1, a tool score, a count that must match, a live page checked for a string (ranked strongest to weakest)
- A retry cap, so it stops and reports instead of spiraling
If a job can't produce those, it doesn't run blind: build the checker first, split the job, or downgrade it to supervised drafts. It explicitly blocks the two classic cheats — the model grading its own homework ("8/10") and invented metrics that sound measurable and aren't.
Includes a TEST-line template for scheduled tasks and how to turn a public scoreboard into a client offer that closes itself ("if it's not 100 on PageSpeed, you don't pay").
Use cases
- Check any long or overnight agent job before you walk away from it
- Diagnose why a previous autonomous run drifted or came back wrong
- Add a real pass/fail TEST line to every scheduled task so none of them fail silently for weeks
- Turn a vague client service into a checkable offer
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 6 days ago
- Passed all security checks, Safe to install