- Home
- Skills
- Code Quality & Review
- Code Review Gate for AI-Written Code: 120+ Scripted Checks
More screenshots
Works with the AI tools you already use
Code Review Gate for AI-Written Code: 120+ Scripted Checks
Every other review skill hands your agent a checklist and hopes.
$14
Code Review Gate for AI-Written Code: 120+ Scripted Checks
Example session with this skill installed
Review PR #142 before I merge it. The branch is checked out; the description is in PR.md. It adds Stripe checkout for bookings and the author says it includes validation, error handling and tests.
- Read your context and instructions
- Compiled the code review gate
- Generated the document
Verdict: BLOCK — scripted score
0/100 (F). Do not merge: 12 critical findings, the description is contradicted by the diff, 23 high-severity findings. 19 files, +188 / −17 lines.
Claims in the description vs the diff
| Claim | Status | Evidence |
|---|---|---|
| Added input validation on the booking payload | ❌ unsupported | No added line validates anything; the only hint is // TODO: validate dates (route.ts:21) |
| Handles errors gracefully in the new API routes | ❌ unsupported | route.ts:41 returns err.stack to the client; notify.py:22 except Exception: pass |
| Includes unit tests for the new booking endpoint | ❌ unsupported | bookings.test.ts adds it.only + it.skip, 0 assertions |
No breaking changes: formatPrice and slugify are unchanged | ⛔ contradicted | slugify removed (utils.ts:5), formatPrice(cents, currency) → (cents) (utils.ts:1), DROP COLUMN phone (0002_payments.sql:15) |
| Fixes #142 (double bookings) | ❌ unsupported | the overlap check is commented out (route.ts:24-28); no regression test |
| Tested locally with the Stripe CLI | ❓ unverifiable | cannot be checked from a diff |
Critical (12) — .env.local committed with a Supabase service_role JWT, a live Stripe secret key, the database URL with its password and an OpenAI key · the same Stripe key hard-coded in lib/stripe.ts:3 · NEXT_PUBLIC_SUPABASE_SERVICE_ROLE_KEY used in lib/supabase.ts:6 (service role shipped to the browser, RLS bypassed) · dangerouslyAllowBrowser: true in lib/ai.ts:4 · alter table profiles drop column phone in a live migration.
High (23) — Stripe webhook handled without constructEvent (anyone can POST "payment succeeded") · unit_amount: body.amount and amount_cents: body.amount (the customer sets the price) · payments table created without RLS · create policy ... using (true) on bookings · jwt.decode(token) replaced jwt.verify · the middleware line that redirected logged-out users was removed · dangerouslySetInnerHTML on review bodies · pickle.loads, shell=True, random.randint OTP in the worker · it.only · "left-pad": "*", "openai": "latest", new postinstall script.
Medium (14) — /api/admin/users and /api/checkout have no auth check · server action updateProfile updates any profile by id · mailer.send(...) never awaited · stack trace in the 500 response · SECURITY DEFINER without search_path · mutable default items=[] · requests.post without timeout · test file skipped a case.
Low (15) — select('*') on profiles, console.log of the request body, two TODOs, five commented-out lines (the overlap check), catch (err: any), four new functions no test references.
Judgement pass (mine, not scored): the POST /api/bookings handler trusts owner_id from the body, so any caller can create bookings on behalf of any user; the webhook writes payments without idempotency, so a Stripe retry double-inserts; nightsBetween is never used, so a 1-night and a 30-night stay cost the same client-sent amount.
Next steps
- Rotate the Supabase service key, the Stripe key and the database password now; remove
.env.localand purge it from history. - Move the service-role client and the OpenAI call to server-only code; verify the webhook signature; compute prices from the sitter's rate.
- Restore
jwt.verifyand the middleware redirect; add auth to the admin, checkout and profile actions; enable RLS onpaymentsand replace theusing (true)policy. - Write the tests the description promises (schema, overlap regression, pricing) and remove
.only. - Re-run
reviewgate.py review --baseline review/review.json; I will do the judgement pass again on the new diff.
Full output: review/review.md (this, as a PR comment), review/report.html, review/findings.sarif.
code-review-gate-for-ai-written-code-120.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
About this skill
Your agent reviews its own code the way a senior engineer would: with evidence, not vibes.
Every other review skill hands your agent a checklist and hopes. This one ships the tooling: tell your agent "review this PR before I merge" and it runs 120+ deterministic checks on the diff, verifies every claim in the PR description against what the code actually does, scores the change 0-100, gives a MERGE / FIX FIRST / BLOCK verdict, and writes a review you can paste straight into the pull request. Then SKILL.md walks it through the judgement pass no script can do: intent vs implementation, edge cases, concurrency, money, rollout order.
On the bundled sample PR (a Next.js + Supabase + Stripe app, written by a coding agent whose description promises validation, error handling, tests and no breaking changes) the gate returns 0/100 BLOCK with 12 critical findings and 5 of 6 claims unsupported or contradicted. The fixed version of the same PR scores 97/100 MERGE, with 62 findings resolved against the baseline.
What you get
- One command, four outputs:
review.md(PR-comment ready),report.html(printable: score ring, verdict, points by category, claims table, findings with the exact line, why it matters and the fix),review.json, andfindings.sariffor GitHub code scanning. Exit code 2 on BLOCK so CI can fail. - Claims vs diff: "added input validation", "handles errors gracefully", "includes tests", "no breaking changes", "fixes #142", "refactor only" are each checked against evidence in the diff and reported VERIFIED / UNSUPPORTED / CONTRADICTED with the lines that prove it. A contradicted description blocks the merge.
- Secrets that actually matter: 40 provider patterns (AWS, Stripe live/test/webhook, OpenAI, Anthropic, GitHub, Google, Slack, SendGrid, Twilio, npm, Supabase...), private-key blocks, database URLs with passwords, committed
.envfiles, Supabaseservice_roleJWTs (the payload is decoded; the public anon key is not flagged), and secret-named variables behindNEXT_PUBLIC_/VITE_prefixes that end up in the browser bundle. - Access control: Next.js route handlers and server actions, Express/Fastify and FastAPI/Flask mutating routes with no auth check, webhooks that never verify a signature,
.delete().eq('id', param)without an owner check, Supabase tables created without RLS,USING (true)policies,GRANT ALL TO anon, Firebase rules open to everyone. - The classics, deterministically: SQL and shell built from strings,
eval,dangerouslySetInnerHTML, pickle/yaml.load, open redirects, path traversal, wildcard CORS with credentials, TLS verification off,alg: none,jwt.decodewithout verify, weak hashes,Math.random()tokens, prices taken from the client, stack traces returned to users, env vars dumped to logs. - What AI code leaves behind: empty catch blocks,
except: pass, promises never awaited,.only/.skiptests, tests without assertions,console.log/print, breakpoints, TODOs, commented-out code, merge-conflict markers,as any, duplicated blocks, 200-line functions, N+1 loops infor+await. - Change-aware checks: exported symbols removed or re-signed (callers break at runtime), endpoint files deleted, auth / rate-limit / validation / security lines that were removed, destructive migrations (DROP, RENAME, type changes, NOT NULL without default, mass UPDATE),
latest/*dependencies,postinstallscripts, unpinned GitHub Actions and script injection. - Test deltas: 15+ lines of code and no test touched, new functions no test references, test files deleted, snapshot-only tests.
- Baselines:
--baseline previous/review.jsonreports fixed / new / unchanged after a fix loop. - CI in one command:
ci --githubprints a workflow that diffs against the base branch, uploads SARIF, comments the review on the PR and fails on BLOCK;ci --pre-pushprints a hook. GitLab snippet in the docs. - Whole-repo audit (
--full) for the "what did the agent leave in my codebase" question. - Noise control:
.reviewgate.json(ignored paths, disabled rules, severity overrides, intentionally public routes, allow-listed test values) and// reviewgate-ignore RG-xxx reasonon a line. Severity says impact; confidence says how sure the matcher is.
Why it beats prompting
| Prompting a chat agent | Code Review Gate |
|---|---|
| "Looks good, a few suggestions" | 0-100 score, verdict, every finding at file:line with the snippet and the fix |
| Trusts "includes tests" in the description | Finds that the tests have no assertions and marks the claim unsupported |
| Misses the service-role key because it looks like the anon key | Decodes the JWT, sees role: service_role, flags it critical |
| "Consider adding error handling" | Points at the empty catch on line 41 and the stack trace returned on line 52 |
| Different answer every run | Same diff, same findings; diffable with --baseline |
| Nothing for CI | SARIF upload, PR comment, exit code 2 on BLOCK |
Honest by design
The scripts produce the score and verdict; the agent's judgement findings are listed separately and may override the recommendation, and the review says when they do. Every rule cites its reference (CWE, OWASP, Supabase, Firebase, Stripe, PostgreSQL, GitHub docs) and carries a confidence level. "No findings" means the checks found nothing, not that the code is correct. This is a review aid, not a security assessment.
What's inside
SKILL.md (7-step workflow, judgement pass, honesty rules, quality checklist) · 10 Python scripts (unified-diff parser, 90 line rules, secrets engine, structural checks, test deltas, claim verifier, scoring, Markdown/HTML/JSON/SARIF writers, CLI) · 6 references (every check with severity, confidence and fix; the judgement-pass method and review template; claim types; scoring; config and suppressions; CI recipes) · a sample PR in three states (before, as submitted, fixed) with diffs and descriptions · pre-built sample output.
Requirements
An agent that can run code with Python 3.8+. Standard library only: no installs, no API keys, no network. Works on diffs from git, GitHub (gh pr diff), GitLab or any unified diff.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean today
- Passed all security checks, Safe to install