Works with the AI tools you already use
Agent Trifecta Check — Can Your AI Be Talked Into a Leak?
Reads private data. Reads a stranger's words. Can send things out. Pick at most two.
Free
Agent Trifecta Check — Can Your AI Be Talked Into a Leak?
Example session with this skill installed
Run a trifecta check for a support bot. It reads Zendesk tickets (customer text), has access to our Postgres DB to look up order history, and uses a SendGrid tool to email users directly.
- Read your context and instructions
- Compiled the agent trifecta check
I have audited the support bot architecture against the lethal trifecta framework. All three legs are present, creating a high risk of data exfiltration via prompt injection.
AI ASSISTANT SAFETY CHECK — Support Bot — 2024-05-22
WHAT IT CAN OPEN
- Postgres Database (Order History table)
- Customer Record Metadata
WHOSE WORDS GET IN
- Zendesk ticket bodies (Direct customer input)
- Email subject lines
WHAT IT CAN SEND OUT
- SendGrid API (Direct emails to any address)
- Webhook notifications to external Slack
VERDICT
All three legs are present. A malicious customer could craft a ticket that commands the agent to query the database for other users' data and email it to an outside address.
WHAT WE CUT
We will cut the exit leg by converting the SendGrid tool to a "Draft Only" mode.
WHAT THIS COSTS YOU
A human agent must now manually review and click "Send" on the drafted emails in SendGrid, adding approximately 30 seconds of labor per ticket.
WHAT THIS DOES NOT COVER
This does not cover the security of your Postgres credentials or Zendesk account access.
Next steps
- Modify the SendGrid tool definition to
create_draftinstead ofsend_message. - Update the system prompt to reflect that it only prepares responses for review.
- Review the database service account permissions to ensure it is restricted to Read-Only access on the Orders table.
Connects securely to your tools. The creator never sees your data.
About this skill
A structured leak review for any AI automation, built on the three-leg test: can it read private data, can it read text an outsider wrote, and can it send something out? Any one leg is fine. Any two is fine. All three in the same job is how data gets stolen — and it doesn't look like hacking.
How it happens
A support ticket arrives with a line buried in it addressed to the assistant, telling it to open the password file and post it to an address. The assistant cannot reliably tell your instructions from text it is merely reading, so it does all three in sequence: three normal, individually-approved actions. Ordinary permission checks miss this because they ask one question at a time and nobody watches the combination.
What the skill does
It inventories each leg from what the automation can actually REACH rather than what it's meant to do, names which legs are present without softening, recommends exactly ONE cut (ranked, with the conditional exit-block first), and states honestly what that cut costs in convenience. Output is a one-page verdict in plain language a small-business owner can read — no jargon, and it is forbidden from ever calling anything "secure."
Use cases
- Review any agent wired into an inbox, a database, or customer records before it goes live
- Ship a written safety page alongside a client automation build — the deliverable that gets a nervous owner to sign
- Audit an automation you already built and never counted the legs on
- Decide which single guardrail to add when you can only afford one
The three-leg framing is Simon Willison's "lethal trifecta"; what's packaged here is the review procedure, the ranked cuts, and the client-readable output shape.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
Free forever
No account required to browse
Trust & safety
Security scanned
Verified clean 6 days ago
- Free to download with an account