agent reliability audit
Turn raw agent traces and tool logs into professional production-readiness audits and remediation reports.
Ship better AI in 30 seconds. Browse 2,000+ expert-built and security scanned skills -> Browse skills
THE AGENSI STORE
31 skills found
Turn raw agent traces and tool logs into professional production-readiness audits and remediation reports.
Detect and ethically author persuasion in copy. Replication-filtered mechanism catalog with per-row ethics gate. Quarantines 13 zombie findings the popular canon still teaches.
A modular governance framework for AI policy, agent risk assessment, human-in-the-loop approvals, and audit trails.
Audit and de-conflict complex agent instruction stacks to fix inconsistent behavior and logic bloat.
Transform ambiguous AI tasks into auditable execution traces with verified evidence and AI-smell detection.
Audit applications against 12-Factor methodology to identify architectural risks and generate cloud-native fix plans.
Reconstruct architecture and map risks in inherited legacy codebases with evidence-based auditing and migration plans.
Automated versioning, provenance stamping, and integrity hashing for AI-generated artifacts.
Bulk-update cover images and thumbnails across your Gumroad catalog from a CSV or a folder of files named by permalink. Pre-flight checks validate each image (format, size, corruption) before it hits the Gumroad API, and you get an audit log of every success and failure. No clicking through the dashboard product by product.
Scan your schemas, seed data, config, and logs for personal data before it leaks. Detects PII-indicating column and key names (email, ssn, phone, address) across SQL, CSV, and JSON, plus PII in the data itself: email addresses, SSN-like numbers, credit-card-like numbers, phone numbers, and PII written into log files. Each finding is flagged with its location and a GDPR-style review note. Heuristic by design: it surfaces what to review, not a compliance guarantee.
Audits Kubernetes manifests, Helm values, deployment logs, and service configs to detect configuration errors and produce safe, reviewable fix plans.
Find the unit tests that pass without testing anything. Flags tests with no assertions, trivial existence-only checks (toBeDefined, assertIsNotNone), tests that assert the exact value they just mocked, snapshot-only tests, tautological assertions (expect(true).toBe(true)), empty placeholders, and over-mocked tests with more setup than assertions. Works on Jest/Vitest and pytest/unittest.
An adversarial senior engineer review gate that audits AI-written code for security gaps and logic errors before shipping.
Run structural QA on your translation files across locales. Flags missing keys, placeholder mismatches ({name}, %s, {{var}}), strings left untranslated and identical to the source, length-overflow risk that breaks UI, terminology drift against a glossary, empty targets, and plural-category gaps. Works on JSON, gettext .po/.pot, and .properties. It checks form, not meaning, so you do not need to speak the target language to use it.
Generate structured, audit-ready evidence logs documenting human review of AI-assisted work products.
Hardens AI prompts and agent workflows against logic errors, tool-misuse, and prompt injection.
A high-ticket offer diagnostic engine that identifies structural conversion leaks and pricing logic failures.
Audits fragile Zapier, Make, n8n, Airtable, Google Sheets, CRM, webhook, API, and script automations for failure points, data-loss risks, weak logging, missing retries, and risky dependencies.
Check your app for the security mistakes that leak data before you launch, explained in plain English. Flags API keys and secrets sitting in your code, a committed .env file, data with no login protecting it, database tables anyone can read, debug mode left on, wide-open sharing (CORS), hardcoded admin passwords, and public storage buckets. Built for non-technical founders shipping AI-built apps: every finding tells you what is wrong, why it matters, and how to fix it.
Deep conversion psychology audit to identify and fix invisible friction in landing pages and product listings.
Designs controlled multi-agent workflows with clear roles, contracts, validation gates, approval steps, memory boundaries, audit trails, error handling, and fallback behavior.
Adversarially audit your agent hooks before you trust them. Catches command injection, secret leakage, over-broad matchers, destructive actions, and blocking-logic mistakes in pre/post-tool-use, prompt, and stop hooks — with a PASS or REVISE verdict and severity-ranked fixes.
Clean messy tabular data with a profile-first workflow, audit logs, and reproducible pandas scripts.
Map your real cloud dependency tree — data plane and control plane — find the single points of failure ranked by blast radius, catch the classic traps like monitoring that dies with the region it watches, and get a prioritized resilience backlog. The audit every outage post-mortem says to do first.