codex grade coding
by Roy Yuen
Turn your AI agent into a senior engineer with strict task classification and verification-driven coding protocols.
New: UPI payments are live. Buyers in India can now pay for skills with UPI in INR -> Browse skills
THE AGENSI STORE
40 skills found
by Roy Yuen
Turn your AI agent into a senior engineer with strict task classification and verification-driven coding protocols.
by loreto
Evaluates AI coding agent platforms across five structural dimensions that determine real-world performance independently of model quality, so teams select on architectural fit rather than benchmark scores.
by Sinu
A risk-aware, evidence-based engineering lifecycle protocol for robust agentic task execution and safety.
by Roy Yuen
Upgrade your AI agent with a senior-level engineering SOP focused on inspection, minimal diffs, and hard verification.
by Julian
A 5-gate pre-flight audit to ensure your AI agent has the context, scope, and safety boundaries needed to code successfully.
by Kevin Cline
Automate data profiling with type detection, statistical analysis, and quality flags saved to a Markdown report.
by Perry
Generate professional, formula-driven DMAIC/DFSS artifacts, Excel tools, and interactive Lean Six Sigma visualizations.
by Shandra
Audits AI agent failures and converts recurring mistakes into durable rules, anti-patterns, regression tests, memory candidates, and improved SKILL.md sections.
An adversarial senior engineer review gate that audits AI-written code for security gaps and logic errors before shipping.
by Matthew King
Audit, score, and improve your AI agent skills for higher quality, lower token costs, better reliability, and marketplace success. Get actionable recommendations for prompts, instructions, tool usage, error handling, and user experience.
Audit local dbt SQL and YAML for missing model tests, source freshness and test gaps, likely-key coverage, missing model descriptions, SELECT *, and raw table references. Get severity-ranked findings plus starter tests YAML for model/key gaps without running dbt or querying a warehouse.
Continuously monitor data pipelines, detect anomalies, and explain root causes before failures impact production.
Builds a complete LLM evaluation framework — quality dimensions, a golden dataset, code-based and model-graded rubric graders, judge calibration, and CI regression rules. Use when the user says build LLM evals, create a golden dataset, or set up LLM-as-judge. Do not use when they want to debug one bad model output, not build a repeatable measurement system.
by Shogun Labs
Battle-tested prompting patterns to eliminate LLM output drift. Sandwich structure, few-shot examples, history limits, retry, and token caps — 6 composable layers for production-grade agent reliability.
by Joker
Data quality diagnosis, dedup strategies, format standardization, anomaly handling, batch processing.
by Joker
Governance framework, data quality, metadata, compliance, data lineage, 2026 trends.
by Joker
Model selection matrix (MJ/Flux/SD), composition rules, 3-level quality gates, 10 style recipes.
by GTDataworks
Run a buyer-readiness check before publishing an AI agent skill package.
by SkillForge
Force your AI agent to apply senior-level architectural judgment and design notes before writing any code.
by Shandra
Tests AI agents, prompts, and agent skills against edge cases, unsafe behavior, output failures, permission risks, escalation gaps, memory leaks, and marketplace-quality weaknesses.
by Kaymue
Pre-merge PR quality gate. 40+ checks: diff size, tests, security, conventional commits, breaking changes, reviewer routing.
by Timoranjes
Transform your AI agent from a code generator into a senior architect that enforces clean design and SOLID principles.
Model quality is table stakes — the harness is where agents win or fail. This designs yours: it writes a structured, testable system prompt (role, tools, boundaries, method, output contract, failure handling) and maps every concern to the right layer — prompt, tool, guardrail, or evaluation — so the pieces reinforce each other instead of fighting.
by PromptWagon
Reviews document sets, source quality, chunking logic, metadata, retrieval coverage, citation traceability, answer grounding, source gaps, stale content, duplicate content, and failure patterns for RAG knowledge-base chatbots. Helps AI, product, support, governance, and engineering teams diagnose common and costly RAG quality problems before deployment or after incidents.