- Home
- Skills
- Agents & Orchestration
- Bounded RAG Surface Design
Bounded RAG Surface Design
Designs bounded RAG pipelines: document ingestion, chunking rules, retrieval thresholds, citations, and abstention logic.
$5
Works with the AI tools you already use
Bounded RAG Surface Design
Example session with this skill installed
Design the bounded RAG surface specification for our Employee Benefits Policy Assistant.
- Subject: BenefitsBot v1.2 answering employee questions on health insurance, 401(k), and parental leave (run-id: policy-bot-rag-001).
- Deciders: HR Operations Director (Rachel Adams) and AI Engineering Lead (Marcus Vance).
- Corpus Scope: 42 official benefits policy PDFs (~1,800 total pages); Markdown summary handbook (150 pages).
- Traffic & Budget: Peak 25 queries/sec during open enrollment; p95 latency budget <= 1,800 ms; max context retrieval budget 4,000 tokens (top-4 chunks).
- Tension: HR leads want the bot to answer edge-case health claims with general advice when exact policy text is missing; Marcus Vance warns that hallucinated medical coverage creates severe corporate liability, mandating strict cosine similarity cutoffs (>= 0.76), mandatory chunk citation tags ([^chunk_id]), and explicit abstention when no verified policy clause matches.
- Constraints: Chunking: 400 tokens with 40-token overlap; metadata must capture policy_id, effective_date, and jurisdiction.
Write the RAG surface specification under docs/.
- Read your context and instructions
- Compiled the bounded rag surface
- Generated the UI component
Wrote docs/architecture/tasks/policy-bot-rag-001/rag-design-task/rag-surface-spec.md. Complete bounded RAG surface specification establishing 400-token chunking rules, metadata schemas, a 0.76 cosine similarity cutoff, mandatory chunk-level citations, and strict liability-free abstention.
---
skill: rag-design-task
---
# RAG Surface Specification: Employee Benefits Policy Assistant [RAG-BEN-001]
## Summary
This specification establishes the bounded Retrieval-Augmented Generation (RAG) surface for `BenefitsBot v1.2` under run ID `policy-bot-rag-001`. It governs semantic retrieval and grounded answer synthesis across 42 official HR benefits policy PDFs (~1,800 pages) and the company handbook. To eliminate employer liability arising from hallucinated health insurance coverage, the design decisively rejects conversational extrapolation for missing policy terms. It enforces deterministic 400-token chunking with 40-token overlap, rich metadata tagging (jurisdiction, effective date, policy ID), a strict cosine similarity cutoff (>= 0.76), top-4 candidate context budgeting (capped at 1,600 tokens), mandatory verifiable chunk citations (`[^chunk_id]`), and fail-safe abstention.
## Detailed Description
Employee benefits inquiries involve legal and financial implications. When employees ask about specific elective surgery coverages or parental leave eligibility dates, the assistant must ground statements exclusively in authoritative HR documents. Speculative or extrapolated answers risk corporate ERISA non-compliance and financial disputes.
Employee Query
│
▼
[ Query Embedding: text-embedding-3-small ]
│
▼
[ Vector Search & Metadata Filter: Qdrant ]
(Filter: active effective_date <= today)
│
├─► Max Cosine Similarity < 0.76 ──► [ ABSTAIN: "Policy not found" ]
│
▼ (Top-4 Chunks >= 0.76, max 1,600 tokens)
[ Grounded Synthesis Prompt (GPT-4o-mini) ]
│
▼
[ Post-Generation Citation Validator ] ──► Verified Answer + HR Contact Link
### Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Zero Hallucinated Coverage (Grounding Fidelity) | Speculative insurance answers create direct company liability under ERISA regulations. | 0.40 | Rachel Adams (HR Operations) |
| Latency Budget (p95 <= 1,800 ms) | High concurrent open enrollment queries must not experience conversational lag. | 0.25 | Intake SLA requirement |
| Citation Traceability | Every factual benefit assertion must link back to an exact page in the canonical PDF policy. | 0.20 | Marcus Vance (AI Lead) |
| Token Cost Efficiency | Top-k retrieval must remain bounded to avoid unnecessary context window inflation. | 0.15 | Cost Governance Policy |
### Comparison
| Design Option | Similarity Threshold | Chunk Size / Overlap | Citation Contract | Residual Hallucination Risk |
|---|---|---|---|---|
| Option A: Permissive Top-K Retrieval | None (always returns top-5) | 1,000 tokens / 100 overlap | General document title | Critical: Synthesizes speculative answers when topics are missing. |
| Option B: Dense Hybrid without Cutoff | Score >= 0.65 | 512 tokens / 50 overlap | Section heading only | Medium: Outdated 2024 policy chunks leak into 2026 open enrollment queries. |
| Option C: Bounded Strict-Cutoff RAG (Chosen) | Score >= 0.76 + Date Filter | 400 tokens / 40 overlap | Deterministic `[^chunk_id]` regex | Minimal: Abstains immediately if authoritative text is absent. |
### Result
Option C is selected. Strict similarity gating combined with temporal metadata filtering ensures only active, highly relevant policy provisions enter the LLM context window.
---
### Required Mechanisms
#### 1. Task Contract & Corpus Boundary [MC-TC-01]
- **Ingestion Corpus**: 42 official benefits policy PDFs and the 150-page Employee Handbook markdown file. External websites or public health portals are strictly excluded.
- **Document Metadata Schema**:
```json
```json
{
"document_id": "BEN-2026-HLTH-01",
"title": "Comprehensive Healthcare Plan Summary 2026",
"jurisdiction": "US-ALL",
"effective_date": "2026-01-01",
"version": "2.0",
"source_uri": "s3://hr-docs/benefits/2026/health_summary.pdf"
}
#### 2. Ingestion & Chunking Specification [MC-IC-01]
- **Chunk Size**: Fixed 400 tokens.
- **Overlap**: Derived as 400 * 0.10 = 40 tokens.
- **Chunk Schema**:
- `chunk_id`: Formatted as `<document_id>_p<page_number>_c<chunk_index>` (e.g. `BEN-2026-HLTH-01_p14_c02`).
- `page_number`: Integer reference for citation mapping.
- `content`: Plain UTF-8 text with headers prepended.
#### 3. Retrieval & Scoring Thresholds [MC-RS-01]
- **Embedding Model**: `text-embedding-3-small` (1536 dimensions, cosine distance).
- **Candidate Retrieval**:
1. Apply Qdrant payload pre-filter: `effective_date <= CURRENT_DATE` and `jurisdiction IN ("US-ALL", user_state)`.
2. Retrieve top-10 nearest neighbors.
3. Filter candidates: discard all chunks with cosine similarity score < 0.76.
4. Select at most top-4 remaining chunks (total context payload <= 1,600 tokens).
#### 4. Generation & Grounding Contract [MC-GG-01]
- **Prompt Directive**:
```markdown
You are an HR Benefits Assistant. Answer the question using ONLY the provided context chunks below.
Every factual statement must end with an exact citation referencing the chunk ID: [^chunk_id].
If the provided context does not explicitly contain the answer, you must output:
"I cannot find this specific coverage in official benefits documentation. Please consult HR directly."
- Validation Oracle: A regex parser asserts all bracketed citations match IDs present in the injected prompt context. Citations of unprovided documents cause immediate answer rejection.
5. Failure & Abstention Semantics [MC-FA-01]
- Abstention Triggers:
- Zero retrieved chunks pass the 0.76 similarity threshold.
- Retrieved chunks contain conflicting policy rules across overlapping jurisdictions.
- Customer Response:
"I cannot find this specific coverage in official benefits documentation. Please contact the HR Benefits Desk at benefits@company.com."
Invariants and Contracts
Strict Similarity Cutoff Invariant [INV-RAG-01]
No chunk with cosine similarity score < 0.76 may be injected into the generation prompt.
If all candidates score below 0.76, the system must immediately return the standard abstention message.
Deterministic Citation Invariant [INV-RAG-02]
Every factual claim in the synthesized response must terminate with an exact chunk citation
in the format `[^<document_id>_p<page>_c<index>]`. Unsupported assertions are stripped.
Temporal Metadata Invariant [INV-RAG-03]
The vector pre-filter must exclude documents whose `effective_date` is in the future or which
have been superseded by a newer version with identical `policy_id`.
Explicit Unknowns
- Handling of state-specific dental insurance variances in Puerto Rico and Hawaii (G-1).
- PDF table extraction fidelity on complex dental tier comparison matrices (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 42 policy PDFs (~1,800 pages) | provided | Intake corpus inventory | Current |
| 25 QPS peak load during open enrollment | provided | Traffic profile | Current |
| Latency budget p95 <= 1,800 ms | provided | Intake constraint | Current |
| Cosine cutoff >= 0.76 | decided | Marcus Vance & Rachel Adams | 2026-09-15 |
| 400 token chunk / 40 token overlap | provided | Request constraint | Current |
| Mandatory chunk-level citations | decided | Architectural invariant INV-RAG-02 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against RAG surface contracts:
- Chunk Geometry: PASS. 400 token chunk size with derived 40 token overlap (10%).
- Score Gating: PASS. Strict 0.76 threshold stops generation on low-confidence matches.
- Citation Compliance: PASS. Regex validator enforces
[^chunk_id]provenance tracing back to source page. - Abstention Safeguard: PASS. Hardcoded fallback prevents speculative coverage answers.
Open Decisions
DEC-RAG-01: Rachel Adams to confirm whether bilingual Spanish translations of benefits policies should share vector collections or use distinct language namespaces (Owner: Rachel Adams).
Next steps
- Ingestion team runs PDF chunking extractor script over all 42 benefits documents.
- Configure Qdrant collection
hr_benefits_2026with payload indices oneffective_dateandjurisdiction. - Execute evaluation benchmark against 100 known benefits queries asserting 100% citation compliance.
bounded-rag-surface-design.tsx
TSX · React component
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill maps one accepted question/claim/answer surface to governed source revisions, derived retrieval representations, access-aware query/retrieval/reranking, selected context, generation/citation and evaluation contracts. It does not choose chunks, embeddings or vector databases from generic defaults.
Use it when
Use when an accepted RAG architecture needs one scoped retrieval-to-generation surface designed under existing source and evaluation authority.
For example: “Our policy assistant answers from the staff handbook. It quotes the right paragraph about half the time, and for anything in a table it makes something up.”
What you get
- RAG Architecture Specification
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/rag-design-task/.
What it will not do
Do not use for cross-system RAG architecture, generic search/vector-index design, one ingestion job, embedding/vector-DB selection, chunk or retrieval tuning, memory/context/prompt design, answering from documents, evaluation execution or framework implementation.
How it works
- Check retrieval is required for this surface.
- Bound the corpus and its access rule.
- Choose the segmentation against the question shape, then measure it.
- Define retrieval, reranking and the evidence budget.
- Fix citation and abstention.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install