Bounded RAG Surface Design

    1

    Designs bounded RAG pipelines: document ingestion, chunking rules, retrieval thresholds, citations, and abstention logic.

    $5

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Bounded RAG Surface Design

    Example session with this skill installed

    Design the bounded RAG surface specification for our Employee Benefits Policy Assistant.

    • Subject: BenefitsBot v1.2 answering employee questions on health insurance, 401(k), and parental leave (run-id: policy-bot-rag-001).
    • Deciders: HR Operations Director (Rachel Adams) and AI Engineering Lead (Marcus Vance).
    • Corpus Scope: 42 official benefits policy PDFs (~1,800 total pages); Markdown summary handbook (150 pages).
    • Traffic & Budget: Peak 25 queries/sec during open enrollment; p95 latency budget <= 1,800 ms; max context retrieval budget 4,000 tokens (top-4 chunks).
    • Tension: HR leads want the bot to answer edge-case health claims with general advice when exact policy text is missing; Marcus Vance warns that hallucinated medical coverage creates severe corporate liability, mandating strict cosine similarity cutoffs (>= 0.76), mandatory chunk citation tags ([^chunk_id]), and explicit abstention when no verified policy clause matches.
    • Constraints: Chunking: 400 tokens with 40-token overlap; metadata must capture policy_id, effective_date, and jurisdiction.

    Write the RAG surface specification under docs/.

    • Read your context and instructions
    • Compiled the bounded rag surface
    • Generated the UI component

    Wrote docs/architecture/tasks/policy-bot-rag-001/rag-design-task/rag-surface-spec.md. Complete bounded RAG surface specification establishing 400-token chunking rules, metadata schemas, a 0.76 cosine similarity cutoff, mandatory chunk-level citations, and strict liability-free abstention.

    ---
    skill: rag-design-task
    ---
    
    # RAG Surface Specification: Employee Benefits Policy Assistant [RAG-BEN-001]
    
    ## Summary
    
    This specification establishes the bounded Retrieval-Augmented Generation (RAG) surface for `BenefitsBot v1.2` under run ID `policy-bot-rag-001`. It governs semantic retrieval and grounded answer synthesis across 42 official HR benefits policy PDFs (~1,800 pages) and the company handbook. To eliminate employer liability arising from hallucinated health insurance coverage, the design decisively rejects conversational extrapolation for missing policy terms. It enforces deterministic 400-token chunking with 40-token overlap, rich metadata tagging (jurisdiction, effective date, policy ID), a strict cosine similarity cutoff (>= 0.76), top-4 candidate context budgeting (capped at 1,600 tokens), mandatory verifiable chunk citations (`[^chunk_id]`), and fail-safe abstention.
    
    ## Detailed Description
    
    Employee benefits inquiries involve legal and financial implications. When employees ask about specific elective surgery coverages or parental leave eligibility dates, the assistant must ground statements exclusively in authoritative HR documents. Speculative or extrapolated answers risk corporate ERISA non-compliance and financial disputes.
    
    

    Employee Query
    │
    ▼
    [ Query Embedding: text-embedding-3-small ]
    │
    ▼
    [ Vector Search & Metadata Filter: Qdrant ]
    (Filter: active effective_date <= today)
    │
    ├─► Max Cosine Similarity < 0.76 ──► [ ABSTAIN: "Policy not found" ]
    │
    ▼ (Top-4 Chunks >= 0.76, max 1,600 tokens)
    [ Grounded Synthesis Prompt (GPT-4o-mini) ]
    │
    ▼
    [ Post-Generation Citation Validator ] ──► Verified Answer + HR Contact Link

    
    ### Criteria and weights
    
    | Criterion | Why it matters here | Weight | Source of the weight |
    |---|---|---|---|
    | Zero Hallucinated Coverage (Grounding Fidelity) | Speculative insurance answers create direct company liability under ERISA regulations. | 0.40 | Rachel Adams (HR Operations) |
    | Latency Budget (p95 <= 1,800 ms) | High concurrent open enrollment queries must not experience conversational lag. | 0.25 | Intake SLA requirement |
    | Citation Traceability | Every factual benefit assertion must link back to an exact page in the canonical PDF policy. | 0.20 | Marcus Vance (AI Lead) |
    | Token Cost Efficiency | Top-k retrieval must remain bounded to avoid unnecessary context window inflation. | 0.15 | Cost Governance Policy |
    
    
    ### Comparison
    
    | Design Option | Similarity Threshold | Chunk Size / Overlap | Citation Contract | Residual Hallucination Risk |
    |---|---|---|---|---|
    | Option A: Permissive Top-K Retrieval | None (always returns top-5) | 1,000 tokens / 100 overlap | General document title | Critical: Synthesizes speculative answers when topics are missing. |
    | Option B: Dense Hybrid without Cutoff | Score >= 0.65 | 512 tokens / 50 overlap | Section heading only | Medium: Outdated 2024 policy chunks leak into 2026 open enrollment queries. |
    | Option C: Bounded Strict-Cutoff RAG (Chosen) | Score >= 0.76 + Date Filter | 400 tokens / 40 overlap | Deterministic `[^chunk_id]` regex | Minimal: Abstains immediately if authoritative text is absent. |
    
    
    ### Result
    
    Option C is selected. Strict similarity gating combined with temporal metadata filtering ensures only active, highly relevant policy provisions enter the LLM context window.
    
    ---
    
    ### Required Mechanisms
    
    #### 1. Task Contract & Corpus Boundary [MC-TC-01]
    - **Ingestion Corpus**: 42 official benefits policy PDFs and the 150-page Employee Handbook markdown file. External websites or public health portals are strictly excluded.
    - **Document Metadata Schema**:
      ```json
    
    ```json
      {
        "document_id": "BEN-2026-HLTH-01",
    
    "title": "Comprehensive Healthcare Plan Summary 2026",
    
        "jurisdiction": "US-ALL",
        "effective_date": "2026-01-01",
        "version": "2.0",
        "source_uri": "s3://hr-docs/benefits/2026/health_summary.pdf"
      }
    
    
    #### 2. Ingestion & Chunking Specification [MC-IC-01]
    - **Chunk Size**: Fixed 400 tokens.
    - **Overlap**: Derived as 400 * 0.10 = 40 tokens.
    - **Chunk Schema**:
      - `chunk_id`: Formatted as `<document_id>_p<page_number>_c<chunk_index>` (e.g. `BEN-2026-HLTH-01_p14_c02`).
      - `page_number`: Integer reference for citation mapping.
      - `content`: Plain UTF-8 text with headers prepended.
    
    #### 3. Retrieval & Scoring Thresholds [MC-RS-01]
    - **Embedding Model**: `text-embedding-3-small` (1536 dimensions, cosine distance).
    - **Candidate Retrieval**:
      1. Apply Qdrant payload pre-filter: `effective_date <= CURRENT_DATE` and `jurisdiction IN ("US-ALL", user_state)`.
      2. Retrieve top-10 nearest neighbors.
      3. Filter candidates: discard all chunks with cosine similarity score < 0.76.
      4. Select at most top-4 remaining chunks (total context payload <= 1,600 tokens).
    
    #### 4. Generation & Grounding Contract [MC-GG-01]
    - **Prompt Directive**:
      ```markdown
      You are an HR Benefits Assistant. Answer the question using ONLY the provided context chunks below.
      Every factual statement must end with an exact citation referencing the chunk ID: [^chunk_id].
      If the provided context does not explicitly contain the answer, you must output:
      "I cannot find this specific coverage in official benefits documentation. Please consult HR directly."
    
    • Validation Oracle: A regex parser asserts all bracketed citations match IDs present in the injected prompt context. Citations of unprovided documents cause immediate answer rejection.
    5. Failure & Abstention Semantics [MC-FA-01]
    • Abstention Triggers:
      1. Zero retrieved chunks pass the 0.76 similarity threshold.
      2. Retrieved chunks contain conflicting policy rules across overlapping jurisdictions.
    • Customer Response:
      "I cannot find this specific coverage in official benefits documentation. Please contact the HR Benefits Desk at benefits@company.com."

    Invariants and Contracts

    Strict Similarity Cutoff Invariant [INV-RAG-01]
      No chunk with cosine similarity score < 0.76 may be injected into the generation prompt.
      If all candidates score below 0.76, the system must immediately return the standard abstention message.
    
    Deterministic Citation Invariant [INV-RAG-02]
      Every factual claim in the synthesized response must terminate with an exact chunk citation
      in the format `[^<document_id>_p<page>_c<index>]`. Unsupported assertions are stripped.
    
    Temporal Metadata Invariant [INV-RAG-03]
      The vector pre-filter must exclude documents whose `effective_date` is in the future or which
      have been superseded by a newer version with identical `policy_id`.
    

    Explicit Unknowns

    • Handling of state-specific dental insurance variances in Puerto Rico and Hawaii (G-1).
    • PDF table extraction fidelity on complex dental tier comparison matrices (G-2).

    Traceability

    ClaimClassificationSourceFreshness
    42 policy PDFs (~1,800 pages)providedIntake corpus inventoryCurrent
    25 QPS peak load during open enrollmentprovidedTraffic profileCurrent
    Latency budget p95 <= 1,800 msprovidedIntake constraintCurrent
    Cosine cutoff >= 0.76decidedMarcus Vance & Rachel Adams2026-09-15
    400 token chunk / 40 token overlapprovidedRequest constraintCurrent
    Mandatory chunk-level citationsdecidedArchitectural invariant INV-RAG-022026-09-15

    Verification

    No validator was supplied, so no command was run.

    Reviewer self-check against RAG surface contracts:

    • Chunk Geometry: PASS. 400 token chunk size with derived 40 token overlap (10%).
    • Score Gating: PASS. Strict 0.76 threshold stops generation on low-confidence matches.
    • Citation Compliance: PASS. Regex validator enforces [^chunk_id] provenance tracing back to source page.
    • Abstention Safeguard: PASS. Hardcoded fallback prevents speculative coverage answers.

    Open Decisions

    • DEC-RAG-01: Rachel Adams to confirm whether bilingual Spanish translations of benefits policies should share vector collections or use distinct language namespaces (Owner: Rachel Adams).

    Next steps

    1. Ingestion team runs PDF chunking extractor script over all 42 benefits documents.
    2. Configure Qdrant collection hr_benefits_2026 with payload indices on effective_date and jurisdiction.
    3. Execute evaluation benchmark against 100 known benefits queries asserting 100% citation compliance.

    bounded-rag-surface-design.tsx

    TSX · React component

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Define structural chunking rules for complex documents like tables.Set retrieval thresholds and reranking logic for context windows.Establish citation formats and fallback abstention messaging.Map governed corpus sources to specific query and answer surfaces.

    About this skill

    What it does

    This skill maps one accepted question/claim/answer surface to governed source revisions, derived retrieval representations, access-aware query/retrieval/reranking, selected context, generation/citation and evaluation contracts. It does not choose chunks, embeddings or vector databases from generic defaults.

    Use it when

    Use when an accepted RAG architecture needs one scoped retrieval-to-generation surface designed under existing source and evaluation authority.

    For example: “Our policy assistant answers from the staff handbook. It quotes the right paragraph about half the time, and for anything in a table it makes something up.”

    What you get

    • RAG Architecture Specification

    Written as Markdown to <your output folder>/architecture/tasks/<run-id>/rag-design-task/.

    What it will not do

    Do not use for cross-system RAG architecture, generic search/vector-index design, one ingestion job, embedding/vector-DB selection, chunk or retrieval tuning, memory/context/prompt design, answering from documents, evaluation execution or framework implementation.

    How it works

    1. Check retrieval is required for this surface.
    2. Bound the corpus and its access rule.
    3. Choose the segmentation against the question shape, then measure it.
    4. Define retrieval, reranking and the evidence budget.
    5. Fix citation and abstention.
    6. Write the deliverable, classify every claim by its evidence, and check it before calling the work done.

    What's in the package

    Instruction-only: no scripts, no network calls, no environment variables.

    • LICENSE.txt
    • SKILL.md
    • agents/openai.yaml
    • assets/output-template-task.md
    • references/domain-rules.md
    • references/operating-rules.md
    • references/output-contract.md

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Passed all security checks, Safe to install

    Listed12 days ago

    What's inside

    Frequently Asked Questions