- Home
- Skills
- Agents & Orchestration
- AI Context Optimization Design
AI Context Optimization Design
Designs dynamic context assembly, token budgeting, caching boundaries, and compaction rules for AI agent runs.
$5
Works with the AI tools you already use
AI Context Optimization Design
Example session with this skill installed
Design runtime context optimization and token budgeting for our Legal Discovery Assistant.
- Subject: LegalReview Assistant v2.1 running on Claude 3.5 Sonnet (200k window, run-id: legal-doc-ctx-001).
- Deciders: Legal AI Architecture Lead (Sarah Chen) and Cost Controller (Marcus Vance).
- Workload & Constraints: Reviews 120-page deposition transcripts (~95,000 tokens) alongside a 15,000-token case playbook and 40 tool logs; target cost ceiling <= $0.45 per review turn; p95 latency <= 4,500 ms.
- Tension: Legal team wants to inject raw deposition pages, exhibits, and past agent turns into every prompt turn; Marcus Vance warns that uncompacted prompts exhaust the $0.45 cost budget after 3 turns and trigger prompt cache misses due to shifting dynamic headers, requiring prefix caching boundaries, recursive summary distillation, and hard token tier ceilings (max 32k active working context).
- Rules: Prompt prefix caching for static playbook (15k tokens); tier-1 essential testimony protected from truncation; rolling turn history compacted when history exceeds 8,000 tokens.
Write the context budget specification under docs/.
- Read your context and instructions
- Compiled the ai context optimization
- Generated the document
Wrote docs/architecture/tasks/legal-doc-ctx-001/context-optimization-design/context-budget-spec.md. Complete context optimization specification establishing prefix caching boundaries, priority token tiers, rolling summarization, and cost ceiling enforcement.
---
skill: context-optimization-design
---
# Context Budget Specification: Legal Discovery Review Agent [CTX-LEG-001]
## Summary
This specification establishes the context optimization and dynamic token budgeting architecture for `LegalReview Assistant v2.1` under run ID `legal-doc-ctx-001` on Claude 3.5 Sonnet. Evaluating 120-page deposition transcripts (~95,000 tokens), it resolves the conflict between complete evidentiary visibility and strict execution cost controls ($0.45 per turn ceiling). It rejects raw, uncompacted full-document injection across multi-turn agent runs, mandating a 32,000-token active working context ceiling, deterministic Anthropic prompt-prefix caching for the static 15,000-token legal playbook, tier-based token prioritization, and rolling extractive transcript summarization.
## Detailed Description
Unconstrained context inflation in document-heavy agent workflows degrades reasoning quality (the "lost-in-the-middle" phenomenon), multiplies token costs, and violates latency SLAs. In multi-turn legal analysis, injecting raw transcripts repeatedly causes linear cost accumulation and repeatedly misses prompt caching due to volatile dynamic timestamps prepended to system instructions.
### Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Execution Cost Containment | Uncompacted 110k token prompts cost ~$0.33/turn in input tokens; 4 turns exceed the $0.45 turn ceiling. | 0.35 | Marcus Vance (Cost Controller) |
| Prompt Cache Hit Ratio | Static playbook text (15k tokens) must hit Anthropic prefix cache with >= 90% frequency. | 0.30 | Claude API Architecture Rule |
| Evidentiary Fidelity (Zero Omission of Key Testimony) | Truncation or compaction must never drop designated Tier-1 critical testimony quotes. | 0.20 | Sarah Chen (Legal AI Lead) |
| Latency Budget Headroom | p95 end-to-end response latency must remain <= 4,500 ms under sustained analytical turns. | 0.15 | SLA intake constraint |
### Comparison
| Candidate Strategy | Active Token Ceiling | Prompt Cache Strategy | Multi-Turn Cost / Turn | Evidence |
|---|---|---|---|---|
| Option A: Raw Ingestion (Legal Proposal) | ~110,000 tokens | Dynamic injection (Cache Miss) | $0.33 input + $0.15 output = $0.48 | Exceeds $0.45 ceiling on turn 1 |
| Option B: Extreme Naive Truncation | 16,000 tokens | No prefix caching | $0.05 input | High risk: Drops critical cross-examination exhibits |
| Option C: Structured Prefix Caching & Tiered Compaction (Chosen) | 32,000 tokens | Pinned 15k prefix with stable breakpoint | $0.06 cached input + $0.05 dynamic = $0.11 | Saves 77% cost; preserves 100% Tier-1 evidence |
### Result
Option C is selected. The prompt template is physically partitioned into a static cached prefix and a dynamic compacted working buffer.
---
### Required Mechanisms
#### 1. Context Layout & Prefix Caching Boundary [MC-CL-01]
To guarantee Anthropic prompt caching hits, volatile elements (timestamps, case IDs) are strictly segregated from the static prefix:
[ Block 1: Static Legal Playbook ] (15,000 tokens)
- 40 standard clause definitions, evidentiary rules, citation format
- Terminates with:
{"type": "text", "text": "...", "cache_control": {"type": "ephemeral"}}
▲ STABLE CACHE BREAKPOINT (Cache Hit: 90% cost reduction on 15k tokens)
────────────────────────────────────────────────────────────────────────
[ Block 2: Tier-1 Extracted Key Deposition Excerpts ] (Up to 10,000 tokens) - Bounded excerpts matching case issue tags; verbatim quote integrity
────────────────────────────────────────────────────────────────────────
[ Block 3: Rolling Compacted Session History ] (Up to 5,000 tokens) - Distilled summaries of turns 1..(N-1); structured JSON observation state
────────────────────────────────────────────────────────────────────────
[ Block 4: Active User Question & Tool Returns ] (Up to 2,000 tokens) - Current turn prompt and latest tool outputs
#### 2. Token Budget Allocation & Priority Tiers [MC-TB-01]
| Priority Tier | Content Type | Hard Token Ceiling | Eviction Policy |
|---|---|---:|---|
| Tier 0 (Pinned) | System Contract & Legal Playbook | 15,000 tokens | Never evicted; immutable cached prefix. |
| Tier 1 (Critical) | Key Testimony & Cross-Examination Quotes | 10,000 tokens | Protected; evicted only if superseding quote verified. |
| Tier 2 (Contextual) | General Transcript Background & Chronology | 4,000 tokens | Compacted via extractive summarization when buffer full. |
| Tier 3 (Volatile) | Multi-Turn Conversation History & Tool Logs | 3,000 tokens | Rolling window; turns older than 3 turns compressed to JSON state. |
| **Total Envelope** | **Combined Active Prompt Context** | **32,000 tokens** | **Hard limit enforced prior to API dispatch.** |
#### 3. Compaction & Distillation Pipeline [MC-CP-01]
- **Trigger**: When conversation history exceeds 8,000 tokens or total active context exceeds 32,000 tokens.
- **Algorithm**:
1. The runtime identifies completed turn pairs older than 2 turns.
2. A local deterministic reducer extracts tool execution statuses, cited document chunk IDs, and legal conclusions into a compact state ledger:
`{"turn": 2, "issue": "BreachOfWarranty", "cited_exhibits": ["EX-04", "EX-12"], "status": "CONFIRMED"}`
3. Raw intermediate tool call logs and model conversational fluff are pruned.
4. Context footprint of old turns drops from ~4,500 tokens to ~350 tokens (92% reduction).
---
### Invariants and Contracts
Prompt Cache Breakpoint Immutability [INV-CTX-01]
The first 15,000 tokens containing the Legal Playbook must remain character-for-character identical
across all turns. Dynamic timestamps, request IDs, and session metadata are strictly forbidden
from being prepended before the cache control breakpoint.
Active Context Hard Ceiling [INV-CTX-02]
The total token count of the assembled prompt passed to Claude 3.5 Sonnet must not exceed 32,000 tokens.
The context assembler must fail closed with an error if compaction fails to bring context below 32,000.
Tier-1 Eviction Prohibition [INV-CTX-03]
The compaction pipeline is prohibited from truncating or paraphrasing Tier-1 verbatim testimony quotes
without explicit user confirmation or automated citation tracking.
## Explicit Unknowns
- Tokenization variance across non-English deposition exhibits translated via external OCR (G-1).
- Anthropic ephemeral prompt cache eviction timing during multi-hour user idle periods (G-2).
## Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Claude 3.5 Sonnet 200k window | provided | Intake specification | Current |
| 120-page transcript (~95k tokens) | provided | Workload intake | Current |
| 15,000-token legal playbook | provided | Intake specification | Current |
| Cost ceiling <= $0.45 per turn | provided | Marcus Vance (Cost SLA) | Current |
| Latency target p95 <= 4,500 ms | provided | Intake constraint | Current |
| 32,000 active token ceiling | decided | Architectural decision INV-CTX-02 | 2026-09-16 |
| 90% cache discount on pinned prefix | observed | Anthropic API pricing tier | 2026-09-16 |
## Verification
No validator was supplied, so no command was run.
Reviewer self-check against context optimization standards:
- **Cacheability Check**: PASS. Static playbook isolated behind `cache_control` breakpoint; dynamic headers removed.
- **Budget Compliance**: PASS. $0.06 cached input + $0.05 dynamic = $0.11/turn, well beneath $0.45 ceiling.
- **Evidence Protection**: PASS. Tier-1 testimony explicitly shielded from lossy summarization.
- **Compaction Safety**: PASS. JSON state ledger preserves legal findings across multi-turn rolling pruning.
## Open Decisions
- `DEC-CTX-01`: Sarah Chen to determine whether exhibits containing complex scanned tables should use markdown table formatting or pre-extracted key-value JSON (Owner: Sarah Chen).
## Next steps
1. Sarah Chen reviews the 15,000-token static playbook markdown to finalize the cache breakpoint structure.
2. Platform team implements the prefix caching wrapper in `services/legal_agent/prompt_assembler.py`.
3. Validate prompt cache hit rate and measure p95 latency across a simulated 10-turn deposition review session.
ai-context-optimization-design.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill defines how one exact AI runtime selects, validates, orders, packs, degrades and optionally compresses/cache-aligns context under supplied limits while preserving authority, provenance and task evidence. It does not own prompt/memory/RAG systems, invent budgets or guarantee quality from fewer tokens.
Use it when
Use when an accepted AI runtime/context contract needs a bounded assembly/overflow optimization policy backed by representative evaluations.
For example: “Our coding assistant costs $14k a month and hits the context limit on large files. Someone suggested we just summarise the conversation history.”
What you get
- Context Optimization Spec
Written as Markdown to <your output folder>/architecture/tasks/<run-id>/context-optimization-design/.
What it will not do
Do not use for prompt architecture/editing, memory or RAG design, model selection, agent/workflow architecture, generic summarization, cache implementation or evaluation execution.
How it works
- Check the budget is actually binding.
- Attribute every token to a component and a purpose.
- Rank components by contribution to the answer, measured.
- Choose a reduction technique per component.
- State what must never be dropped, and what happens at the limit.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-task.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install