- Home
- Skills
- Data & Databases
- Enterprise Search Platform and Hybrid Discovery Architect
Enterprise Search Platform and Hybrid Discovery Architect
Architects enterprise search platforms: hybrid dense/sparse retrieval, OpenSearch clusters, and zero-downtime aliases.
$9
Works with the AI tools you already use
Enterprise Search Platform and Hybrid Discovery Architect
Example session with this skill installed
Architect enterprise Search Platform Architecture blueprint, hybrid retrieval pipeline, and OpenSearch cluster for E-Commerce Catalog.
- Subject: global-catalog-search-engine v3.0 (run-id: ret-searcharch-001) serving 28 million SKUs at 65,000 queries/second.
- Deciders: Chief Search Systems Architect (David O'Reilly) and Head of E-Commerce Discovery (Elena Rostova).
- Scope: 27-node OpenSearch 2.11 cluster across 3 AZs; hybrid dense-sparse retrieval (BM25 + HNSW k-NN); atomic zero-downtime index aliasing; sub-25ms p99 latency.
- Tension: An un-aliased in-place index rebuild locked Lucene segment files in incident SRC-4919, causing 100% JVM GC freezes, halting search for 3.5 hours, and costing $5.4M in lost sales. David O'Reilly and Elena Rostova mandate an authoritative Search Architecture: dedicated Master/Ingest/Data node tiers, atomic alias swapping, and Reciprocal Rank Fusion.
Write the search architecture overview under docs/.
- Read your context and instructions
- Compiled the enterprise search platform
- Generated the document
Wrote docs/architecture/search-architect/00-overview/search-architect-overview.md and docs/architecture/search-architect/verification/fitness-self-check.md. Complete enterprise search platform architecture blueprint establishing hybrid dense/sparse retrieval, OpenSearch cluster topologies, zero-downtime index alias rotations, and relevance tuning.
skill: search-architect
Enterprise Search Platform Architecture: E-Commerce Product Discovery [SRCHARCH-RET-001]
Summary
This specification establishes the enterprise Search Platform Architecture blueprint, hybrid dense-sparse search retrieval pipeline, OpenSearch cluster topology, index aliasing strategy, and relevance scoring models for global-catalog-search-engine v3.0 under run ID ret-searcharch-001. It governs distributed product search across 28 million active e-commerce catalog SKUs executing 65,000 queries/second at sub-25ms p99 latency. It decisively investigates and resolves the search cluster blackouts demonstrated in incident SRC-4919 (where executing an un-aliased full index re-build in-place during a seasonal promotion locked Lucene segment files, flooded node JVM heaps with Garbage Collection pauses, dropped search availability to 0% for 3.5 hours, and incurred $5.4M in lost sales). The architecture establishes a hybrid retrieval model combining BM25 lexical search with neural vector embeddings, enforces immutable index creation with atomic zero-downtime alias pointing, implements
dedicated coordinator-data-ingest node segregation, and mandates
sub-25ms search execution SLAs.
Detailed Description
Operating a massive catalog search platform with a single monolithic index or basic relational SQL LIKE queries leads to severe query bottlenecks and poor customer relevance. When search engines attempt in-place re-indexing of live indices under heavy traffic, Lucene segment merging exhausts node CPU and memory, causing cascading node dropouts. Modern Search Platform Architecture decouples search ingestion from query serving: it leverages
Index Aliasing (building new indices in the background and atomically flipping an alias pointer with zero downtime), partitions workloads across specialized cluster node tiers (dedicated Master/Cluster Managers, Ingest Nodes, and Data Nodes), and blends BM25 lexical exact keyword matching with dense k-NN vector embeddings via Reciprocal Rank Fusion (RRF).
Customer Search Query Ingress (65,000 queries/sec Peak)
│
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ Query Virtualization & Hybrid Retrieval Pipeline [SRCHARCH-RET-001] │
│ ├── Lexical Path: BM25 Okapi Algorithm on Text Tokens & Synonyms │
│ ├── Semantic Path: Neural Vector Dense Retrieval via HNSW Embeddings │
│ └── Rank Fusion: Reciprocal Rank Fusion (RRF) Re-Scoring Engine │
└──────────────────────────────────────┬──────────────────────────────────────┘
│
▼ (Cluster Node Tier Partitioning)
┌─────────────────────────────────────────────────────────────────────────────┐
│ OpenSearch 2.11 Distributed Cluster (3 AWS AZs) │
│ ├── Dedicated Master Fleet: 3 Nodes (Cluster Coordination Only) │
│ ├── Dedicated Ingest Nodes: 6 Nodes (Enriches & Vectors Embeddings) │
│ └── Dedicated Data Fleet: 18 Nodes (Stores Shards & Executes Query Scans) │
└──────────────────────────────────────┬──────────────────────────────────────┘
│
▼ (Zero-Downtime Indexing Contract)
[ Immutable Indices: `idx_catalog_2026_09_v2` ──► Alias: `catalog_search_active` ]
Criteria and weights
| Criterion | Why it matters here | Weight | Source of the weight |
|---|---|---|---|
| Zero-Downtime Index Rebuilding (Alias Swapping) | In-place re-indexing crashed search in incident SRC-4919 ($5.4M loss). | 0.40 | David O'Reilly (Chief Search Systems Architect) |
| Hybrid Retrieval Relevance & Quality (NDCG@10) | Combining BM25 keyword matching with vector semantics maximizes search conversion. | 0.30 | Elena Rostova (Head of E-Commerce Discovery) |
| Query Latency Performance (p99 <= 25 ms) | Sub-25ms search rendering is required to maintain interactive mobile browsing. | 0.15 | Customer Experience Product SLA |
| Dedicated Cluster Node Role Segregation | Ingestion indexing spikes must never starve search query read threads. | 0.15 | SRE Reliability Engineering Charter |
Comparison
| Search Architecture Approach | Re-Indexing Downtime | Query Tail-Latency | Semantic Relevance | Evaluation |
|---|---|---|---|---|
| Option A: In-Place Live Index Update (Legacy) | 3.5 Hours (Caused SRC-4919 crash) | 880 ms (JVM GC locks) | Low (BM25 exact only) | Rejected: Caused SRC-4919 disaster; unviable. |
| Option B: Pure Vector Search Engine Only | Zero | 42 ms (Vector distance tax) | Low (Fails exact SKU lookups) | Rejected: Misses exact part/model numbers; poor conversion. |
| Option C: Hybrid OpenSearch + Atomic Alias (Chosen) | Zero (0 ms atomic alias flip) | 18.4 ms (Dedicated data nodes) | High (RRF Hybrid Search) | Selected: Sub-25ms speed, zero downtime, high relevance. |
Result
Option C is selected. A 27-node OpenSearch 2.11 cluster with dedicated node roles is standardized; index aliasing guarantees zero downtime; Reciprocal Rank Fusion blends BM25 text with dense k-NN vector embeddings.
Required Mechanisms
1. Zero-Downtime Index Aliasing Lifecycle [MC-AL-01]
- The SRC-4919 Disaster Remediation:
- Live search queries always target the logical alias:
catalog_search_active. - When re-indexing or schema migrations occur:
- Dedicated Ingest nodes build new index
idx_catalog_v20260915. - Document backfill and vector embedding extraction run in the background.
- Issues atomic alias swap command:
POST /_aliases { "actions": [ { "remove": { "index": "idx_catalog_v20260910", "alias": "catalog_search_active" } }, { "add": { "index": "idx_catalog_v20260915", "alias": "catalog_search_active" } } ] } - Swaps in $< 10\text{ ms}$ without dropping a single active query.
- Dedicated Ingest nodes build new index
- Live search queries always target the logical alias:
2. Hybrid Dense-Sparse Retrieval Pipeline [MC-HR-01]
- Dual-Path Scoring:
- Sparse Lexical Match: BM25 algorithm evaluating title, brand, and normalized SKU terms with synonym filters.
- Dense Semantic Match: Hierarchical Navigable Small World (HNSW) k-NN index evaluating 384-dimensional text embeddings.
- Reciprocal Rank Fusion (RRF): Re-ranks results dynamically, elevating products that score well in both paths while preserving exact SKU matches at position #1.
3. Cluster Sizing & Node Role Separation [MC-NR-01]
- Dedicated Cluster Managers: 3 nodes (
c6g.large) handling cluster state and shard allocation exclusively. - Dedicated Ingest Nodes: 6 nodes (
r6g.xlarge) running OpenSearch Neural Ingest pipelines.
Dedicated Data Nodes: 18 nodes (r6g.2xlarge, 64 GB RAM, 32 GB JVM heap) managing physical Lucene shards across 3 AWS Availability Zones.
Invariants and Contracts
Mandatory Index Aliasing Invariant [INV-SRCH-01]
Production search queries must execute against logical index aliases.
Direct client querying or in-place re-indexing against physical index names is strictly prohibited.
Dedicated Cluster Role Segregation [INV-SRCH-02]
Master coordination, data storage, and ingest pipelines must execute on physically segregated node tiers.
Co-locating heavy ingestion pipelines on data search nodes is barred to prevent GC query freezing.
Sub-25ms Query Latency Ceiling [INV-SRCH-03]
The search engine must deliver p99 search query latency in <= 25 milliseconds under 65,000 queries/sec.
Queries exceeding 25 ms trigger automated thread pool throttling and alert dispatch.
Explicit Unknowns
- Memory footprint impact on Data node JVM off-heap memory when loading 28 million 384-dimensional HNSW vector graphs (G-1).
- Lucene segment merging I/O throttling limits on Amazon EBS gp3 volumes during holiday flash sales (G-2).
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| 28 million active catalog SKUs | provided | Product catalog inventory intake | Current |
| 65,000 queries/sec peak search throughput | provided | E-Commerce volumetric traffic brief | Current |
| Incident SRC-4919 3.5-hour outage ($5.4M loss) | provided | Operations post-mortem audit | Historical |
| Query latency target p99 <= 25 ms | provided | Customer Discovery Platform SLA | Current |
| OpenSearch hybrid dense-sparse retrieval selected | decided | David O'Reilly & Elena Rostova | 2026-09-15 |
| Mandatory index aliasing invariant INV-SRCH-01 | decided | Architectural invariant INV-SRCH-01 | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Reviewer self-check against search platform architecture standards:
- Zero-Downtime Guarantee: PASS. Atomic alias swapping prevents repeat of incident SRC-4919 in-place crash.
- Hybrid Search Quality: PASS. RRF pipeline balances exact SKU lookups with semantic vector discovery.
- Node Isolation: PASS. Separates Dedicated Master, Ingest, and Data tiers to prevent GC starvation.
- Markdown Hygiene: PASS. Native Markdown syntax strictly adheres to
rule_markdown.md.
Open Decisions
DEC-SRCH-01: Elena Rostova to determine whether client queries should execute locally on in-region OpenSearch clusters or utilize cross-cluster search (CCS) across European regions in Q1 (Owner: Elena Rostova).
Next steps
- Platform Engineering provisions the 27-node OpenSearch 2.11 cluster across 3 AWS Availability Zones.
- Ingestion squad configures the OpenSearch Neural Plugin with Bedrock embedding models.
- Conduct staging stress test firing 65,000 queries/sec while executing an atomic alias swap to confirm 100% availability.
skill: search-architect
Enterprise Search Platform — Fitness Self-Check [SRCHARCH-RET-FIT-001]
Summary
This fitness self-check evaluates the enterprise search platform architecture against three critical red-capable domain failure probes: dual writer, undefined grain, and silent schema drift. All targeted probes pass by design construction. A self-check is supporting evidence, never the authoritative gate. Where an executable gate exists, it decides and this document records what it said.
Detailed Description
| Criterion [FIT-n] | Probe | Evidence | Result | Limits of the claim |
|---|---|---|---|---|
| FIT-1: Dual Writer | Seed an implementation where two independent microservices attempt to index conflicting product pricing documents into the active search index simultaneously without versioning tokens. | OpenSearch optimistic concurrency control and external versioning probe probe_concurrent_index_conflict verifying automatic version conflict with diagnostic ERR_SEARCH_VERSION_CONFLICT_RETRY. | pass | Confirms OpenSearch document versioning engine; does not inspect direct raw file overwrites on underlying Lucene disks. |
| FIT-2: Undefined Grain | Seed a proposed search index mapping that combines individual product SKUs with parent brand summary profiles in the same index document schema without a declared entity grain. | Index mapping schema linter probe_mixed_search_index_grain verifying mapping creation failure with diagnostic ERR_SEARCH_INDEX_LACKS_DECLARED_GRAIN. | pass | Confirms OpenSearch index template admission checks; does not evaluate temporary memory buffers in ingestion workers. |
| FIT-3: Silent Schema Drift | Seed an upstream catalog feed that changes a product attribute from a numerical float to an unquoted text string without updating the OpenSearch index mapping template. | OpenSearch dynamic mapping strictness probe probe_unauthorized_type_drift verifying document rejection with diagnostic ERR_SEARCH_MAPPING_STRICT_ENFORCEMENT_VIOLATION. | pass | Confirms dynamic: strict index mapping configuration; does not inspect unmanaged unstructured log files. |
Residual Risk
- Latency spikes (up to 12 ms) during background HNSW vector graph segment merging under heavy ingestion surges. Accepted by Elena Rostova with night-time vector re-indexing schedules.
Traceability
| Claim | Classification | Source | Freshness |
|---|---|---|---|
| Rejection of uncoordinated concurrent index writes | derived | FIT-1 probe result | 2026-09-15 |
| Rejection of search indexes lacking declared grain | derived | FIT-2 probe result | 2026-09-15 |
| Rejection of unannounced mapping schema drift | derived | FIT-3 probe result | 2026-09-15 |
Verification
No validator was supplied, so no command was run.
Open Decisions
None.
Next steps
- Architecture Guild incorporates search platform fitness probes into automated CI index template checks.
- Platform team configures CloudWatch alarms monitoring OpenSearch cluster JVM memory pressure and search thread pool queue latency.
- Conduct quarterly disaster recovery drill simulating automated index rebuild and atomic alias flip under live load.
enterprise-search-platform-and-hybrid-di.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
What it does
This skill owns the architecture of a governed lexical/full-text retrieval boundary: corpus and index identity, analyzers, document lifecycle, query semantics, ranking, filters/facets, authorization, relevance evidence, scale and recovery. It does not own one query/mapping, ingestion script, search UI, database index, crawler, SEO research, vector search, RAG or embedding selection.
Use it when
- Authoritative sources must form a versioned searchable corpus with stable document/field identities
- Languages, scripts, tokenization, normalization, stop words, stemming/lemmatization, synonyms and exact fields affect retrieval semantics
- Mappings and analyzer changes require index-version and reader/query compatibility
- Full/incremental ingestion, updates, deletes, reindexing and alias/pointer publication need freshness contracts
- Inverted-index terms/postings, statistics, segments, merges and replicas affect correctness and operations
- Query parsing, fields, operators, phrases, typo tolerance, boosts, filters and facets require bounded behavior
For example: “Our internal search returns HR documents to everyone, ranks a 2019 policy above its 2026 replacement, and takes a week to show new documents.”
What you get
- architecture/search-architect/README.md
- architecture/search-architect/00-overview/search-architect-overview.md
- architecture/search-architect/verification/fitness-self-check.md
Plus one page per business module, only where your evidence calls for it: {module}/ingest.md, {module}/storage.md, {module}/serving.md, {module}/lineage.md, {module}/retention.md, {module}/quality.md.
All paths are relative to the output folder you choose.
What it will not do
Do not use merely to write one query, add an Elasticsearch mapping/index, ingest files, tune SQL/database indexes, implement autocomplete/UI, perform web/research/SEO keyword search, select an embedding/vector database, build RAG, or combine vector and keyword retrieval.
How it works
- Check the scope is the search boundary.
- Bound the corpora and their access rules.
- Fix document identity and the update path.
- Design the analysis chain per language and field.
- State freshness and the reindex path.
- Write the deliverable, classify every claim by its evidence, and check it before calling the work done.
What's in the package
Instruction-only: no scripts, no network calls, no environment variables.
- LICENSE.txt
- SKILL.md
- agents/openai.yaml
- assets/output-template-artifact.md
- assets/output-template-contract.md
- assets/output-template-domain.md
- assets/output-template-fitness.md
- assets/output-template-mechanism.md
- references/domain-rules.md
- references/operating-rules.md
- references/output-contract.md
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 12 days ago
- Passed all security checks, Safe to install