More screenshots

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseOpenClawOpenClaw+21 more

    Repair Multi-Region Failover Races

    1

    Diagnose and repair multi-region failover, routing authority, stale regional writers, and recovery coordination.

    $8.99

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    Security scanned
    Repair Multi-Region Failover Races

    Repair Multi-Region Failover Races

    Example session with this skill installed

    We are seeing 'stale write' errors after failing over from us-east-1 to us-west-2. It seems the old primary in east is still processing queued messages. How do we fence it?

    • Read your context and instructions
    • Compiled the repair multi-region failover
    • Generated the document

    I have analyzed the failover race condition between us-east-1 and us-west-2. The issue stems from a lack of generation identity in your worker heartbeats, allowing the 'degraded' region to continue processing the task queue without verifying its current authority status.

    Regional Write Fencing Implementation

    To resolve this, we must implement a monotonic generation counter (fence token) stored in a global consensus layer (like DynamoDB or Consul). Every write must include this token, and the storage layer must reject writes with a lower generation than the current global maximum.

    // Proposed Fencing Logic for Task Workers
    interface TaskMetadata {
      id: string;
      generation: number; // The fencing token
      region: string;
    }
    
    async function processTask(task: TaskMetadata) {
      const currentAuthority = await authorityService.getGlobalGeneration();
      
      // 1. Pre-flight check
      if (task.generation < currentAuthority.generation) {
        console.error(`Fencing violation: Task gen ${task.generation} < Current ${currentAuthority.generation}`);
        return; // Drop task, it belongs to a stale region
      }
    
      // 2. Atomic Write with Condition Expression
      try {
        await db.update({
          TableName: 'RegionalTasks',
          Key: { id: task.id },
          UpdateExpression: 'SET val = :val, gen = :gen',
          // Ensure the write only happens if no newer generation has taken over
          ConditionExpression: 'attribute_not_exists(gen) OR gen <= :gen',
          ExpressionAttributeValues: {
            ':val': task.data,
            ':gen': task.generation
          }
        });
      } catch (err) {
        if (err.code === 'ConditionalCheckFailedException') {
          // Logic for handling a write that arrived after a newer region took over
          await handleStaleWrite(task);
        }
      }
    }
    

    Deterministic Test Plan

    1. Initialize Authority: Set global generation to 10.
    2. Simulate Lag: Start a worker in us-east-1 with generation: 10.
    3. Trigger Failover: Update global generation to 11 (simulating us-west-2 takeover).
    4. Execute Stale Write: Attempt the us-east-1 write.
    5. Verify Invariant: Assert the database rejects the write and the us-east-1 worker enters a 'dormant' state.

    Next steps

    • Audit the authorityService to ensure it uses a strongly consistent read for generation checks.
    • Update the CI/CD pipeline to increment the global generation counter during regional traffic shifts.
    • Implement a 'drain' signal for workers that detects generation mismatches via heartbeats.

    repair-multi-region-failover-races.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Identify authority gaps during active/passive regional transitions.Implement fencing tokens to prevent stale regional writes.Design deterministic tests for cross-region race conditions.Audit health signals to prevent routing loops during outages.

    About this skill

    Repair Multi-Region Failover Races targets production and repository defects around diagnose and repair multi-region failover, routing authority, stale regional writers, and recovery coordination.. It examines active/active versus active/passive ownership, routing and failover authority, regional write fencing or generation identity, replication lag and read/write locality, health signals and traffic withdrawal, background workers during failover, session/cache affinity across regions, recovery, failback, and observability, then establishes one explicit repair contract and focused deterministic validation.

    The skill works from repository access or partial evidence such as snippets, logs, configuration, traces, and architecture notes. It preserves existing product/security policy, avoids invented execution claims, and keeps production systems, real credentials, and deployment outside the validation path.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 1 month ago

    • Passed all security checks, Safe to install

    Listed1 month ago

    What's inside

    Frequently Asked Questions