- Home
- Skills
- Agents & Orchestration
- Fable Workflow
Works with the AI tools you already use
Fable Workflow
Makes any agent surface the unknowns in a vague spec before it builds, log its assumptions, and verify the result.
Free
Fable Workflow
Example session with this skill installed
Add rate limiting to our web API. The product owner's spec, verbatim: "Users shouldn't be able to spam our API. Limit it to 100 requests per minute." Deliver a Python rate limiter plus a short note. I can't answer questions right now, so use the fable workflow and log what you decide.
- Read your context and instructions
- Compiled the fable workflow
Summary
The spec says "100 requests per minute" but leaves 8 decisions open. I can't ask you right now, so I surfaced them, picked a default for each with the reason, built it, and verified it. The defaults most likely to need your call are marked ⚠️.
UNKNOWNS (decided before building)
| # | Unknown | Default chosen | Why |
|---|---|---|---|
| 1 | ⚠️ Scope: per IP, user, API key, or global? | Per key (caller passes the key) | A global limit lets one user lock out everyone |
| 2 | Window algorithm | Sliding-window log | A fixed window allows 200 requests across a minute boundary |
| 3 | ⚠️ Storage: in-memory or shared? | In-memory | With N workers this becomes N×100. Use Redis before scaling out |
| 4 | Over-limit response | 429 + Retry-After | Standard behavior that clients can back off on |
| 5 | Burst tolerance | None beyond 100/60s | The spec gives no burst allowance |
| 6 | Concurrency | One lock around check-and-append | Avoids a race where two requests both see 99 |
| 7 | Clock | time.monotonic | Wall-clock jumps (NTP) would reset or extend windows |
| 8 | Memory growth | evict_idle() drops keys idle for at least one window | Otherwise every IP ever seen stays in memory |
| + | Do rejected requests consume budget? | No | Otherwise a client that keeps retrying can never get back under the limit |
Implementation
import threading, time
from collections import deque
class SlidingWindowLimiter:
"""Per-key sliding-window log: at most `limit` requests in any rolling `window` seconds."""
def __init__(self, limit=100, window=60.0, clock=time.monotonic):
self.limit, self.window, self.clock = limit, window, clock
self._hits: dict[str, deque] = {}
self._lock = threading.Lock()
def check(self, key: str) -> tuple[bool, float]:
"""Return (allowed, retry_after_seconds). Rejected requests don't consume budget."""
now = self.clock()
with self._lock:
q = self._hits.setdefault(key, deque())
while q and now - q[0] >= self.window:
q.popleft()
if len(q) < self.limit:
q.append(now)
return True, 0.0
return False, self.window - (now - q[0])
def evict_idle(self) -> None:
now = self.clock()
with self._lock:
for k in [k for k, q in self._hits.items() if not q or now - q[-1] >= self.window]:
del self._hits[k]
In your handler, when the result is allowed=False, return 429 with Retry-After: ceil(retry_after).
Verification (ran it, didn't just assert it)
t = [0.0]
rl = SlidingWindowLimiter(100, 60, clock=lambda: t[0])
assert all(rl.check("a")[0] for _ in range(100))
ok, retry = rl.check("a"); assert not ok and retry == 60
assert rl.check("b")[0] # keys are independent
t[0] = 59.9; assert not rl.check("a")[0] # no fixed-window boundary burst
t[0] = 60.0; assert rl.check("a")[0] # oldest hit expired
t[0] = 200; rl.evict_idle(); assert rl._hits == {}
Output: all checks pass
ASSUMPTIONS log
- Assumed a single process. Multi-worker deployments need a shared store, or the limit is per worker.
- Assumed the key is the authenticated user ID when there is one, otherwise the client IP.
- Memory cost is O(limit) timestamps per active key, about 100 floats each. That's fine at this limit, but a token bucket is cheaper if the limit grows to 10k+.
Next steps
- Confirm #1 (scope) and #3 (storage). Those two change the architecture.
- If you run more than one worker, I'll swap the deque for a Redis sorted set with the same interface.
- Quiz-back before merge: what happens to request #101 at t=59.9s, and why?
Connects securely to your tools. The creator never sees your data.
About this skill
Capable models can explore a huge solution space on their own, so the bottleneck is no longer the model. It is whether your map matches the territory before the model starts moving. Every place your spec is silent is an unknown: a decision the model will otherwise guess at silently. Fable Workflow makes the model surface those unknowns first, decide them explicitly, and only then build.
It packages a Fable-style working method into one portable SKILL.md that works with Opus, Sonnet, Haiku, Fable, or any other tool-using model.
The loop
- Unhobble. Counting, enumeration, and precise lookups go to a tool or a script, not to memory.
- Find the unknowns before building. Blind-spot pass, interview-me (questions that would change the architecture come first), N genuinely different variants for taste calls, and references-as-maps.
- Build, logging deviations. A running ASSUMPTIONS list records every unknown the model hit and the choice it made.
- Verify. Run it, or write the smallest check that fails if the logic breaks. A good plan is not a correct answer. When verification fails, it runs a bounded correction loop with explicit exits.
- Stay in the loop. Before merge, the model quizzes you on what changed so you still own the work.
What's in the ZIP
SKILL.md: the method itself.prompts.md: copy-paste prompts (blind-spot pass, interview, variants, plus a lite variant for small local models).standing-orders.md: per-answer discipline and a pre-send gate.loop-engineering.mdandcompletion-gate.md: the correction loop, and a goal ledger (scripts/goals.py) that refuses "done" without a verify command and its result.hooks/: optional Claude Code hooks. One injects a single task-matched line on non-trivial prompts. The other is a Stop gate that flags edited-but-never-run code (advisory by default, blocking withFABLE_STRICT=1).integrations/: a Cursor rule and anAGENTS.mdversion for Codex, Aider, Zed, Antigravity, and others.
Benchmark (small, and stated as such)
We ran an A/B test on one deliberately under-specified task ("Limit our API to 100 requests per minute", which hides about 8 architecture-changing unknowns) across 8 models, cloud and local, with and without the skill. Each run was scored /100 for thinking plus answer (no skill → with skill):
- Haiku 4.5: 66 → 81
- Sonnet 5: 91 → 98
- Opus 4.8: 95 → 98
- Fable 5: 96 → 100
- llama3:8b (local): 8 → 27
The skill is a reasoning amplifier, not a coding amplifier. Thinking quality rose for every model. Answer quality rose only when the model could already code the plan. This was a directional pilot: single trials, judged by a model. Local-model deltas swung widely between reruns. The full methodology and every caveat are in the repo's benchmark/RESULTS.md.
Best for
- Vague specs and "just make it work" tickets
- Unfamiliar codebases or new domains
- Design, API-shape, or output-format calls where you'll "know it when you see it"
- Any build where you want a written trail of the model's assumptions
Not for one-line fixes. The method says so itself.
Open source (MIT): https://github.com/joey114132/fable-workflow-skill
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
Free forever
No account required to browse
Trust & safety
Security scanned
Verified clean 2 days ago
- Free to download with an account