arazvan-ec/xmarks/tree/main/mods/newspeak
關於這個 mod
flywheel 🎡
一個 Claude Code 外掛,把臨時的「憑感覺寫程式」變成有紀律、能自我驗證的 AI 輔助開發 迴圈。它把 obra/superpowers、EveryInc/compound-engineering-plugin、karpathy/autoresearch、addyosmani/agent-skills 和 gszhangwei/open-spdd 的實務整合成一套一致的系統。
這個儲存庫 就是 外掛,透過 xmarks 市集(.claude-plugin/marketplace.json)提供。
安裝(Claude Code)
/plugin marketplace add arazvan-ec/xmarks
/plugin install flywheel@xmarks
接著執行 /reload-plugins,再執行 /flywheel:help。如果要讓 flywheel 在儲存庫中 自動啟用,請參閱 docs/add-flywheel-to-a-repo.md。
僅限 Claude Code —— claude.ai 聊天應用程式使用不同的 Skills 系統,不會執行 Claude Code 外掛。
Claude Code web —— 網頁工作階段不會自動安裝市集外掛,因此
/plugin install和settings.json中的市集金鑰都不會讓/flywheel:*出現在那裡。請改為在目標儲存庫中使用scripts/install-vendored.sh內嵌 flywheel 一次;之後每個介面都能使用/flywheel-help、/flywheel-loop、…。請參閱 docs/add-flywheel-to-a-repo.md。這個缺口也曾影響 flywheel 本身:沒有已註冊的代理程式時,它自己的開發迴圈無法遵循它所規定的+delegate路由。install-vendored.sh --agents-only是守門器允許的唯一自我目標 —— 它會把agents/*.md註冊到.claude/agents/,除此之外什麼都不做(沒有 skills,也沒有 hooks);若已發佈副本和已註冊副本在任一方向不一致,scripts/check-agent-parity.sh會讓 CI 失敗。它的兄弟選項--hooks-only會把這個儲存庫自己的 hooks 註冊到.claude/settings.json,指向scripts/;scripts/check-hook-parity.sh現在也會確認hooks/hooks.json宣告的每個 hook 都已在那裡註冊。
這個想法:兩層迴圈
外層迴圈(開發週期) —— 一項工作會依序通過六個閘門階段:
spec → plan → work → verify → review → compound
內層迴圈(在 work 裡) —— 一個緊湊的 寫失敗測試 → 實作 → 執行 → 觀察 → 修正 週期;只有客觀檢查變成綠燈後,才會宣告「完成」。
不會因為「看起來沒問題」就往前走:verify 會執行真正的應用程式/測試,而每次完成的週期都會把可重用的知識存進帳本,替下一個週期預先準備。
📚 第一次接觸迴圈概念?請參閱 docs/getting-started-with-loops.md ——其中介紹四種迴圈類型(回合式、目標式、時間式、主動式),以及 flywheel 如何對應到它們。
⏱️ 想排程或無人值守地執行 flywheel?請參閱 docs/proactive-loops.md ——其中說明如何把
/flywheel:verify/review與/loop、/schedule例程、/goal和工作流程組合起來。
第二根支柱:代理原生執行階段(v0.15.0)
上面的迴圈會 建置 軟體。process/run 這對指令也讓 flywheel 能夠
運作 軟體 ——把儲存庫變成 代理原生:Claude 是執行階段的一等公民,不是外掛上去的附加物。與其為重複的領域操作(「分析一輛車」、「替潛在客戶評分」、「匯入報告」)撰寫靜態後端函式,不如定義一份 程序契約,讓 Claude 執行它。
/flywheel:process <desc>會建立.claude/flywheel/processes/<slug>.md:操作始終遵循的 固定規則、輸出結構,以及結果要在哪裡 持久化 ——遵循儲存庫在.claude/flywheel/DATA.md中一次宣告的自身資料策略(例如透過儲存庫的 client 使用 Postgres),絕不套用 flywheel 強加的資料儲存區。/flywheel:run <slug> [input]會 以後端身分執行契約:遵循規則,只在契約允許的地方運用判斷,把結果寫入資料儲存區並證明已落地(具冪等性,且已讀回驗證),然後 讓契約成熟 ——追加一項以證據為基礎的修訂,讓下一次執行更精準。固定規則加上會自我改進的提示,正是所要求的形式。
完整願景與汽車範例:docs/research/agent-native-processes.md。
指令
| 指令 | 功能 |
| --- | --- |
| /flywheel:help | 新手導覽與指令地圖。 |
| /flywheel:loop <feature> | 從頭到尾執行完整週期,在階段之間設閘門。 |
| /flywheel:brainstorm <idea> | 在撰寫規格前,把模糊想法整理成已同意的需求。 |
| /flywheel:spec <feature> | 撰寫 REASONS 規格契約與可由機器檢查的成功指標。 |
| /flywheel:plan <spec-slug> | 把規格轉成有順序的工作,每項工作都有自己的檢查。 |
| /flywheel:work <task> | 以內層的反覆迭代直到綠燈迴圈實作。 |
| /flywheel:debug <symptom> | 系統化除錯:重現 → 假設 → 隔離 → 修正 → 回歸測試。 |
| /flywheel:verify | 客觀的 PASS/FAIL 閘門 ——透過 verifier 代理程式執行真正的應用程式/測試。 |
| /flywheel:review <ref> | 依 diff 類型分派多位專家審查(文件 diff ≠ 完整分派),再彙整結果。 |
| /flywheel:compound | 把本週期的決定、陷阱與模式追加到帳本。 |
| /flywheel:recall <query> | 隨需搜尋帳本 ——低成本列出符合的學習內容,依要求展開一項。 |
| /flywheel:route <task> | 在把工作委派到計畫之外前,依 route-tiers.txt 推薦工具、新工作階段、子代理程式或目前工作階段,以及使用的模型和 effort。 |
| /flywheel:ship <title> | 以乾淨的 commit 加上 push 與 PR 結束週期。 |
| /flywheel:process <desc> | 定義 代理原生程序 ——可重用的提示契約(固定規則+輸出結構+持久化),讓 Claude 作為後端執行重複的領域操作。 |
| /flywheel:run <slug> [input] | 把已定義的程序作為執行階段執行 ——遵循規則、把結果持久化到儲存庫的資料儲存區,再根據執行結果讓契約成熟。 |
| /flywheel:autoloop <goal> ⚡ | 由自主指標驅動的迴圈 ——持續迭代,直到達成指標或用盡預算。 |
| /flywheel:sync <spec-slug> ⚡ | 協調規格與程式碼之間的漂移(雙向)。 |
| /flywheel:update [vendored\|marketplace] | 更新 flywheel 本身 ——自動判斷 marketplace 或 vendored 安裝,也可把模式作為引數傳入。 |
代理程式
verifier——執行應用程式/測試,並以證據回傳客觀的 PASS/FAIL。reviewer-correctness、reviewer-security、reviewer-performance——由/flywheel:review平行分派的對抗式專家審查器。測試套件並未測試這個分派扇出本身(P32、v0.60.0):評估執行器是子代理程式,而子代理程式不能再建立子代理程式,因此這個儲存庫從未有任何分級執行真的分派reviewer-*。skills/review/evals/覆蓋剩下的部分 ——哪些審查器是由 diff 推導出來的(從成品讀取,而非報告文字),以及報告是否說明專家並未執行 ——真正的分派只有一個分支,由頂層工作階段手動執行,並記錄在該套件的 README 中。evaluator——由/flywheel:autoloop分派的獨立交叉檢查器,用於含糊的保留/捨棄結果,以及在宣告目標達成前重新執行指標指令,而不是相信工作代理程式的自我報告。executor——階段路由中的低成本層級(Haiku、low effort):在自己的上下文中完成一項已完整規格化的機械式計畫工作,若任務沒有決定就回傳ESCALATE: <reason>,而不是自行創造決策。當/flywheel:work將任務路由到haiku/low+delegate時分派它。
依角色路由模型(v0.9.0):機械式的 verifier 使用 Haiku(它執行指令並回報證據);需要大量判斷的 reviewer-* 使用 Sonnet。可透過代理程式的 model: frontmatter(例如高風險審查使用 opus)覆寫單一代理程式,或用 CLAUDE_CODE_SUBAGENT_MODEL 一次覆寫全部。自 v0.39.0 起,每個代理程式也會固定其 effort: ——機械式的 verifier/evaluator/executor 使用 low,對抗式的 reviewer-* 使用 high。
每項工作的 check: 都會被執行,而不是只讀取(v0.70.0):每項計畫工作都必須帶有一個 - check: ——plan-route.sh 會拒絕沒有它的計畫——過去這個欄位只檢查是否存在、從未執行,因此「T4 已完成」其實是模型根據通常已經可執行的指令替自己評分。bash scripts/check-task-closure.sh 會執行它,並為每項工作印出一列:PASS(符合允許清單的指令、已執行且退出碼為 0)、FAIL(退出碼非 0)、UNRUNNABLE(check 中沒有允許清單指令——已命名並計數,但從未執行,也不會被視為綠燈),或 PENDING(週期保留帳本,而這項工作在其中沒有狀態轉換列,所以尚未開始——迴圈會在工作前的核准閘門提交計畫,評分時就會把未開始視為故障)。列數等於工作數,因此計數在哪裡停止對不上會立刻看出遺漏。可執行內容由 scripts/task-closure-allow.txt 鎖定 ——含 shell 運算子的區段會在查詢允許清單前被拒絕,通過的內容會以不經 shell 的 argv 執行,因此允許清單前綴不能再附加第二個指令。這個檔案是安全邊界,不只是便利清單:計畫是任何 PR 都能寫入的儲存庫內容,若從其中執行任意 shell,就會把 /flywheel:verify 暴露在 gate.sh 信任存放區本想封住的風險中。截止點前加入的計畫屬於語料 ——只回報、不判失敗,也不會為了買到綠燈而重寫(P18、P48)。閘門評分的是指令退出碼,不是周圍的文字,因此從現在起撰寫的 check: 會陳述收尾時成立的條件。
每項計畫工作的階段路由:模型+effort(P27、v0.39.0):計畫不是單一價格。/flywheel:plan 會依 3 層規則,為每項工作路由 route: <model>/<effort>[+delegate] ——T1 是機械且規格完整(haiku/low+delegate,由 executor 代理程式執行),T2 是一般測試優先工作(sonnet/medium),T3 是判斷工作,也永遠是風險最高的一步(opus/high)——而路由表是計畫閘門核准內容的一部分,所以模型切換是你簽署的決定,不是任務中途才做的決定。/flywheel:work 會遵循該路由,第二次變紅時升高一層,不會廉價地原地打轉,並在狀態轉換列記錄 route(升級時另記 route_escalated_from)。bash scripts/plan-route.sh <plan.md> 會檢查路由並印出層級摘要;若最具風險的步驟低於最高層級,它會 讓計畫失敗(層級定義放在 scripts/route-tiers.txt,因此重新調整它就是修改資料)。這裡刻意不宣稱節省了多少 token ——檢查器只計算工作數,任何成本差異都由下面的 cost proxy 承擔(P18/P23)。
迴圈內的原子 commit(P28、v0.40.0):週期不再把工作留到最後才儲存。/flywheel:work 會在每項工作的檢查變成綠燈時立刻 commit 並 push ——git commit -m "<subject>" -- <the task's paths>(使用 pathspec,因此 commit 只包含該工作的路徑,不包含原本已 dirty 的內容),再加上不強制的 git push -u origin <branch>;/flywheel:debug 也會對修正與回歸測試做同樣的事。這兩種形式正是 P21 授權事先核准的內容,因此不需要新的權限範圍。/flywheel:ship 接著只提交剩下的內容(帳本、規格、文件),而且 絕不 squash、rebase 或 amend 工作歷史 ——要改寫歷史必須由你提出要求。所有地方的 commit 都採 fail open:沒有儲存庫、沒有 remote、沒有可提交的內容,或 push 被拒絕,都只回報一次,迴圈仍會繼續;若位於預設分支,會先建立功能分支,而不是提交到 main。
**Token 紀律
安裝
請先查看作者 README,確認 marketplace 與外掛名稱;指令可能隨儲存庫結構而變動。
claude plugin marketplace add arazvan-ec/xmarks claude plugin install newspeak
原文 / README
flywheel 🎡
A Claude Code plugin that turns ad-hoc "vibe coding" into a disciplined, self-verifying loop for AI-assisted development. It distills the best practices from obra/superpowers, EveryInc/compound-engineering-plugin, karpathy/autoresearch, addyosmani/agent-skills, and gszhangwei/open-spdd into one coherent system.
This repository is the plugin, served through the xmarks marketplace (.claude-plugin/marketplace.json).
Install (Claude Code)
/plugin marketplace add arazvan-ec/xmarks
/plugin install flywheel@xmarks
Then /reload-plugins and run /flywheel:help. To make flywheel auto-activate in a repo, see docs/add-flywheel-to-a-repo.md.
Claude Code only — the claude.ai chat app uses a different Skills system and does not run Claude Code plugins.
Claude Code web — web sessions do not auto-install marketplace plugins, so neither
/plugin installnor thesettings.jsonmarketplace keys make/flywheel:*appear there. Instead, vendor flywheel into the target repo once withscripts/install-vendored.sh; the commands then work on every surface as/flywheel-help,/flywheel-loop, … See docs/add-flywheel-to-a-repo.md. The same gap applied to flywheel itself: with no registered agents its own dev loop could not honor the+delegateroutes it prescribes.install-vendored.sh --agents-onlyis the one self-target the guard allows — it registersagents/*.mdinto.claude/agents/and nothing else (no skills, no hooks), andscripts/check-agent-parity.shfails CI if the shipped and registered copies ever disagree in either direction. Its sibling--hooks-onlyregisters this repo's own hooks into.claude/settings.jsonpointing atscripts/, andscripts/check-hook-parity.shnow also asserts that every hookhooks/hooks.jsondeclares is registered there.
The idea: two nested loops
Outer loop (development cycle) — one unit of work flows through six gated phases:
spec → plan → work → verify → review → compound
Inner loop (inside work) — a tight write failing test → implement → run → observe → fix cycle that never declares "done" until an objective check is green.
Nothing advances on "seems right": verify runs the real app/tests, and every finished cycle deposits reusable knowledge into a ledger that primes the next one.
📚 New to loops as a concept? See docs/getting-started-with-loops.md — the four loop types (turn-based, goal-based, time-based, proactive) and how flywheel maps onto them.
⏱️ Want to run flywheel on a schedule or unattended? See docs/proactive-loops.md — composing
/flywheel:verify/reviewwith/loop,/scheduleroutines,/goal, and workflows.
The second pillar: an agent-native runtime (v0.15.0)
The loop above builds software. The process/run pair lets flywheel also
operate it — turning the repo agent-native:
Claude is a first-class part of the runtime, not a bolt-on. Instead of writing a
static backend function for a recurring domain operation ("analyze a car", "score
a lead", "ingest a report"), you define a process contract and let Claude run it.
/flywheel:process <desc>scaffolds.claude/flywheel/processes/<slug>.md: the fixed rules the operation always follows, its output schema, and where results persist — following the repo's own data strategy declared once in.claude/flywheel/DATA.md(e.g. Postgres via the repo's client), never a datastore flywheel imposes./flywheel:run <slug> [input]executes the contract as the backend: follow the rules, apply judgment only where the contract allows, write the result to the datastore and prove it landed (idempotent, read-back verified), then mature the contract — appending one evidence-based refinement so the next run is sharper. Fixed rules + a self-improving prompt, exactly as asked.
Full vision + the worked car example: docs/research/agent-native-processes.md.
Commands
| Command | What it does |
| --- | --- |
| /flywheel:help | Onboarding + command map. |
| /flywheel:loop <feature> | Run the whole cycle end to end, gating between phases. |
| /flywheel:brainstorm <idea> | Sharpen a fuzzy idea into agreed requirements before the spec. |
| /flywheel:spec <feature> | Write a REASONS spec-contract + a machine-checkable success metric. |
| /flywheel:plan <spec-slug> | Turn the spec into ordered tasks, each with its own check. |
| /flywheel:work <task> | Implement with the inner iterate-until-green loop. |
| /flywheel:debug <symptom> | Systematic debugging: reproduce → hypothesis → isolate → fix → regression test. |
| /flywheel:verify | Objective PASS/FAIL gate — runs the real app/tests (via the verifier agent). |
| /flywheel:review <ref> | Multi-specialist review routed by diff type (docs diff ≠ full fan-out), synthesized. |
| /flywheel:compound | Append this cycle's decisions, gotchas, and patterns to the ledger. |
| /flywheel:recall <query> | On-demand ledger search — list matching learnings cheaply, expand one on request. |
| /flywheel:route <task> | Before delegating work outside a plan: recommends tool, fresh session, subagent or here, at which model and effort, from route-tiers.txt. |
| /flywheel:ship <title> | Clean commit + push + PR to close out the cycle. |
| /flywheel:process <desc> | Define an agent-native process — a reusable prompt-contract (fixed rules + output schema + persistence) for a recurring domain operation Claude runs as the backend. |
| /flywheel:run <slug> [input] | Execute a defined process as the runtime — follow its rules, persist the result to the repo's datastore, then mature the contract from the run. |
| /flywheel:autoloop <goal> ⚡ | Autonomous metric-driven loop — iterate hands-off until a metric is met or a budget is spent. |
| /flywheel:sync <spec-slug> ⚡ | Reconcile drift between a spec and the code (bidirectional). |
| /flywheel:update [vendored\|marketplace] | Update flywheel itself — autodetects marketplace vs vendored install, or takes the mode as an argument. |
Agents
verifier— runs the app/tests and returns an objective PASS/FAIL with evidence.reviewer-correctness,reviewer-security,reviewer-performance— adversarial specialist reviewers dispatched in parallel by/flywheel:review. The fan-out itself is untested by the suites (P32, v0.60.0): an eval executor is a subagent, a subagent cannot spawn subagents, so no graded run this repo has ever made dispatched areviewer-*at all.skills/review/evals/covers what is left — which reviewers a diff draws (read from an artifact, not from the report's prose) and whether the report says the specialists did not run — and real dispatch has exactly one arm, run by hand from a top-level session, documented in that suite's README.evaluator— independent cross-check dispatched by/flywheel:autoloopon ambiguous keep/discard results and before it declares its target met; re-runs the metric command itself instead of trusting the working agent's self-report.executor— the cheap tier of stage routing (Haiku, low effort): does one fully-specified mechanical plan task in its own context and returnsESCALATE: <reason>rather than inventing a decision the task left open. Dispatched by/flywheel:workfor a task routedhaiku/low+delegate.
Model routing by role (v0.9.0): the mechanical verifier runs on Haiku (it runs commands and reports evidence); the judgment-heavy reviewer-* run on Sonnet. Override any agent via its model: frontmatter (e.g. a reviewer → opus for high-stakes reviews), or all at once with CLAUDE_CODE_SUBAGENT_MODEL. Since v0.39.0 each agent also pins its effort: — low for the mechanical verifier/evaluator/executor, high for the adversarial reviewer-*.
A task's check: is executed, not read (v0.70.0): every plan task must carry a - check: — plan-route.sh rejects a plan without one — and until now that field was linted for presence and never run, so "T4 is done" was a model grading its own work over a field that usually already held a runnable command. bash scripts/check-task-closure.sh runs it and prints one row per task: PASS (a command matched the allowlist, ran, exited 0), FAIL (it exited non-zero), UNRUNNABLE (no allowlisted command in the check — named and counted, never executed and never read as green), or PENDING (the cycle keeps a ledger and this task has no transition line in it, so it has not started — the loop commits a plan at its approval gate, before the work, and grading then would report not-started as broken). Rows equal tasks, so a dropped item is visible exactly where the count stops reconciling. What it may execute is pinned by scripts/task-closure-allow.txt — a span carrying a shell operator is refused before the allowlist is consulted, and what survives runs as argv with no shell, so an allowlisted prefix cannot append a second command. That file is a security boundary rather than a convenience list: a plan is repo content any PR can write, so running arbitrary shell out of it would hand /flywheel:verify the exposure gate.sh's trust store exists to close. Plans added before the cutoff are corpus — reported, never failed, never rewritten to buy a green (P18, P48). The gate grades the command's exit code, not the prose around it, so a check: written from here on states the condition that holds at close.
Stage routing: model + effort per plan task (P27, v0.39.0): a plan is not one price. /flywheel:plan routes every task with route: <model>/<effort>[+delegate] from a 3-tier rubric — T1 mechanical and fully specified (haiku/low+delegate, run by the executor agent), T2 ordinary test-first work (sonnet/medium), T3 judgment, and always the riskiest step (opus/high) — and the routing table is part of what the plan gate approves, so a model switch is a decision you signed, not one taken mid-task. /flywheel:work honors the route, escalates one tier on the second red instead of grinding cheap, and records route (plus route_escalated_from on an escalation) on the transition line. bash scripts/plan-route.sh <plan.md> lints the routes and prints the tier summary; it fails a plan whose riskiest step runs below the top tier (scripts/route-tiers.txt holds the tiers, so retuning them is a data edit). The payoff is deliberately not claimed in tokens — the linter counts tasks, and any cost delta rides on the cost proxies below (P18/P23).
Atomic commits inside the loop (P28, v0.40.0): a cycle no longer saves its work for the end. /flywheel:work commits each task the moment its check goes green and pushes it — git commit -m "<subject>" -- <the task's paths> (pathspec, so the commit holds that task and nothing that was already dirty) plus a force-free git push -u origin <branch>; /flywheel:debug does the same for a fix and its regression test. Both forms are exactly what the P21 grant already pre-approves, so the discipline costs no new permission surface. /flywheel:ship then commits only what is left (ledger, spec, docs) and never squashes, rebases or amends the task history — rewriting it is yours to ask for. Committing fails open everywhere: no repo, no remote, nothing to commit, or a rejected push is reported once and the loop carries on; on the default branch it creates a feature branch first rather than committing to main.
Token discipline (v0.12.0): /flywheel:autoloop treats its iteration budget as a hard stop and recommends piloting on a small budget before scaling; /flywheel:help points to /usage, /goal, and /workflows for spend visibility. See skills/autoloop/SKILL.md.
Delegation triggers (v0.13.0): /flywheel:work names advisory thresholds for handing off to a fresh-context subagent — reading 4+ files, touching 2+ non-trivial files, or ~20 tool calls deep without converging — to keep each turn's context lean.
Live progress (v0.16.0, two-tier since v0.30.0): every process run (/flywheel:run) and dev cycle (/flywheel:loop/work) materializes its steps as visible tasks in the host task system — states updated at every transition — and keeps per-execution telemetry at .claude/flywheel/runs/<slug>/<date>.jsonl (one appended JSON line per transition) plus an HTML report at …/<date>.html, rendered from the JSONL only at gates and at close and republished to a stable artifact URL. Output tokens are the expensive ones: a transition costs one line, never a regenerated page. Chat stays reserved for gates, blockers, and the final summary. Fail-open: reporting never blocks execution.
Cycle cost, measured (P23): each transition line carries a cost object — bytes_out, bytes_in, tool_calls, elapsed_s — and the rendered report ends with a cost block. These are proxies, labelled as such everywhere they appear, never token counts: a session cannot observe its own token usage, so recording one would put unverifiable evidence in the ledger (P18) — scripts/run-cost.sh warns if it finds a tokens key. bash scripts/run-cost.sh <run.jsonl> [baseline.jsonl] totals a run and prints the per-field delta against a baseline, so "this made the loop cheaper" becomes a number. Transitions from before the schema are reported as unmeasured, never counted as zero — otherwise every old run would look free. bytes_in (P40a) floors read volume (charged once per read, nothing for conversation) and is reported UNMEASURED, never 0, in old runs to prevent fabricating improvements. bytes_in and tool_calls are recorded as they happen (P44): scripts/read-meter.sh sits on PostToolUse and appends one line per call — the bytes of tool_response that entered context — so a transition line reads bytes_in, tool_calls and elapsed_s with read-meter.sh --since <previous ts> (or --since first on a cycle's opening transition, which has no previous ts — every run before P45 omitted elapsed_s on line 1 because a commit-time delta has no previous commit) instead of recalling a total nothing kept, which is why both fields were absent from every line in the repo's history. A write tool's response echoes the file it changed without that text reaching context, so the call counts and its bytes do not, and no meter at all reports UNMEASURED rather than a zero. bash scripts/check-telemetry.sh gates conformance (every runs JSONL line holds valid proxies) and coverage (every spec slug is instrumented unless exempted with a reason in scripts/telemetry-baseline.txt), failing the build if either rule breaks.
State it keeps (in the project you use it on)
.claude/flywheel/specs/<slug>.md— REASONS specs and.plan.mdplans..claude/flywheel/processes/<slug>.md— agent-native process contracts (fixed rules + output schema + persistence + an append-only improvement log), created by/flywheel:processand matured by/flywheel:run..claude/flywheel/DATA.md— the repo's data-persistence strategy (Store / Access / Schema / Conventions) that every/flywheel:runwrites through, so results land the way the repo already stores them..claude/flywheel/LEARNINGS.md— the compounding ledger. Typed entries (## <type>: <title>+ a greppable<!-- fw: … -->metadata line;type∈decision/gotcha/pattern/bugfix/fixture) let theSessionStarthook inject only a relevance-scored, budgeted subset (branch/files/recency, default top 12,FLYWHEEL_LEARNINGS_INJECTto override) instead of a blind reload;/flywheel:recall <query>reaches the rest on demand. Created by/flywheel:compound. Older free-prose entries still load, as always-eligible low-priority entries.fixtureentries (v0.21.0) capture how to set up the world — the recipe to build a valid stub for a domain entity, seed the datastore, or stand up a test harness — the costliest thing a session otherwise re-derives./flywheel:workoffers to record one when it spends real effort building test data, and/flywheel:spec+workprime from any that match the task's entities before the rediscovery.- Evidence-gated (v0.25.0): flywheel gates knowledge the way it gates code. Each entry carries
evidence=— what proved it (a test, a run/PR, acommand → result). A lesson that can't point to a proof is writtenevidence=unverifiedexplicitly, and the SessionStart injection +/flywheel:recallflag those so a wrong-but-plausible conclusion can never masquerade as proven context./flywheel:compoundrecords only what a cycle actually proved.
.claude/flywheel/runs/<slug>/<date>.jsonl+.html— per-execution telemetry for process runs and dev cycles (v0.16.0): one JSONL line appended per state transition; the HTML report (task ledger + states, gates, unit telemetry, verdict) rendered from the JSONL only at gates and close (v0.30.0).
Read-priming hook (advisory)
Before reading a file, a PreToolUse hook greps the ledger's files= metadata for that path and, if any typed entry names it, injects a short "prior learnings touch this file" note into context via the hook's additionalContext field (v0.18.0 — plain stdout is transcript-only and never reaches the model) — cheap context ahead of an expensive read. A bash pre-filter skips the python parser entirely for the no-match majority. It never blocks the read (unlike claude-mem's File Read Gate) and fails silently (no ledger, no match, or no python3) so it can never slow down or break a read.
Approval-coherent permissions
flywheel has exactly two deliberate approval gates, and both are conversational: the spec sign-off and the plan approval. The harness's tool-permission layer knows nothing about them — so without help it re-asks "allow?" for actions the approval already implied. Two allow-only PreToolUse hooks close the gap; everything they don't match keeps the normal permission flow, and (docs-guaranteed) a hook "allow" can never override a deny/ask rule you wrote yourself. Both are fail-open by contract: they never deny, never ask, never block; malformed input or a missing python3 just falls back to the ordinary prompt.
- State writes (v0.27.0,
Write|Edit|MultiEdit|NotebookEdit): a write whose target resolves inside<project>/.claude/flywheel/is auto-allowed — specs, plans, the ledger, process contracts, run reports. Repo code and.claude/settings.jsonare out of scope. Paths arerealpath-resolved before the containment check, so..traversal, prefix siblings (.claude/flywheel-evil/) and symlinks planted inside the state dir that point elsewhere get no grant. - Loop-advancing git (v0.28.0,
Bash): one plaingit add,git commit,git stash(bare/push/pop/list), or a force-freegit push [-u] origin <branch>where<branch>is the current, non-default branch — checked live against the repo. Any shell metacharacter outside single quotes (chaining, pipes, redirects,$(…)/backticks even inside double quotes) disqualifies the whole command, sogit commit -m "x" && anythingnever rides the grant while a quoted-m "fix: A & B"passes. Global git flags (-C,-c,--git-dir), foreign remotes, refspecs,--force*,stash drop/clearand every other verb stay prompted.
The commands the plugin cannot know — your test/metric command, your DATA.md datastore write path — get their grant at the gate that approves them: /flywheel:spec and /flywheel:process offer at sign-off (never write unasked) to append the matching narrow rule (e.g. Bash(npm test:*)) to the project's .claude/settings.json permissions.allow, committed with the spec/contract; /flywheel:sync flags signed pre-v0.28.0 specs/contracts that lack their rule as drift.
Progress toolbar, enforced (P56, v0.74.0)
While a .claude/flywheel/specs/<slug>.plan.md that the current branch touches (vs its base, or uncommitted) still has a task with no transition line, the final reply of each turn must open with <🟢|⏸️|🔴|🏁> <done>/<total> <bar> · ▶ <item> · «<what, in the owner's words>». scripts/toolbar.sh enforces it as two hooks: UserPromptSubmit injects the format and the live count (slug done/total, open task ids) as context every prompt, and Stop blocks (exit 2) a final reply whose first line is not the toolbar or whose total is not the plan's. Mid-turn notes are exempt. stop_hook_active never re-traps, and anything it cannot read is a no-op.
A shipped cycle says what it learned (P69, v0.83.0)
scripts/compound-due.sh is a Stop hook: a runs/<slug>/*.jsonl this branch touches that has a phase: ship line and no phase: compound line blocks the turn (exit 2). /flywheel:compound appends that line with entries: N; a cycle that proved nothing durable records entries: 0 with a reason. Before it, 14 of 16 shipped runs in this repo carried no compound line.
Feedback about flywheel reaches flywheel (P70, v0.83.0)
A lesson about the plugin, learned in a repo that only uses it, used to stay in that repo's ledger. When /flywheel:compound writes one outside this repo, it offers scripts/upstream-issue.sh, which renders the entry as a prefilled flywheel feedback issue URL (label flywheel-feedback) with only flywheel paths kept. Nothing is sent until a human reviews the form and submits it. The same form is open to anyone using the plugin. The flow-audit process reads the open issues as input and gives each one a disposition.
A delegated review says it started (P57, v0.75.0)
A review sent to another session (create_session) that posts nothing reads the same whether it found nothing, could not post, or is still running. skills/review/references/delegated-review.md is the child's prompt template: post a 🔎 Review started … fw-review-start comment on the PR first, through the GitHub MCP tools (the container has no gh), then post the findings as one review with inline comments, or a "no findings" comment. The delegation guard's REVIEW family asks when a delegated review prompt lacks fw-review-start, naming the template.
Mods (P72, v0.84.0+)
Claude Code mods ship from this marketplace as separate, opt-in plugins under
mods/<name>/, each with its own version and tests (scripts/check-mods.sh).
Claude Code web (cloud sessions): marketplace plugins are not installed there, mods included. Load mods by folder instead, with
CLAUDE_CODE_PLUGIN_DIRS, verified in a cloud container on CLI 2.1.289. In the environment's settings (cloud environment
menu → Edit), add to the setup script git clone --depth 1 https://github.com/arazvan-ec/xmarks /opt/xmarks, and set the
environment variable CLAUDE_CODE_PLUGIN_DIRS=/opt/xmarks/mods/resource-committee:/opt/xmarks/mods/big-brother-token
(absolute paths, :-separated). New sessions load them. A project's own settings.json cannot set this variable. Step by step, and what each mod measures: docs/mods-in-the-cloud.md.
| Mod | What it does | Install |
| --- | --- | --- |
| resource-committee | Assigns every turn its model and effort: sonnet/medium by default, opus/high for judgment, sonnet/low for mechanical work, classified once per prompt by Haiku so the cache survives. /committee shows the decision or pins haiku/sonnet/opus/auto; /committee stats sums each session's tally. Subagents keep their own model. | /plugin install resource-committee@xmarks |
| big-brother-token | Live bytes read, session cost and context % in the status line; a toast for any single read over 8 KB and for a Write that rewrites a file already read; /ministry gives the dossier by tool and Bash bytes by command. Stores each session's summary, with its cost. | /plugin install big-brother-token@xmarks |
| memory-hole | A prompt carrying a list (2+ numbered items or 3+ bullets, code fences ignored) holds Edit/Write/NotebookEdit until the list is written to a .plan.md, the scratchpad, or tasks.md/todo.md. /memory-hole release lets it go, logged. | /plugin install memory-hole@xmarks |
| ventanilla-unica | /ventanilla [base] runs scripts/sweep.sh in the background and stamps each gate in a pane as its line arrives; a toast gives the verdict. A sweep the model runs through Bash is stamped too. /ventanilla status answers in text. | /plugin install ventanilla-unica@xmarks |
| social-credit | A citizen score in the status line, kept across sessions: +1 per Edit, −5 for a Write over a file already read, ±3 for test-first on scripts/, +1 per commit, −10 per SKIP_*= or --no-verify. Below 80 a re-education section joins the system prompt. /social-credit lists acts; amnesty resets. | /plugin install social-credit@xmarks |
| thought-police | At Stop, a reply whose fenced code repeats 8+ lines written this turn (Write/Edit) is blocked once with the reason: report what and where, the diff is in git. Never blocks twice in a row. | /plugin install thought-police@xmarks |
| newspeak | For a prompt carrying a list, names the items with no success criterion (English or Spanish markers: so that, must, passes, para que, en verde, a number with a unit…) in a context note asking the model to get one first. Never blocks. Also counts multi-item asks written in prose, without nudging. /newspeak shows both. | /plugin install newspeak@xmarks |
| ration-book | A per-session read ration: p75 of the last 10 sessions once 5 are recorded (floor 100 KB), else 400 KB. Coupons left in the status line, a toast at 80%, and at 100% Read/Grep/Glob/WebFetch/WebSearch are held (Bash and writes never). /ration grant <KB> adds coupons, announced. | /plugin install ration-book@xmarks |
| black-market | Ledgers every SKIP_*=<reason>, --no-verify and Release-Exception: the session uses, with its reason; one without a reason (1, empty) raises a toast. /black-market audit totals the standing permits in scripts/*allow*.txt, *baseline*.txt, invocation-budget.txt and the git log's trailers. | /plugin install black-market@xmarks |
| citizen-file | /expediente [slug] opens a pane on a cycle's run record (.claude/flywheel/runs/<slug>/*.jsonl, newest by default): each transition with its route, escalations marked, then transitions, bytes in and escalations totalled. list and status answer in text. | /plugin install citizen-file@xmarks |
| telescreen | When a Read/Edit/Write touches a file a LEARNINGS.md entry cites (files=), a band above the prompt shows the newest such lesson, with Hide. No match, no band. /telescreen counts lessons loaded and slogans shown. | /plugin install telescreen@xmarks |
| supervisor | Every N prompts (default 5, /supervisor every <n>) the turn is asked to name what closed with its evidence; a reply naming no file:line, backticked command or commit gets a toast. Never blocks, spawns nothing. /supervisor counts checks and evidenced replies. | /plugin install supervisor@xmarks |
| general-strike | Tracks check commands (test, check, sweep, pytest, jest, go test…): the 3rd consecutive failure of the same one, with an edit since the last, holds Edit/Write/NotebookEdit and tells the next prompt to run /flywheel:debug. Reads and Bash keep running. Ends on a pass, on opening flywheel:debug, or /strike end. | /plugin install general-strike@xmarks |
Deterministic completion gate (opt-in)
Drop an executable .claude/flywheel/gate.sh in your project with your verification command (e.g. npm test && npm run lint). While it exists and you've trusted it, flywheel's Stop hook runs it whenever Claude tries to finish and blocks finishing if it fails — so nothing is declared "done" with checks red.
Trust it first (v0.20.0): because the gate is a repo file that runs automatically, a PR could plant a malicious one — so an unrecognized gate is not executed. The first time it's seen, the hook prints the one command to trust it (a content hash stored outside the repo, so a PR can't self-authorize); editing the gate revokes trust until you re-consent. It is a no-op when absent, skips re-running when the git-tracked working tree is byte-for-byte the last-passing state (a git-derived signature over tracked + untracked content; non-git state like env vars or ignored files isn't observed, and it only ever skips a re-run — a changed tree always re-runs), bounded to a few consecutive blocks per failing tree with a persisted bypass (so a red gate never re-traps you), and fails open on internal errors.
One suite run per cycle (v0.30.0): /flywheel:verify uses the project's gate.sh as its suite command when present, and after a PASS runs gate.sh seal — recording the passing tree's signature so the Stop hook cache-hits instead of re-running the same suite minutes later. Seal accepts the caller's evidence, never creates it: it refuses an untrusted gate (a planted gate can't arrive pre-passed) and covers exactly one tree signature — any change re-runs.
Auto-update runs pinned code (v0.61.0)
install-vendored.sh --auto-update writes .github/workflows/flywheel-update.yml into your repo: a thin caller of flywheel's reusable workflow that refreshes your vendored copy and opens a PR. That workflow runs in your CI, with contents: write + pull-requests: write, on a weekly cron, so what it executes is a trust boundary — the only Critical in docs/research/pillar2-threat-model.md. It used to git clone this repo's main and bash the result, which handed arbitrary code execution in your repo to anyone who could push here.
Pinning the caller's uses: alone does not fix that: a pinned workflow that still clones a moving branch executes the branch. So both halves are pinned — the caller names a full commit SHA and passes that same commit as a flywheel_sha input, and the reusable workflow checks that commit out, refusing before any vendored bash runs if the input is absent or is not a full 40-hex SHA. bash scripts/check-supply-chain-pin.sh asserts both halves independently, so deleting either one turns it red on its own; exceptions live in scripts/supply-chain-pin-allow.txt with a reason, and a stale entry fails. There is no github.job_workflow_sha to read the pin from — the field does not exist, which is why the caller passes it explicitly.
This buys one specific thing: a compromised flywheel release can no longer reach your repo by itself. It does not make a release you choose to adopt safe. The pin is now the update lever — nothing bumps it for you, so the weekly job no longer opens update PRs on its own. Discovery still works by the routes that have a human in them: the SessionStart hook polls this repo's plugin.json and prints ⬆️ flywheel X.Y.Z is available, and /flywheel:update refreshes on demand. What is gone is precisely the unattended part — which was the part executing unreviewed code. See upgrades/v0.61.0.md; repos installed before 0.61.0 carry an @main caller whose weekly job now fails until they re-vendor.
Skill evals (manual release gate — not in CI)
Skills are prompts, so structural checks can't catch a behavioral regression — verify starting to rationalize a FAIL into a PASS, work skipping the red step, or process emitting a contract with no fixed rules. The skills where that hurts most carry behavioral evals in the skill-creator format, each with evals.json (realistic prompts + objective assertions), fixtures whose ground truth is known, and benchmarks/<date>/ holding the committed evidence of the last graded iteration:
- Pillar 1 —
skills/verify/evals/,skills/work/evals/andskills/loop/evals/: planted-bug mini-repos underevals/fixtures/, graded by a committedcheck.sh(last-lineVERDICT:regex over the savedreport.md/transcript.md;.check-logred-before-the-implementation-changed — aRESULT=FAILat the pristineIMPL_SHAwith no changed-sha entry before it — then a final PASS, plus a behaviour probe and an independent suite re-run).loop(P29, v0.41.0) is the only suite that runs a whole cycle, and therefore the only placework's cycle telemetry exists to be graded at all: its grader parses the transition JSONL, rejects atokenskey, resolves every recordedcommitagainst git withcat-file(a plausible sha nobody made fails), rejects a commit that mixes.pysource with.claude/flywheel/state, and runs a second eval in a non-git workdir where the correct outcome is nocommitfield and none invented. A third eval (P30, v0.42.0) grades the other half — a cycle that cannot pass, because its work item demands a tax-inclusive total and that the existing ex-tax test keep passing. An honest run may stop at any gate and leave the suite green or red, so neither route nor colour is asserted; what is asserted is what every cheap way out must touch — the pre-existing assertion stays verbatim, the telemetry records the blockage, and no transition claims averdict: PASS. - Pillar 2 —
skills/process/evals/andskills/run/evals/: a versioned mini target repo (DATA.md + the trivialplate-auditcontract + a seeded datastore, its setup recipe captured as atype=fixturelearning) and a committed grader scriptcheck.shthat greps the artifacts the run left behind.
All six graders share one contract — bash skills/<name>/evals/check.sh <id> <workdir>, one PASS:/FAIL: line per expectation, exit 0 only if all pass, exit 2 on an unknown id — and each mechanizes exactly the expectations in its own evals.json. Two CI gates keep them honest (P26), because both defects they exist to catch were found by accident rather than by a check:
scripts/test-eval-graders.shruns every grader against an untouched fixture and requires a red — the question "can this assertion even fail?" is what exposed the hollowruneval-2 grader — and runs the pillar-1 graders against an ideal outcome to require a green, so a grader that can never pass is caught too. Those ideal outcomes are committed assets underskills/<name>/evals/solutions/<solution>/, applied byfixture-scratch.sh --solution(P33), so the gate grades exactly what a person can reproduce by hand. Each solution declares its fixture and eval ids in aMANIFEST—cart-feature/cart.pyandcart-bugfix/cart.pyare byte-identical, so a clean patch apply proves nothing about which fixture it was written for — and pins the fixture's tree digest inBASED-ON, sincegit applydetects drift only in a patch's own context lines. Edits to files the fixture already has arepatch/entries, never whole-file overlay copies: a copy oftest_cart.pywould shadow itsKATA_HARNESSguard, the thing that makes.check-logtrustworthy, and the green arm would stay green after the fixture moved. Solutions live outsidefixtures/because they are the answer key. Pillar 2's green side is still not built (that would reimplement what it grades); its green evidence is the committed benchmarks.scripts/check-fixture-leaks.shfails the build when a fixture file contains the assertion vocabulary (VERDICT:,baseline-sha,IMPL_SHA, "the eval asserts", "pristine", "eval fixture", …). Legitimate hits —run-tests.shmust name the.check-logit writes — are allowlisted per path and per pattern with a reason inscripts/fixture-leak-allow.txt; stale entries fail. Ground truth belongs inskills/<name>/evals/README.md, which is never copied into a workdir.
When to run: manually, before bumping the version on any release whose diff touches one of those skills. They are deliberately not in CI — one iteration costs roughly 300–800k tokens.
How to run one iteration (from a Claude Code session on this repo):
-
Instantiate the eval's fixture into a scratch workdir with
bash scripts/fixture-scratch.sh <skill> <eval-id> --keep, which prints the path. It readsevals.json, so it applies whichever schema the suite uses — copy thefileslist (pillar 1) or run the eval's ownsetup(pillar 2) — and--print-promptemits the prompt with{{WORKDIR}}already substituted. Never point a run at the fixture template itself.Doing it by hand is a trap worth naming, because this step used to say
W=$(mktemp -d): thesetupcommands arecp -r <fixture> "$W", which nests the fixture into$W/<fixture-name>/when$Walready exists, and every grader then fails on a workdir that is correct in every other respect. -
Spawn a fresh-context subagent per run, told to read
skills/<name>/SKILL.mdand execute it as if the user had invoked the eval's prompt, with every repo-relative path resolved against the workdir. Pillar 2 briefs add eval mode: gates are pre-approved and the host task system / artifact publishing are unavailable, so the skill's fail-open paths apply — the telemetry report still gets written. A baseline (no-skill) run per eval is optional: it measures the skill's value, while the release gate only needs the with-skill regression signal. -
Grade with the skill's committed grader —
bash skills/<name>/evals/check.sh <id> "$W"for all six skills now (one PASS/FAIL line per expectation, exit 0 = all green). Pillar 2: exportFW_EVAL_DATE=<run date>when regrading a workdir on a later day. Pillar 1verify: the grader readsreport.mdandtranscript.mdfrom the workdir root — save them there, or pointFW_EVAL_REPORT/FW_EVAL_TRANSCRIPTat them. Never grade by re-deriving regexes by hand: that is how a vacuous assertion survives.While designing a fixture — before any executor run exists to grade — chain the whole loop in one command instead of writing a scratch script:
bash scripts/fixture-scratch.sh <skill> <id> --solution <name> --suite --check. It instantiates, applies the reference solution, runs the suite the grader would run (KATA_HARNESS=1, neverrun-tests.sh, which appends to the.check-logbeing graded), grades, prints onePASS:/FAIL:line per step, and tears the workdir down.--digestprints the value a new solution'sBASED-ONneeds. -
Aggregate into
benchmark.json/benchmark.md(skill-creator schema) and commit them underskills/<name>/evals/benchmarks/<date>/as the release evidence.
A regression (with-skill pass rate below the committed benchmark, or any planted-bug eval rationalized into a PASS) blocks the release until the skill text is fixed.
The baseline arm (value study) — documented, not yet run
The release gate needs only the with-skill arm: it answers "did this skill text regress?". A baseline arm — the same eval run by a subagent that is not given the skill — answers a different and unanswered question: "is this behavior the skill's, or would a strong model do it anyway?". Run it as a deliberate study, never as part of a release:
- Same fixture instantiation as step 1 above, into a separate workdir per arm so the two never share state.
- Spawn the baseline subagent with the eval's prompt and no reference to
skills/<name>/SKILL.md— no summary of it, no paraphrase. Everything else (eval-mode briefing, pre-approved gates, path resolution) stays identical, or the comparison measures the briefing instead of the skill. - Grade both arms with the same grader and record them as two
configurationvalues (with_skill,without_skill) in onebenchmark.json, as the pillar-1 benchmarks do. - Report the delta per assertion, not just per eval — a tie on pass rate can still hide which specific contract the baseline broke.
Status: not run for process/run. Their committed iteration is with-skill only (~237k tokens), so their 49 green assertions are a regression baseline and not evidence that the skills beat an unaided model. Cost of closing that gap is roughly 120k tokens per skill. Until it is run, do not cite pillar-2 pass rates as skill value.
work's katas are regression-only, and that is now a measured conclusion rather than a caveat. Three iterations tried to make them discriminate and all three tied at 100%: the v0.31.0 run (prompt named ./run-tests.sh), the P25 run (prompt de-hinted, but the fixture README still stated the grading rule verbatim), and the 2026-07-30 run with that leak removed and the fixture reading like an ordinary repo — where the baseline still wrote the regression test first, ran it red against a pristine cart.py, and only then fixed it. A strong model does test-first on a kata this small whether or not the skill says so. That is a fact about the task, not a defect in the skill, and citing the 100% as skill value would be the kind of unverifiable claim these evals exist to remove. What the suite still earns its cost for: if a future edit to work/SKILL.md stops inducing the red step, the with-skill arm drops below 8/8 and that is a real regression signal. Evidence: skills/work/evals/benchmarks/2026-07-30/benchmark.json.
Repo layout
The plugin lives at the repo root: .claude-plugin/ (manifest + marketplace), skills/, agents/, hooks/, scripts/. Setup guides are in docs/, and upgrades/ holds the per-version, AI-authored migration notes that /flywheel:update executes in installed repos (CI requires one per release).
Design research and the improvement backlog live in docs/research/ — see improvement-proposals.md for the living roadmap (P1–P6).

