ClaudeMods
☰
JA
● 0 人がオンライン ・閲覧 0 回
スポンサー作品を投稿
GitHub リポジトリ · 投稿者 tommy5dollar

effort-router

各タスクに必要な推論の労力を選択します。セッション自身のモデルがタスクを判断し、そのモデルで各レベルが何ができるかを知っており、ルーターは確信した場合にのみ動作します。労力設定と一致しない場合、タスク実行前に確認を求めます。サブエージェントは、起動したエージェントによって選択された独自のレベルを取得します。/route report は、労力がどこに使われたかを示します。

tommy5dollar@tommy5dollar

tommy5dollar/claude-mods/tree/main/effort-router

翻訳済み

この mod について

effort-router

タスクが明確になった後、セッションの推論労力をタスクと照合し、変更前に確認を求め、その設定を保持する Claude Code mod です。各サブエージェントは、それを起動したエージェントによって選択された独自のレベルを取得します。

これは、Anthropic の Claude Code の使用: 労力の使い方 (Thariq Shihipar、2026 年 9 月 25 日)] のガイダンスに従っています。この記事では、労力は検証とエッジケースのテストをもたらし、より良いアプローチではないことがわかりました。ルーターは、これらの原則と、各レベルがそのモデルで何ができるかについての情報源となるメモをセッションのモデルに提供し、判断させます。レベル名が Opus、Sonnet、Fable で異なる意味を持つため、特定の種類のタスクを固定レベルにマッピングすることはありません。

Claude Code 2.1.287 以降 (Claude Mods) が必要です。

ルール

労力ピッカーのレベルがデフォルトです。各プロンプトの後、タスクが明確になるまで、セッション自身のモデルが会話を見て、タスクに必要なレベルと、その確信度を命名します。その後:

  • まだ明確なタスクがない、または十分に確信がない場合: そのターンはピッカーのレベルで実行され、次のプロンプトが再度チェックされます。
  • ルーターがピッカーのレベルを命名した場合: そのターンが実行され、そのレベルがセッションに保持されます。
  • ルーターが異なるレベルを命名した場合: そのターンはルーターのレベルで実行され、そのレベルがセッションに保持されます。プロンプトの上の帯が一度開き、その理由と、ルーティングを停止して設定に戻るボタンが表示されます。

いずれにしても、一度レベルが保持されるとセッションは決定され、ルーターは自動的にチェックを停止します。帯では、Reassess now (または /route) がこれまでの会話について再度尋ねます。Reassess with my next prompt (または /route next) は設定に戻り、次のメッセージを送信するときに再度チェックします。これは、作業を別の方向に誘導しようとしているときに使用するものです。

ask の同意があれば、異なるレベルが代わりにあなたの回答を待ちます。Claude 自身の質問カードは、「労力ルーター: マルチプラットフォーム金融統合。中程度の労力ではなく高労力を使用しますか?」と尋ね、その間のレベルもオプションとして提供します。Use high は高労力でターンを実行し、高労力を保持します。Keep medium は中程度の労力で実行し、中程度の労力を保持します。質問を閉じると、ターンはピッカーのレベルで実行され、ルーターは未決定のままで、後のプロンプトで再度尋ねることができます。

状態

フッターは、ネイティブモデルと労力ピッカーのすぐ隣に、ルーターの状態を表示します。使用中のレベルは省略されています。数ピクセル離れた労力ピッカーがすでに表示しているためです。

| フッター | 意味 | 帯のボタン (フッターを押す) | | --- | --- | --- | | undecided (薄暗い) | まだ何も決定されていません。設定が適用され、送信する各プロンプトがチェックされます。待つ必要はありません | Assess now、Stop routing (back to medium) | | deciding… | 現在チェックが実行中です。数秒後に完了するとターンが開始されます | | | high? | ask の同意があれば: 質問が開いており、ターンはあなたの回答を待ちます | Assess now、Stop routing (back to medium) | | using high | 決定済み: すべてのメインスレッドリクエストは高労力で実行され、ルーターは自動的にチェックを停止します | Reassess now、Reassess with my next prompt、Stop routing (back to medium) | | no decision (薄暗い) | ルーターは確信することなくプロンプトを使い果たしました。設定はセッションの残りの部分に適用されます | Start routing | | off (薄暗い) | ルーティングを停止したため、ピッカーが担当します | Start routing |

Stop routing はこのセッションのルーターをオフにします。これ以上チェックは行われず、サブエージェントはルーティングされず、労力設定 (ボタンに記載) が再度適用されます。/route は引き続き回答し、Start routing または /route on は再度開始します。新しいセッションは通常通りルーティングされます。

ルーターがレベルを強制している間に労力ピッカーを自分で変更すると、同じことが起こります。ルーティングが停止し、そのリクエストから新しいレベルが使用されます。デスクトップでは、ルーターが別のレベルを実行している間、ピッカーは設定を表示し続けるため、すでに表示されているレベルを選択しても何も変わりません。代わりに Stop routing を使用してください。

<!-- スクリーンショット: ゲージとネイティブピッカーの横に「undecided」と表示されたフッター --> <!-- スクリーンショット: 「Effort router: ... Use high effort instead of medium?」という質問カードと、フッターに「high?」と表示されている -->

フッターの状態は単なるボタンです。それを押すと、プロンプトの上のルーターの帯が開きます。Effort router: using high for this session (bug fix in existing code). The crash needs tracing through the parser, but the fix is local. のような行、次にボタンと Hide (ホットキー x) が表示されます。括弧内の単語は、チェックがタスクと見なしたもの、その後の文は、そのレベルを選択した理由です。ボタンには 1、2 と番号が付けられています。いずれかの操作で帯が閉じ、フッターを再度押しても閉じます。ask の同意があれば、帯は自動的に開くことはありません。質問カードで同意します。

auto (デフォルト) の同意があれば、ルーターは尋ねません。ピッカーのレベルと異なるレベルを保持する場合、帯は一度自動的に開き、Effort router: changed from medium to high for this session (<task>). <why> と表示され、OK (変更が有効になる)、Go back to xhigh (前のレベルが設定でなかった場合)、および Stop routing (back to medium) が表示されます。フッターを押すと、ask の下と同じ帯が表示されます。

フッターはドロップダウンではなくボタンです。デスクトップアプリはフッターに Select をサイレントにドロップするためです。これは描画されず、エラーも報告されません (2.1.286 アプリでライブ検証済み。テストキットはこれを受け入れるため、キットではこれを検出できません)。スペースがなくなるとフッターは … で切り詰められるため、ラベルは短く保たれます。

ルーターは手動で選択したレベルを設定することはありません。それはネイティブの労力ピッカーの役割です。特定のレベルで実行するには、ルーターをオフにしてピッカーを使用します。ルーターがオフの場合、すべてのリクエストはピッカーのレベルで送信されます。

なぜ質問が Claude 自身の質問カードなのか

ルーターは、リクエスト自体からのみピッカーのレベルを知ることができます。リクエストが到着したときの turn.step フックの e.effort は、エンジンのレベルです。他に何も表示されません。設定リストには労力行がなく、セッションの最初のプロンプトはリクエストが存在する前に送信されます。したがって、比較と質問は、ターンのリクエストが送信されようとしているときに発生します。

通常のプロミスを待つ turn.step フックは、約 10 秒後に破棄され、リクエストはそれなしで送信されます。エンジンへの呼び出し ($.ui.ask。これは Claude 自身の AskUserQuestion カードを描画します) は、その制限にカウントされません。デスクトップ 2.1.286 でライブ検証済み: リクエストは回答が来るまで 23.6 秒間保持され、その後選択されたレベルで送信されました。帯のボタンでは待つものがなくなるため、質問はカードである必要があります。

決定方法

  • プロンプトが実行される前。 ルーターが決定している間、送信する各プロンプトは、ターンが開始される前に 1 回のチェックを待ちます。これはセッションごとに最大 decideWithin 回発生します。チェックに classifyTimeoutMs (15 秒) より長くかかったり、失敗したりした場合、ターンはピッカーのレベルで実行され、/route status が理由を示します。
  • チェックはセッションのモデルで実行されます。 作業するモデルがタスクを判断します。これは、小さなモデルよりも判断が優れており、レベルを正しく設定することによる節約がそれに比例して拡大するためです。2 番目のプロンプト以降、チェックは会話のフォークです。セッション自身の要求 (システムプロンプト、ツール、CLAUDE.md、メモリ、および会話全体) に 1 つの質問が追加され、セッションのプロンプトキャッシュから提供されます。Opus 5.5 で 72k トークンの会話で測定: 1.6 秒、会話全体がキャッシュから読み取られ、約 2.8k の新しい入力トークンと 40 の出力トークンで、約 3 セントです。フォークはセッションが最後に使用した労力で実行されます。
  • 最初のプロンプトは別の呼び出しです。 セッションが何も送信する前にフォークするリクエストはなく、mod は Claude Code のシステムプロンプトとツールでそれを作成することはできません。したがって、最初のチェックは、CLAUDE.md ファイル、ルール、メモリ (Claude Code が会話に渡すもの)、およびプロンプトを含む同じモデルへの 1 回の呼び出しであり、モデルのデフォルトの労力で実行されます。これはキャッシュされません。大量の指示を含む約 13k トークンで、Opus 5.5 で約 5 セント、Fable 5.1 で約 13 セント、セッションごとに 1 回 (測定値 1.4 秒) です。firstCheckInstructions: false はプロンプトのみを送信します。ターンがまだ実行中に送信されたプロンプトはチェックされません (このケースはライブテストされていません)。次のプロンプトがチェックされます。
  • 確信度。 各チェックは、各レベルが正しい確率を与えます。たとえば、medium 10%, high 50%, xhigh 40% です。ルーターは、チェックがあなたのレベルが一方の方向に間違っているとどれだけ確信しているかを尋ねます。ここでは、中程度が低すぎる可能性が 90% です。それが confidence (デフォルトは 0.7) をクリアすると、スプレッドの中央、つまり十分である可能性が少なくとも同じくらい高い最低レベルに移動します。ここでは高です。したがって、高と xhigh の間で迷っているチェックでも、中程度から高に移動させ、間違っていると確信しているレベルに留まらせることはありません。あなたのレベルが正しいと十分に確信している場合、それを保持します。バーの下では何も変更せず、次のプロンプトの後に再度チェックします。そのプロンプトにはより多くの会話が含まれます。各チェックのスプレッド、信頼度、結果は支出台帳に保持されるため、自信のある動きが保持されたり覆されたりする頻度からバーを設定できます。showChecks をオンにすると、各チェックの後に 1 行表示されます。
  • xhigh まで。 ルーターは highestLevel まで、デフォルトでは xhigh を選択し、max はチェックに提供されません。3 つのモデルすべてで、max が xhigh を上回ることはめったになく、考えすぎる可能性があります。highestLevel を max に設定すると許可されます。
  • 中間のレベル。 ルーターのレベルがあなたのレベルから 2 つ以上離れている場合、質問は中間のレベルも提供します。中程度から、xhigh を選択するチェックは、Use xhigh、Use high、または Keep medium を尋ねます。
  • 別のチェックモデル。 classifierModel を別のサポートされているモデル (opus、sonnet、または fable) に設定すると、会話の短縮コピーを読み取るための個別の呼び出しが行われます。Haiku は使用されません。他の名前はセッションのモデルにフォールバックします。
  • 最初のリクエストで比較。 読み取りのレベルは、ターンの最初のリクエストを待ちます。そこでピッカーのレベルがわかっており、上記のルールが適用されます。比較は、ルーターが何も変更する前の、ルーターに到達したときのレベルを使用するため、常にピッカーのレベルです。
  • あなたの回答もカウントされます。 Claude の多肢選択問題 (メインスレッドの AskUserQuestion) への回答は、次のチェックが確認する会話の一部です。ルーターは回答が Claude に戻る前に再度チェックし、ターンの次のリクエストがルールを適用し、回答された質問はプロンプトと同様に予算にカウントされます。セッションのモデルでは、そのチェックはあなたの回答を含むフォークであり、ターン途中のフォークは他のフォークと同様にキャッシュを読み取ります。
  • 読み取る内容。 フォークは、セッションのモデルが見るものとまったく同じものを見ます。個別の呼び出し (最初のプロンプト、または別の classifierModel) は、プロンプト全体、Claude の質問とあなたの回答、Claude の返信 (最後のものはそれほどではない) を切り詰めて読み取り、その他のツール呼び出しは名前のみで、classifierMaxChars (24,000) に制限されます。上限を超えると、最初のプロンプト (元のタスク)、次に最新の行、Claude の返信の前のプロンプトと回答が保持されます。/route status は、最後の読み取りで送信された量を示します。
  • タスクがある前のみ未決定。 モデルは、挨拶、たとえば「最新のコードをプルする」などのハウスキーピング、または作業前の質問などのオープニングフィラーに対してのみ「未決定」と回答します。実際のタスクを述べると、詳細は未定であっても、そのタスクに最も必要なレベルを選択します。プロンプトの作業例は、会話 (フィラー、範囲の絞り込み、番号付きの質問への短い返信、回答済みの質問) の読み方を教え、それらのどれもレベルを命名しません。
  • 最新のやり取りが最も重要。 後からの明確化は以前の要求を上書きし、短い返信はそれが回答する質問に対して読み取られます。「支払い再試行ロジックのリファクタリング」という質問を却下し、Claude の「1. 完全な書き換えか 2. 定数のみを抽出するか?」という質問に「2」と回答した場合、次の読み取りでは低と命名されます。

インストール

まず作者の README で marketplace とプラグイン名を確認してください。コマンドはリポジトリの構成によって変わる場合があります。

claude plugin marketplace add tommy5dollar/claude-mods
claude plugin install effort-router
原文 / README

effort-router

A Claude Code mod that checks your session's reasoning effort against the task once the task is clear, asks before changing it, then holds it. Each subagent gets its own level, chosen by the agent that launches it.

It follows Anthropic's guidance in Using Claude Code: Spending your effort (Thariq Shihipar, 25 September 2026). The article found that effort buys verification and edge-case testing, not a better approach. The router gives your session's model those principles and sourced notes on what each level can do on that model, then lets it judge. It never maps a kind of task to a fixed level, because level names mean different things on Opus, Sonnet and Fable.

Requires Claude Code 2.1.287 or later (Claude Mods).

The rule

Your effort picker's level is the default. After each prompt, until the task is clear, your session's own model looks at the conversation and names the level the task needs, with how sure it is. Then:

  • No clear task yet, or not sure enough: the turn runs at your picker's level, and the next prompt is checked again.
  • The router names your picker's level: the turn runs, and that level is kept for the session.
  • The router names a different level: the turn runs at the router's level, which is kept for the session. The band above the prompt opens once to say so, with why, and a button to stop routing and go back to your setting.

Either way, once a level is kept the session is decided and the router stops checking by itself. In the band, Reassess now (or /route) asks again about the conversation so far. Reassess with my next prompt (or /route next) goes back to your setting and checks again when you send your next message, which is the one to use when you're about to steer the work somewhere else.

With consent ask, a different level waits for your answer instead. Claude's own question card asks "Effort router: Multi-platform finance integration. Use high effort instead of medium?", with any levels in between as options too. Use high runs the turn at high and keeps high. Keep medium runs it at medium and keeps medium. If you dismiss the question, the turn runs at your picker's level, the router stays undecided, and a later prompt can ask again.

The states

The footer, right beside the native model and effort pickers, shows the router's state. It leaves out the level in use, because the effort picker a few pixels away already shows it.

| Footer | What it means | Band buttons (press the footer) | | --- | --- | --- | | undecided (dim) | Nothing decided yet. Your setting applies, and each prompt you send is checked. Nothing to wait for | Assess now, Stop routing (back to medium) | | deciding… | A check is running now. The turn starts when it's done, in a few seconds | | | high? | With consent ask: the question is open and the turn waits for your answer | Assess now, Stop routing (back to medium) | | using high | Decided: every main-thread request runs at high, and the router stops checking by itself | Reassess now, Reassess with my next prompt, Stop routing (back to medium) | | no decision (dim) | The router ran out of prompts without being sure. Your setting applies for the rest of the session | Start routing | | off (dim) | You stopped routing, so the picker is in charge | Start routing |

Stop routing turns the router off for this session: no more checks, subagents aren't routed and your effort setting (named in the button) applies again. /route still answers, and Start routing or /route on starts it again. New sessions are routed as usual.

Changing the effort picker yourself while the router has a level in force does the same: routing stops and your new level is used from that request on. In Desktop the picker keeps showing your setting while the router runs another level, so picking the level it already shows changes nothing there. Use Stop routing instead.

<!-- screenshot: footer showing "undecided" beside the gauge and the native pickers --> <!-- screenshot: the question card "Effort router: ... Use high effort instead of medium?" with the footer reading "high?" -->

The footer state is a plain button. Pressing it opens the router's band above the prompt: a line such as Effort router: using high for this session (bug fix in existing code). The crash needs tracing through the parser, but the fix is local., then the buttons and Hide (hotkey x). The words in brackets are what the check took the task to be, and the sentence after is why it chose that level. The buttons are numbered 1, 2. Any action closes the band, and pressing the footer again closes it too. With consent ask the band never opens by itself: the question card is where you agree.

With consent auto (the default) the router doesn't ask. When it keeps a level that differs from your picker's, the band opens by itself once, reading Effort router: changed from medium to high for this session (<task>). <why>, with OK (the change stands), Go back to xhigh when the level before wasn't your setting, and Stop routing (back to medium). Pressing the footer shows the same band as under ask.

The footer is a button, not a dropdown, because the Desktop app silently drops a Select in the footer: it is not drawn, and nothing reports an error (verified live on the 2.1.286 app; the test kit accepts it, so the kit cannot catch this). The footer truncates with … when space runs out, so the label stays short.

The router never sets a level you pick by hand: that is what the native effort picker is for. To run at a specific level, turn the router off and use the picker; with the router off, every request goes out at the picker's level.

Why the question is Claude's own question card

The router can only learn your picker's level from the request itself. e.effort in the turn.step hook, as the request arrives, is the engine's level for it. Nothing else shows it: the config list has no effort row, and a session's first prompt is submitted before any request exists. So the comparison, and the question, happen when the turn's request is about to go out.

A turn.step hook that waits on an ordinary promise is abandoned after about 10 seconds, and the request goes out without it. A call into the engine ($.ui.ask, which draws Claude's own AskUserQuestion card) doesn't count against that limit. Verified live on Desktop 2.1.286: the request was held for 23.6 seconds until the answer came, then went out at the chosen level. A button in the band would leave nothing to wait on, so the question has to be the card.

How it decides

  • Before your prompt runs. While the router is deciding, each prompt you send waits for one check before the turn starts. It happens at most decideWithin times per session. If the check takes longer than classifyTimeoutMs (15 s) or fails, the turn runs at the picker's level and /route status shows why.
  • Checks run on your session's model. The model you chose to work in judges the task, because it judges better than a small model and the savings from getting the level right scale with it. From the second prompt on, a check is a fork of the conversation: the session's own request (system prompt, tools, CLAUDE.md, memory and the whole conversation) with one question added, served from the session's prompt cache. Measured on Opus 5.5 with a 72k-token conversation: 1.6 s, the whole conversation read from cache, about 2.8k fresh input tokens and 40 output tokens, so about 3 cents. The fork runs at the effort the session last used.
  • The first prompt is a separate call. Before the session has sent anything there is no request to fork, and a mod can't build one with Claude Code's system prompt and tools. So the first check is one call to the same model with your CLAUDE.md files, rules and memory (as Claude Code hands them to the conversation) and your prompt, at the model's default effort. It isn't cached: about 13k tokens with a large set of instructions, so roughly 5 cents on Opus 5.5 or 13 cents on Fable 5.1, once per session (1.4 s measured). firstCheckInstructions: false sends only the prompt. A prompt sent while a turn is still running isn't checked (that case hasn't been tested live); the next prompt is.
  • How sure it is. Each check gives every level a probability of being the right one, for example medium 10%, high 50%, xhigh 40%. The router asks how sure the check is that the level you're on is wrong in one direction: here 90% that medium is too low. When that clears confidence (0.7 by default), it moves to the middle of the spread, the lowest level at least as likely as not to be enough. Here that's high. So a check torn between high and xhigh still moves you off medium, to high, rather than leaving you on the one level it's sure is wrong. When it's sure enough your level is right, it keeps it. Below the bar it changes nothing and checks again after your next prompt, which by then carries more of the conversation. Every check's spread, confidence and outcome is kept in the spend ledger, so the bar can be set from how often a confident move was kept or overruled. Turn on showChecks to see a line after each check.
  • Up to xhigh. The router picks up to highestLevel, xhigh by default, and the checks aren't offered max: on all three models max rarely beats xhigh and can overthink. Set highestLevel to max to allow it.
  • The levels in between. When the router's level is two or more away from yours, the question offers the levels in between too: from medium, a check that picks xhigh asks Use xhigh, Use high or Keep medium.
  • Another check model. Set classifierModel to another supported model (opus, sonnet or fable) for separate calls that read a shortened copy of the conversation. Haiku is never used: any other name falls back to your session's model.
  • Compared at the first request. The read's level waits for the turn's first request, where your picker's level is known, and the rule above applies there. The comparison uses the level as it reached the router, before the router changes anything, so it is always your picker's.
  • Your answers count too. Answers to Claude's multiple-choice questions (AskUserQuestion on the main thread) are part of the conversation the next check sees. The router checks again before the answers go back to Claude, the turn's next request applies the rule, and answered questions count toward the budget like a prompt. On the session's model that check is a fork carrying your answers, and a fork mid-turn reads the cache like any other.
  • What it reads. A fork sees exactly what the session's model sees. A separate call (the first prompt, or another classifierModel) reads your prompts in full, Claude's questions with your answers, Claude's replies truncated (the last one less so) and other tool calls as names only, capped at classifierMaxChars (24,000). Over the cap it keeps your first prompt (the original task), then the newest lines, your prompts and answers before Claude's replies. /route status says how much the last read sent.
  • Undecided only before there is a task. The model answers "undecided" only for opening filler: greetings, housekeeping such as "pull the latest code", or questions before any work. Once you state a real task it picks the level that task most likely needs, even while the details are open. The prompt's worked examples teach reading the conversation (filler, a narrowed scope, a short reply to a numbered question, answered questions), and none of them names a level.
  • The latest exchange counts most. A later clarification overrides an earlier ask, and a short reply is read against the question it answers. If you dismissed the question for "refactor the payment retry logic" and then answer Claude's "1. full rewrite or 2. just extract the constant?" with "2", the next read names low.
  • Decided is decided. Once a level is kept, every later main-thread request runs at it and the router stops reading. Subagents get their own level (below). When the kept level isn't the picker's, the terminal also runs /effort <level> once the session is idle, so the native picker label matches. In the Desktop app the picker belongs to the app, so its label stays where you set it; trust the footer. A kept level survives claude --resume. Choosing Keep medium keeps medium even if you move the picker later; /route off hands control back to the picker.
  • It stops after decideWithin prompts (6 by default), counted from the start of the session. If nothing is decided by then, the router turns off with the reason no clear task after 6 prompts, or not sure enough after 6 prompts, last check high at 65% when the checks named a level they weren't sure of, after asking any question still waiting. It never calls the model again on its own.
  • Existing sessions are left alone. The first time the router sees a session that already has decideWithin or more prompts in it, or more than skipAboveTokens (20,000) tokens of conversation (a long chat from before the router was installed, say), it starts off with the reason session started before the router: no question, no model calls. Fewer earlier prompts count toward the budget. A resumed session with saved router state keeps that state.
  • /route asks now. It reads the whole conversation in any state, ignoring the budget. If the answer is the level already in use, it says so and changes nothing. If it is your picker's level while a different one is kept, it keeps the picker's level without asking. Otherwise the question card opens straight away ("Use low effort instead of high?", naming the level in force now, with any levels in between), and your answer is kept. Add a hint to steer it: /route this is a security review, /route keep it quick. The hint is weighed strongly and kept for later reads until a level is kept. The band's Assess now (Reassess once decided) is the same as bare /route: it asks the session's model, so it takes a couple of seconds and uses your plan like any request. The footer reads checking… while it runs, then the band shows what it found ("high still fits (bug fix, 90% sure). Nothing changed.") until you hide it.
  • Consent ask. If you'd rather approve each change: a level other than your picker's waits on the question card. Set it in /plugin configure, or with EFFORT_ROUTER_CONSENT=ask. Headless runs (-p) have no one to answer, so leave them on auto.

Models

The router supports the current models: Fable 5.1, Opus 5.5 and Sonnet 5.5. Level names don't mean the same amount of thinking on each, and each responds to effort differently. In Claude Code, Opus 5.5 and Sonnet 5.5 default to medium and Fable 5.1 to high. Opus 5.5 gains most from low to medium and little above high, while Sonnet 5.5 gains a lot at every step. Routing one like another would be a mistake. Each has a notes file in rules/models/ on what each level can do there: what it's good for, what it misses, and measured gains and costs. Every check carries the notes for the session's model (or the subagent's) after the routing rules, as the main guide to the level. /route rules prints them. The evidence behind each note, with sources, is in rules/models/research-2026-10.md.

The notes guide the level instead of fixed rules because of an eval on 4 October 2026. With rules that tied kinds of task to levels, all three models gave almost the same answers and ignored their notes: Sonnet kept picking max for autonomous work, though its evidence says xhigh. Without those rules, each model's answers moved the way its evidence predicts (TESTING.md, "Prompt variants").

On any other model the router stands aside: the footer reads off, /route status says which models it works with, and no checks run. Its state is kept, so switching back with /model picks up where it was. A new model needs a new version of the router.

Subagents

Each subagent gets its own level, chosen by the agent that launches it.

  • Its parent decides. When Claude launches a subagent, the launch waits for one fork of the parent's conversation, asked which level the subagent needs, with its brief (capped at classifierMaxChars). The parent knows the task and why it's delegating this part, which a brief alone often doesn't say. Then the subagent starts, and every request it makes carries that level. Measured on Opus 5.5: 2.3 to 3.3 s, the parent's conversation read from cache, about 2 cents. With another classifierModel it's a separate call that reads the brief alone.
  • On its own model. The check is told which model the subagent runs on (the Agent call's model, else its definition's, else the parent's) and gets that model's notes. A subagent has no user in the loop, and the check always picks a level. Your rules and your organisation's rules apply here too.
  • Haiku agents are left alone. A subagent on Haiku (the built-in Explore agent runs there) or on another model the router doesn't support isn't checked. Haiku takes no effort setting anyway.
  • An agent's own effort: wins. If the agent's definition sets an effort, the router doesn't read its brief and leaves its requests alone, so the engine applies the definition's level. It looks for the definition by its frontmatter name: in the project's .claude/agents/*.md, then your ~/.claude/agents/*.md, and in the agents key of policy, project and user settings. The first definition with that name decides, as it does for the engine: a project definition without effort: still beats a user one with it. Definitions are scanned once per session. /route status shows such an agent as low: <description> (set by its agent definition).
  • Forks and failures take the parent's level. A fork shares its parent's context, so it skips the read. If a read fails, times out (classifyTimeoutMs) or returns something unusable, the subagent also takes its parent's level. That is the main thread's level in use, or for a subagent launched by another subagent, that subagent's level. With no level anywhere, its requests are left alone.
  • It runs even when the main thread is left alone. In an existing session the router leaves the main thread alone, but each new subagent brief is a fresh, whole task, so subagents are still routed. The same holds after the router turns itself off with no clear task. When you turn the router off yourself (/route off or Stop routing), subagents go back to the picker's level too, and /route on brings their routed levels back.
  • Seeing it. /route status lists this session's routed subagents, newest first (the last 10), with level, description, agent type and why. The debug log has one line per routed launch. Nothing is added to the footer or to the parent's conversation.

Claude can't set a subagent's effort itself today: the Agent tool takes a model but no effort, so without the router every subagent runs at the session's level unless its agent definition sets one. Set routeSubagents to false to go back to that (the main thread's level in use, as before 0.7.0).

Where the effort went

/route report shows what your requests spent at each level over the last 7 days. /route report session, month or all cover other spans. For example:

Effort for the last 7 days (since 2026-09-28): 412 requests in 9 sessions, 610k output tokens.
By level:
  low: 120 requests, 31k output tokens (avg 258)
  medium: 260 requests, 410k output tokens (avg 1.6k)
  high: 32 requests, 169k output tokens (avg 5.3k)
Changed by the router: 74 requests
  subagents, medium → low: 44 requests, 9.9k output tokens (avg 225, vs 1.6k for those left at medium)
  main conversation, medium → high: 30 requests, 160k output tokens (avg 5.3k, vs 1.6k for those left at medium)
The router's own checks: 61 (9 of a first prompt, 14 of a conversation, 38 for subagents), using 3.1k output and 1.20M input tokens.
By repo (output tokens): payments 400k, web 210k.
No "saved" figure: the router lowers easy tasks and raises hard ones, so these averages can't show what a changed request would have cost.
  • What it records. Every model request in every session with the router installed (0.9.0 on), on the main thread and in subagents, with the router on or off. For each one it keeps the level the request arrived at (your picker's, or the level a subagent would have inherited), the level it went out at, and its tokens as the API reported them. Requests are summed per day into one small JSON file per session, in ~/.claude/effort-router/spend/. The file is written when a turn ends, and nothing leaves your machine.
  • What it shows. Requests and output tokens per level, with the average per request. The requests the router moved, by thread and direction, each beside the average request left at the level it came from. Requests whose agent definition set their level. The router's own checks, by kind, so its cost is in the same report. Each session check's level, confidence and outcome is kept in the file too, for setting the confidence bar later. Over more than one session, output by repo.
  • Why output tokens. Output (thinking plus the answer) is what effort changes most. Input is recorded too.
  • Why there is no "saved" figure. The router lowers easy tasks and raises hard ones. A lowered request is small partly because its task was small, so comparing it with the average medium request would overstate the saving, and the same comparison would overstate what a raised request cost extra. Only running the same task at both levels can say what a request would have cost at its old level. The report gives the measured numbers side by side and leaves that estimate out.

Policy

The shipped rules (rules/default.md) are principles, not a table of levels:

  • Effort buys verification, edge-case testing and independent judgement, not a better approach.
  • Weigh how much is hidden (edge cases, existing code, money, several external systems, concurrency, security), whether you're in the loop, how well specified the task is, and how big it is.
  • Pick the level that does the work well on this model without paying for thinking it won't use.

They come from the article above and Anthropic's effort docs. What each level can do comes from the model notes.

Commands

| Command | What it does | | --- | --- | | /route | Runs the router now over the whole conversation, in any state, and asks if its level differs | | /route <hint> | The same, with a hint for the classifier (/route this is a security review) | | /route status | Shows the state and why, the consent mode, automatic reads used of the budget, classifier calls and how long the last read took, how much transcript it sent, the last verdict (with the raw reply and when), the last error, and this session's routed subagents | | /route report [session\|week\|month\|all] | Shows where the effort went: requests and output tokens per level, what the router moved, its own reads, and output by repo. The last 7 days by default | | /route off | Turns the router off and restores the picker's earlier level | | /route on | Turns the router back on: deciding over the whole conversation, with a fresh budget | | /route next | Back to deciding at your setting, with a fresh budget: your next prompt is checked before its turn starts (the band's Reassess with my next prompt) | | /route rules | Prints the effective rules and which layers contributed | | /route rules init [user\|project] | Writes a starter rules file that keeps the defaults | | /route rules critique | Asks Sonnet to critique your custom rules |

/route is registered with $.command.register, so it shows in the typeahead. State is per session. /route decide is a hidden alias of bare /route. There is no command to set a level: turn the router off and use the effort picker.

Options

Set them in /config, or under pluginConfigs["effort-router@tommy-mods"].options in settings.json.

| Option | Default | Meaning | | --- | --- | --- | | consent | auto | When the router wants a different level from your picker: auto uses the router's level without asking and shows it once in the band, with a button to stop routing. ask holds the turn and asks (Use the router's level, a level in between or Keep yours) | | decideWithin | 6 | Prompts (and answered questions) the router reads automatically, counted from the session's start | | classifyTimeoutMs | 15000 | How long a prompt waits for the check before it runs anyway | | classifierMaxChars | 24000 | Most transcript characters a separate check sends | | classifierModel | session | session: your session's own model, as a fork from the second prompt. Or another supported model (opus, sonnet, fable) for separate checks. Haiku is never used | | highestLevel | xhigh | The highest level the router picks. max allows max, which rarely beats xhigh on the current models | | confidence | 0.7 | How sure (0 to 1) a check must be that your current level is wrong in one direction (or right) before the router acts. 0 acts on any check | | showChecks | false | Print a line after each automatic check: its spread, how sure it was and what the router did | | skipAboveTokens | 20000 | A session first seen with more conversation than this keeps your effort setting | | firstCheckInstructions | true | Send your CLAUDE.md files, rules and memory with the first prompt's check. false: the prompt only | | syncPicker | true | Run /effort <level> so the terminal's picker label matches | | routeSubagents | true | Give each subagent its own level, chosen at launch by the agent starting it. false: subagents run at the main thread's level | | footerControl | button | button makes the footer state a button that opens the band. label draws plain text, and /route is the control | | rules | empty | Rules text for your user layer. A rules file takes precedence |

The environment variable EFFORT_ROUTER_CONSENT=ask|auto overrides consent, which helps in headless runs (-p has no one to answer a question, so use auto there). The older names still work there: apply and none mean auto, while confirm and band mean ask.

Customising the rules

The rules are plain markdown, layered from the bottom up:

  1. The shipped defaults (rules/default.md)
  2. Your organisation's rules from managed (policy) settings, if it sets any
  3. Your rules: ~/.claude/effort-router.md, or the rules option in your user settings
  4. The project's rules: <project root>/.claude/effort-router.md (commit it), or the rules option in project settings

A line that is exactly $defaults pulls in everything beneath that layer. Text after the line adds to the rules, and later rules win. Text before it goes first. A file with no $defaults line replaces everything beneath it. Missing or empty files change nothing, and HTML comments are ignored.

$defaults

- This is a payments codebase. Never pick below high: money movement needs verification.

The files are re-read on every classification, so edits apply without a reload. An unreadable file is skipped. The frame around the rules (undecided only before a task, the worked examples, reply in JSON) is fixed, so no rules file can break the parser.

For organisations

An organisation can set routing rules centrally in managed settings (managed-settings.json), as it does for other Claude Code policy:

{
  "pluginConfigs": {
    "effort-router@tommy-mods": {
      "options": {
        "rules": "$defaults\n\n- Code under payments/ or ledger/ is never routed below high.\n- Infrastructure changes (terraform/, k8s/) are high.",
        "rulesMode": "enforce",
        "allowOff": false
      }
    }
  }
}
  • rulesMode: "extend" (the default) layers the org rules over the shipped defaults. Users and projects can add to them with $defaults, or replace them.
  • rulesMode: "enforce" makes the org layer final. Personal and project rules are ignored, and /route rules init says so.
  • routeSubagents: false turns subagent routing off for everyone, whatever their own setting.
  • allowOff: false stops users turning the router off, so the organisation's routing always applies. /route off refuses, the band has no Stop routing, and a session saved as off comes back deciding. When the budget runs out with nothing suggested, the router idles as deciding (no more reads) instead of turning off. /route still works.

A top-level "effortRouter": { "rules": ..., "rulesMode": ..., "allowOff": ..., "routeSubagents": ... } object works too. The router reads these four settings only from the policy source, so a user cannot claim enforce for themselves.

Install

claude plugin marketplace add tommy5dollar/claude-mods
claude plugin install effort-router@tommy-mods

There's nothing to configure: every option has a default. /plugin configure lists the options as not yet set, which just means the defaults apply. Change one only when you want something different (see Options).

For development, run claude --plugin-dir ./effort-router.

Known limits

  • In the Desktop app the native effort picker never changes: the app owns it and nothing a mod can call sets it. The requests still go out at the routed level; trust the footer label.
  • In the terminal the router can't set the picker label directly either. It runs /effort <level> when the session is idle, which prints a line in the transcript, and the footer shows the true level until then. Headless (-p) runs skip the sync because its output would replace the run's printed result. The per-request override still applies there.
  • Each prompt waits for the check while the router is deciding (at most decideWithin prompts per session): about 1.5 s on Opus 5.5, longer on a model that thinks more by default. If the check times out (classifyTimeoutMs), that turn goes at the picker's level; a late answer is ignored.
  • The first prompt's check can't share the session's prompt cache: the engine offers no way to fork before the first response, and a separate call can't carry Claude Code's system prompt or tools. It pays for your instructions and the prompt once per session. Claude Code also never reuses the cache for the conversation part of a session's first request, so the first fork after a one-request first turn pays for about 13k tokens again.
  • On Bedrock, Google Cloud or an LLM gateway, Claude Code clears the cached conversation when effort changes (Anthropic's docs). A locked level that differs from your setting changes effort once, so expect one uncached request there. With an API key or a subscription the cache is kept.
  • The confidence bar (0.7) is a starting guess, not calibrated yet. The ledger keeps every check's confidence and outcome for that.
  • A fork answers at the effort the session last used, and the router can't change it.
  • A request whose picker level is a number rather than a named level, or a model that takes no effort, can't be compared, so the waiting verdict waits for the next request that can.
  • The router changes effort only, never the model. A request to a model that takes no effort is left alone.
  • Each automatic check is one call per prompt while deciding, for at most decideWithin prompts. Nothing more is spent once a level is locked or the budget is spent, except when you run /route. /route report shows what the checks cost.
  • Once decided, the router does not notice a change of phase on its own (for example "now verify it" after an implementation). Run /route (or the band's Reassess now), optionally with a hint, or /route next to be asked with your next prompt.
  • A definition's effort: is respected for user and project agent files and for the agents key in settings, not for plugin agents (<plugin>:<name>), which can't be located reliably. Those are routed from their brief, which replaces any effort their definition sets. A mod's $.agent.register({ effort }) is ignored by the engine itself (an engine bug), and the router routes those agents too.
  • An agent file added or edited mid-session is seen from the next session: definitions are scanned once per session.
  • Workflow agents that don't launch through the Agent tool raise no agent.spawn, so they keep the main thread's level.
  • Subagent levels are kept in memory only. After a restart or resume, a subagent still running from before takes the main thread's level.
  • The router adds no note about the chosen level to the system prompt, because changing a cached prompt section would break the prompt cache. The /effort echo tells the model instead, and it is appended to the transcript, so the cache holds.
  • The spend report starts at 0.9.0: sessions from before it aren't in it. A request with no reported usage (failed or interrupted) isn't counted. Days are UTC.
  • The band and footer draw in the terminal and the Desktop app. VS Code and -p run the hooks without the UI. Under ask in -p the question has no one to answer, so requests stay at the picker's level; set consent (or EFFORT_ROUTER_CONSENT) to auto there.

Development

bun test                            # pure policy: trimming, parsing, rule layering, /route grammar, subagent reads, the spend ledger and report
claude plugin test .                # engine kit: band, buttons, footer button, turn.step, /route, org layers, agent.spawn, the ledger saved and reported
claude plugin validate . --strict
bun run eval                        # opt-in: the real classifier over eval/fixtures.ts (see below)

bun run eval sends each fixture to the real model through claude -p --safe-mode (no plugins or hooks, no tools), with the system prompt and input the router builds from rules/default.md: for the session read, a first check on the session's model (--model, default opus) at its default effort with that model's notes from rules/models/, and for a subagent's read, a separate call on the same model with that model's notes and the brief. It can't reproduce the forks that later checks and subagent reads make. It parses the reply with the router's own parser and prints each verdict, the pass rate and every miss. --runs 3 repeats each fixture (the model is not deterministic), --model sonnet runs the session set as a Sonnet session, --effort sets the effort and --confidence the bar, --set session or --set subagent runs one set and --only <text> filters fixtures by name. It uses your Claude Code login, and each fixture costs one small model call.

TESTING.md lists the live checks.

Licence

MIT

関連作品