ClaudeMods
☰
ZH-CN
● 0 人在线 · 浏览 0 次
赞助提交作品
GitHub 仓库 · 发布者 tommy5dollar

effort-router

根据每个任务所需的工作量进行选择。您会话的模型会根据该模型判断任务,了解每个级别在该模型上可以做什么,并且路由器只有在确信后才会执行操作。当它与您的工作量设置不一致时,它会在任务运行前进行询问。子代理会获得自己的级别,由启动它们的代理选择。/route report 显示了您的工作量去向。

tommy5dollar@tommy5dollar

tommy5dollar/claude-mods/tree/main/effort-router

已翻译

关于这个 mod

effort-router

一个 Claude Code 模组,它在任务明确后检查您会话的推理工作量与任务是否匹配,在更改前进行询问,然后保持该设置。每个子代理都会获得自己的级别,由启动它的代理选择。

它遵循 Anthropic 在 使用 Claude Code:投入您的精力(Thariq Shihipar,2026 年 9 月 25 日)] 中的指导。文章发现,投入精力可以带来验证和边缘案例测试,而不是更好的方法。路由器将这些原则和关于每个级别在该模型上可以做什么的来源说明提供给您会话的模型,然后让它进行判断。它从不将某种任务映射到固定级别,因为级别名称在 Opus、Sonnet 和 Fable 上意味着不同的东西。

需要 Claude Code 2.1.287 或更高版本(Claude Mods)。

规则

您的工作量选择器的级别是默认值。在每次提示后,直到任务明确,您会话的模型会查看对话并命名任务所需的级别,以及它有多确定。然后:

  • 任务尚不明确,或不够确定: 该回合以您选择器的级别运行,并再次检查下一个提示。
  • 路由器命名您选择器的级别: 该回合运行,并且该级别会话保持不变。
  • 路由器命名不同的级别: 该回合以路由器的级别运行,该级别会话保持不变。提示上方的横幅会打开一次,说明原因,并提供一个按钮来停止路由并返回您的设置。

无论哪种方式,一旦级别确定,会话就已决定,路由器会自行停止检查。在横幅中,Reassess now(或 /route)会再次询问到目前为止的对话。Reassess with my next prompt(或 /route next)会返回您的设置,并在您发送下一条消息时再次检查,这是当您打算将工作引导到其他地方时使用的。

经同意 ask,不同的级别会等待您的答案。Claude 自己的问题卡片会询问“工作量路由器:多平台金融集成。使用高工作量而不是中等工作量?”并提供中间的任何级别作为选项。Use high 以高工作量运行该回合并保持高工作量。Keep medium 以中等工作量运行并保持中等工作量。如果您关闭问题,该回合以您选择器的级别运行,路由器保持未决定状态,并且稍后的提示可能会再次询问。

状态

页脚,紧邻原生模型和工作量选择器,显示路由器的状态。它省略了正在使用的级别,因为几像素之外的工作量选择器已经显示了它。

| 页脚 | 含义 | 横幅按钮(按下页脚) | | --- | --- | --- | | undecided (暗淡) | 尚未决定。您的设置适用,并且您发送的每个提示都会被检查。无需等待 | Assess now,Stop routing (back to medium) | | deciding… | 正在运行检查。回合将在几秒钟内完成时开始 | | | high? | 经同意 ask:问题已打开,回合等待您的答案 | Assess now,Stop routing (back to medium) | | using high | 已决定:每个主线程请求都以高工作量运行,路由器自行停止检查 | Reassess now,Reassess with my next prompt,Stop routing (back to medium) | | no decision (暗淡) | 路由器在没有确定的情况下用完了提示。您的设置适用于会话的其余部分 | Start routing | | off (暗淡) | 您停止了路由,因此选择器负责 | Start routing |

Stop routing 会关闭此会话的路由器:不再进行检查,子代理不再路由,并且您的工作量设置(按钮中命名)再次适用。/route 仍然回答,Start routing 或 /route on 再次启动它。新会话照常路由。

当路由器强制执行某个级别时,您自己更改工作量选择器也会执行相同的操作:路由停止,并且从该请求开始使用您的新级别。在桌面版中,当路由器运行另一个级别时,选择器会继续显示您的设置,因此选择它已经显示的级别不会改变任何内容。请改用 Stop routing。

<!-- 截图:页脚显示“未决定”在仪表和原生选择器旁边 --> <!-- 截图:问题卡片“工作量路由器:...使用高工作量而不是中等工作量?”页脚显示“高?” -->

页脚状态是一个普通按钮。按下它会打开提示上方的路由器横幅:一行,例如 Effort router: using high for this session (bug fix in existing code). The crash needs tracing through the parser, but the fix is local.,然后是按钮和 Hide(热键 x)。括号中的词是检查认为的任务,后面的句子是它选择该级别的原因。按钮编号为 1,2。任何操作都会关闭横幅,再次按下页脚也会关闭它。经同意 ask,横幅永远不会自行打开:问题卡片是您同意的地方。

经同意 auto(默认),路由器不会询问。当它保持与您选择器不同的级别时,横幅会自行打开一次,显示 Effort router: changed from medium to high for this session (<task>). <why>,并带有 OK(更改生效)、Go back to xhigh(如果之前的级别不是您的设置)和 Stop routing (back to medium)。按下页脚会显示与 ask 下相同的横幅。

页脚是一个按钮,而不是下拉菜单,因为桌面应用程序会在页脚中静默放置一个 Select:它不会被绘制,也没有任何错误报告(在 2.1.286 应用程序上实时验证;测试套件接受它,因此套件无法捕获此问题)。当空间不足时,页脚会用 … 截断,因此标签保持简短。

路由器从不设置您手动选择的级别:这是原生工作量选择器的用途。要以特定级别运行,请关闭路由器并使用选择器;关闭路由器后,每个请求都以选择器的级别发出。

为什么问题是 Claude 自己的问题卡片

路由器只能从请求本身了解您选择器的级别。e.effort 在 turn.step 钩子中,当请求到达时,是引擎的级别。没有其他东西显示它:配置列表没有工作量行,并且会话的第一个提示在任何请求存在之前提交。因此,比较和问题发生在回合的请求即将发出时。

一个等待普通 promise 的 turn.step 钩子在大约 10 秒后被放弃,请求在没有它的情况下发出。对引擎的调用($.ui.ask,它绘制 Claude 自己的 AskUserQuestion 卡片)不计入该限制。在桌面版 2.1.286 上实时验证:请求被 удержи 23.6 秒,直到答案到来,然后以选定的级别发出。横幅中的按钮将没有任何可等待的,因此问题必须是卡片。

它是如何决定的

  • 在您的提示运行之前。 当路由器正在决定时,您发送的每个提示都会等待一次检查,然后回合开始。每个会话最多发生 decideWithin 次。如果检查时间超过 classifyTimeoutMs(15 秒)或失败,该回合以选择器的级别运行,并且 /route status 显示原因。
  • 检查在您会话的模型上运行。 您选择工作的模型会判断任务,因为它比小型模型判断得更好,并且从正确设置级别中获得的节省会随之扩大。从第二个提示开始,检查是对话的一个分支:会话自己的请求(系统提示、工具、CLAUDE.md、内存和整个对话)添加了一个问题,从会话的提示缓存中提供。在 Opus 5.5 上测量,72k 令牌对话:1.6 秒,整个对话从缓存中读取,大约 2.8k 新输入令牌和 40 个输出令牌,因此大约 3 美分。分支以会话上次使用的工作量运行。
  • 第一个提示是单独的调用。 在会话发送任何内容之前,没有要分支的请求,并且模组无法使用 Claude Code 的系统提示和工具构建一个。因此,第一次检查是对同一模型的一次调用,其中包含您的 CLAUDE.md 文件、规则和内存(如 Claude Code 将它们交给对话)以及您的提示,以模型的默认工作量。它没有缓存:大约 13k 令牌,包含大量指令,因此在 Opus 5.5 上大约 5 美分,或在 Fable 5.1 上大约 13 美分,每个会话一次(测量 1.4 秒)。firstCheckInstructions: false 仅发送提示。在回合仍在运行时发送的提示不会被检查(这种情况尚未进行实时测试);下一个提示会被检查。
  • 它有多确定。 每次检查都会为每个级别提供正确的概率,例如 medium 10%, high 50%, xhigh 40%。路由器会询问检查有多确定您所处的级别在一个方向上是错误的:这里有 90% 的可能性中等工作量太低。当它清除 confidence(默认为 0.7)时,它会移动到分布的中间,即最低级别至少与不足够的可能性一样大。这里是高。因此,在高和 xhigh 之间犹豫不决的检查仍然会将您从中等工作量移开,移到高工作量,而不是让您停留在它确定是错误的级别上。当它足够确定您的级别是正确的时,它会保持不变。低于阈值时,它不会改变任何内容,并在您的下一个提示后再次检查,届时该提示会包含更多对话。每次检查的分布、置信度和结果都保存在支出分类账中,因此可以根据自信的移动被保留或推翻的频率来设置阈值。打开 showChecks 以在每次检查后查看一行。
  • 最高到 xhigh。 路由器选择最高到 highestLevel,默认为 xhigh,并且不提供 max 检查:在所有三个模型上,max 很少能胜过 xhigh,并且可能会过度思考。将 highestLevel 设置为 max 以允许它。
  • 中间级别。 当路由器的级别与您的级别相差两个或更多时,问题也会提供中间级别:从中等工作量开始,选择 xhigh 的检查会询问 Use xhigh、Use high 或 Keep medium。
  • 另一个检查模型。 将 classifierModel 设置为另一个受支持的模型(opus、sonnet 或 fable),用于读取对话缩短副本的单独调用。Haiku 从不使用:任何其他名称都会回退到您会话的模型。
  • 在第一个请求时进行比较。 读取级别等待回合的第一个请求,在那里您的选择器级别是已知的,并且上述规则适用于那里。比较使用到达路由器时的级别,在路由器更改任何内容之前,因此它始终是您的选择器级别。
  • 您的答案也算数。 对 Claude 的多项选择问题(主线程上的 AskUserQuestion)的答案是下一次检查看到的对话的一部分。路由器在答案返回 Claude 之前再次检查,回合的下一个请求应用规则,并且已回答的问题计入预算,就像提示一样。在会话的模型上,该检查是一个包含您的答案的分支,并且回合中的分支像任何其他分支一样读取缓存。
  • 它读取什么。 分支看到的与会话模型看到的完全相同。单独的调用(第一个提示,或另一个 classifierModel)会完整读取您的提示,Claude 的问题和您的答案,Claude 的回复被截断(最后一个截断较少),以及其他工具调用仅作为名称,上限为 classifierMaxChars(24,000)。超过上限时,它会保留您的第一个提示(原始任务),然后是最新行、您的提示和答案,然后是 Claude 的回复。/route status 说明上次读取发送了多少。
  • 只有在没有任务之前才未决定。 模型仅对开场白回答“未决定”:问候语、诸如“拉取最新代码”之类的内务管理,或在任何工作之前的问题。一旦您陈述了一个真正的任务,它就会选择该任务最可能需要的级别,即使细节尚未确定。提示的示例教导阅读对话(填充、缩小范围、对编号问题的简短回复、已回答的问题),并且它们都没有命名级别。
  • 最新交流最重要。 后来的澄清会覆盖早期的询问,简短的回复会根据它回答的问题进行阅读。如果您驳回了“重构支付重试逻辑”的问题,然后用“2”回答 Claude 的“1. 完全重写还是 2. 只提取常量?”,则下一次读取会命名为低。

安装

请先查看作者 README 确认 marketplace 和插件名称;命令可能随仓库结构改变。

claude plugin marketplace add tommy5dollar/claude-mods
claude plugin install effort-router
原文 / README

effort-router

A Claude Code mod that checks your session's reasoning effort against the task once the task is clear, asks before changing it, then holds it. Each subagent gets its own level, chosen by the agent that launches it.

It follows Anthropic's guidance in Using Claude Code: Spending your effort (Thariq Shihipar, 25 September 2026). The article found that effort buys verification and edge-case testing, not a better approach. The router gives your session's model those principles and sourced notes on what each level can do on that model, then lets it judge. It never maps a kind of task to a fixed level, because level names mean different things on Opus, Sonnet and Fable.

Requires Claude Code 2.1.287 or later (Claude Mods).

The rule

Your effort picker's level is the default. After each prompt, until the task is clear, your session's own model looks at the conversation and names the level the task needs, with how sure it is. Then:

  • No clear task yet, or not sure enough: the turn runs at your picker's level, and the next prompt is checked again.
  • The router names your picker's level: the turn runs, and that level is kept for the session.
  • The router names a different level: the turn runs at the router's level, which is kept for the session. The band above the prompt opens once to say so, with why, and a button to stop routing and go back to your setting.

Either way, once a level is kept the session is decided and the router stops checking by itself. In the band, Reassess now (or /route) asks again about the conversation so far. Reassess with my next prompt (or /route next) goes back to your setting and checks again when you send your next message, which is the one to use when you're about to steer the work somewhere else.

With consent ask, a different level waits for your answer instead. Claude's own question card asks "Effort router: Multi-platform finance integration. Use high effort instead of medium?", with any levels in between as options too. Use high runs the turn at high and keeps high. Keep medium runs it at medium and keeps medium. If you dismiss the question, the turn runs at your picker's level, the router stays undecided, and a later prompt can ask again.

The states

The footer, right beside the native model and effort pickers, shows the router's state. It leaves out the level in use, because the effort picker a few pixels away already shows it.

| Footer | What it means | Band buttons (press the footer) | | --- | --- | --- | | undecided (dim) | Nothing decided yet. Your setting applies, and each prompt you send is checked. Nothing to wait for | Assess now, Stop routing (back to medium) | | deciding… | A check is running now. The turn starts when it's done, in a few seconds | | | high? | With consent ask: the question is open and the turn waits for your answer | Assess now, Stop routing (back to medium) | | using high | Decided: every main-thread request runs at high, and the router stops checking by itself | Reassess now, Reassess with my next prompt, Stop routing (back to medium) | | no decision (dim) | The router ran out of prompts without being sure. Your setting applies for the rest of the session | Start routing | | off (dim) | You stopped routing, so the picker is in charge | Start routing |

Stop routing turns the router off for this session: no more checks, subagents aren't routed and your effort setting (named in the button) applies again. /route still answers, and Start routing or /route on starts it again. New sessions are routed as usual.

Changing the effort picker yourself while the router has a level in force does the same: routing stops and your new level is used from that request on. In Desktop the picker keeps showing your setting while the router runs another level, so picking the level it already shows changes nothing there. Use Stop routing instead.

<!-- screenshot: footer showing "undecided" beside the gauge and the native pickers --> <!-- screenshot: the question card "Effort router: ... Use high effort instead of medium?" with the footer reading "high?" -->

The footer state is a plain button. Pressing it opens the router's band above the prompt: a line such as Effort router: using high for this session (bug fix in existing code). The crash needs tracing through the parser, but the fix is local., then the buttons and Hide (hotkey x). The words in brackets are what the check took the task to be, and the sentence after is why it chose that level. The buttons are numbered 1, 2. Any action closes the band, and pressing the footer again closes it too. With consent ask the band never opens by itself: the question card is where you agree.

With consent auto (the default) the router doesn't ask. When it keeps a level that differs from your picker's, the band opens by itself once, reading Effort router: changed from medium to high for this session (<task>). <why>, with OK (the change stands), Go back to xhigh when the level before wasn't your setting, and Stop routing (back to medium). Pressing the footer shows the same band as under ask.

The footer is a button, not a dropdown, because the Desktop app silently drops a Select in the footer: it is not drawn, and nothing reports an error (verified live on the 2.1.286 app; the test kit accepts it, so the kit cannot catch this). The footer truncates with … when space runs out, so the label stays short.

The router never sets a level you pick by hand: that is what the native effort picker is for. To run at a specific level, turn the router off and use the picker; with the router off, every request goes out at the picker's level.

Why the question is Claude's own question card

The router can only learn your picker's level from the request itself. e.effort in the turn.step hook, as the request arrives, is the engine's level for it. Nothing else shows it: the config list has no effort row, and a session's first prompt is submitted before any request exists. So the comparison, and the question, happen when the turn's request is about to go out.

A turn.step hook that waits on an ordinary promise is abandoned after about 10 seconds, and the request goes out without it. A call into the engine ($.ui.ask, which draws Claude's own AskUserQuestion card) doesn't count against that limit. Verified live on Desktop 2.1.286: the request was held for 23.6 seconds until the answer came, then went out at the chosen level. A button in the band would leave nothing to wait on, so the question has to be the card.

How it decides

  • Before your prompt runs. While the router is deciding, each prompt you send waits for one check before the turn starts. It happens at most decideWithin times per session. If the check takes longer than classifyTimeoutMs (15 s) or fails, the turn runs at the picker's level and /route status shows why.
  • Checks run on your session's model. The model you chose to work in judges the task, because it judges better than a small model and the savings from getting the level right scale with it. From the second prompt on, a check is a fork of the conversation: the session's own request (system prompt, tools, CLAUDE.md, memory and the whole conversation) with one question added, served from the session's prompt cache. Measured on Opus 5.5 with a 72k-token conversation: 1.6 s, the whole conversation read from cache, about 2.8k fresh input tokens and 40 output tokens, so about 3 cents. The fork runs at the effort the session last used.
  • The first prompt is a separate call. Before the session has sent anything there is no request to fork, and a mod can't build one with Claude Code's system prompt and tools. So the first check is one call to the same model with your CLAUDE.md files, rules and memory (as Claude Code hands them to the conversation) and your prompt, at the model's default effort. It isn't cached: about 13k tokens with a large set of instructions, so roughly 5 cents on Opus 5.5 or 13 cents on Fable 5.1, once per session (1.4 s measured). firstCheckInstructions: false sends only the prompt. A prompt sent while a turn is still running isn't checked (that case hasn't been tested live); the next prompt is.
  • How sure it is. Each check gives every level a probability of being the right one, for example medium 10%, high 50%, xhigh 40%. The router asks how sure the check is that the level you're on is wrong in one direction: here 90% that medium is too low. When that clears confidence (0.7 by default), it moves to the middle of the spread, the lowest level at least as likely as not to be enough. Here that's high. So a check torn between high and xhigh still moves you off medium, to high, rather than leaving you on the one level it's sure is wrong. When it's sure enough your level is right, it keeps it. Below the bar it changes nothing and checks again after your next prompt, which by then carries more of the conversation. Every check's spread, confidence and outcome is kept in the spend ledger, so the bar can be set from how often a confident move was kept or overruled. Turn on showChecks to see a line after each check.
  • Up to xhigh. The router picks up to highestLevel, xhigh by default, and the checks aren't offered max: on all three models max rarely beats xhigh and can overthink. Set highestLevel to max to allow it.
  • The levels in between. When the router's level is two or more away from yours, the question offers the levels in between too: from medium, a check that picks xhigh asks Use xhigh, Use high or Keep medium.
  • Another check model. Set classifierModel to another supported model (opus, sonnet or fable) for separate calls that read a shortened copy of the conversation. Haiku is never used: any other name falls back to your session's model.
  • Compared at the first request. The read's level waits for the turn's first request, where your picker's level is known, and the rule above applies there. The comparison uses the level as it reached the router, before the router changes anything, so it is always your picker's.
  • Your answers count too. Answers to Claude's multiple-choice questions (AskUserQuestion on the main thread) are part of the conversation the next check sees. The router checks again before the answers go back to Claude, the turn's next request applies the rule, and answered questions count toward the budget like a prompt. On the session's model that check is a fork carrying your answers, and a fork mid-turn reads the cache like any other.
  • What it reads. A fork sees exactly what the session's model sees. A separate call (the first prompt, or another classifierModel) reads your prompts in full, Claude's questions with your answers, Claude's replies truncated (the last one less so) and other tool calls as names only, capped at classifierMaxChars (24,000). Over the cap it keeps your first prompt (the original task), then the newest lines, your prompts and answers before Claude's replies. /route status says how much the last read sent.
  • Undecided only before there is a task. The model answers "undecided" only for opening filler: greetings, housekeeping such as "pull the latest code", or questions before any work. Once you state a real task it picks the level that task most likely needs, even while the details are open. The prompt's worked examples teach reading the conversation (filler, a narrowed scope, a short reply to a numbered question, answered questions), and none of them names a level.
  • The latest exchange counts most. A later clarification overrides an earlier ask, and a short reply is read against the question it answers. If you dismissed the question for "refactor the payment retry logic" and then answer Claude's "1. full rewrite or 2. just extract the constant?" with "2", the next read names low.
  • Decided is decided. Once a level is kept, every later main-thread request runs at it and the router stops reading. Subagents get their own level (below). When the kept level isn't the picker's, the terminal also runs /effort <level> once the session is idle, so the native picker label matches. In the Desktop app the picker belongs to the app, so its label stays where you set it; trust the footer. A kept level survives claude --resume. Choosing Keep medium keeps medium even if you move the picker later; /route off hands control back to the picker.
  • It stops after decideWithin prompts (6 by default), counted from the start of the session. If nothing is decided by then, the router turns off with the reason no clear task after 6 prompts, or not sure enough after 6 prompts, last check high at 65% when the checks named a level they weren't sure of, after asking any question still waiting. It never calls the model again on its own.
  • Existing sessions are left alone. The first time the router sees a session that already has decideWithin or more prompts in it, or more than skipAboveTokens (20,000) tokens of conversation (a long chat from before the router was installed, say), it starts off with the reason session started before the router: no question, no model calls. Fewer earlier prompts count toward the budget. A resumed session with saved router state keeps that state.
  • /route asks now. It reads the whole conversation in any state, ignoring the budget. If the answer is the level already in use, it says so and changes nothing. If it is your picker's level while a different one is kept, it keeps the picker's level without asking. Otherwise the question card opens straight away ("Use low effort instead of high?", naming the level in force now, with any levels in between), and your answer is kept. Add a hint to steer it: /route this is a security review, /route keep it quick. The hint is weighed strongly and kept for later reads until a level is kept. The band's Assess now (Reassess once decided) is the same as bare /route: it asks the session's model, so it takes a couple of seconds and uses your plan like any request. The footer reads checking… while it runs, then the band shows what it found ("high still fits (bug fix, 90% sure). Nothing changed.") until you hide it.
  • Consent ask. If you'd rather approve each change: a level other than your picker's waits on the question card. Set it in /plugin configure, or with EFFORT_ROUTER_CONSENT=ask. Headless runs (-p) have no one to answer, so leave them on auto.

Models

The router supports the current models: Fable 5.1, Opus 5.5 and Sonnet 5.5. Level names don't mean the same amount of thinking on each, and each responds to effort differently. In Claude Code, Opus 5.5 and Sonnet 5.5 default to medium and Fable 5.1 to high. Opus 5.5 gains most from low to medium and little above high, while Sonnet 5.5 gains a lot at every step. Routing one like another would be a mistake. Each has a notes file in rules/models/ on what each level can do there: what it's good for, what it misses, and measured gains and costs. Every check carries the notes for the session's model (or the subagent's) after the routing rules, as the main guide to the level. /route rules prints them. The evidence behind each note, with sources, is in rules/models/research-2026-10.md.

The notes guide the level instead of fixed rules because of an eval on 4 October 2026. With rules that tied kinds of task to levels, all three models gave almost the same answers and ignored their notes: Sonnet kept picking max for autonomous work, though its evidence says xhigh. Without those rules, each model's answers moved the way its evidence predicts (TESTING.md, "Prompt variants").

On any other model the router stands aside: the footer reads off, /route status says which models it works with, and no checks run. Its state is kept, so switching back with /model picks up where it was. A new model needs a new version of the router.

Subagents

Each subagent gets its own level, chosen by the agent that launches it.

  • Its parent decides. When Claude launches a subagent, the launch waits for one fork of the parent's conversation, asked which level the subagent needs, with its brief (capped at classifierMaxChars). The parent knows the task and why it's delegating this part, which a brief alone often doesn't say. Then the subagent starts, and every request it makes carries that level. Measured on Opus 5.5: 2.3 to 3.3 s, the parent's conversation read from cache, about 2 cents. With another classifierModel it's a separate call that reads the brief alone.
  • On its own model. The check is told which model the subagent runs on (the Agent call's model, else its definition's, else the parent's) and gets that model's notes. A subagent has no user in the loop, and the check always picks a level. Your rules and your organisation's rules apply here too.
  • Haiku agents are left alone. A subagent on Haiku (the built-in Explore agent runs there) or on another model the router doesn't support isn't checked. Haiku takes no effort setting anyway.
  • An agent's own effort: wins. If the agent's definition sets an effort, the router doesn't read its brief and leaves its requests alone, so the engine applies the definition's level. It looks for the definition by its frontmatter name: in the project's .claude/agents/*.md, then your ~/.claude/agents/*.md, and in the agents key of policy, project and user settings. The first definition with that name decides, as it does for the engine: a project definition without effort: still beats a user one with it. Definitions are scanned once per session. /route status shows such an agent as low: <description> (set by its agent definition).
  • Forks and failures take the parent's level. A fork shares its parent's context, so it skips the read. If a read fails, times out (classifyTimeoutMs) or returns something unusable, the subagent also takes its parent's level. That is the main thread's level in use, or for a subagent launched by another subagent, that subagent's level. With no level anywhere, its requests are left alone.
  • It runs even when the main thread is left alone. In an existing session the router leaves the main thread alone, but each new subagent brief is a fresh, whole task, so subagents are still routed. The same holds after the router turns itself off with no clear task. When you turn the router off yourself (/route off or Stop routing), subagents go back to the picker's level too, and /route on brings their routed levels back.
  • Seeing it. /route status lists this session's routed subagents, newest first (the last 10), with level, description, agent type and why. The debug log has one line per routed launch. Nothing is added to the footer or to the parent's conversation.

Claude can't set a subagent's effort itself today: the Agent tool takes a model but no effort, so without the router every subagent runs at the session's level unless its agent definition sets one. Set routeSubagents to false to go back to that (the main thread's level in use, as before 0.7.0).

Where the effort went

/route report shows what your requests spent at each level over the last 7 days. /route report session, month or all cover other spans. For example:

Effort for the last 7 days (since 2026-09-28): 412 requests in 9 sessions, 610k output tokens.
By level:
  low: 120 requests, 31k output tokens (avg 258)
  medium: 260 requests, 410k output tokens (avg 1.6k)
  high: 32 requests, 169k output tokens (avg 5.3k)
Changed by the router: 74 requests
  subagents, medium → low: 44 requests, 9.9k output tokens (avg 225, vs 1.6k for those left at medium)
  main conversation, medium → high: 30 requests, 160k output tokens (avg 5.3k, vs 1.6k for those left at medium)
The router's own checks: 61 (9 of a first prompt, 14 of a conversation, 38 for subagents), using 3.1k output and 1.20M input tokens.
By repo (output tokens): payments 400k, web 210k.
No "saved" figure: the router lowers easy tasks and raises hard ones, so these averages can't show what a changed request would have cost.
  • What it records. Every model request in every session with the router installed (0.9.0 on), on the main thread and in subagents, with the router on or off. For each one it keeps the level the request arrived at (your picker's, or the level a subagent would have inherited), the level it went out at, and its tokens as the API reported them. Requests are summed per day into one small JSON file per session, in ~/.claude/effort-router/spend/. The file is written when a turn ends, and nothing leaves your machine.
  • What it shows. Requests and output tokens per level, with the average per request. The requests the router moved, by thread and direction, each beside the average request left at the level it came from. Requests whose agent definition set their level. The router's own checks, by kind, so its cost is in the same report. Each session check's level, confidence and outcome is kept in the file too, for setting the confidence bar later. Over more than one session, output by repo.
  • Why output tokens. Output (thinking plus the answer) is what effort changes most. Input is recorded too.
  • Why there is no "saved" figure. The router lowers easy tasks and raises hard ones. A lowered request is small partly because its task was small, so comparing it with the average medium request would overstate the saving, and the same comparison would overstate what a raised request cost extra. Only running the same task at both levels can say what a request would have cost at its old level. The report gives the measured numbers side by side and leaves that estimate out.

Policy

The shipped rules (rules/default.md) are principles, not a table of levels:

  • Effort buys verification, edge-case testing and independent judgement, not a better approach.
  • Weigh how much is hidden (edge cases, existing code, money, several external systems, concurrency, security), whether you're in the loop, how well specified the task is, and how big it is.
  • Pick the level that does the work well on this model without paying for thinking it won't use.

They come from the article above and Anthropic's effort docs. What each level can do comes from the model notes.

Commands

| Command | What it does | | --- | --- | | /route | Runs the router now over the whole conversation, in any state, and asks if its level differs | | /route <hint> | The same, with a hint for the classifier (/route this is a security review) | | /route status | Shows the state and why, the consent mode, automatic reads used of the budget, classifier calls and how long the last read took, how much transcript it sent, the last verdict (with the raw reply and when), the last error, and this session's routed subagents | | /route report [session\|week\|month\|all] | Shows where the effort went: requests and output tokens per level, what the router moved, its own reads, and output by repo. The last 7 days by default | | /route off | Turns the router off and restores the picker's earlier level | | /route on | Turns the router back on: deciding over the whole conversation, with a fresh budget | | /route next | Back to deciding at your setting, with a fresh budget: your next prompt is checked before its turn starts (the band's Reassess with my next prompt) | | /route rules | Prints the effective rules and which layers contributed | | /route rules init [user\|project] | Writes a starter rules file that keeps the defaults | | /route rules critique | Asks Sonnet to critique your custom rules |

/route is registered with $.command.register, so it shows in the typeahead. State is per session. /route decide is a hidden alias of bare /route. There is no command to set a level: turn the router off and use the effort picker.

Options

Set them in /config, or under pluginConfigs["effort-router@tommy-mods"].options in settings.json.

| Option | Default | Meaning | | --- | --- | --- | | consent | auto | When the router wants a different level from your picker: auto uses the router's level without asking and shows it once in the band, with a button to stop routing. ask holds the turn and asks (Use the router's level, a level in between or Keep yours) | | decideWithin | 6 | Prompts (and answered questions) the router reads automatically, counted from the session's start | | classifyTimeoutMs | 15000 | How long a prompt waits for the check before it runs anyway | | classifierMaxChars | 24000 | Most transcript characters a separate check sends | | classifierModel | session | session: your session's own model, as a fork from the second prompt. Or another supported model (opus, sonnet, fable) for separate checks. Haiku is never used | | highestLevel | xhigh | The highest level the router picks. max allows max, which rarely beats xhigh on the current models | | confidence | 0.7 | How sure (0 to 1) a check must be that your current level is wrong in one direction (or right) before the router acts. 0 acts on any check | | showChecks | false | Print a line after each automatic check: its spread, how sure it was and what the router did | | skipAboveTokens | 20000 | A session first seen with more conversation than this keeps your effort setting | | firstCheckInstructions | true | Send your CLAUDE.md files, rules and memory with the first prompt's check. false: the prompt only | | syncPicker | true | Run /effort <level> so the terminal's picker label matches | | routeSubagents | true | Give each subagent its own level, chosen at launch by the agent starting it. false: subagents run at the main thread's level | | footerControl | button | button makes the footer state a button that opens the band. label draws plain text, and /route is the control | | rules | empty | Rules text for your user layer. A rules file takes precedence |

The environment variable EFFORT_ROUTER_CONSENT=ask|auto overrides consent, which helps in headless runs (-p has no one to answer a question, so use auto there). The older names still work there: apply and none mean auto, while confirm and band mean ask.

Customising the rules

The rules are plain markdown, layered from the bottom up:

  1. The shipped defaults (rules/default.md)
  2. Your organisation's rules from managed (policy) settings, if it sets any
  3. Your rules: ~/.claude/effort-router.md, or the rules option in your user settings
  4. The project's rules: <project root>/.claude/effort-router.md (commit it), or the rules option in project settings

A line that is exactly $defaults pulls in everything beneath that layer. Text after the line adds to the rules, and later rules win. Text before it goes first. A file with no $defaults line replaces everything beneath it. Missing or empty files change nothing, and HTML comments are ignored.

$defaults

- This is a payments codebase. Never pick below high: money movement needs verification.

The files are re-read on every classification, so edits apply without a reload. An unreadable file is skipped. The frame around the rules (undecided only before a task, the worked examples, reply in JSON) is fixed, so no rules file can break the parser.

For organisations

An organisation can set routing rules centrally in managed settings (managed-settings.json), as it does for other Claude Code policy:

{
  "pluginConfigs": {
    "effort-router@tommy-mods": {
      "options": {
        "rules": "$defaults\n\n- Code under payments/ or ledger/ is never routed below high.\n- Infrastructure changes (terraform/, k8s/) are high.",
        "rulesMode": "enforce",
        "allowOff": false
      }
    }
  }
}
  • rulesMode: "extend" (the default) layers the org rules over the shipped defaults. Users and projects can add to them with $defaults, or replace them.
  • rulesMode: "enforce" makes the org layer final. Personal and project rules are ignored, and /route rules init says so.
  • routeSubagents: false turns subagent routing off for everyone, whatever their own setting.
  • allowOff: false stops users turning the router off, so the organisation's routing always applies. /route off refuses, the band has no Stop routing, and a session saved as off comes back deciding. When the budget runs out with nothing suggested, the router idles as deciding (no more reads) instead of turning off. /route still works.

A top-level "effortRouter": { "rules": ..., "rulesMode": ..., "allowOff": ..., "routeSubagents": ... } object works too. The router reads these four settings only from the policy source, so a user cannot claim enforce for themselves.

Install

claude plugin marketplace add tommy5dollar/claude-mods
claude plugin install effort-router@tommy-mods

There's nothing to configure: every option has a default. /plugin configure lists the options as not yet set, which just means the defaults apply. Change one only when you want something different (see Options).

For development, run claude --plugin-dir ./effort-router.

Known limits

  • In the Desktop app the native effort picker never changes: the app owns it and nothing a mod can call sets it. The requests still go out at the routed level; trust the footer label.
  • In the terminal the router can't set the picker label directly either. It runs /effort <level> when the session is idle, which prints a line in the transcript, and the footer shows the true level until then. Headless (-p) runs skip the sync because its output would replace the run's printed result. The per-request override still applies there.
  • Each prompt waits for the check while the router is deciding (at most decideWithin prompts per session): about 1.5 s on Opus 5.5, longer on a model that thinks more by default. If the check times out (classifyTimeoutMs), that turn goes at the picker's level; a late answer is ignored.
  • The first prompt's check can't share the session's prompt cache: the engine offers no way to fork before the first response, and a separate call can't carry Claude Code's system prompt or tools. It pays for your instructions and the prompt once per session. Claude Code also never reuses the cache for the conversation part of a session's first request, so the first fork after a one-request first turn pays for about 13k tokens again.
  • On Bedrock, Google Cloud or an LLM gateway, Claude Code clears the cached conversation when effort changes (Anthropic's docs). A locked level that differs from your setting changes effort once, so expect one uncached request there. With an API key or a subscription the cache is kept.
  • The confidence bar (0.7) is a starting guess, not calibrated yet. The ledger keeps every check's confidence and outcome for that.
  • A fork answers at the effort the session last used, and the router can't change it.
  • A request whose picker level is a number rather than a named level, or a model that takes no effort, can't be compared, so the waiting verdict waits for the next request that can.
  • The router changes effort only, never the model. A request to a model that takes no effort is left alone.
  • Each automatic check is one call per prompt while deciding, for at most decideWithin prompts. Nothing more is spent once a level is locked or the budget is spent, except when you run /route. /route report shows what the checks cost.
  • Once decided, the router does not notice a change of phase on its own (for example "now verify it" after an implementation). Run /route (or the band's Reassess now), optionally with a hint, or /route next to be asked with your next prompt.
  • A definition's effort: is respected for user and project agent files and for the agents key in settings, not for plugin agents (<plugin>:<name>), which can't be located reliably. Those are routed from their brief, which replaces any effort their definition sets. A mod's $.agent.register({ effort }) is ignored by the engine itself (an engine bug), and the router routes those agents too.
  • An agent file added or edited mid-session is seen from the next session: definitions are scanned once per session.
  • Workflow agents that don't launch through the Agent tool raise no agent.spawn, so they keep the main thread's level.
  • Subagent levels are kept in memory only. After a restart or resume, a subagent still running from before takes the main thread's level.
  • The router adds no note about the chosen level to the system prompt, because changing a cached prompt section would break the prompt cache. The /effort echo tells the model instead, and it is appended to the transcript, so the cache holds.
  • The spend report starts at 0.9.0: sessions from before it aren't in it. A request with no reported usage (failed or interrupted) isn't counted. Days are UTC.
  • The band and footer draw in the terminal and the Desktop app. VS Code and -p run the hooks without the UI. Under ask in -p the question has no one to answer, so requests stay at the picker's level; set consent (or EFFORT_ROUTER_CONSENT) to auto there.

Development

bun test                            # pure policy: trimming, parsing, rule layering, /route grammar, subagent reads, the spend ledger and report
claude plugin test .                # engine kit: band, buttons, footer button, turn.step, /route, org layers, agent.spawn, the ledger saved and reported
claude plugin validate . --strict
bun run eval                        # opt-in: the real classifier over eval/fixtures.ts (see below)

bun run eval sends each fixture to the real model through claude -p --safe-mode (no plugins or hooks, no tools), with the system prompt and input the router builds from rules/default.md: for the session read, a first check on the session's model (--model, default opus) at its default effort with that model's notes from rules/models/, and for a subagent's read, a separate call on the same model with that model's notes and the brief. It can't reproduce the forks that later checks and subagent reads make. It parses the reply with the router's own parser and prints each verdict, the pass rate and every miss. --runs 3 repeats each fixture (the model is not deterministic), --model sonnet runs the session set as a Sonnet session, --effort sets the effort and --confidence the bar, --set session or --set subagent runs one set and --only <text> filters fixtures by name. It uses your Claude Code login, and each fixture costs one small model call.

TESTING.md lists the live checks.

Licence

MIT

其他同名作品

更多类似作品