tommy5dollar/claude-mods/tree/main/effort-router
effort-router
각 작업에 필요한 추론 노력을 선택합니다. 세션의 자체 모델이 작업을 판단하고, 해당 모델에서 각 레벨이 무엇을 할 수 있는지 알며, 라우터는 확신할 때만 작동합니다. 노력 설정과 일치하지 않으면 작업 실행 전에 묻습니다. 하위 에이전트는 시작한 에이전트가 선택한 자체 레벨을 얻습니다. /route report는 노력이 어디에 사용되었는지 보여줍니다.
이 mod 소개
effort-router
작업이 명확해지면 세션의 추론 노력을 작업과 비교하고, 변경하기 전에 묻고, 그 설정을 유지하는 Claude Code 모드입니다. 각 하위 에이전트는 이를 시작한 에이전트가 선택한 자체 레벨을 얻습니다.
이는 Anthropic의 Claude Code 사용: 노력 투입하기 (Thariq Shihipar, 2026년 9월 25일)] 지침을 따릅니다. 이 기사는 노력이 검증과 엣지 케이스 테스트를 가져오지만 더 나은 접근 방식은 아니라는 것을 발견했습니다. 라우터는 이러한 원칙과 각 레벨이 해당 모델에서 무엇을 할 수 있는지에 대한 출처 정보를 세션의 모델에 제공한 다음 판단하게 합니다. 레벨 이름이 Opus, Sonnet, Fable에서 다른 의미를 가지므로 특정 유형의 작업을 고정 레벨에 매핑하지 않습니다.
Claude Code 2.1.287 이상 (Claude Mods)이 필요합니다.
규칙
노력 선택기의 레벨이 기본값입니다. 각 프롬프트 후, 작업이 명확해질 때까지 세션의 자체 모델이 대화를 보고 작업에 필요한 레벨과 그 확신도를 명명합니다. 그런 다음:
- 아직 명확한 작업이 없거나 충분히 확신하지 못하는 경우: 해당 턴은 선택기의 레벨로 실행되고 다음 프롬프트가 다시 확인됩니다.
- 라우터가 선택기의 레벨을 명명하는 경우: 해당 턴이 실행되고 해당 레벨이 세션에 유지됩니다.
- 라우터가 다른 레벨을 명명하는 경우: 해당 턴은 라우터의 레벨로 실행되고 해당 레벨이 세션에 유지됩니다. 프롬프트 위의 밴드가 한 번 열려 이유를 설명하고, 라우팅을 중지하고 설정으로 돌아가는 버튼이 표시됩니다.
어떤 경우든, 레벨이 유지되면 세션이 결정되고 라우터는 스스로 확인을 중지합니다. 밴드에서 Reassess now (또는 /route)는 지금까지의 대화에 대해 다시 묻습니다. Reassess with my next prompt (또는 /route next)는 설정으로 돌아가 다음 메시지를 보낼 때 다시 확인합니다. 이는 작업을 다른 곳으로 유도하려는 경우에 사용합니다.
ask 동의가 있으면 다른 레벨이 대신 답변을 기다립니다. Claude 자신의 질문 카드는 "노력 라우터: 다중 플랫폼 금융 통합. 중간 노력 대신 높은 노력을 사용하시겠습니까?"라고 묻고, 그 사이의 모든 레벨도 옵션으로 제공합니다. Use high은 높은 노력으로 턴을 실행하고 높은 노력을 유지합니다. Keep medium은 중간 노력으로 실행하고 중간 노력을 유지합니다. 질문을 닫으면 턴은 선택기의 레벨로 실행되고 라우터는 미결정 상태로 유지되며, 나중에 프롬프트에서 다시 물을 수 있습니다.
상태
푸터는 기본 모델 및 노력 선택기 바로 옆에 라우터의 상태를 표시합니다. 사용 중인 레벨은 몇 픽셀 떨어진 노력 선택기가 이미 표시하고 있으므로 생략됩니다.
| 푸터 | 의미 | 밴드 버튼 (푸터 누르기) |
| --- | --- | --- |
| undecided (흐리게) | 아직 결정된 것이 없습니다. 설정이 적용되고 보내는 각 프롬프트가 확인됩니다. 기다릴 필요가 없습니다 | Assess now、Stop routing (back to medium) |
| deciding… | 현재 확인 중입니다. 몇 초 후에 완료되면 턴이 시작됩니다 | |
| high? | ask 동의: 질문이 열려 있고 턴은 답변을 기다립니다 | Assess now、Stop routing (back to medium) |
| using high | 결정됨: 모든 메인 스레드 요청은 높은 노력으로 실행되고 라우터는 스스로 확인을 중지합니다 | Reassess now、Reassess with my next prompt、Stop routing (back to medium) |
| no decision (흐리게) | 라우터가 확신하지 못한 채 프롬프트를 모두 사용했습니다. 설정은 세션의 나머지 부분에 적용됩니다 | Start routing |
| off (흐리게) | 라우팅을 중지했으므로 선택기가 담당합니다 | Start routing |
Stop routing는 이 세션의 라우터를 끕니다. 더 이상 확인하지 않고, 하위 에이전트는 라우팅되지 않으며, 노력 설정(버튼에 명시됨)이 다시 적용됩니다. /route는 여전히 답변하고, Start routing 또는 /route on는 다시 시작합니다. 새 세션은 평소와 같이 라우팅됩니다.
라우터가 레벨을 강제하고 있는 동안 노력 선택기를 직접 변경하면 동일한 작업이 수행됩니다. 라우팅이 중지되고 해당 요청부터 새 레벨이 사용됩니다. 데스크톱에서는 라우터가 다른 레벨을 실행하는 동안 선택기가 설정을 계속 표시하므로, 이미 표시된 레벨을 선택해도 아무것도 변경되지 않습니다. 대신 Stop routing을 사용하십시오.
푸터 상태는 일반 버튼입니다. 푸터를 누르면 프롬프트 위에 라우터 밴드가 열립니다. Effort router: using high for this session (bug fix in existing code). The crash needs tracing through the parser, but the fix is local.와 같은 줄, 그리고 버튼과 Hide (단축키 x)이 표시됩니다. 괄호 안의 단어는 확인이 작업으로 간주한 것이고, 그 뒤의 문장은 해당 레벨을 선택한 이유입니다. 버튼은 1、2으로 번호가 매겨져 있습니다. 어떤 작업이든 밴드를 닫고, 푸터를 다시 눌러도 닫힙니다. ask 동의가 있으면 밴드는 스스로 열리지 않습니다. 질문 카드에서 동의합니다.
auto (기본값) 동의가 있으면 라우터는 묻지 않습니다. 선택기의 레벨과 다른 레벨을 유지할 때 밴드는 한 번 스스로 열리며, Effort router: changed from medium to high for this session (<task>). <why>이 표시되고, OK (변경 사항 유지), Go back to xhigh (이전 레벨이 설정이 아니었던 경우), 그리고 Stop routing (back to medium)이 표시됩니다. 푸터를 누르면 ask 아래와 동일한 밴드가 표시됩니다.
푸터는 드롭다운이 아닌 버튼입니다. 데스크톱 앱이 푸터에 Select을 자동으로 배치하기 때문입니다. 이는 그려지지 않으며 오류도 보고되지 않습니다(2.1.286 앱에서 실시간 확인; 테스트 키트는 이를 허용하므로 키트에서 이를 감지할 수 없습니다). 공간이 부족하면 푸터는 …로 잘리므로 레이블은 짧게 유지됩니다.
라우터는 수동으로 선택한 레벨을 설정하지 않습니다. 이는 기본 노력 선택기의 역할입니다. 특정 레벨에서 실행하려면 라우터를 끄고 선택기를 사용하십시오. 라우터가 꺼져 있으면 모든 요청은 선택기의 레벨로 전송됩니다.
질문이 Claude 자신의 질문 카드인 이유
라우터는 요청 자체에서만 선택기의 레벨을 알 수 있습니다. 요청이 도착할 때 turn.step 후크의 e.effort은 엔진의 레벨입니다. 다른 어떤 것도 이를 표시하지 않습니다. 구성 목록에는 노력 행이 없으며, 세션의 첫 번째 프롬프트는 어떤 요청도 존재하기 전에 제출됩니다. 따라서 비교와 질문은 턴의 요청이 나갈 때 발생합니다.
일반적인 프로미스를 기다리는 turn.step 후크는 약 10초 후에 중단되고, 요청은 그것 없이 나갑니다. 엔진 호출($.ui.ask, Claude 자신의 AskUserQuestion 카드를 그리는)은 해당 제한에 포함되지 않습니다. 데스크톱 2.1.286에서 실시간 확인: 요청은 답변이 올 때까지 23.6초 동안 보류되었고, 그 다음 선택된 레벨로 나갔습니다. 밴드의 버튼은 기다릴 것이 없으므로 질문은 카드여야 합니다.
결정 방법
- 프롬프트가 실행되기 전. 라우터가 결정하는 동안, 보내는 각 프롬프트는 턴이 시작되기 전에 한 번의 확인을 기다립니다. 이는 세션당 최대
decideWithin번 발생합니다. 확인에classifyTimeoutMs(15초)보다 오래 걸리거나 실패하면, 턴은 선택기의 레벨로 실행되고/route status가 이유를 표시합니다. - 확인은 세션의 모델에서 실행됩니다. 작업할 모델이 작업을 판단합니다. 이는 작은 모델보다 판단력이 뛰어나고, 레벨을 올바르게 설정함으로써 얻는 절감 효과가 그에 비례하여 커지기 때문입니다. 두 번째 프롬프트부터 확인은 대화의 포크입니다. 세션 자체의 요청(시스템 프롬프트, 도구, CLAUDE.md, 메모리 및 전체 대화)에 질문 하나가 추가되어 세션의 프롬프트 캐시에서 제공됩니다. Opus 5.5에서 72k 토큰 대화로 측정: 1.6초, 전체 대화가 캐시에서 읽히고 약 2.8k의 새로운 입력 토큰과 40개의 출력 토큰으로 약 3센트입니다. 포크는 세션이 마지막으로 사용한 노력으로 실행됩니다.
- 첫 번째 프롬프트는 별도의 호출입니다. 세션이 아무것도 보내기 전에 포크할 요청이 없으며, 모드는 Claude Code의 시스템 프롬프트와 도구로 이를 구축할 수 없습니다. 따라서 첫 번째 확인은 CLAUDE.md 파일, 규칙 및 메모리(Claude Code가 대화에 전달하는 대로)와 프롬프트를 포함하는 동일한 모델에 대한 한 번의 호출이며, 모델의 기본 노력으로 실행됩니다. 이는 캐시되지 않습니다. 많은 지침이 포함된 약 13k 토큰으로, Opus 5.5에서 약 5센트 또는 Fable 5.1에서 약 13센트이며, 세션당 한 번(측정값 1.4초)입니다.
firstCheckInstructions: false은 프롬프트만 보냅니다. 턴이 아직 실행 중일 때 보낸 프롬프트는 확인되지 않습니다(이 경우는 실시간 테스트되지 않았습니다). 다음 프롬프트가 확인됩니다. - 확신도. 각 확인은 각 레벨이 올바른 확률을 제공합니다. 예를 들어
medium 10%, high 50%, xhigh 40%입니다. 라우터는 확인이 현재 레벨이 한 방향으로 잘못되었다고 얼마나 확신하는지 묻습니다. 여기서는 중간이 너무 낮을 확률이 90%입니다. 그것이confidence(기본값 0.7)를 통과하면, 확산의 중간, 즉 충분할 가능성이 적어도 그만큼 높은 가장 낮은 레벨로 이동합니다. 여기서는 높음입니다. 따라서 높음과 xhigh 사이에서 갈등하는 확인도 중간에서 높음으로 이동시키며, 잘못되었다고 확신하는 레벨에 머무르게 하지 않습니다. 현재 레벨이 올바르다고 충분히 확신하면 이를 유지합니다. 기준치 미만에서는 아무것도 변경하지 않고 다음 프롬프트 후에 다시 확인합니다. 그러면 그 프롬프트에는 더 많은 대화가 포함됩니다. 각 확인의 확산, 신뢰도 및 결과는 지출 원장에 보관되므로, 확신 있는 이동이 유지되거나 기각된 빈도에 따라 기준치를 설정할 수 있습니다.showChecks을 켜면 각 확인 후에 한 줄이 표시됩니다. - xhigh까지. 라우터는
highestLevel까지, 기본적으로 xhigh를 선택하며, max는 확인에 제공되지 않습니다. 세 가지 모델 모두에서 max가 xhigh를 능가하는 경우는 드물고 과도하게 생각할 수 있습니다.highestLevel를max으로 설정하여 허용합니다. - 중간 레벨. 라우터의 레벨이 사용자 레벨에서 두 개 이상 떨어져 있는 경우, 질문은 중간 레벨도 제공합니다. 중간에서 xhigh를 선택하는 확인은
Use xhigh、Use high또는Keep medium를 묻습니다. - 다른 확인 모델.
classifierModel을 다른 지원되는 모델(opus、sonnet또는fable)로 설정하면 대화의 단축된 사본을 읽는 별도의 호출이 이루어집니다. Haiku는 사용되지 않습니다. 다른 이름은 세션의 모델로 폴백됩니다. - 첫 번째 요청에서 비교. 읽기 레벨은 턴의 첫 번째 요청을 기다립니다. 거기서 선택기의 레벨을 알 수 있으며, 위의 규칙이 적용됩니다. 비교는 라우터가 아무것도 변경하기 전, 라우터에 도달했을 때의 레벨을 사용하므로 항상 선택기의 레벨입니다.
- 귀하의 답변도 중요합니다. Claude의 객관식 질문(메인 스레드의 AskUserQuestion)에 대한 답변은 다음 확인이 볼 대화의 일부입니다. 라우터는 답변이 Claude로 돌아가기 전에 다시 확인하고, 턴의 다음 요청이 규칙을 적용하며, 답변된 질문은 프롬프트와 마찬가지로 예산에 포함됩니다. 세션의 모델에서 해당 확인은 귀하의 답변을 포함하는 포크이며, 턴 중간의 포크는 다른 포크와 마찬가지로 캐시를 읽습니다.
- 읽는 내용. 포크는 세션의 모델이 보는 것과 정확히 동일한 것을 봅니다. 별도의 호출(첫 번째 프롬프트 또는 다른
classifierModel)은 귀하의 프롬프트 전체, Claude의 질문과 귀하의 답변, Claude의 답변(마지막 것은 덜 잘림)을 잘라내어 읽고, 다른 도구 호출은 이름만으로,classifierMaxChars(24,000)로 제한됩니다. 제한을 초과하면 첫 번째 프롬프트(원래 작업), 그 다음 최신 줄, Claude의 답변 전의 귀하의 프롬프트와 답변이 유지됩니다./route status은 마지막 읽기에서 보낸 양을 나타냅니다. - 작업이 있기 전까지만 미결정. 모델은 인사말, "최신 코드 가져오기"와 같은 하우스키핑, 또는 작업 전 질문과 같은 시작 필러에 대해서만 "미결정"이라고 답변합니다. 실제 작업을 명시하면 세부 사항이 미정인 상태에서도 해당 작업에 가장 필요한 레벨을 선택합니다. 프롬프트의 예시는 대화(필러, 범위 좁히기, 번호가 매겨진 질문에 대한 짧은 답변, 답변된 질문)를 읽는 방법을 가르치며, 그 중 어느 것도 레벨을 명명하지 않습니다.
- 최신 교류가 가장 중요합니다. 나중의 명확화는 이전 요청을 무효화하며, 짧은 답변은 그것이 답변하는 질문에 대해 읽힙니다. "결제 재시도 로직 리팩토링" 질문을 기각하고 Claude의 "1. 전체 재작성 또는 2. 상수만 추출?" 질문에 "2"라고 답변하면 다음 읽기에서는 낮음으로 명명됩니다.
설치
먼저 작성자의 README에서 marketplace와 플러그인 이름을 확인하세요. 저장소 구조에 따라 명령어가 달라질 수 있습니다.
claude plugin marketplace add tommy5dollar/claude-mods claude plugin install effort-router
원문 / README
effort-router
A Claude Code mod that checks your session's reasoning effort against the task once the task is clear, asks before changing it, then holds it. Each subagent gets its own level, chosen by the agent that launches it.
It follows Anthropic's guidance in Using Claude Code: Spending your effort (Thariq Shihipar, 25 September 2026). The article found that effort buys verification and edge-case testing, not a better approach. The router gives your session's model those principles and sourced notes on what each level can do on that model, then lets it judge. It never maps a kind of task to a fixed level, because level names mean different things on Opus, Sonnet and Fable.
Requires Claude Code 2.1.287 or later (Claude Mods).
The rule
Your effort picker's level is the default. After each prompt, until the task is clear, your session's own model looks at the conversation and names the level the task needs, with how sure it is. Then:
- No clear task yet, or not sure enough: the turn runs at your picker's level, and the next prompt is checked again.
- The router names your picker's level: the turn runs, and that level is kept for the session.
- The router names a different level: the turn runs at the router's level, which is kept for the session. The band above the prompt opens once to say so, with why, and a button to stop routing and go back to your setting.
Either way, once a level is kept the session is decided and the router stops checking by itself. In the band, Reassess now (or /route) asks again about the conversation so far. Reassess with my next prompt (or /route next) goes back to your setting and checks again when you send your next message, which is the one to use when you're about to steer the work somewhere else.
With consent ask, a different level waits for your answer instead. Claude's own question card asks "Effort router: Multi-platform finance integration. Use high effort instead of medium?", with any levels in between as options too. Use high runs the turn at high and keeps high. Keep medium runs it at medium and keeps medium. If you dismiss the question, the turn runs at your picker's level, the router stays undecided, and a later prompt can ask again.
The states
The footer, right beside the native model and effort pickers, shows the router's state. It leaves out the level in use, because the effort picker a few pixels away already shows it.
| Footer | What it means | Band buttons (press the footer) |
| --- | --- | --- |
| undecided (dim) | Nothing decided yet. Your setting applies, and each prompt you send is checked. Nothing to wait for | Assess now, Stop routing (back to medium) |
| deciding… | A check is running now. The turn starts when it's done, in a few seconds | |
| high? | With consent ask: the question is open and the turn waits for your answer | Assess now, Stop routing (back to medium) |
| using high | Decided: every main-thread request runs at high, and the router stops checking by itself | Reassess now, Reassess with my next prompt, Stop routing (back to medium) |
| no decision (dim) | The router ran out of prompts without being sure. Your setting applies for the rest of the session | Start routing |
| off (dim) | You stopped routing, so the picker is in charge | Start routing |
Stop routing turns the router off for this session: no more checks, subagents aren't routed and your effort setting (named in the button) applies again. /route still answers, and Start routing or /route on starts it again. New sessions are routed as usual.
Changing the effort picker yourself while the router has a level in force does the same: routing stops and your new level is used from that request on. In Desktop the picker keeps showing your setting while the router runs another level, so picking the level it already shows changes nothing there. Use Stop routing instead.
The footer state is a plain button. Pressing it opens the router's band above the prompt: a line such as Effort router: using high for this session (bug fix in existing code). The crash needs tracing through the parser, but the fix is local., then the buttons and Hide (hotkey x). The words in brackets are what the check took the task to be, and the sentence after is why it chose that level. The buttons are numbered 1, 2. Any action closes the band, and pressing the footer again closes it too. With consent ask the band never opens by itself: the question card is where you agree.
With consent auto (the default) the router doesn't ask. When it keeps a level that differs from your picker's, the band opens by itself once, reading Effort router: changed from medium to high for this session (<task>). <why>, with OK (the change stands), Go back to xhigh when the level before wasn't your setting, and Stop routing (back to medium). Pressing the footer shows the same band as under ask.
The footer is a button, not a dropdown, because the Desktop app silently drops a Select in the footer: it is not drawn, and nothing reports an error (verified live on the 2.1.286 app; the test kit accepts it, so the kit cannot catch this). The footer truncates with … when space runs out, so the label stays short.
The router never sets a level you pick by hand: that is what the native effort picker is for. To run at a specific level, turn the router off and use the picker; with the router off, every request goes out at the picker's level.
Why the question is Claude's own question card
The router can only learn your picker's level from the request itself. e.effort in the turn.step hook, as the request arrives, is the engine's level for it. Nothing else shows it: the config list has no effort row, and a session's first prompt is submitted before any request exists. So the comparison, and the question, happen when the turn's request is about to go out.
A turn.step hook that waits on an ordinary promise is abandoned after about 10 seconds, and the request goes out without it. A call into the engine ($.ui.ask, which draws Claude's own AskUserQuestion card) doesn't count against that limit. Verified live on Desktop 2.1.286: the request was held for 23.6 seconds until the answer came, then went out at the chosen level. A button in the band would leave nothing to wait on, so the question has to be the card.
How it decides
- Before your prompt runs. While the router is deciding, each prompt you send waits for one check before the turn starts. It happens at most
decideWithintimes per session. If the check takes longer thanclassifyTimeoutMs(15 s) or fails, the turn runs at the picker's level and/route statusshows why. - Checks run on your session's model. The model you chose to work in judges the task, because it judges better than a small model and the savings from getting the level right scale with it. From the second prompt on, a check is a fork of the conversation: the session's own request (system prompt, tools, CLAUDE.md, memory and the whole conversation) with one question added, served from the session's prompt cache. Measured on Opus 5.5 with a 72k-token conversation: 1.6 s, the whole conversation read from cache, about 2.8k fresh input tokens and 40 output tokens, so about 3 cents. The fork runs at the effort the session last used.
- The first prompt is a separate call. Before the session has sent anything there is no request to fork, and a mod can't build one with Claude Code's system prompt and tools. So the first check is one call to the same model with your CLAUDE.md files, rules and memory (as Claude Code hands them to the conversation) and your prompt, at the model's default effort. It isn't cached: about 13k tokens with a large set of instructions, so roughly 5 cents on Opus 5.5 or 13 cents on Fable 5.1, once per session (1.4 s measured).
firstCheckInstructions: falsesends only the prompt. A prompt sent while a turn is still running isn't checked (that case hasn't been tested live); the next prompt is. - How sure it is. Each check gives every level a probability of being the right one, for example
medium 10%, high 50%, xhigh 40%. The router asks how sure the check is that the level you're on is wrong in one direction: here 90% that medium is too low. When that clearsconfidence(0.7 by default), it moves to the middle of the spread, the lowest level at least as likely as not to be enough. Here that's high. So a check torn between high and xhigh still moves you off medium, to high, rather than leaving you on the one level it's sure is wrong. When it's sure enough your level is right, it keeps it. Below the bar it changes nothing and checks again after your next prompt, which by then carries more of the conversation. Every check's spread, confidence and outcome is kept in the spend ledger, so the bar can be set from how often a confident move was kept or overruled. Turn onshowChecksto see a line after each check. - Up to xhigh. The router picks up to
highestLevel, xhigh by default, and the checks aren't offered max: on all three models max rarely beats xhigh and can overthink. SethighestLeveltomaxto allow it. - The levels in between. When the router's level is two or more away from yours, the question offers the levels in between too: from medium, a check that picks xhigh asks
Use xhigh,Use highorKeep medium. - Another check model. Set
classifierModelto another supported model (opus,sonnetorfable) for separate calls that read a shortened copy of the conversation. Haiku is never used: any other name falls back to your session's model. - Compared at the first request. The read's level waits for the turn's first request, where your picker's level is known, and the rule above applies there. The comparison uses the level as it reached the router, before the router changes anything, so it is always your picker's.
- Your answers count too. Answers to Claude's multiple-choice questions (AskUserQuestion on the main thread) are part of the conversation the next check sees. The router checks again before the answers go back to Claude, the turn's next request applies the rule, and answered questions count toward the budget like a prompt. On the session's model that check is a fork carrying your answers, and a fork mid-turn reads the cache like any other.
- What it reads. A fork sees exactly what the session's model sees. A separate call (the first prompt, or another
classifierModel) reads your prompts in full, Claude's questions with your answers, Claude's replies truncated (the last one less so) and other tool calls as names only, capped atclassifierMaxChars(24,000). Over the cap it keeps your first prompt (the original task), then the newest lines, your prompts and answers before Claude's replies./route statussays how much the last read sent. - Undecided only before there is a task. The model answers "undecided" only for opening filler: greetings, housekeeping such as "pull the latest code", or questions before any work. Once you state a real task it picks the level that task most likely needs, even while the details are open. The prompt's worked examples teach reading the conversation (filler, a narrowed scope, a short reply to a numbered question, answered questions), and none of them names a level.
- The latest exchange counts most. A later clarification overrides an earlier ask, and a short reply is read against the question it answers. If you dismissed the question for "refactor the payment retry logic" and then answer Claude's "1. full rewrite or 2. just extract the constant?" with "2", the next read names low.
- Decided is decided. Once a level is kept, every later main-thread request runs at it and the router stops reading. Subagents get their own level (below). When the kept level isn't the picker's, the terminal also runs
/effort <level>once the session is idle, so the native picker label matches. In the Desktop app the picker belongs to the app, so its label stays where you set it; trust the footer. A kept level survivesclaude --resume. ChoosingKeep mediumkeeps medium even if you move the picker later;/route offhands control back to the picker. - It stops after
decideWithinprompts (6 by default), counted from the start of the session. If nothing is decided by then, the router turns off with the reasonno clear task after 6 prompts, ornot sure enough after 6 prompts, last check high at 65%when the checks named a level they weren't sure of, after asking any question still waiting. It never calls the model again on its own. - Existing sessions are left alone. The first time the router sees a session that already has
decideWithinor more prompts in it, or more thanskipAboveTokens(20,000) tokens of conversation (a long chat from before the router was installed, say), it startsoffwith the reasonsession started before the router: no question, no model calls. Fewer earlier prompts count toward the budget. A resumed session with saved router state keeps that state. /routeasks now. It reads the whole conversation in any state, ignoring the budget. If the answer is the level already in use, it says so and changes nothing. If it is your picker's level while a different one is kept, it keeps the picker's level without asking. Otherwise the question card opens straight away ("Use low effort instead of high?", naming the level in force now, with any levels in between), and your answer is kept. Add a hint to steer it:/route this is a security review,/route keep it quick. The hint is weighed strongly and kept for later reads until a level is kept. The band'sAssess now(Reassessonce decided) is the same as bare/route: it asks the session's model, so it takes a couple of seconds and uses your plan like any request. The footer readschecking…while it runs, then the band shows what it found ("high still fits (bug fix, 90% sure). Nothing changed.") until you hide it.- Consent
ask. If you'd rather approve each change: a level other than your picker's waits on the question card. Set it in/plugin configure, or withEFFORT_ROUTER_CONSENT=ask. Headless runs (-p) have no one to answer, so leave them onauto.
Models
The router supports the current models: Fable 5.1, Opus 5.5 and Sonnet 5.5. Level names don't mean the same amount of thinking on each, and each responds to effort differently. In Claude Code, Opus 5.5 and Sonnet 5.5 default to medium and Fable 5.1 to high. Opus 5.5 gains most from low to medium and little above high, while Sonnet 5.5 gains a lot at every step. Routing one like another would be a mistake. Each has a notes file in rules/models/ on what each level can do there: what it's good for, what it misses, and measured gains and costs. Every check carries the notes for the session's model (or the subagent's) after the routing rules, as the main guide to the level. /route rules prints them. The evidence behind each note, with sources, is in rules/models/research-2026-10.md.
The notes guide the level instead of fixed rules because of an eval on 4 October 2026. With rules that tied kinds of task to levels, all three models gave almost the same answers and ignored their notes: Sonnet kept picking max for autonomous work, though its evidence says xhigh. Without those rules, each model's answers moved the way its evidence predicts (TESTING.md, "Prompt variants").
On any other model the router stands aside: the footer reads off, /route status says which models it works with, and no checks run. Its state is kept, so switching back with /model picks up where it was. A new model needs a new version of the router.
Subagents
Each subagent gets its own level, chosen by the agent that launches it.
- Its parent decides. When Claude launches a subagent, the launch waits for one fork of the parent's conversation, asked which level the subagent needs, with its brief (capped at
classifierMaxChars). The parent knows the task and why it's delegating this part, which a brief alone often doesn't say. Then the subagent starts, and every request it makes carries that level. Measured on Opus 5.5: 2.3 to 3.3 s, the parent's conversation read from cache, about 2 cents. With anotherclassifierModelit's a separate call that reads the brief alone. - On its own model. The check is told which model the subagent runs on (the Agent call's model, else its definition's, else the parent's) and gets that model's notes. A subagent has no user in the loop, and the check always picks a level. Your rules and your organisation's rules apply here too.
- Haiku agents are left alone. A subagent on Haiku (the built-in Explore agent runs there) or on another model the router doesn't support isn't checked. Haiku takes no effort setting anyway.
- An agent's own
effort:wins. If the agent's definition sets an effort, the router doesn't read its brief and leaves its requests alone, so the engine applies the definition's level. It looks for the definition by its frontmattername:in the project's.claude/agents/*.md, then your~/.claude/agents/*.md, and in theagentskey of policy, project and user settings. The first definition with that name decides, as it does for the engine: a project definition withouteffort:still beats a user one with it. Definitions are scanned once per session./route statusshows such an agent aslow: <description> (set by its agent definition). - Forks and failures take the parent's level. A fork shares its parent's context, so it skips the read. If a read fails, times out (
classifyTimeoutMs) or returns something unusable, the subagent also takes its parent's level. That is the main thread's level in use, or for a subagent launched by another subagent, that subagent's level. With no level anywhere, its requests are left alone. - It runs even when the main thread is left alone. In an existing session the router leaves the main thread alone, but each new subagent brief is a fresh, whole task, so subagents are still routed. The same holds after the router turns itself off with no clear task. When you turn the router off yourself (
/route offorStop routing), subagents go back to the picker's level too, and/route onbrings their routed levels back. - Seeing it.
/route statuslists this session's routed subagents, newest first (the last 10), with level, description, agent type and why. The debug log has one line per routed launch. Nothing is added to the footer or to the parent's conversation.
Claude can't set a subagent's effort itself today: the Agent tool takes a model but no effort, so without the router every subagent runs at the session's level unless its agent definition sets one. Set routeSubagents to false to go back to that (the main thread's level in use, as before 0.7.0).
Where the effort went
/route report shows what your requests spent at each level over the last 7 days. /route report session, month or all cover other spans. For example:
Effort for the last 7 days (since 2026-09-28): 412 requests in 9 sessions, 610k output tokens.
By level:
low: 120 requests, 31k output tokens (avg 258)
medium: 260 requests, 410k output tokens (avg 1.6k)
high: 32 requests, 169k output tokens (avg 5.3k)
Changed by the router: 74 requests
subagents, medium → low: 44 requests, 9.9k output tokens (avg 225, vs 1.6k for those left at medium)
main conversation, medium → high: 30 requests, 160k output tokens (avg 5.3k, vs 1.6k for those left at medium)
The router's own checks: 61 (9 of a first prompt, 14 of a conversation, 38 for subagents), using 3.1k output and 1.20M input tokens.
By repo (output tokens): payments 400k, web 210k.
No "saved" figure: the router lowers easy tasks and raises hard ones, so these averages can't show what a changed request would have cost.
- What it records. Every model request in every session with the router installed (0.9.0 on), on the main thread and in subagents, with the router on or off. For each one it keeps the level the request arrived at (your picker's, or the level a subagent would have inherited), the level it went out at, and its tokens as the API reported them. Requests are summed per day into one small JSON file per session, in
~/.claude/effort-router/spend/. The file is written when a turn ends, and nothing leaves your machine. - What it shows. Requests and output tokens per level, with the average per request. The requests the router moved, by thread and direction, each beside the average request left at the level it came from. Requests whose agent definition set their level. The router's own checks, by kind, so its cost is in the same report. Each session check's level, confidence and outcome is kept in the file too, for setting the confidence bar later. Over more than one session, output by repo.
- Why output tokens. Output (thinking plus the answer) is what effort changes most. Input is recorded too.
- Why there is no "saved" figure. The router lowers easy tasks and raises hard ones. A lowered request is small partly because its task was small, so comparing it with the average medium request would overstate the saving, and the same comparison would overstate what a raised request cost extra. Only running the same task at both levels can say what a request would have cost at its old level. The report gives the measured numbers side by side and leaves that estimate out.
Policy
The shipped rules (rules/default.md) are principles, not a table of levels:
- Effort buys verification, edge-case testing and independent judgement, not a better approach.
- Weigh how much is hidden (edge cases, existing code, money, several external systems, concurrency, security), whether you're in the loop, how well specified the task is, and how big it is.
- Pick the level that does the work well on this model without paying for thinking it won't use.
They come from the article above and Anthropic's effort docs. What each level can do comes from the model notes.
Commands
| Command | What it does |
| --- | --- |
| /route | Runs the router now over the whole conversation, in any state, and asks if its level differs |
| /route <hint> | The same, with a hint for the classifier (/route this is a security review) |
| /route status | Shows the state and why, the consent mode, automatic reads used of the budget, classifier calls and how long the last read took, how much transcript it sent, the last verdict (with the raw reply and when), the last error, and this session's routed subagents |
| /route report [session\|week\|month\|all] | Shows where the effort went: requests and output tokens per level, what the router moved, its own reads, and output by repo. The last 7 days by default |
| /route off | Turns the router off and restores the picker's earlier level |
| /route on | Turns the router back on: deciding over the whole conversation, with a fresh budget |
| /route next | Back to deciding at your setting, with a fresh budget: your next prompt is checked before its turn starts (the band's Reassess with my next prompt) |
| /route rules | Prints the effective rules and which layers contributed |
| /route rules init [user\|project] | Writes a starter rules file that keeps the defaults |
| /route rules critique | Asks Sonnet to critique your custom rules |
/route is registered with $.command.register, so it shows in the typeahead. State is per session. /route decide is a hidden alias of bare /route. There is no command to set a level: turn the router off and use the effort picker.
Options
Set them in /config, or under pluginConfigs["effort-router@tommy-mods"].options in settings.json.
| Option | Default | Meaning |
| --- | --- | --- |
| consent | auto | When the router wants a different level from your picker: auto uses the router's level without asking and shows it once in the band, with a button to stop routing. ask holds the turn and asks (Use the router's level, a level in between or Keep yours) |
| decideWithin | 6 | Prompts (and answered questions) the router reads automatically, counted from the session's start |
| classifyTimeoutMs | 15000 | How long a prompt waits for the check before it runs anyway |
| classifierMaxChars | 24000 | Most transcript characters a separate check sends |
| classifierModel | session | session: your session's own model, as a fork from the second prompt. Or another supported model (opus, sonnet, fable) for separate checks. Haiku is never used |
| highestLevel | xhigh | The highest level the router picks. max allows max, which rarely beats xhigh on the current models |
| confidence | 0.7 | How sure (0 to 1) a check must be that your current level is wrong in one direction (or right) before the router acts. 0 acts on any check |
| showChecks | false | Print a line after each automatic check: its spread, how sure it was and what the router did |
| skipAboveTokens | 20000 | A session first seen with more conversation than this keeps your effort setting |
| firstCheckInstructions | true | Send your CLAUDE.md files, rules and memory with the first prompt's check. false: the prompt only |
| syncPicker | true | Run /effort <level> so the terminal's picker label matches |
| routeSubagents | true | Give each subagent its own level, chosen at launch by the agent starting it. false: subagents run at the main thread's level |
| footerControl | button | button makes the footer state a button that opens the band. label draws plain text, and /route is the control |
| rules | empty | Rules text for your user layer. A rules file takes precedence |
The environment variable EFFORT_ROUTER_CONSENT=ask|auto overrides consent, which helps in headless runs (-p has no one to answer a question, so use auto there). The older names still work there: apply and none mean auto, while confirm and band mean ask.
Customising the rules
The rules are plain markdown, layered from the bottom up:
- The shipped defaults (
rules/default.md) - Your organisation's rules from managed (policy) settings, if it sets any
- Your rules:
~/.claude/effort-router.md, or therulesoption in your user settings - The project's rules:
<project root>/.claude/effort-router.md(commit it), or therulesoption in project settings
A line that is exactly $defaults pulls in everything beneath that layer. Text after the line adds to the rules, and later rules win. Text before it goes first. A file with no $defaults line replaces everything beneath it. Missing or empty files change nothing, and HTML comments are ignored.
$defaults
- This is a payments codebase. Never pick below high: money movement needs verification.
The files are re-read on every classification, so edits apply without a reload. An unreadable file is skipped. The frame around the rules (undecided only before a task, the worked examples, reply in JSON) is fixed, so no rules file can break the parser.
For organisations
An organisation can set routing rules centrally in managed settings (managed-settings.json), as it does for other Claude Code policy:
{
"pluginConfigs": {
"effort-router@tommy-mods": {
"options": {
"rules": "$defaults\n\n- Code under payments/ or ledger/ is never routed below high.\n- Infrastructure changes (terraform/, k8s/) are high.",
"rulesMode": "enforce",
"allowOff": false
}
}
}
}
rulesMode: "extend"(the default) layers the org rules over the shipped defaults. Users and projects can add to them with$defaults, or replace them.rulesMode: "enforce"makes the org layer final. Personal and project rules are ignored, and/route rules initsays so.routeSubagents: falseturns subagent routing off for everyone, whatever their own setting.allowOff: falsestops users turning the router off, so the organisation's routing always applies./route offrefuses, the band has noStop routing, and a session saved as off comes back deciding. When the budget runs out with nothing suggested, the router idles asdeciding(no more reads) instead of turning off./routestill works.
A top-level "effortRouter": { "rules": ..., "rulesMode": ..., "allowOff": ..., "routeSubagents": ... } object works too. The router reads these four settings only from the policy source, so a user cannot claim enforce for themselves.
Install
claude plugin marketplace add tommy5dollar/claude-mods
claude plugin install effort-router@tommy-mods
There's nothing to configure: every option has a default. /plugin configure lists the options as not yet set, which just means the defaults apply. Change one only when you want something different (see Options).
For development, run claude --plugin-dir ./effort-router.
Known limits
- In the Desktop app the native effort picker never changes: the app owns it and nothing a mod can call sets it. The requests still go out at the routed level; trust the footer label.
- In the terminal the router can't set the picker label directly either. It runs
/effort <level>when the session is idle, which prints a line in the transcript, and the footer shows the true level until then. Headless (-p) runs skip the sync because its output would replace the run's printed result. The per-request override still applies there. - Each prompt waits for the check while the router is deciding (at most
decideWithinprompts per session): about 1.5 s on Opus 5.5, longer on a model that thinks more by default. If the check times out (classifyTimeoutMs), that turn goes at the picker's level; a late answer is ignored. - The first prompt's check can't share the session's prompt cache: the engine offers no way to fork before the first response, and a separate call can't carry Claude Code's system prompt or tools. It pays for your instructions and the prompt once per session. Claude Code also never reuses the cache for the conversation part of a session's first request, so the first fork after a one-request first turn pays for about 13k tokens again.
- On Bedrock, Google Cloud or an LLM gateway, Claude Code clears the cached conversation when effort changes (Anthropic's docs). A locked level that differs from your setting changes effort once, so expect one uncached request there. With an API key or a subscription the cache is kept.
- The confidence bar (0.7) is a starting guess, not calibrated yet. The ledger keeps every check's confidence and outcome for that.
- A fork answers at the effort the session last used, and the router can't change it.
- A request whose picker level is a number rather than a named level, or a model that takes no effort, can't be compared, so the waiting verdict waits for the next request that can.
- The router changes effort only, never the model. A request to a model that takes no effort is left alone.
- Each automatic check is one call per prompt while deciding, for at most
decideWithinprompts. Nothing more is spent once a level is locked or the budget is spent, except when you run/route./route reportshows what the checks cost. - Once decided, the router does not notice a change of phase on its own (for example "now verify it" after an implementation). Run
/route(or the band'sReassess now), optionally with a hint, or/route nextto be asked with your next prompt. - A definition's
effort:is respected for user and project agent files and for theagentskey in settings, not for plugin agents (<plugin>:<name>), which can't be located reliably. Those are routed from their brief, which replaces any effort their definition sets. A mod's$.agent.register({ effort })is ignored by the engine itself (an engine bug), and the router routes those agents too. - An agent file added or edited mid-session is seen from the next session: definitions are scanned once per session.
- Workflow agents that don't launch through the Agent tool raise no
agent.spawn, so they keep the main thread's level. - Subagent levels are kept in memory only. After a restart or resume, a subagent still running from before takes the main thread's level.
- The router adds no note about the chosen level to the system prompt, because changing a cached prompt section would break the prompt cache. The
/effortecho tells the model instead, and it is appended to the transcript, so the cache holds. - The spend report starts at 0.9.0: sessions from before it aren't in it. A request with no reported usage (failed or interrupted) isn't counted. Days are UTC.
- The band and footer draw in the terminal and the Desktop app. VS Code and
-prun the hooks without the UI. Underaskin-pthe question has no one to answer, so requests stay at the picker's level; setconsent(orEFFORT_ROUTER_CONSENT) toautothere.
Development
bun test # pure policy: trimming, parsing, rule layering, /route grammar, subagent reads, the spend ledger and report
claude plugin test . # engine kit: band, buttons, footer button, turn.step, /route, org layers, agent.spawn, the ledger saved and reported
claude plugin validate . --strict
bun run eval # opt-in: the real classifier over eval/fixtures.ts (see below)
bun run eval sends each fixture to the real model through claude -p --safe-mode (no plugins or hooks, no tools), with the system prompt and input the router builds from rules/default.md: for the session read, a first check on the session's model (--model, default opus) at its default effort with that model's notes from rules/models/, and for a subagent's read, a separate call on the same model with that model's notes and the brief. It can't reproduce the forks that later checks and subagent reads make. It parses the reply with the router's own parser and prints each verdict, the pass rate and every miss. --runs 3 repeats each fixture (the model is not deterministic), --model sonnet runs the session set as a Sonnet session, --effort sets the effort and --confidence the bar, --set session or --set subagent runs one set and --only <text> filters fixtures by name. It uses your Claude Code login, and each fixture costs one small model call.
TESTING.md lists the live checks.
Licence
MIT
