KilimcininKorOglu/claude-code-mods/tree/main/plugins/cache-warm
cache-warm
設定したウィンドウの間、1時間プロンプトキャッシュを温かく保つ、またはalwaysモード下のすべてのセッションで無制限に保つ。アイドルの各区間ごとに1つのキャッシュ共有フォーク、再開後に1つのウォームアップメッセージ、有料のコールド書き込み後にウィンドウを有効化、キャッシュ状態、コールド価格、このセッションのコール数を表示します。
この mod について
cache-warm
Claude Codeは会話をプロンプトキャッシュに1時間保持します。それ以上離れると、次のメッセージはキャッシュ書き込み率で全コンテキストを再書き込みします。Claude Fable 5.1では200kトークンあたり$4.00ですが、ウォームキャッシュから同じトークンを読み込むには$0.05です。このmodは選んだウィンドウの間キャッシュを温かく保ち、コールドキャッシュのコストを表示します。
Karan Bansal(karanb192/claude-code-mods)のcache-taxモッドの動作に従いますが、その送信ガードなしです。コードは新規です。
機能
キャッシュを温かく保つ。 /cache-warmは6時間のウィンドウを有効化します。その中で、メインループの最後のモデル要求から50分後、modはセッション自身のトランスクリプト上でツールレスの$.model.forkを1つ送信します。サーバーはそれをキャッシュから応答し、1時間をリフレッシュします。新しい要求ごとにpingが遅くなるため、活動的に使用しているセッションは全くpingを送信しません。セッションの最初の要求、および/clearまたは圧縮後の最初の要求は、pingも設定します。そのため、最初のターン内で1時間を超えて実行するツール呼び出しでも、ターンの中ほどでそのpingを取得します(2.1.281で測定:pingはsleep 110実行中に出ていき、ターンは正常に終了)。
alwaysの下で無制限に実行。 /cache-warm alwaysはウィンドウではありません。pingは50分ごとにセッションが生存している限り出ていき、/cache-warm offのみがそれを終了します。スイッチはmodの独自の$.store内の1つのグローバルキーなので、すべてのプロジェクトの後続セッションは起動時と/clear後に同じループを開始します。
- 各ターンの終わりは最後の要求の時刻をセッションのidの下の
$.storeに保持し、各pingはそこに独自の読み込みも保持します。実行中の会話にロードされたモジュール(/reload-plugins、更新)はそのため、2つの後の方からタイムリーにpingします。pingsのみで温かく保たれるセッションは、最後のターンがそのキャッシュより古い状態です。両方が1時間以上古い場合、キャッシュは消え、ループはコールドpingの料金を支払う代わりに次のターンを待ちます。 - セッションのトランスクリプトはこれに使用されません。
/reload-pluginsは独自の行をそこに書き込み、ファイルの最後の書き込みは要求に見えるためです(測定:最後のターンから90秒後のリロードはpingを90秒遅く設定するはずです)。時刻が未保持のセッション、古いバージョンを実行したセッションのみが、そのトランスクリプトの最後の書き込みを1回読み込みます(~/.claude/projects/<directory>/<session id>.jsonl、CLAUDE_CONFIG_DIRが設定されている場合はその下)。トランスクリプトはセッションが開始したディレクトリの下で検索されます。シェルのcdが移動したディレクトリではなく、見つからないディレクトリはトランスクリプト行で名前が付けられます。 - キャッシュが消えているのが見つかるpingはこのループを終了しません。そのpingが支払った書き込みが新しいキャッシュです。modはそれをトランスクリプト行で言及し、セッションの集計に書き込みをカウント、続行します。ウォームpingはコンテキストを読み込み率で読み込みます。約$0.05で200kトークンなので、アイドルの1日のpingsのコストは約$1.40です。
キャッシュが消えた時に停止。 これはウィンドウの終わりがある場合に適用されます。alwaysには適用されません。ウォームpingはコンテキストを読み込み、独自の数トークンのみを書き込みます。pingが何も読み込まない、または読み込みの10分の1以上を書き込む場合、キャッシュはすでに消えており、ping自身が書き込みの料金を支払ったため、modは停止し、理由を表示します。APIがフォークにエラーで応答する場合にも停止し(行はそのステータスと種類を名前付けます)、フォークが応答前に切断された場合にも停止します。エンジンがフォークするものがない場合、再開処理がその最初の応答の前のように、ウィンドウは停止しません。次の応答を待ち、そのターンが再度pingを有効化し、行はキャッシュがいつまで保持されるかを言及します。テキストなしの応答でもキャッシュを読み込むため、pingとしてカウントされます。always下で、そのような失敗はそのターンのみループを停止します。次のターンが再度開始するため、セッションはスイッチが何も実行されていない状態を保有しません。
有料のコールド書き込み後に自動有効化。 ターンが20kトークンより大きいコンテキストの少なくとも半分を再書き込みする場合、modはそのコールド書き込みをカウント、より長いものがすでに有効化されていない限り6時間のウィンドウを有効化します。always下ではエンドレスループがすでにそのキャッシュを保つため、6時間のウィンドウは有効化されません。
状態を表示。 /cache-statusはモデル、ウォームまたはコールド、コンテキストサイズ、コールド価格、ウィンドウ、ブレイクイーブン、このセッションのコールド書き込みを出力します。
コールドキャッシュに送信するメッセージは決して停止または遅延されません。キャッシュがlapseしたセッションを再開すると、その最初のメッセージの価格の1行を取得します。Claude Codeはキャッシュをトランスクリプトの最後の応答から日付けします。pingはそれを決して書き込まないため、その行は最後のpingがmodがセッション用に保持した間、キャッシュを1時間以内に読み込んだ間は省略されます。
再開後にウォームアップメッセージを送信。 再開処理はそれ独自の最初の応答の前にはフォークできません。$.model.forkはnothing-to-forkと応答します(ヘッドレスのclaude --resumeで測定)。そのためpingは出ていくことができず、30分後に閉じて開いたセッションは1時間でキャッシュを失うことになります。何かを書き込まない限り。ウィンドウまたはalwaysが実行されている間にインタラクティブセッションが再開され、そのコンテキストが50k以上のトークン、そのキャッシュが保持されている場合、modは再開から3秒後に独自の/cache-warm:sendコマンドを通してメッセージを1つ送信します。
/cache-warm:send This message was sent by the cache-warm plugin, not by the person. The session was resumed, and a resumed session can keep its prompt cache warm only after a reply. Do not run a tool or continue a task. Reply with the single word: warm
これは実際のターンです。キャッシュを読み込み、モデルが1単語で答え、ペアが会話に留まり、その終わりはpingを再度有効化します。キャッシュがすでに消えている場合は何も送信されません(次のメッセージはいずれにせよ同じ書き込みの料金を支払います)、-p実行内、または3秒以内にメッセージを送信した場合。他のmodsはそれをあらゆるターンのように見ます。task-pokeはタスクが開いている間、それの後で続行プロンプトを送信し、desk-notifyはターン終了通知を表示します。2.1.283で116kトークンの再開されたインタラクティブセッションで測定されました。メッセージは再開時に出ていき、モデルはwarmと答え、次のpingはフォークして117kトークンを読み込みました。再開は独自にプレフィックスの一部を破壊する可能性があります(そのセッションは116kの42kを再書き込みしました。別のセッション、最後のpingから40分後に再開、4k)。ウォームアップターンはその再開で最初のメッセージの代わりに書き込みの料金を支払います。3秒は2つの開始フックの後の方から数えられます。classic.SessionStartとsession.start、これらは固定順序で解決されないため。マーケットプレイスのすべてのmodがロードされて、session.startはclassic.SessionStartの4秒後に解決されました(2.1.285で測定)。最初のみで開始された待機は開始していないセッションを検出、何も送信しません。
コマンド
/cache-warm keep warm for six hours
/cache-warm 90m keep warm for a window of your own (also 2h30m)
/cache-warm always keep the cache warm with no end, in every session of every project
/cache-warm 6h every 2m ping every two minutes; a test setting, floor 1m, forgotten after this window
/cache-warm status the status line text
/cache-warm off stop, forget the window, and turn always off
/cache-status the card
/cache-warm:send <text> the keep-warm message the mod sends after a resume; its body is the text alone
表示内容
ウィンドウが有効化されている間またはストップ後、プロンプト下のステータスライン:
cache-warm: 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
cache-warm: stopped: the ping read 0 and wrote 180k tokens ($3.60), the cache was already gone
lastパートは最後の要求を名前付けます。これはキャッシュを読み込んだもの、pingまたはメインループのターン、どちらか後の方です。last ping read 200k $0.05またはlast turn read 250k $0.07、両方ではありません。ターンの数字はターン全体をカバーします。すべての要求が合計されます。括弧内の時刻はその答えが来た時刻です。現地時間で。早い日からのものは日と月を持ちます。(22 Sep 23:10)として。記録はセッションのidの下の$.storeに保持されるため、/reload-pluginsまたは更新はすぐに再度表示されます。
ウィンドウが実行されている間、インタラクティブセッションは1分ごと行をリドロー、左時間とnextpingの時刻はターン間とping間でカウントダウン、lastパートは行に留まります。pingの前の最後の分はping nowと読みます。分は丸められ、行は1分ごとにドロー、pingのフォークが出ている間、pinging…と読みます。エンジンがフォークするものがない場合、行はカウントダウン代わりに言及します。
cache-warm: always · no ping before the next reply · cache holds until 19:16
サイドバーが開いている間、pings試行ごとの1つのストリームエントリ。 それはサイドバーのログファイルにも保持されます(~/.claude/sidebar/<project>-<date>.log)。後で読み込み戻すことができるため。pingが出たかどうか何をしたかを見ます。サイドバー閉じると、同じテキストはトランスクリプト行です。
ping sent · read 901k · wrote 0 · $0.45
ping found the cache gone · read 0 · wrote 180k · $3.60
ping not sent: the conversation has no reply to fork yet; the ping waits for the next reply, and the cache holds until 19:16
ping failed: the ping failed, the API answered 529 (overloaded)
keep-warm message sent: the session was resumed and its cache holds until 19:16; a resumed session pings only after a reply
sidebarが開いている場合、ステータスラインはそこにcache windowセクションとして移動。セッション用に留まり、すべての変更で書き直されます。ステータスラインはクリアに留まります。左時間のみ(またはalways)が色付けされます。ウィンドウが1pingサイクルより前に終わるとき、黄;保有される間は緑;最初のターンを待つ間は薄くなります。後のping詳細は薄く、ストップウィンドウはそのstopped:フロントを赤で、理由をデフォルト色で表示します。サイドバーなしで、ステータスラインは上記として描画されます。
ストップ理由は1ターン留まります。次のターンでセクションはアイドル行を代わりに表示します。薄い、支払われたN cold writes paid $X除いて、これは黄。そのように窓ペインは今の測定を表示します。終わったウィンドウの最後の文ではなく。理由はトランスクリプトに留まり、ウィンドウが実行されていないときのステータスラインはクリアです。
cache window
off · 2 cold writes paid $6.30 · context 315k tokens
時間が切れたウィンドウはあなたの次のメッセージで再度有効化されます。終わった方と同じ限り、同じpingサイクルで、アイドル行はそれが待つ間に言及します。
cache window
off · 6h again at your next message · 1 cold write paid $4.01 · context 201k tokens
cache-warm: the 6h window ran out; this message arms another one. /cache-warm off stops it.
時間が切れたウィンドウのみ戻ります。pingが停止したもの、ではありません。キャッシュはすでにそこで消え、次のメッセージのコールド書き込みが独自の6hウィンドウを有効化します。/cache-warm offは戻りを待つウィンドウを忘れます。
ウィンドウの下、セクションはの2番目、薄い行を保有します。最後のトランスクリプト行、短縮、コールド書き込みのコストで黄。ウィンドウ行はキャッシュがどのくらい保持されるかを言及します。2番目の行はmodが最後にしたことを言及します。
cache window
6h left · ping in 50m
cold write 201k tokens paid ($4.01)
/cache-statusのカード:
claude-fable-5-1
state warm, 42m left
context 200,502 tokens
cold cost $4.01 to re-write it (warm turn $0.05)
keep warm on, 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
break-even up to 80 pings at the read rate cost one cold write, about 2d 18h of idle at one ping per 50m
session 1 cold write paid, $4.01
1つのトランスクリプト行、モデルに送信されず、コールド書き込みがウィンドウを有効化、または再開がコールドで開始されたとき。ウォッチャーがオフの間(ウィンドウなし、always off)、行はサイドバーのストリームに行く代わり、トランスクリプトに何も書き込みません。クローズされたサイドバーはそれをドロップします。実行ウォッチャーが書き込む行は変わりません。
価格
テーブル i
インストール
まず作者の README で marketplace とプラグイン名を確認してください。コマンドはリポジトリの構成によって変わる場合があります。
claude plugin marketplace add KilimcininKorOglu/claude-code-mods claude plugin install cache-warm
原文 / README
cache-warm
Claude Code keeps your conversation in a prompt cache for one hour. Step away for longer, and the next message re-writes the whole context at the cache-write rate: on Claude Fable 5.1 that is $4.00 for 200k tokens, against $0.05 to read the same tokens from a warm cache. This mod keeps the cache warm for a window you choose, and shows you what a cold cache would cost.
It follows the behaviour of the cache-tax mod by Karan Bansal (karanb192/claude-code-mods), without its send guard. The code is new.
What it does
Keeps the cache warm. /cache-warm arms a six-hour window. Inside it, 50 minutes after the main loop's last model request, the mod sends one tool-less $.model.fork over the session's own transcript. The server answers it from the cache, and that refreshes the hour. Every new request pushes the ping later, so a session you are actively using sends no ping at all. The first request of a session, and the first after /clear or a compaction, sets the ping as well, so even a tool call that runs past the hour inside the first turn gets its ping in the middle of the turn (measured on 2.1.281: the ping went out while sleep 110 ran, and the turn ended normally).
Runs with no end under always. /cache-warm always is not a window: the ping goes out every 50 minutes for as long as the session lives, and only /cache-warm off ends it. The switch is one global key in the mod's own $.store, so every later session of every project starts the same loop at its start and after /clear.
- Each turn's end keeps the last request's time in
$.storeunder the session's id, and each ping keeps its own read there too. A module loaded into a running conversation (/reload-plugins, an update) therefore pings on time from the later of the two; a session kept warm by pings alone has a last turn older than its cache. When both are more than an hour old, the cache is gone, and the loop waits for the next turn instead of paying for a cold ping. - The session's transcript is not used for this, because
/reload-pluginswrites a line of its own there and the file's last write would then look like a request (measured: a reload 90 seconds after the last turn would have set the ping 90 seconds late). Only a session with no time kept yet, one that ran an older version, reads the last write of its transcript (~/.claude/projects/<directory>/<session id>.jsonl, underCLAUDE_CONFIG_DIRwhen it is set) once. The transcript is looked up under the directory the session started in, never the one a shellcdmoved to, and one it cannot find is named in a transcript line. - A ping that finds the cache gone does not end this loop. The write that ping paid for is the new cache: the mod says so in a transcript line, counts the write in the session's tally and keeps going. A warm ping reads the context at the read rate, about $0.05 for 200k tokens, so an idle day of pings costs about $1.40.
Stops when the cache is gone. This applies to a window with an end, not to always. A warm ping reads the context and writes only its own few tokens. When a ping reads nothing, or writes a tenth of what it read or more, the cache was already gone and the ping itself paid for the write, so the mod stops and shows why. It also stops when the API answers the fork with an error (the line names its status and kind), and when the fork is cut before it replies. When the engine has nothing to fork, as in a resumed process before its first reply, the window does not stop: it waits for the next reply, whose turn arms the ping again, and the line says until when the cache holds. A reply without text still read the cache, so it counts as a ping. Under always such a failure stops the loop for that turn only: the next turn starts it again, so the session never holds the switch while nothing runs.
Arms itself after a paid cold write. When a turn re-writes at least half of a context larger than 20k tokens, the mod counts that cold write and arms a six-hour window, unless a longer one is already armed. Under always no six-hour window is armed, because the endless loop already keeps that cache.
Shows the state. /cache-status prints the model, warm or cold, the context size, the cold price, the window, the break-even and this session's cold writes.
A message you send to a cold cache is never stopped or delayed. A resumed session whose cache has lapsed gets one line with the price of its first message. Claude Code dates the cache from the transcript's last reply, which a ping never writes, so that line is left out while the last ping the mod kept for the session read the cache within the hour.
Sends a keep-warm message after a resume. A resumed process cannot fork before its own first reply: $.model.fork answers nothing-to-fork (measured with a headless claude --resume). So no ping can go out, and a session you closed and opened again 30 minutes later would lose its cache at the hour unless you wrote something. When an interactive session is resumed while a window or always runs, its context is 50k tokens or more and its cache still holds, the mod therefore sends one message three seconds after the resume, through its own /cache-warm:send command:
/cache-warm:send This message was sent by the cache-warm plugin, not by the person. The session was resumed, and a resumed session can keep its prompt cache warm only after a reply. Do not run a tool or continue a task. Reply with the single word: warm
This is a real turn: it reads the cache, the model answers one word, the pair stays in the conversation, and its end arms the ping again. Nothing is sent when the cache is already gone (your next message pays the same write anyway), in a -p run, or when you sent a message within those three seconds. Other mods see it like any other turn: task-poke may send its continue prompt after it while tasks are open, and desk-notify shows its turn-end notification. Measured on 2.1.283 in a resumed interactive session of 116k tokens: the message went out at the resume, the model answered warm, and the next ping forked and read 117k tokens. A resume can break part of the prefix on its own (that session re-wrote 42k of the 116k; another, resumed 40 minutes after its last ping, 4k); the keep-warm turn pays that write at the resume instead of your first message. The three seconds count from the later of the two start hooks, classic.SessionStart and session.start, because they settle in no fixed order: with every mod of the marketplace loaded, session.start settled four seconds after classic.SessionStart (measured on 2.1.285), and a wait started by the first alone found a session that had not started and sent nothing.
Commands
/cache-warm keep warm for six hours
/cache-warm 90m keep warm for a window of your own (also 2h30m)
/cache-warm always keep the cache warm with no end, in every session of every project
/cache-warm 6h every 2m ping every two minutes; a test setting, floor 1m, forgotten after this window
/cache-warm status the status line text
/cache-warm off stop, forget the window, and turn always off
/cache-status the card
/cache-warm:send <text> the keep-warm message the mod sends after a resume; its body is the text alone
What it shows
A status line under the prompt while a window is armed or after a stop:
cache-warm: 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
cache-warm: stopped: the ping read 0 and wrote 180k tokens ($3.60), the cache was already gone
The last part names the last request that read the cache, a ping or a main-loop turn, whichever came later: last ping read 200k $0.05 or last turn read 250k $0.07, never both. A turn's figures cover the whole turn, every request summed. The time in brackets is when that answer came, in local time; one from an earlier day carries its day and month, as (22 Sep 23:10). The record is kept in $.store under the session's id, so /reload-plugins or an update shows it again at once.
While a window runs, an interactive session redraws the line every minute, so the time left and the time to the next ping count down between turns and pings, and the last part stays on the line. The last minute before a ping reads ping now, because minutes are rounded and the line is drawn once a minute; while the ping's fork is out it reads pinging…. When the engine has nothing to fork, the line says so instead of counting down:
cache-warm: always · no ping before the next reply · cache holds until 19:16
One stream entry per ping attempt while the sidebar is open. It is also kept in the sidebar's log file (~/.claude/sidebar/<project>-<date>.log), so you can read back later whether a ping went out and what it did; with the sidebar closed the same text is a transcript line:
ping sent · read 901k · wrote 0 · $0.45
ping found the cache gone · read 0 · wrote 180k · $3.60
ping not sent: the conversation has no reply to fork yet; the ping waits for the next reply, and the cache holds until 19:16
ping failed: the ping failed, the API answered 529 (overloaded)
keep-warm message sent: the session was resumed and its cache holds until 19:16; a resumed session pings only after a reply
With the sidebar open, the status line moves there as a cache window section that stays for the session and is rewritten at every change, and the status line stays clear. Only the time left (or always) is coloured: yellow when the window ends sooner than one ping period, green while it holds, faint while it waits for the first turn. The ping details after it are faint, and a stopped window shows its stopped: front in red with the reason in the default colour. Without the sidebar, the status line is drawn as above.
A stop reason stays for one turn. At the next turn the section shows the idle line instead: faint, except for a paid N cold writes paid $X, which is yellow. That way the pane shows a measurement of now, not the last sentence of a window that ended. The reason stays in the transcript, and the status line is empty while no window runs:
cache window
off · 2 cold writes paid $6.30 · context 315k tokens
A window that runs out of time is armed again by your next message, as long as the one that ended and with the same ping period, and the idle line says so while it waits:
cache window
off · 6h again at your next message · 1 cold write paid $4.01 · context 201k tokens
cache-warm: the 6h window ran out; this message arms another one. /cache-warm off stops it.
Only a window that ran out of time comes back. One that a ping stopped does not: the cache is already gone there, and the cold write of your next message arms its own 6h window. /cache-warm off forgets a window waiting to come back.
Under the window the section holds a second, faint line: the last transcript line, shortened, with the cost of a cold write in yellow. The window line says how long the cache is kept; the second line says what the mod did last:
cache window
6h left · ping in 50m
cold write 201k tokens paid ($4.01)
The card of /cache-status:
claude-fable-5-1
state warm, 42m left
context 200,502 tokens
cold cost $4.01 to re-write it (warm turn $0.05)
keep warm on, 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
break-even up to 80 pings at the read rate cost one cold write, about 2d 18h of idle at one ping per 50m
session 1 cold write paid, $4.01
One transcript line, not sent to the model, when a cold write arms the window or a resume starts cold. While the watcher is off (no window, always off), the line goes to the sidebar's stream instead and writes nothing to the transcript; a closed sidebar drops it. The line a running watcher writes is unchanged.
Prices
The table in hooks/pricing.ts holds the cache-read, 1-hour cache-write and output rates of every model on the Anthropic pricing page, read in September 2026. A model id takes the first family it contains, so claude-opus-4-1 is priced as Opus 4.1 ($1.50 / $30 / $75) and claude-opus-4-8 as Opus 4.8 ($0.50 / $10 / $25). claude-opus-5-5 also contains opus-5, so its own row comes first: Opus 5.5 is $0.20 / $8 / $20, cheaper than Opus 5. Sonnet 5.5 has its own row at the Sonnet 5 rates, $0.20 / $4 / $10. A ping is priced in full: the cache read, its cache write, its uncached input at the base rate (half the 1-hour write rate) and its output. An unknown model shows n/a.
Fast mode bills Opus 5.5, Opus 5 and Opus 4.8 at their own base rates ($8 and $10 input), with the cache multipliers applied on top. The mod prices Opus 5.5 at $0.40 / $16 / $40 and Opus 5 and 4.8 at $1 / $20 / $50 while the fastMode setting, which /fast writes, is on. It reads the settings at the session's start and at the end of each main-loop turn, so a /fast counts from the next turn. With fastModePerSessionOptIn set to true, every session starts with fast mode off, so the standard rates apply. Any other model keeps its standard rates, and the card mentions fast mode only when the rates changed:
claude-opus-5-5 · fast mode rates (the fastMode setting)
On a subscription the dollars are a yardstick, not your bill. How a cache read counts against the 5-hour and weekly limits is not documented.
Install
claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install cache-warm@kilimcininkoroglu-mods
Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.
To load it from a local checkout for one session:
claude --plugin-dir plugins/cache-warm
After installing
- Restart Claude Code.
- Disable every other keep-warm mod, for example
claude plugin disable cache-tax@claude-code-mods. Two keep-warm mods in one session send two pings per idle stretch. - Check once that a ping reads your cache, as "Prove it on your own session" below describes.
- To keep the cache warm with no end, run
/cache-warm alwaysonce. The switch is global: every later session of every project starts the loop by itself, and/cache-warm offends it for good. Without it, a window is armed only by/cache-warmor after a paid cold write.
What it can reach
Validated with claude plugin validate on Claude Code 2.1.288:
❯ ./register.ts hooks: session.start, classic.SessionStart, prompt.submit, command.run{command=cache-warm}, command.run{command=cache-status}, turn.step, turn.complete, session.compact
❯ ./register.ts calls: $.clock.after (via arm, scheduleKeepWarm), $.clock.every, $.clock.now, $.command.register (via registerCommands), $.command.run (via keepWarmAfterResume), $.env.get (via seedFromTranscript), $.fs.exists (via seedFromTranscript), $.fs.stat (via seedFromTranscript), $.model.fork (via forkPing), $.prompt.submit (via keepWarmAfterResume), $.session.id, $.session.model, $.session.root (via seedFromTranscript), $.session.usage, $.settings.read (via readFast), $.sidebar.set (via logEvent, toSidebar, toStream), $.store.delete (via prune, pruneRequests, startEndless, startWindow, stop), $.store.get, $.store.keys (via prune, pruneRequests), $.store.set (via afterTurn, keepLastRead, startWindow, warmCommand), $.ui.log (via logEvent, seedFromTranscript, toStream), $.ui.status (via showStatusAt)
❯ ./register.ts env writes: nothing
❯ ./register.ts env reads: CLAUDE_CONFIG_DIR, HOME
Reach L2: it drives Claude.
1. Reads: the time of each main-loop model request; the token counts and model id of each turn and of each ping; the live context size; the origin of each message, to arm a window again; the resume fields Claude Code computes for settings hooks; the session id and model; the last write time of the session's own transcript file, once when the module loads into a running conversation; the fastMode and fastModePerSessionOptIn settings, at the session's start and at each turn's end; its own $.store. It never reads a prompt's text, a file's content or a tool result.
2. Runs: one $.model.fork per idle stretch while a window or the always loop runs, 50 minutes after the last request unless the test setting is used (floor 1 minute); never while off; a ping that found the cache gone ends a window with an end, and under always the loop carries on; after a resume of an interactive session whose cache still holds, one keep-warm message through /cache-warm:send (a plugin prompt when the engine refuses the command), which is a real turn
3. Sends: the fork, an API request over the session's own transcript with a fixed one-line prompt, and after a resume the fixed keep-warm message as a turn of the conversation
4. Persists: in $.store, the window end, the ping period and the last main-loop request's time and the last ping or turn read (tokens, cost, time) under this session's id, and the global always switch, which the endless loop needs no window key beside; this session's ended window is deleted at stop and at its next start, another session's window one week after it ended, another session's request time and last read once they are an hour old; the cold-write tally lives in memory and ends with the session
5. Hostile input: the only text it parses is the argument of /cache-warm, matched against a duration pattern and three words; the fork's prompt is a constant, so nothing crafted can reach it
Prove it on your own session
The mock-clock tests prove the timer and the scoring, not that a fork reads the main cache. One ping proves that. In a warm session:
> Reply with one word: ready
> /cache-warm 1h every 1m
After a minute the status line should read last ping read <close to your context> $.... A stopped: the ping read ... line means the fork did not share the cache, and the mod has already stopped. /cache-warm off ends the test.
Limits
- The 50-minute ping assumes the 1-hour cache tier, which the main conversation uses.
- A warm ping only proves the cache was warm at that moment. A model or effort switch, an edited CLAUDE.md or a changed tool list breaks the prefix no matter the time, and your next message pays.
- A ping's output cannot be capped; a model at high effort may think before it answers. The status line prices what the ping really billed.
- The resume logic is covered by hook tests that raise
classic.SessionStartwith the resume fields, and by a live check of a resumed interactive session in tmux. Whether a resume keeps the whole prefix is out of the mod's hands: it re-wrote 42k of 116k tokens in one measured session and 4k in another. - The cold-write tally is per session and lives in memory.
/clearempties it. - Fast mode is read from the
fastModesetting, the saved preference, not from the request. The mod does not see Claude Code fall back to standard speed within a session (a fast mode rate-limit cooldown, usage credits that ran out, an organization that turned fast mode off). Those turns bill standard rates while the mod prices them as fast. - Whether a ping, a
$.model.fork, runs at fast speed while the session does has not been measured; the mod prices it at the session's rates.
Development
make install # eslint, typescript-eslint, typescript
make lint # complexity limit 10, the build fails above it
make typecheck # needs .claude/types/ from /plugin-types
make validate
make test # claude plugin test
