ClaudeMods
☰
ZH-TW
● 0 人在線上 · 瀏覽 0 次
贊助提交作品
GitHub 儲存庫 · 發布者 KilimcininKorOglu

cache-warm

為你設定的時間段保持 1 小時提示快取熱狀態,或在 always 模式下每個工作階段都保持(無時間限制),每個空閒階段一個快取共享分叉,恢復後一條保溫訊息,在付費冷寫入後為你啟用一個時間段,並顯示快取狀態、冷價格和本工作階段的冷寫入次數。

KilimcininKorOglu@KilimcininKorOglu

KilimcininKorOglu/claude-code-mods/tree/main/plugins/cache-warm

已翻譯

關於這個 mod

cache-warm

Claude Code 將你的對話保存在提示快取中 1 小時。如果你離開更長時間,下一條訊息將以快取寫入率重寫整個上下文:在 Claude Fable 5.1 上,這是 200k 權杖 $4.00,而從熱快取讀取相同權杖僅需 $0.05。此 mod 為你選擇的時間段保持快取熱狀態,並顯示冷快取的成本。

它遵循 Karan Bansal (karanb192/claude-code-mods) 的 cache-tax mod 的行為,但沒有其傳送保護。程式碼是全新的。

它的功能

保持快取熱狀態。 /cache-warm 啟用一個 6 小時的時間段。在其中,在主迴圈最後一次模型要求後 50 分鐘,mod 透過工作階段自己的記錄發送一個無工具的 $.model.fork。伺服器從快取中回覆,這會重新整理 1 小時。每個新要求都會推遲 ping,因此你主動使用的工作階段根本不會發送 ping。工作階段的第一個要求,以及 /clear 或壓縮後的第一個要求,也會設定 ping,所以即使工具呼叫在第一個回合內超過 1 小時,它也會在回合中間獲得 ping(在 2.1.281 上測量:ping 在 sleep 110 執行時發出,回合正常結束)。

在 always 下無限執行。 /cache-warm always 不是一個時間段:ping 每 50 分鐘發出一次,只要工作階段活躍,直到 /cache-warm off 結束。開關是 mod 自己 $.store 中的一個全域鍵,所以每個專案的每個後續工作階段在其開始和 /clear 後啟動相同的迴圈。

  • 每個回合的結束在 $.store 的工作階段 ID 下保持最後一次要求的時間,每個 ping 也在那裡保持自己的讀取時間。因此,一個載入到正在執行的對話中的模組(/reload-plugins、更新)會從兩者的較晚者準時 ping;僅由 ping 保溫的工作階段的最後回合比其快取更舊。當兩者都超過 1 小時舊時,快取已消失,迴圈等待下一個回合而不是為冷 ping 付費。
  • 工作階段的記錄不用於此,因為 /reload-plugins 在其中寫入自己的一行,檔案的最後寫入時間看起來像要求(測量:最後一個回合後 90 秒的重新載入會將 ping 設定為晚 90 秒)。只有沒有保持時間的工作階段,即執行較舊版本的工作階段,才會讀取一次其記錄的最後寫入(~/.claude/projects/<directory>/<session id>.jsonl,當設定時在 CLAUDE_CONFIG_DIR 下)。記錄在工作階段開始的目錄下查找,從不是 shell cd 移動到的目錄,找不到的會在記錄行中命名。
  • 找到快取已消失的 ping 不會結束此迴圈。該 ping 支付的寫入是新快取:mod 在記錄行中說明這一點,在工作階段的統計中計算寫入並繼續。熱 ping 以讀取率讀取上下文,200k 權杖約 $0.05,所以空閒一天的 ping 成本約 $1.40。

快取消失時停止。 這適用於有結束時間的時間段,不適用於 always。溫 ping 讀取上下文,僅寫入自己的幾個權杖。當 ping 什麼都沒讀到,或寫入讀取內容的十分之一或更多時,快取已經消失,ping 本身支付了寫入,所以 mod 停止並顯示原因。當 API 用錯誤回覆分叉時也會停止(行命名其狀態和類型),以及當分叉在回覆前被切斷時。當引擎沒有分叉時,如在恢復流程的第一個回覆前,時間段不會停止:它等待下一個回覆,其回合重新啟用 ping,行說明快取的保持時間。沒有文字的回覆仍然讀取了快取,所以它計為 ping。在 always 下這樣的失敗僅在那個回合停止迴圈:下一個回合重新啟動它,所以工作階段在執行時從不持有開關。

在付費冷寫入後啟用自己。 當一個回合重寫大小超過 20k 權杖的上下文的至少一半時,mod 計算該冷寫入並啟用一個 6 小時的時間段,除非已經啟用了更長的。在 always 下不啟用 6 小時的時間段,因為無限迴圈已經保持該快取。

顯示狀態。 /cache-status 列印模型(熱或冷)、上下文大小、冷價格、時間段、平衡點和本工作階段的冷寫入。

發送給冷快取的訊息永遠不會被停止或延遲。快取已過期的恢復工作階段會獲得一行,顯示其第一條訊息的價格。Claude Code 從記錄的最後回覆對快取進行日期計算,ping 從不寫入,所以在 mod 為工作階段保持的最後 ping 在 1 小時內讀取快取時,該行被省略。

恢復後發送保溫訊息。 恢復流程無法在其自己的第一個回覆前分叉:$.model.fork 回覆 nothing-to-fork(使用無頭 claude --resume 測量)。所以沒有 ping 可以發出,如果你 30 分鐘後關閉並重新開啟工作階段,除非你寫了什麼,否則會在 1 小時處失去快取。當交互式工作階段在時間段或 always 執行時被恢復,其上下文 50k 權杖或更多且其快取仍持有時,mod 因此在恢復後 3 秒鐘發送一條訊息,透過其自己的 /cache-warm:send 命令:

/cache-warm:send 此訊息由 cache-warm 外掛發送,而非個人。工作階段已恢復,恢復的工作階段僅在回覆後才能保持其提示快取熱狀態。不要執行工具或繼續工作。用單個單詞回覆:warm

這是一個真實的回合:它讀取快取,模型回答一個單詞,兩者保留在對話中,其結束再次啟用 ping。當快取已經消失時(你的下一條訊息無論如何都會支付相同的寫入)、在 -p 執行中或你在這三秒內發送訊息時,不會發送任何東西。其他 mod 像任何其他回合一樣看待它:task-poke 在工作打開時可能在其後發送繼續提示,desk-notify 顯示其回合結束通知。在 2.1.283 上測量在一個 116k 權杖的恢復交互式工作階段中:訊息在恢復時發出,模型回答 warm,下一個 ping 分叉並讀取 117k 權杖。恢復可以自己打破前置詞的一部分(該工作階段重寫了 116k 中的 42k;另一個,在最後一個 ping 後 40 分鐘恢復,4k);保溫回合在恢復時支付該寫入,而不是你的第一條訊息。三秒是從兩個啟動鉤點的較晚者計算的,classic.SessionStart 和 session.start,因為它們沒有固定順序:載入了市場上每個 mod 時,session.start 在 classic.SessionStart 後四秒進行(在 2.1.285 上測量),僅由第一個啟動的等待發現了一個尚未啟動的工作階段並未發送任何東西。

命令

/cache-warm               保持溫度 6 小時
/cache-warm 90m           保持溫度你自己的時間段(也可以 2h30m)
/cache-warm always        保持快取溫度無限制,在每個專案的每個工作階段中
/cache-warm 6h every 2m   每兩分鐘 ping 一次;測試設定,下限 1 分鐘,在此時間段後被遺忘
/cache-warm status        狀態行文字
/cache-warm off           停止,忘記時間段,並關閉 always
/cache-status             卡片
/cache-warm:send <text>   mod 在恢復後發送的保溫訊息;其正文僅為文字

它的顯示內容

時間段啟用或停止後提示下的狀態行:

cache-warm: 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
cache-warm: stopped: the ping read 0 and wrote 180k tokens ($3.60), the cache was already gone

last 部分命名最後讀取快取的要求,ping 或主迴圈回合,後者較晚:last ping read 200k $0.05 或 last turn read 250k $0.07,從不兩者。回合的數字覆蓋整個回合,所有要求和。括號中的時間是該答覆何時到達,本機時間;來自較早日期的那一天帶有其日期和月份,如 (22 Sep 23:10)。記錄保存在 $.store 下的工作階段 ID 中,所以 /reload-plugins 或更新立即再次顯示它。

當時間段執行時,交互式工作階段每分鐘重繪該行,因此剩餘時間和下一個 ping 的時間在回合和 ping 之間倒計時,last 部分保留在行上。ping 前的最後一分鐘讀取 ping now,因為分鐘四捨五入且行每分鐘繪製一次;當 ping 的分叉在外面時,它讀取 pinging…。當引擎沒有分叉時,行說明這個而不是倒計時:

cache-warm: always · no ping before the next reply · cache holds until 19:16

每個 ping 嘗試一個流條目當邊欄開啟時。它也保存在邊欄的日誌檔案中(~/.claude/sidebar/<project>-<date>.log),所以你以後可以回讀 ping 是否發出以及它做了什麼;當邊欄關閉時相同的文字是記錄行:

ping sent · read 901k · wrote 0 · $0.45
ping found the cache gone · read 0 · wrote 180k · $3.60
ping not sent: the conversation has no reply to fork yet; the ping waits for the next reply, and the cache holds until 19:16
ping failed: the ping failed, the API answered 529 (overloaded)
keep-warm message sent: the session was resumed and its cache holds until 19:16; a resumed session pings only after a reply

當 sidebar 開啟時,狀態行移到那裡作為一個 cache window 部分,為工作階段停留並在每次更改時重寫,狀態行保持清晰。僅剩餘時間(或 always)被著色:當時間段比一個 ping 週期更早結束時為黃色,保持時為綠色,等待第一個回合時為淺色。ping 詳情在其後為淺色,停止的時間段顯示其 stopped: 前面為紅色,原因為預設顏色。沒有邊欄時,狀態行如上所述繪製。

停止原因停留一個回合。在下一個回合處,部分顯示空閒行:淺色,除了已支付的 N cold writes paid $X 為黃色。這樣窗格顯示現在的測量,而不是已結束時間段的最後一句。原因保留在記錄中,當沒有時間段執行時狀態行為空:

cache window
off · 2 cold writes paid $6.30 · context 315k tokens

當時間段用盡時,只要已結束的時間和相同的 ping 週期,你的下一條訊息會再次啟用它,空閒行在等待時說明這一點:

cache window
off · 6h again at your next message · 1 cold write paid $4.01 · context 201k tokens

cache-warm: the 6h window ran out; this message arms another one. /cache-warm off stops it.

僅當時間段用盡時才回來。ping 停止的時間段不:快取已在那裡消失,下一條訊息的冷寫入啟用其自己的 6 小時視窗。/cache-warm off 忘記等待回來的時間段。

在時間段下,部分保持第二條淺色行:最後的記錄行、縮短的,具有黃色冷寫入成本。時間段行說明快取的保持時間;第二行說明 mod 最後做了什麼:

cache window
6h left · ping in 50m
cold write 201k tokens paid ($4.01)

/cache-status 的卡片:

claude-fable-5-1
state       warm, 42m left
context     200,502 tokens
cold cost   $4.01 to re-write it (warm turn $0.05)
keep warm   on, 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
break-even  up to 80 pings at the read rate cost one cold write, about 2d 18h of idle at one ping per 50m
session     1 cold write paid, $4.01

一條記錄行,未發送給模型,當冷寫入啟用時間段或恢復開始冷時。當觀察器關閉(無時間段、always 關閉)時,行轉到邊欄的流,不向記錄寫入;關閉的邊欄刪除它。正在執行的觀察器寫入的行是不變的。

價格

表格 i

安裝

請先查看作者 README,確認 marketplace 與外掛名稱;指令可能隨儲存庫結構而變動。

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install cache-warm
原文 / README

cache-warm

Claude Code keeps your conversation in a prompt cache for one hour. Step away for longer, and the next message re-writes the whole context at the cache-write rate: on Claude Fable 5.1 that is $4.00 for 200k tokens, against $0.05 to read the same tokens from a warm cache. This mod keeps the cache warm for a window you choose, and shows you what a cold cache would cost.

It follows the behaviour of the cache-tax mod by Karan Bansal (karanb192/claude-code-mods), without its send guard. The code is new.

What it does

Keeps the cache warm. /cache-warm arms a six-hour window. Inside it, 50 minutes after the main loop's last model request, the mod sends one tool-less $.model.fork over the session's own transcript. The server answers it from the cache, and that refreshes the hour. Every new request pushes the ping later, so a session you are actively using sends no ping at all. The first request of a session, and the first after /clear or a compaction, sets the ping as well, so even a tool call that runs past the hour inside the first turn gets its ping in the middle of the turn (measured on 2.1.281: the ping went out while sleep 110 ran, and the turn ended normally).

Runs with no end under always. /cache-warm always is not a window: the ping goes out every 50 minutes for as long as the session lives, and only /cache-warm off ends it. The switch is one global key in the mod's own $.store, so every later session of every project starts the same loop at its start and after /clear.

  • Each turn's end keeps the last request's time in $.store under the session's id, and each ping keeps its own read there too. A module loaded into a running conversation (/reload-plugins, an update) therefore pings on time from the later of the two; a session kept warm by pings alone has a last turn older than its cache. When both are more than an hour old, the cache is gone, and the loop waits for the next turn instead of paying for a cold ping.
  • The session's transcript is not used for this, because /reload-plugins writes a line of its own there and the file's last write would then look like a request (measured: a reload 90 seconds after the last turn would have set the ping 90 seconds late). Only a session with no time kept yet, one that ran an older version, reads the last write of its transcript (~/.claude/projects/<directory>/<session id>.jsonl, under CLAUDE_CONFIG_DIR when it is set) once. The transcript is looked up under the directory the session started in, never the one a shell cd moved to, and one it cannot find is named in a transcript line.
  • A ping that finds the cache gone does not end this loop. The write that ping paid for is the new cache: the mod says so in a transcript line, counts the write in the session's tally and keeps going. A warm ping reads the context at the read rate, about $0.05 for 200k tokens, so an idle day of pings costs about $1.40.

Stops when the cache is gone. This applies to a window with an end, not to always. A warm ping reads the context and writes only its own few tokens. When a ping reads nothing, or writes a tenth of what it read or more, the cache was already gone and the ping itself paid for the write, so the mod stops and shows why. It also stops when the API answers the fork with an error (the line names its status and kind), and when the fork is cut before it replies. When the engine has nothing to fork, as in a resumed process before its first reply, the window does not stop: it waits for the next reply, whose turn arms the ping again, and the line says until when the cache holds. A reply without text still read the cache, so it counts as a ping. Under always such a failure stops the loop for that turn only: the next turn starts it again, so the session never holds the switch while nothing runs.

Arms itself after a paid cold write. When a turn re-writes at least half of a context larger than 20k tokens, the mod counts that cold write and arms a six-hour window, unless a longer one is already armed. Under always no six-hour window is armed, because the endless loop already keeps that cache.

Shows the state. /cache-status prints the model, warm or cold, the context size, the cold price, the window, the break-even and this session's cold writes.

A message you send to a cold cache is never stopped or delayed. A resumed session whose cache has lapsed gets one line with the price of its first message. Claude Code dates the cache from the transcript's last reply, which a ping never writes, so that line is left out while the last ping the mod kept for the session read the cache within the hour.

Sends a keep-warm message after a resume. A resumed process cannot fork before its own first reply: $.model.fork answers nothing-to-fork (measured with a headless claude --resume). So no ping can go out, and a session you closed and opened again 30 minutes later would lose its cache at the hour unless you wrote something. When an interactive session is resumed while a window or always runs, its context is 50k tokens or more and its cache still holds, the mod therefore sends one message three seconds after the resume, through its own /cache-warm:send command:

/cache-warm:send This message was sent by the cache-warm plugin, not by the person. The session was resumed, and a resumed session can keep its prompt cache warm only after a reply. Do not run a tool or continue a task. Reply with the single word: warm

This is a real turn: it reads the cache, the model answers one word, the pair stays in the conversation, and its end arms the ping again. Nothing is sent when the cache is already gone (your next message pays the same write anyway), in a -p run, or when you sent a message within those three seconds. Other mods see it like any other turn: task-poke may send its continue prompt after it while tasks are open, and desk-notify shows its turn-end notification. Measured on 2.1.283 in a resumed interactive session of 116k tokens: the message went out at the resume, the model answered warm, and the next ping forked and read 117k tokens. A resume can break part of the prefix on its own (that session re-wrote 42k of the 116k; another, resumed 40 minutes after its last ping, 4k); the keep-warm turn pays that write at the resume instead of your first message. The three seconds count from the later of the two start hooks, classic.SessionStart and session.start, because they settle in no fixed order: with every mod of the marketplace loaded, session.start settled four seconds after classic.SessionStart (measured on 2.1.285), and a wait started by the first alone found a session that had not started and sent nothing.

Commands

/cache-warm               keep warm for six hours
/cache-warm 90m           keep warm for a window of your own (also 2h30m)
/cache-warm always        keep the cache warm with no end, in every session of every project
/cache-warm 6h every 2m   ping every two minutes; a test setting, floor 1m, forgotten after this window
/cache-warm status        the status line text
/cache-warm off           stop, forget the window, and turn always off
/cache-status             the card
/cache-warm:send <text>   the keep-warm message the mod sends after a resume; its body is the text alone

What it shows

A status line under the prompt while a window is armed or after a stop:

cache-warm: 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
cache-warm: stopped: the ping read 0 and wrote 180k tokens ($3.60), the cache was already gone

The last part names the last request that read the cache, a ping or a main-loop turn, whichever came later: last ping read 200k $0.05 or last turn read 250k $0.07, never both. A turn's figures cover the whole turn, every request summed. The time in brackets is when that answer came, in local time; one from an earlier day carries its day and month, as (22 Sep 23:10). The record is kept in $.store under the session's id, so /reload-plugins or an update shows it again at once.

While a window runs, an interactive session redraws the line every minute, so the time left and the time to the next ping count down between turns and pings, and the last part stays on the line. The last minute before a ping reads ping now, because minutes are rounded and the line is drawn once a minute; while the ping's fork is out it reads pinging…. When the engine has nothing to fork, the line says so instead of counting down:

cache-warm: always · no ping before the next reply · cache holds until 19:16

One stream entry per ping attempt while the sidebar is open. It is also kept in the sidebar's log file (~/.claude/sidebar/<project>-<date>.log), so you can read back later whether a ping went out and what it did; with the sidebar closed the same text is a transcript line:

ping sent · read 901k · wrote 0 · $0.45
ping found the cache gone · read 0 · wrote 180k · $3.60
ping not sent: the conversation has no reply to fork yet; the ping waits for the next reply, and the cache holds until 19:16
ping failed: the ping failed, the API answered 529 (overloaded)
keep-warm message sent: the session was resumed and its cache holds until 19:16; a resumed session pings only after a reply

With the sidebar open, the status line moves there as a cache window section that stays for the session and is rewritten at every change, and the status line stays clear. Only the time left (or always) is coloured: yellow when the window ends sooner than one ping period, green while it holds, faint while it waits for the first turn. The ping details after it are faint, and a stopped window shows its stopped: front in red with the reason in the default colour. Without the sidebar, the status line is drawn as above.

A stop reason stays for one turn. At the next turn the section shows the idle line instead: faint, except for a paid N cold writes paid $X, which is yellow. That way the pane shows a measurement of now, not the last sentence of a window that ended. The reason stays in the transcript, and the status line is empty while no window runs:

cache window
off · 2 cold writes paid $6.30 · context 315k tokens

A window that runs out of time is armed again by your next message, as long as the one that ended and with the same ping period, and the idle line says so while it waits:

cache window
off · 6h again at your next message · 1 cold write paid $4.01 · context 201k tokens

cache-warm: the 6h window ran out; this message arms another one. /cache-warm off stops it.

Only a window that ran out of time comes back. One that a ping stopped does not: the cache is already gone there, and the cold write of your next message arms its own 6h window. /cache-warm off forgets a window waiting to come back.

Under the window the section holds a second, faint line: the last transcript line, shortened, with the cost of a cold write in yellow. The window line says how long the cache is kept; the second line says what the mod did last:

cache window
6h left · ping in 50m
cold write 201k tokens paid ($4.01)

The card of /cache-status:

claude-fable-5-1
state       warm, 42m left
context     200,502 tokens
cold cost   $4.01 to re-write it (warm turn $0.05)
keep warm   on, 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
break-even  up to 80 pings at the read rate cost one cold write, about 2d 18h of idle at one ping per 50m
session     1 cold write paid, $4.01

One transcript line, not sent to the model, when a cold write arms the window or a resume starts cold. While the watcher is off (no window, always off), the line goes to the sidebar's stream instead and writes nothing to the transcript; a closed sidebar drops it. The line a running watcher writes is unchanged.

Prices

The table in hooks/pricing.ts holds the cache-read, 1-hour cache-write and output rates of every model on the Anthropic pricing page, read in September 2026. A model id takes the first family it contains, so claude-opus-4-1 is priced as Opus 4.1 ($1.50 / $30 / $75) and claude-opus-4-8 as Opus 4.8 ($0.50 / $10 / $25). claude-opus-5-5 also contains opus-5, so its own row comes first: Opus 5.5 is $0.20 / $8 / $20, cheaper than Opus 5. Sonnet 5.5 has its own row at the Sonnet 5 rates, $0.20 / $4 / $10. A ping is priced in full: the cache read, its cache write, its uncached input at the base rate (half the 1-hour write rate) and its output. An unknown model shows n/a.

Fast mode bills Opus 5.5, Opus 5 and Opus 4.8 at their own base rates ($8 and $10 input), with the cache multipliers applied on top. The mod prices Opus 5.5 at $0.40 / $16 / $40 and Opus 5 and 4.8 at $1 / $20 / $50 while the fastMode setting, which /fast writes, is on. It reads the settings at the session's start and at the end of each main-loop turn, so a /fast counts from the next turn. With fastModePerSessionOptIn set to true, every session starts with fast mode off, so the standard rates apply. Any other model keeps its standard rates, and the card mentions fast mode only when the rates changed:

claude-opus-5-5 · fast mode rates (the fastMode setting)

On a subscription the dollars are a yardstick, not your bill. How a cache read counts against the 5-hour and weekly limits is not documented.

Install

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install cache-warm@kilimcininkoroglu-mods

Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.

To load it from a local checkout for one session:

claude --plugin-dir plugins/cache-warm

After installing

  1. Restart Claude Code.
  2. Disable every other keep-warm mod, for example claude plugin disable cache-tax@claude-code-mods. Two keep-warm mods in one session send two pings per idle stretch.
  3. Check once that a ping reads your cache, as "Prove it on your own session" below describes.
  4. To keep the cache warm with no end, run /cache-warm always once. The switch is global: every later session of every project starts the loop by itself, and /cache-warm off ends it for good. Without it, a window is armed only by /cache-warm or after a paid cold write.

What it can reach

Validated with claude plugin validate on Claude Code 2.1.288:

❯ ./register.ts hooks: session.start, classic.SessionStart, prompt.submit, command.run{command=cache-warm}, command.run{command=cache-status}, turn.step, turn.complete, session.compact
❯ ./register.ts calls: $.clock.after (via arm, scheduleKeepWarm), $.clock.every, $.clock.now, $.command.register (via registerCommands), $.command.run (via keepWarmAfterResume), $.env.get (via seedFromTranscript), $.fs.exists (via seedFromTranscript), $.fs.stat (via seedFromTranscript), $.model.fork (via forkPing), $.prompt.submit (via keepWarmAfterResume), $.session.id, $.session.model, $.session.root (via seedFromTranscript), $.session.usage, $.settings.read (via readFast), $.sidebar.set (via logEvent, toSidebar, toStream), $.store.delete (via prune, pruneRequests, startEndless, startWindow, stop), $.store.get, $.store.keys (via prune, pruneRequests), $.store.set (via afterTurn, keepLastRead, startWindow, warmCommand), $.ui.log (via logEvent, seedFromTranscript, toStream), $.ui.status (via showStatusAt)
❯ ./register.ts env writes: nothing
❯ ./register.ts env reads: CLAUDE_CONFIG_DIR, HOME

Reach L2: it drives Claude.

1. Reads:    the time of each main-loop model request; the token counts and model id of each turn and of each ping; the live context size; the origin of each message, to arm a window again; the resume fields Claude Code computes for settings hooks; the session id and model; the last write time of the session's own transcript file, once when the module loads into a running conversation; the fastMode and fastModePerSessionOptIn settings, at the session's start and at each turn's end; its own $.store. It never reads a prompt's text, a file's content or a tool result.
2. Runs:     one $.model.fork per idle stretch while a window or the always loop runs, 50 minutes after the last request unless the test setting is used (floor 1 minute); never while off; a ping that found the cache gone ends a window with an end, and under always the loop carries on; after a resume of an interactive session whose cache still holds, one keep-warm message through /cache-warm:send (a plugin prompt when the engine refuses the command), which is a real turn
3. Sends:    the fork, an API request over the session's own transcript with a fixed one-line prompt, and after a resume the fixed keep-warm message as a turn of the conversation
4. Persists: in $.store, the window end, the ping period and the last main-loop request's time and the last ping or turn read (tokens, cost, time) under this session's id, and the global always switch, which the endless loop needs no window key beside; this session's ended window is deleted at stop and at its next start, another session's window one week after it ended, another session's request time and last read once they are an hour old; the cold-write tally lives in memory and ends with the session
5. Hostile input: the only text it parses is the argument of /cache-warm, matched against a duration pattern and three words; the fork's prompt is a constant, so nothing crafted can reach it

Prove it on your own session

The mock-clock tests prove the timer and the scoring, not that a fork reads the main cache. One ping proves that. In a warm session:

> Reply with one word: ready
> /cache-warm 1h every 1m

After a minute the status line should read last ping read <close to your context> $.... A stopped: the ping read ... line means the fork did not share the cache, and the mod has already stopped. /cache-warm off ends the test.

Limits

  • The 50-minute ping assumes the 1-hour cache tier, which the main conversation uses.
  • A warm ping only proves the cache was warm at that moment. A model or effort switch, an edited CLAUDE.md or a changed tool list breaks the prefix no matter the time, and your next message pays.
  • A ping's output cannot be capped; a model at high effort may think before it answers. The status line prices what the ping really billed.
  • The resume logic is covered by hook tests that raise classic.SessionStart with the resume fields, and by a live check of a resumed interactive session in tmux. Whether a resume keeps the whole prefix is out of the mod's hands: it re-wrote 42k of 116k tokens in one measured session and 4k in another.
  • The cold-write tally is per session and lives in memory. /clear empties it.
  • Fast mode is read from the fastMode setting, the saved preference, not from the request. The mod does not see Claude Code fall back to standard speed within a session (a fast mode rate-limit cooldown, usage credits that ran out, an organization that turned fast mode off). Those turns bill standard rates while the mod prices them as fast.
  • Whether a ping, a $.model.fork, runs at fast speed while the session does has not been measured; the mod prices it at the session's rates.

Development

make install     # eslint, typescript-eslint, typescript
make lint        # complexity limit 10, the build fails above it
make typecheck   # needs .claude/types/ from /plugin-types
make validate
make test        # claude plugin test

更多類似作品