ClaudeMods
☰
ZH-TW
● 0 人在線上 · 瀏覽 0 次
贊助提交作品
GitHub 儲存庫 · 發布者 krika2810

cache-keepalive

一個 Claude Code 函式鉤子外掛,用來保持 1 小時的提示快取溫熱:閒置約 55 分鐘後,透過 $.model.fork 重播快取前綴,讓下一輪避免完整快取寫入,並設有 ping 上限與安全措施。

已翻譯

關於這個 mod

cache-keepalive 是一個 Claude Code mod(函式鉤子外掛),適用於 Claude Code 2.1.289,讓你暫時離開時,1 小時的提示快取不會冷掉。

運作方式:

  • 掛接 turn.step,記錄每次主要執行緒模型請求的送出時間,只統計回應回報讀取或寫入快取的請求。
  • 主要執行緒工作階段結束時啟動倒數,預設在該工作階段最後一次請求送出後 55 分鐘結束。
  • 如果沒有新的工作階段開始,就呼叫 $.model.fork,送出很短的「reply ok」提示,完全重送主要執行緒的上一個請求,讓 API 從快取提供整個前綴,並在不完整寫入快取的情況下重新開始該項目的 60 分鐘生命週期。
  • 這個 fork 不屬於你的對話;它的工具呼叫會被拒絕,提示與回覆也永遠不會被快取。

安全措施:ping 上限(maxPings,預設 6,0 = 不限)、在 59+ 分鐘時跳過太晚的計時器、沒有快取命中時停止並顯示 toast、忽略子代理請求、不在工作階段中途 ping,以及失敗的請求不算預熱。

成本:一次 ping 按一次快取讀取,加上 fork 未快取的輸入/輸出(包括 thinking)計費。大多數模型的快取讀取費用是基礎輸入的 0.1 倍(部分模型更低)。沒有 ping 時,過期後的第一輪會以基礎輸入的 2 倍重寫整段對話,因此大型對話通常使用 keep-alive 比較便宜;但小型對話或 ping 次數很多時,成本可能更高。

用法:/cache-keepalive(狀態)、off、on、now。狀態列顯示 cache keep-alive: armed (n/6)。

透過 /config 或 pluginConfigs 設定選項:idleMinutes(預設 55,限制在 1-58)和 maxPings(預設 6)。

載入:claude --plugin-dir ./keep-cache-warm。驗證/測試:claude plugin validate . 和 claude plugin test .。授權採用 MIT。

安裝

請先查看作者 README,確認 marketplace 與外掛名稱;指令可能隨儲存庫結構而變動。

claude plugin marketplace add krika2810/keep-cache-warm
claude plugin install cache-keepalive
原文 / README

keep-cache-warm

cache-keepalive: a Claude Code mod (function-hooks plugin) that keeps the 1-hour prompt cache from going cold while you step away.

Compatibility: Claude Code mods and function hooks are early access, and their API can change between Claude Code releases. This repository was tested against Claude Code 2.1.289. After upgrading, load the mod once so Claude Code rewrites its type definitions in .claude-plugin/types/. Then re-run claude plugin validate . and claude plugin test ..

How it works

  • A cache entry's 1-hour lifetime starts when the request that read or wrote it is sent, not when the response or the turn finishes. A 10-minute response leaves only 50 minutes. So the mod hooks turn.step and records when each main-thread model request is sent. It counts only requests whose response reports cache reads or writes.
  • When a main-thread turn finishes, the mod starts a countdown. It ends 55 minutes (the default) after the turn's last request was sent.
  • If no new turn starts before then, the mod calls $.model.fork with a tiny "reply ok" prompt. The fork resends the main thread's last request exactly (same model, system prompt, tool definitions and messages) with that prompt appended. The API serves the whole prefix from the cache, and the read restarts the entry's 60-minute lifetime without paying for another full cache write.
  • The fork is not a turn of your conversation. Its tools are declared, so the prefix matches the cache, but every tool call it attempts is denied. Its own prompt and reply are never cached. The reply goes only to the mod, which discards it.
  • Each successful ping restarts the countdown from when the ping was sent. A new turn from you resets everything.

Safeguards

  • Ping cap: it stops after maxPings pings in a row (default 6, about 5.5 hours) so an abandoned session doesn't keep spending. Set it to 0 for no limit.
  • Late timer: if a ping would go out 59 or more minutes after the last request that touched the cache was sent, it is skipped. This covers a laptop that slept, or a turn whose last request was sent near the deadline. Rebuilding a cache that has already expired costs a full cache write.
  • No cache hit: if a ping reads 0 tokens from the cache and writes some, the mod stops until your next turn and shows a toast. This happens when the session uses the 5-minute TTL or the model was switched.
  • Subagents: their requests and turns are ignored. They send their own prefix and don't warm the main thread's.
  • No ping mid-turn: no ping is sent while a turn is running.
  • Failed requests: a request that got no response doesn't count as warming the cache.

Cost

A ping is not free. It is billed as one cache read of the conversation, plus the fork's uncached input (its short prompt) and its output, including any thinking. A cache read costs 0.1× the base input price on most models, and less on some: 0.05× on Claude Opus 5.5 and 0.025× on Claude Fable 5.1.

Without the ping, the first turn after the cache expires writes the whole conversation to the cache again. With the 1-hour TTL, that write costs 2× the base input price. For a large conversation, a keep-alive read is usually much cheaper than that rewrite. The fork's own input and output costs don't shrink with the conversation, though, so on a small conversation, or after many pings in a row, keeping the cache warm can cost more than letting it expire. That is why maxPings caps the run.

Usage

/cache-keepalive           # status: minutes until the next ping, pings sent
/cache-keepalive off       # stop for this session
/cache-keepalive on        # resume
/cache-keepalive now       # ping immediately

The status line shows cache keep-alive: armed (n/6) while a countdown is running.

Options (/config or pluginConfigs in settings)

| Field | Default | Meaning | | --- | --- | --- | | idleMinutes | 55 | Minutes after the last cache-touching request before a ping (clamped to 1–58, so a ping is always due before the 59-minute cutoff) | | maxPings | 6 | Pings in a row before giving up; 0 means no limit |

Load it

claude --plugin-dir ./keep-cache-warm

Check it

claude plugin validate .
claude plugin test .

License

MIT

更多類似作品