krika2810/keep-cache-warm
cache-keepalive
1時間のプロンプトキャッシュを温かく保つ Claude Code の関数フックプラグイン。約55分のアイドル後に $.model.fork でキャッシュ済みプレフィックスを再送し、次のターンで全体を書き直すのを避ける。ping 上限と安全策も備える。
この mod について
cache-keepalive は Claude Code 2.1.289 向けの Claude Code mod(関数フックプラグイン)で、席を外している間も1時間のプロンプトキャッシュを冷やさない。
仕組み:
turn.stepにフックして、メインスレッドのモデルリクエストが送信された時刻を記録する。レスポンスがキャッシュの読み取りまたは書き込みを報告したリクエストだけを数える。- メインスレッドのターンが終わると、最後のリクエスト送信からデフォルト55分後に終わるカウントダウンを始める。
- 新しいターンが始まらなければ、短い「reply ok」プロンプトを付けて
$.model.forkを呼ぶ。メインスレッドの最後のリクエストをそのまま再送するため、API はプレフィックス全体をキャッシュから返し、全体を書き直さずにエントリの60分の有効期間を延長できる。 - fork は会話には含まれない。ツール呼び出しは拒否され、プロンプトと応答もキャッシュされない。
安全策:ping の上限(maxPings、デフォルト6、0 = 無制限)、59+分での遅いタイマーのスキップ、キャッシュヒットがない場合の toast 付き停止、サブエージェントのリクエストの無視、ターン中の ping 禁止、失敗したリクエストを保温として数えない処理。
コスト:ping はキャッシュ読み取り1回と、fork のキャッシュされない入出力(thinking を含む)として課金される。多くのモデルではキャッシュ読み取りは基礎入力の0.1倍で、一部のモデルではさらに安い。ping がなければ、有効期限後の最初のターンで会話全体を基礎入力の2倍で書き直す。そのため大きな会話では通常 keep-alive の方が安いが、小さな会話や ping 回数が多い場合は高くなることがある。
使い方:/cache-keepalive(状態)、off、on、now。状態行には cache keep-alive: armed (n/6) と表示される。
/config または pluginConfigs で設定できるオプションは、idleMinutes(デフォルト55、1-58に制限)と maxPings(デフォルト6)。
読み込み:claude --plugin-dir ./keep-cache-warm。検証/テストは claude plugin validate . と claude plugin test .。MIT ライセンス。
インストール
まず作者の README で marketplace とプラグイン名を確認してください。コマンドはリポジトリの構成によって変わる場合があります。
claude plugin marketplace add krika2810/keep-cache-warm claude plugin install cache-keepalive
原文 / README
keep-cache-warm
cache-keepalive: a Claude Code mod (function-hooks plugin) that keeps the 1-hour prompt cache from going cold while you step away.
Compatibility: Claude Code mods and function hooks are early access, and their API can change between Claude Code releases. This repository was tested against Claude Code 2.1.289. After upgrading, load the mod once so Claude Code rewrites its type definitions in
.claude-plugin/types/. Then re-runclaude plugin validate .andclaude plugin test ..
How it works
- A cache entry's 1-hour lifetime starts when the request that read or wrote it is sent, not when the response or the turn finishes. A 10-minute response leaves only 50 minutes. So the mod hooks
turn.stepand records when each main-thread model request is sent. It counts only requests whose response reports cache reads or writes. - When a main-thread turn finishes, the mod starts a countdown. It ends 55 minutes (the default) after the turn's last request was sent.
- If no new turn starts before then, the mod calls
$.model.forkwith a tiny "replyok" prompt. The fork resends the main thread's last request exactly (same model, system prompt, tool definitions and messages) with that prompt appended. The API serves the whole prefix from the cache, and the read restarts the entry's 60-minute lifetime without paying for another full cache write. - The fork is not a turn of your conversation. Its tools are declared, so the prefix matches the cache, but every tool call it attempts is denied. Its own prompt and reply are never cached. The reply goes only to the mod, which discards it.
- Each successful ping restarts the countdown from when the ping was sent. A new turn from you resets everything.
Safeguards
- Ping cap: it stops after
maxPingspings in a row (default 6, about 5.5 hours) so an abandoned session doesn't keep spending. Set it to0for no limit. - Late timer: if a ping would go out 59 or more minutes after the last request that touched the cache was sent, it is skipped. This covers a laptop that slept, or a turn whose last request was sent near the deadline. Rebuilding a cache that has already expired costs a full cache write.
- No cache hit: if a ping reads 0 tokens from the cache and writes some, the mod stops until your next turn and shows a toast. This happens when the session uses the 5-minute TTL or the model was switched.
- Subagents: their requests and turns are ignored. They send their own prefix and don't warm the main thread's.
- No ping mid-turn: no ping is sent while a turn is running.
- Failed requests: a request that got no response doesn't count as warming the cache.
Cost
A ping is not free. It is billed as one cache read of the conversation, plus the fork's uncached input (its short prompt) and its output, including any thinking. A cache read costs 0.1× the base input price on most models, and less on some: 0.05× on Claude Opus 5.5 and 0.025× on Claude Fable 5.1.
Without the ping, the first turn after the cache expires writes the whole conversation to the cache again. With the 1-hour TTL, that write costs 2× the base input price. For a large conversation, a keep-alive read is usually much cheaper than that rewrite. The fork's own input and output costs don't shrink with the conversation, though, so on a small conversation, or after many pings in a row, keeping the cache warm can cost more than letting it expire. That is why maxPings caps the run.
Usage
/cache-keepalive # status: minutes until the next ping, pings sent
/cache-keepalive off # stop for this session
/cache-keepalive on # resume
/cache-keepalive now # ping immediately
The status line shows cache keep-alive: armed (n/6) while a countdown is running.
Options (/config or pluginConfigs in settings)
| Field | Default | Meaning |
| --- | --- | --- |
| idleMinutes | 55 | Minutes after the last cache-touching request before a ping (clamped to 1–58, so a ping is always due before the 59-minute cutoff) |
| maxPings | 6 | Pings in a row before giving up; 0 means no limit |
Load it
claude --plugin-dir ./keep-cache-warm
Check it
claude plugin validate .
claude plugin test .

