krika2810/keep-cache-warm
cache-keepalive
一个 Claude Code 函数钩子插件,用来保持 1 小时的提示缓存处于热状态:空闲约 55 分钟后,它通过 $.model.fork 重放缓存前缀,让下一轮避免完整缓存写入,并设有 ping 上限和安全措施。
关于这个 mod
cache-keepalive 是一个 Claude Code mod(函数钩子插件),适用于 Claude Code 2.1.289,可在你暂时离开时让 1 小时的提示缓存保持有效。
工作方式:
- 挂接
turn.step,记录每次主线程模型请求的发送时间,只统计响应报告缓存读取或写入的请求。 - 主线程工作阶段结束时,启动倒计时,默认在该工作阶段最后一次请求发送后 55 分钟结束。
- 如果没有新的工作阶段开始,就调用
$.model.fork,发送一个很短的“reply ok”提示,准确重发主线程的最后一个请求,让 API 从缓存提供整个前缀,并在不进行完整缓存写入的情况下重新开始该条目的 60 分钟生命周期。 - 这个 fork 不属于你的对话;它的工具调用会被拒绝,提示和回复也永远不会被缓存。
安全措施:ping 上限(maxPings,默认 6,0 = 不限)、在 59+ 分钟时跳过延后的计时器、没有缓存命中时停止并显示 toast、忽略子代理请求、不在工作阶段中途 ping,以及失败的请求不算作预热。
成本:一次 ping 按一次缓存读取,加上 fork 未缓存的输入/输出(包括 thinking)计费。在大多数模型上,缓存读取的费用是基础输入的 0.1 倍(部分模型更低)。没有 ping 时,过期后的第一轮会以基础输入的 2 倍重写整个对话,因此对于大型对话,keep-alive 通常更便宜;但对于小型对话或 ping 次数很多时,成本可能更高。
用法:/cache-keepalive(状态)、off、on、now。状态行显示 cache keep-alive: armed (n/6)。
通过 /config 或 pluginConfigs 设置选项:idleMinutes(默认 55,限制在 1-58)和 maxPings(默认 6)。
加载:claude --plugin-dir ./keep-cache-warm。验证/测试:claude plugin validate . 和 claude plugin test .。授权采用 MIT。
安装
请先查看作者 README 确认 marketplace 和插件名称;命令可能随仓库结构改变。
claude plugin marketplace add krika2810/keep-cache-warm claude plugin install cache-keepalive
原文 / README
keep-cache-warm
cache-keepalive: a Claude Code mod (function-hooks plugin) that keeps the 1-hour prompt cache from going cold while you step away.
Compatibility: Claude Code mods and function hooks are early access, and their API can change between Claude Code releases. This repository was tested against Claude Code 2.1.289. After upgrading, load the mod once so Claude Code rewrites its type definitions in
.claude-plugin/types/. Then re-runclaude plugin validate .andclaude plugin test ..
How it works
- A cache entry's 1-hour lifetime starts when the request that read or wrote it is sent, not when the response or the turn finishes. A 10-minute response leaves only 50 minutes. So the mod hooks
turn.stepand records when each main-thread model request is sent. It counts only requests whose response reports cache reads or writes. - When a main-thread turn finishes, the mod starts a countdown. It ends 55 minutes (the default) after the turn's last request was sent.
- If no new turn starts before then, the mod calls
$.model.forkwith a tiny "replyok" prompt. The fork resends the main thread's last request exactly (same model, system prompt, tool definitions and messages) with that prompt appended. The API serves the whole prefix from the cache, and the read restarts the entry's 60-minute lifetime without paying for another full cache write. - The fork is not a turn of your conversation. Its tools are declared, so the prefix matches the cache, but every tool call it attempts is denied. Its own prompt and reply are never cached. The reply goes only to the mod, which discards it.
- Each successful ping restarts the countdown from when the ping was sent. A new turn from you resets everything.
Safeguards
- Ping cap: it stops after
maxPingspings in a row (default 6, about 5.5 hours) so an abandoned session doesn't keep spending. Set it to0for no limit. - Late timer: if a ping would go out 59 or more minutes after the last request that touched the cache was sent, it is skipped. This covers a laptop that slept, or a turn whose last request was sent near the deadline. Rebuilding a cache that has already expired costs a full cache write.
- No cache hit: if a ping reads 0 tokens from the cache and writes some, the mod stops until your next turn and shows a toast. This happens when the session uses the 5-minute TTL or the model was switched.
- Subagents: their requests and turns are ignored. They send their own prefix and don't warm the main thread's.
- No ping mid-turn: no ping is sent while a turn is running.
- Failed requests: a request that got no response doesn't count as warming the cache.
Cost
A ping is not free. It is billed as one cache read of the conversation, plus the fork's uncached input (its short prompt) and its output, including any thinking. A cache read costs 0.1× the base input price on most models, and less on some: 0.05× on Claude Opus 5.5 and 0.025× on Claude Fable 5.1.
Without the ping, the first turn after the cache expires writes the whole conversation to the cache again. With the 1-hour TTL, that write costs 2× the base input price. For a large conversation, a keep-alive read is usually much cheaper than that rewrite. The fork's own input and output costs don't shrink with the conversation, though, so on a small conversation, or after many pings in a row, keeping the cache warm can cost more than letting it expire. That is why maxPings caps the run.
Usage
/cache-keepalive # status: minutes until the next ping, pings sent
/cache-keepalive off # stop for this session
/cache-keepalive on # resume
/cache-keepalive now # ping immediately
The status line shows cache keep-alive: armed (n/6) while a countdown is running.
Options (/config or pluginConfigs in settings)
| Field | Default | Meaning |
| --- | --- | --- |
| idleMinutes | 55 | Minutes after the last cache-touching request before a ping (clamped to 1–58, so a ping is always due before the 59-minute cutoff) |
| maxPings | 6 | Pings in a row before giving up; 0 means no limit |
Load it
claude --plugin-dir ./keep-cache-warm
Check it
claude plugin validate .
claude plugin test .

