krika2810/keep-cache-warm
cache-keepalive
1시간 프롬프트 캐시를 따뜻하게 유지하는 Claude Code 함수 훅 플러그인입니다. 약 55분 유휴 후 $.model.fork로 캐시된 접두사를 다시 보내 다음 턴의 전체 캐시 쓰기를 피하며, ping 제한과 안전장치를 제공합니다.
이 mod 소개
cache-keepalive는 Claude Code 2.1.289용 Claude Code 모드(함수 훅 플러그인)로, 자리를 비운 동안에도 1시간 프롬프트 캐시가 식지 않게 합니다.
작동 방식:
turn.step에 훅을 걸어 주 스레드 모델 요청이 전송된 시각을 기록합니다. 응답에서 캐시 읽기나 쓰기가 보고된 요청만 셉니다.- 주 스레드 턴이 끝나면 마지막 요청이 전송된 시각에서 기본 55분 뒤에 끝나는 카운트다운을 시작합니다.
- 새 턴이 시작되지 않으면 짧은 “reply ok” 프롬프트로
$.model.fork를 호출하고 주 스레드의 마지막 요청을 그대로 다시 보냅니다. 그러면 API가 전체 접두사를 캐시에서 제공하고, 전체 캐시 쓰기 없이 항목의 60분 수명을 다시 시작합니다. - fork는 대화에 포함되지 않습니다. 도구 호출은 거부되고 프롬프트와 답변도 캐시되지 않습니다.
안전장치: ping 상한(maxPings, 기본 6, 0 = 무제한), 59+분에 도달한 늦은 타이머 건너뛰기, 캐시 적중이 없으면 toast와 함께 중지, 서브에이전트 요청 무시, 턴 중간 ping 금지, 실패한 요청은 예열로 계산하지 않기.
비용: ping은 캐시 읽기 1회와 fork의 캐시되지 않은 입출력(thinking 포함)으로 청구됩니다. 대부분의 모델에서 캐시 읽기는 기본 입력의 0.1배이며 일부는 더 저렴합니다. ping이 없으면 만료 후 첫 턴에서 전체 대화를 기본 입력의 2배로 다시 씁니다. 따라서 큰 대화에서는 보통 keep-alive가 더 저렴하지만, 작은 대화나 ping이 여러 번 발생하면 더 비쌀 수 있습니다.
사용법: /cache-keepalive(상태), off, on, now. 상태 줄에는 cache keep-alive: armed (n/6)가 표시됩니다.
/config 또는 pluginConfigs에서 idleMinutes(기본 55, 1-58로 제한)와 maxPings(기본 6)을 설정할 수 있습니다.
로드: claude --plugin-dir ./keep-cache-warm. 검증/테스트: claude plugin validate . 및 claude plugin test .. MIT 라이선스입니다.
설치
먼저 작성자의 README에서 marketplace와 플러그인 이름을 확인하세요. 저장소 구조에 따라 명령어가 달라질 수 있습니다.
claude plugin marketplace add krika2810/keep-cache-warm claude plugin install cache-keepalive
원문 / README
keep-cache-warm
cache-keepalive: a Claude Code mod (function-hooks plugin) that keeps the 1-hour prompt cache from going cold while you step away.
Compatibility: Claude Code mods and function hooks are early access, and their API can change between Claude Code releases. This repository was tested against Claude Code 2.1.289. After upgrading, load the mod once so Claude Code rewrites its type definitions in
.claude-plugin/types/. Then re-runclaude plugin validate .andclaude plugin test ..
How it works
- A cache entry's 1-hour lifetime starts when the request that read or wrote it is sent, not when the response or the turn finishes. A 10-minute response leaves only 50 minutes. So the mod hooks
turn.stepand records when each main-thread model request is sent. It counts only requests whose response reports cache reads or writes. - When a main-thread turn finishes, the mod starts a countdown. It ends 55 minutes (the default) after the turn's last request was sent.
- If no new turn starts before then, the mod calls
$.model.forkwith a tiny "replyok" prompt. The fork resends the main thread's last request exactly (same model, system prompt, tool definitions and messages) with that prompt appended. The API serves the whole prefix from the cache, and the read restarts the entry's 60-minute lifetime without paying for another full cache write. - The fork is not a turn of your conversation. Its tools are declared, so the prefix matches the cache, but every tool call it attempts is denied. Its own prompt and reply are never cached. The reply goes only to the mod, which discards it.
- Each successful ping restarts the countdown from when the ping was sent. A new turn from you resets everything.
Safeguards
- Ping cap: it stops after
maxPingspings in a row (default 6, about 5.5 hours) so an abandoned session doesn't keep spending. Set it to0for no limit. - Late timer: if a ping would go out 59 or more minutes after the last request that touched the cache was sent, it is skipped. This covers a laptop that slept, or a turn whose last request was sent near the deadline. Rebuilding a cache that has already expired costs a full cache write.
- No cache hit: if a ping reads 0 tokens from the cache and writes some, the mod stops until your next turn and shows a toast. This happens when the session uses the 5-minute TTL or the model was switched.
- Subagents: their requests and turns are ignored. They send their own prefix and don't warm the main thread's.
- No ping mid-turn: no ping is sent while a turn is running.
- Failed requests: a request that got no response doesn't count as warming the cache.
Cost
A ping is not free. It is billed as one cache read of the conversation, plus the fork's uncached input (its short prompt) and its output, including any thinking. A cache read costs 0.1× the base input price on most models, and less on some: 0.05× on Claude Opus 5.5 and 0.025× on Claude Fable 5.1.
Without the ping, the first turn after the cache expires writes the whole conversation to the cache again. With the 1-hour TTL, that write costs 2× the base input price. For a large conversation, a keep-alive read is usually much cheaper than that rewrite. The fork's own input and output costs don't shrink with the conversation, though, so on a small conversation, or after many pings in a row, keeping the cache warm can cost more than letting it expire. That is why maxPings caps the run.
Usage
/cache-keepalive # status: minutes until the next ping, pings sent
/cache-keepalive off # stop for this session
/cache-keepalive on # resume
/cache-keepalive now # ping immediately
The status line shows cache keep-alive: armed (n/6) while a countdown is running.
Options (/config or pluginConfigs in settings)
| Field | Default | Meaning |
| --- | --- | --- |
| idleMinutes | 55 | Minutes after the last cache-touching request before a ping (clamped to 1–58, so a ping is always due before the 59-minute cutoff) |
| maxPings | 6 | Pings in a row before giving up; 0 means no limit |
Load it
claude --plugin-dir ./keep-cache-warm
Check it
claude plugin validate .
claude plugin test .

