KilimcininKorOglu/claude-code-mods/tree/main/plugins/cache-warm
cache-warm
설정한 기간 동안 1시간 프롬프트 캐시를 따뜻하게 유지하거나, always 모드 아래 모든 세션에서 무제한 유지, 유휴 각 구간당 캐시 공유 포크 1개, 재개 후 하나의 보온 메시지, 유료 콜드 쓰기 후 기간 작동, 캐시 상태, 콜드 가격 및 이 세션의 콜드 쓰기 수를 표시합니다.
이 mod 소개
cache-warm
Claude Code는 대화를 프롬프트 캐시에 1시간 보관합니다. 더 오래 떨어지면 다음 메시지는 캐시 쓰기 요금으로 전체 컨텍스트를 다시 씁니다. Claude Fable 5.1의 경우 200k 토큰당 $4.00이지만, 따뜻한 캐시에서 동일한 토큰을 읽으려면 $0.05입니다. 이 mod는 선택한 기간 동안 캐시를 따뜻하게 유지하고 콜드 캐시의 비용을 표시합니다.
Karan Bansal(karanb192/claude-code-mods)의 cache-tax mod의 동작을 따르지만 전송 보호는 없습니다. 코드는 새로 작성되었습니다.
기능
캐시를 따뜻하게 유지합니다. /cache-warm은 6시간 기간을 활성화합니다. 그 안에서 메인 루프의 마지막 모델 요청으로부터 50분 후, mod는 세션 자신의 트랜스크립트에서 도구 없는 $.model.fork 하나를 보냅니다. 서버는 캐시에서 응답하고 1시간을 새로 고칩니다. 새로운 요청마다 ping이 뒤로 밀리므로 활발하게 사용 중인 세션은 ping을 전송하지 않습니다. 세션의 첫 요청, /clear 또는 압축 후의 첫 요청도 ping을 설정하므로 첫 턴 내에서 1시간을 초과하여 실행되는 도구 호출도 턴 중간에 ping을 받습니다(2.1.281에서 측정: ping은 sleep 110 실행 중에 나갔고 턴은 정상적으로 종료됨).
always 아래에서 무한 실행. /cache-warm always는 기간이 아닙니다. ping은 세션이 존재하는 동안 50분마다 나가며 /cache-warm off만 종료합니다. 스위치는 mod의 자체 $.store의 하나의 글로벌 키이므로 모든 프로젝트의 이후 모든 세션은 시작 시와 /clear 후 동일한 루프를 시작합니다.
- 각 턴의 끝은 세션의 id 아래
$.store에서 마지막 요청의 시간을 유지하고, 각 ping도 자신의 읽음을 거기에 유지합니다. 실행 중인 대화에 로드된 모듈(/reload-plugins, 업데이트)은 따라서 두 시점 중 나중 것으로부터 적시에 ping합니다. ping만으로 따뜻하게 유지되는 세션은 마지막 턴이 캐시보다 오래됩니다. 둘 다 1시간 이상 오래되면 캐시가 사라지고 루프는 콜드 ping 비용을 지불하는 대신 다음 턴을 기다립니다. - 세션의 트랜스크립트는 이에 사용되지 않습니다.
/reload-plugins는 자체 라인을 거기에 쓰고 파일의 마지막 쓰기는 요청처럼 보이기 때문입니다(측정: 마지막 턴으로부터 90초 후 다시 로드하면 ping을 90초 늦게 설정할 것입니다). 아직 시간이 없는 세션, 이전 버전을 실행한 세션만 한 번 해당 트랜스크립트의 마지막 쓰기를 읽습니다(~/.claude/projects/<directory>/<session id>.jsonl,CLAUDE_CONFIG_DIR이 설정된 경우 그 아래). 트랜스크립트는 세션이 시작된 디렉토리에서 검색되며, 셸cd가 이동한 디렉토리에서는 검색되지 않으며, 찾을 수 없는 디렉토리는 트랜스크립트 라인에서 이름이 지정됩니다. - 캐시가 없는 것을 찾는 ping은 이 루프를 종료하지 않습니다. 해당 ping이 지불한 쓰기가 새로운 캐시입니다. mod는 트랜스크립트 라인에 그렇게 말하고 쓰기를 세션 합계에 계산하고 계속합니다. 따뜻한 ping은 컨텍스트를 읽기 요금으로 읽습니다. 200k 토큰에 약 $0.05이므로 유휴 하루의 ping 비용은 약 $1.40입니다.
캐시가 없을 때 중지. 이는 기간의 끝이 있는 경우에만 적용됩니다. always에는 적용되지 않습니다. 따뜻한 ping은 컨텍스트를 읽고 자신의 몇 토큰만 씁니다. ping이 아무것도 읽지 않거나 읽은 것의 10분의 1 이상을 쓰면 캐시가 이미 없었고 ping 자체가 쓰기를 지불했으므로 mod는 중지하고 이유를 표시합니다. API가 오류로 포크에 응답할 때도 중지합니다(라인이 상태와 종류를 이름 지정함). 포크가 응답 전에 끊길 때도 중지합니다. 재개 처리가 자신의 첫 응답 전과 같이 엔진에 포크할 것이 없을 때, 기간은 중지하지 않습니다. 다음 응답을 기다리고 해당 턴이 ping을 다시 활성화하면, 라인은 캐시가 언제까지 유지되는지를 말합니다. 텍스트 없는 응답도 캐시를 읽으므로 ping으로 계산됩니다. always 아래에서 그러한 실패는 그 턴만 루프를 중지합니다. 다음 턴이 다시 시작하므로 세션은 아무것도 실행되지 않는 동안 스위치를 유지하지 않습니다.
유료 콜드 쓰기 후 자동 활성화. 턴이 20k 토큰보다 큰 컨텍스트의 절반 이상을 다시 쓰면 mod는 해당 콜드 쓰기를 계산하고 더 긴 것이 이미 활성화되지 않은 한 6시간 기간을 활성화합니다. always 아래에서는 무한 루프가 이미 해당 캐시를 유지하므로 6시간 기간은 활성화되지 않습니다.
상태를 표시합니다. /cache-status는 모델, 따뜻함 또는 콜드, 컨텍스트 크기, 콜드 가격, 기간, 손익분기점 및 이 세션의 콜드 쓰기를 인쇄합니다.
콜드 캐시에 보내는 메시지는 절대 중지되거나 지연되지 않습니다. 캐시가 lapsed된 세션을 재개하면 첫 메시지의 가격에 대한 한 줄을 얻습니다. Claude Code는 트랜스크립트의 마지막 응답에서 캐시를 날짜 지정합니다. ping은 절대 여기에 쓰지 않으므로 마지막 ping이 mod가 세션을 위해 유지한 동안 1시간 이내에 캐시를 읽은 동안 그 라인은 생략됩니다.
재개 후 보온 메시지를 전송합니다. 재개된 프로세스는 자신의 첫 응답 전에 포크할 수 없습니다. $.model.fork는 nothing-to-fork로 응답합니다(헤드리스 claude --resume으로 측정). 그래서 ping이 나갈 수 없고 30분 후 닫혔다가 다시 열린 세션은 무언가를 쓰지 않는 한 1시간에 캐시를 잃을 것입니다. 기간 또는 always가 실행되는 동안 인터랙티브 세션이 재개되고 컨텍스트가 50k 토큰 이상, 캐시가 여전히 유지되는 경우, mod는 재개로부터 3초 후 자신의 /cache-warm:send 명령을 통해 한 메시지를 전송합니다.
/cache-warm:send This message was sent by the cache-warm plugin, not by the person. The session was resumed, and a resumed session can keep its prompt cache warm only after a reply. Do not run a tool or continue a task. Reply with the single word: warm
이는 실제 턴입니다. 캐시를 읽고 모델이 한 단어로 답하고 페어가 대화에 남고 끝이 ping을 다시 활성화합니다. 캐시가 이미 없을 때 아무것도 전송되지 않습니다(다음 메시지는 어쨌든 동일한 쓰기 비용을 지불합니다). -p 실행 내에서 또는 3초 이내에 메시지를 보낸 경우. 다른 mod는 그것을 모든 턴처럼 봅니다. task-poke는 작업이 열려있는 동안 그 후 계속 프롬프트를 보낼 수 있고, desk-notify는 턴 종료 알림을 표시합니다. 2.1.283에서 116k 토큰의 재개된 인터랙티브 세션에서 측정됨. 메시지는 재개 시 나갔고 모델은 warm이라고 답했으며 다음 ping은 포크되어 117k 토큰을 읽었습니다. 재개는 독립적으로 접두사의 일부를 깨뜨릴 수 있습니다(해당 세션은 116k 중 42k를 다시 썼습니다. 다른 하나는 마지막 ping으로부터 40분 후 재개됨, 4k). 보온 턴은 첫 메시지 대신 재개 시 쓰기를 지불합니다. 3초는 두 시작 훅의 나중 것으로부터 계산됩니다. classic.SessionStart 및 session.start, 고정 순서로 해결되지 않으므로. 마켓플레이스의 모든 mod가 로드되면 session.start는 classic.SessionStart로부터 4초 후 해결됩니다(2.1.285에서 측정). 첫 번째만으로 시작된 대기는 시작되지 않은 세션을 발견하고 아무것도 보내지 않습니다.
명령
/cache-warm keep warm for six hours
/cache-warm 90m keep warm for a window of your own (also 2h30m)
/cache-warm always keep the cache warm with no end, in every session of every project
/cache-warm 6h every 2m ping every two minutes; a test setting, floor 1m, forgotten after this window
/cache-warm status the status line text
/cache-warm off stop, forget the window, and turn always off
/cache-status the card
/cache-warm:send <text> the keep-warm message the mod sends after a resume; its body is the text alone
표시 항목
기간이 활성화되거나 중지된 후 프롬프트 아래의 상태 라인:
cache-warm: 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
cache-warm: stopped: the ping read 0 and wrote 180k tokens ($3.60), the cache was already gone
last 부분은 마지막 요청을 이름 지정합니다. 이는 캐시를 읽은 것, ping 또는 메인 루프 턴, 어느 것이든 나중 것입니다. last ping read 200k $0.05 또는 last turn read 250k $0.07, 둘 다 아닙니다. 턴의 수치는 전체 턴을 포함합니다. 모든 요청이 합산됩니다. 괄호 내의 시간은 해당 답변이 나온 시간입니다. 현지 시간; 더 이른 날의 것은 날짜와 월을 포함합니다. (22 Sep 23:10)로. 기록은 세션의 id 아래 $.store에 유지되므로 /reload-plugins 또는 업데이트가 즉시 다시 표시합니다.
기간이 실행되는 동안 인터랙티브 세션은 1분마다 라인을 다시 그리므로 남은 시간과 다음 ping까지의 시간은 턴과 ping 사이에 카운트다운되고 last 부분은 라인에 남습니다. ping 전 마지막 분은 ping now로 읽습니다. 분은 반올림되고 라인은 1분마다 그려집니다. ping의 포크가 나갈 때 pinging…로 읽습니다. 엔진에 포크할 것이 없을 때 라인은 카운트다운 대신 말합니다.
cache-warm: always · no ping before the next reply · cache holds until 19:16
사이드바가 열려있는 동안 ping 시도당 1개의 스트림 항목. 또한 사이드바의 로그 파일에 유지됩니다(~/.claude/sidebar/<project>-<date>.log). 나중에 읽을 수 있으므로 ping이 나갔는지 무엇을 했는지 확인할 수 있습니다. 사이드바가 닫혀있으면 동일한 텍스트가 트랜스크립트 라인입니다.
ping sent · read 901k · wrote 0 · $0.45
ping found the cache gone · read 0 · wrote 180k · $3.60
ping not sent: the conversation has no reply to fork yet; the ping waits for the next reply, and the cache holds until 19:16
ping failed: the ping failed, the API answered 529 (overloaded)
keep-warm message sent: the session was resumed and its cache holds until 19:16; a resumed session pings only after a reply
sidebar가 열려있는 경우 상태 라인이 거기로 이동합니다. cache window 섹션으로 세션용으로 유지되고 모든 변경 시 다시 작성됩니다. 상태 라인은 명확하게 유지됩니다. 남은 시간만(또는 always)이 색상 표시됩니다. 기간이 하나의 ping 기간보다 빨리 끝날 때 노란색; 유지되는 동안 녹색; 첫 턴을 기다리는 동안 흐립니다. 그 뒤 ping 세부 정보는 흐립니다. 중지된 기간은 stopped: 앞을 빨간색으로, 이유를 기본 색상으로 표시합니다. 사이드바 없이 상태 라인은 위와 같이 그려집니다.
중지 이유는 1턴 유지됩니다. 다음 턴에서 섹션은 아이들 라인을 대신 표시합니다. 흐림, 지불된 N cold writes paid $X 제외하고 이는 노란색입니다. 이렇게 창은 지난 기간의 마지막 문이 아니라 지금의 측정을 표시합니다. 이유는 트랜스크립트에 유지되고 기간이 실행 중이 아닐 때 상태 라인은 명확합니다.
cache window
off · 2 cold writes paid $6.30 · context 315k tokens
시간이 다한 기간은 다음 메시지에서 다시 활성화됩니다. 끝난 것과 동일한 한, 동일한 ping 기간과 함께, 대기 중인 동안 아이들 라인이 말합니다.
cache window
off · 6h again at your next message · 1 cold write paid $4.01 · context 201k tokens
cache-warm: the 6h window ran out; this message arms another one. /cache-warm off stops it.
시간이 다한 기간만 되돌아옵니다. ping이 중지한 것은 아닙니다. 캐시는 이미 거기서 없고 다음 메시지의 콜드 쓰기가 자신의 6h 기간을 활성화합니다. /cache-warm off는 대기 중인 기간을 잊습니다.
기간 아래에서 섹션은 두 번째, 흐린 라인을 유지합니다. 마지막 트랜스크립트 라인, 단축, 콜드 쓰기의 비용 노란색. 기간 라인은 캐시가 얼마나 유지되는지를 말합니다. 두 번째 라인은 mod가 마지막에 한 것을 말합니다.
cache window
6h left · ping in 50m
cold write 201k tokens paid ($4.01)
/cache-status 카드:
claude-fable-5-1
state warm, 42m left
context 200,502 tokens
cold cost $4.01 to re-write it (warm turn $0.05)
keep warm on, 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
break-even up to 80 pings at the read rate cost one cold write, about 2d 18h of idle at one ping per 50m
session 1 cold write paid, $4.01
1개의 트랜스크립트 라인, 모델에 전송되지 않음, 콜드 쓰기가 기간을 활성화하거나 재개가 콜드로 시작될 때. 감시자가 오프인 동안(기간 없음, always off) 라인은 대신 사이드바의 스트림으로 가고 트랜스크립트에 아무것도 쓰지 않습니다. 닫힌 사이드바는 그것을 드롭합니다. 실행 감시자가 쓰는 라인은 변경되지 않습니다.
가격
테이블 i
설치
먼저 작성자의 README에서 marketplace와 플러그인 이름을 확인하세요. 저장소 구조에 따라 명령어가 달라질 수 있습니다.
claude plugin marketplace add KilimcininKorOglu/claude-code-mods claude plugin install cache-warm
원문 / README
cache-warm
Claude Code keeps your conversation in a prompt cache for one hour. Step away for longer, and the next message re-writes the whole context at the cache-write rate: on Claude Fable 5.1 that is $4.00 for 200k tokens, against $0.05 to read the same tokens from a warm cache. This mod keeps the cache warm for a window you choose, and shows you what a cold cache would cost.
It follows the behaviour of the cache-tax mod by Karan Bansal (karanb192/claude-code-mods), without its send guard. The code is new.
What it does
Keeps the cache warm. /cache-warm arms a six-hour window. Inside it, 50 minutes after the main loop's last model request, the mod sends one tool-less $.model.fork over the session's own transcript. The server answers it from the cache, and that refreshes the hour. Every new request pushes the ping later, so a session you are actively using sends no ping at all. The first request of a session, and the first after /clear or a compaction, sets the ping as well, so even a tool call that runs past the hour inside the first turn gets its ping in the middle of the turn (measured on 2.1.281: the ping went out while sleep 110 ran, and the turn ended normally).
Runs with no end under always. /cache-warm always is not a window: the ping goes out every 50 minutes for as long as the session lives, and only /cache-warm off ends it. The switch is one global key in the mod's own $.store, so every later session of every project starts the same loop at its start and after /clear.
- Each turn's end keeps the last request's time in
$.storeunder the session's id, and each ping keeps its own read there too. A module loaded into a running conversation (/reload-plugins, an update) therefore pings on time from the later of the two; a session kept warm by pings alone has a last turn older than its cache. When both are more than an hour old, the cache is gone, and the loop waits for the next turn instead of paying for a cold ping. - The session's transcript is not used for this, because
/reload-pluginswrites a line of its own there and the file's last write would then look like a request (measured: a reload 90 seconds after the last turn would have set the ping 90 seconds late). Only a session with no time kept yet, one that ran an older version, reads the last write of its transcript (~/.claude/projects/<directory>/<session id>.jsonl, underCLAUDE_CONFIG_DIRwhen it is set) once. The transcript is looked up under the directory the session started in, never the one a shellcdmoved to, and one it cannot find is named in a transcript line. - A ping that finds the cache gone does not end this loop. The write that ping paid for is the new cache: the mod says so in a transcript line, counts the write in the session's tally and keeps going. A warm ping reads the context at the read rate, about $0.05 for 200k tokens, so an idle day of pings costs about $1.40.
Stops when the cache is gone. This applies to a window with an end, not to always. A warm ping reads the context and writes only its own few tokens. When a ping reads nothing, or writes a tenth of what it read or more, the cache was already gone and the ping itself paid for the write, so the mod stops and shows why. It also stops when the API answers the fork with an error (the line names its status and kind), and when the fork is cut before it replies. When the engine has nothing to fork, as in a resumed process before its first reply, the window does not stop: it waits for the next reply, whose turn arms the ping again, and the line says until when the cache holds. A reply without text still read the cache, so it counts as a ping. Under always such a failure stops the loop for that turn only: the next turn starts it again, so the session never holds the switch while nothing runs.
Arms itself after a paid cold write. When a turn re-writes at least half of a context larger than 20k tokens, the mod counts that cold write and arms a six-hour window, unless a longer one is already armed. Under always no six-hour window is armed, because the endless loop already keeps that cache.
Shows the state. /cache-status prints the model, warm or cold, the context size, the cold price, the window, the break-even and this session's cold writes.
A message you send to a cold cache is never stopped or delayed. A resumed session whose cache has lapsed gets one line with the price of its first message. Claude Code dates the cache from the transcript's last reply, which a ping never writes, so that line is left out while the last ping the mod kept for the session read the cache within the hour.
Sends a keep-warm message after a resume. A resumed process cannot fork before its own first reply: $.model.fork answers nothing-to-fork (measured with a headless claude --resume). So no ping can go out, and a session you closed and opened again 30 minutes later would lose its cache at the hour unless you wrote something. When an interactive session is resumed while a window or always runs, its context is 50k tokens or more and its cache still holds, the mod therefore sends one message three seconds after the resume, through its own /cache-warm:send command:
/cache-warm:send This message was sent by the cache-warm plugin, not by the person. The session was resumed, and a resumed session can keep its prompt cache warm only after a reply. Do not run a tool or continue a task. Reply with the single word: warm
This is a real turn: it reads the cache, the model answers one word, the pair stays in the conversation, and its end arms the ping again. Nothing is sent when the cache is already gone (your next message pays the same write anyway), in a -p run, or when you sent a message within those three seconds. Other mods see it like any other turn: task-poke may send its continue prompt after it while tasks are open, and desk-notify shows its turn-end notification. Measured on 2.1.283 in a resumed interactive session of 116k tokens: the message went out at the resume, the model answered warm, and the next ping forked and read 117k tokens. A resume can break part of the prefix on its own (that session re-wrote 42k of the 116k; another, resumed 40 minutes after its last ping, 4k); the keep-warm turn pays that write at the resume instead of your first message. The three seconds count from the later of the two start hooks, classic.SessionStart and session.start, because they settle in no fixed order: with every mod of the marketplace loaded, session.start settled four seconds after classic.SessionStart (measured on 2.1.285), and a wait started by the first alone found a session that had not started and sent nothing.
Commands
/cache-warm keep warm for six hours
/cache-warm 90m keep warm for a window of your own (also 2h30m)
/cache-warm always keep the cache warm with no end, in every session of every project
/cache-warm 6h every 2m ping every two minutes; a test setting, floor 1m, forgotten after this window
/cache-warm status the status line text
/cache-warm off stop, forget the window, and turn always off
/cache-status the card
/cache-warm:send <text> the keep-warm message the mod sends after a resume; its body is the text alone
What it shows
A status line under the prompt while a window is armed or after a stop:
cache-warm: 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
cache-warm: stopped: the ping read 0 and wrote 180k tokens ($3.60), the cache was already gone
The last part names the last request that read the cache, a ping or a main-loop turn, whichever came later: last ping read 200k $0.05 or last turn read 250k $0.07, never both. A turn's figures cover the whole turn, every request summed. The time in brackets is when that answer came, in local time; one from an earlier day carries its day and month, as (22 Sep 23:10). The record is kept in $.store under the session's id, so /reload-plugins or an update shows it again at once.
While a window runs, an interactive session redraws the line every minute, so the time left and the time to the next ping count down between turns and pings, and the last part stays on the line. The last minute before a ping reads ping now, because minutes are rounded and the line is drawn once a minute; while the ping's fork is out it reads pinging…. When the engine has nothing to fork, the line says so instead of counting down:
cache-warm: always · no ping before the next reply · cache holds until 19:16
One stream entry per ping attempt while the sidebar is open. It is also kept in the sidebar's log file (~/.claude/sidebar/<project>-<date>.log), so you can read back later whether a ping went out and what it did; with the sidebar closed the same text is a transcript line:
ping sent · read 901k · wrote 0 · $0.45
ping found the cache gone · read 0 · wrote 180k · $3.60
ping not sent: the conversation has no reply to fork yet; the ping waits for the next reply, and the cache holds until 19:16
ping failed: the ping failed, the API answered 529 (overloaded)
keep-warm message sent: the session was resumed and its cache holds until 19:16; a resumed session pings only after a reply
With the sidebar open, the status line moves there as a cache window section that stays for the session and is rewritten at every change, and the status line stays clear. Only the time left (or always) is coloured: yellow when the window ends sooner than one ping period, green while it holds, faint while it waits for the first turn. The ping details after it are faint, and a stopped window shows its stopped: front in red with the reason in the default colour. Without the sidebar, the status line is drawn as above.
A stop reason stays for one turn. At the next turn the section shows the idle line instead: faint, except for a paid N cold writes paid $X, which is yellow. That way the pane shows a measurement of now, not the last sentence of a window that ended. The reason stays in the transcript, and the status line is empty while no window runs:
cache window
off · 2 cold writes paid $6.30 · context 315k tokens
A window that runs out of time is armed again by your next message, as long as the one that ended and with the same ping period, and the idle line says so while it waits:
cache window
off · 6h again at your next message · 1 cold write paid $4.01 · context 201k tokens
cache-warm: the 6h window ran out; this message arms another one. /cache-warm off stops it.
Only a window that ran out of time comes back. One that a ping stopped does not: the cache is already gone there, and the cold write of your next message arms its own 6h window. /cache-warm off forgets a window waiting to come back.
Under the window the section holds a second, faint line: the last transcript line, shortened, with the cost of a cold write in yellow. The window line says how long the cache is kept; the second line says what the mod did last:
cache window
6h left · ping in 50m
cold write 201k tokens paid ($4.01)
The card of /cache-status:
claude-fable-5-1
state warm, 42m left
context 200,502 tokens
cold cost $4.01 to re-write it (warm turn $0.05)
keep warm on, 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
break-even up to 80 pings at the read rate cost one cold write, about 2d 18h of idle at one ping per 50m
session 1 cold write paid, $4.01
One transcript line, not sent to the model, when a cold write arms the window or a resume starts cold. While the watcher is off (no window, always off), the line goes to the sidebar's stream instead and writes nothing to the transcript; a closed sidebar drops it. The line a running watcher writes is unchanged.
Prices
The table in hooks/pricing.ts holds the cache-read, 1-hour cache-write and output rates of every model on the Anthropic pricing page, read in September 2026. A model id takes the first family it contains, so claude-opus-4-1 is priced as Opus 4.1 ($1.50 / $30 / $75) and claude-opus-4-8 as Opus 4.8 ($0.50 / $10 / $25). claude-opus-5-5 also contains opus-5, so its own row comes first: Opus 5.5 is $0.20 / $8 / $20, cheaper than Opus 5. Sonnet 5.5 has its own row at the Sonnet 5 rates, $0.20 / $4 / $10. A ping is priced in full: the cache read, its cache write, its uncached input at the base rate (half the 1-hour write rate) and its output. An unknown model shows n/a.
Fast mode bills Opus 5.5, Opus 5 and Opus 4.8 at their own base rates ($8 and $10 input), with the cache multipliers applied on top. The mod prices Opus 5.5 at $0.40 / $16 / $40 and Opus 5 and 4.8 at $1 / $20 / $50 while the fastMode setting, which /fast writes, is on. It reads the settings at the session's start and at the end of each main-loop turn, so a /fast counts from the next turn. With fastModePerSessionOptIn set to true, every session starts with fast mode off, so the standard rates apply. Any other model keeps its standard rates, and the card mentions fast mode only when the rates changed:
claude-opus-5-5 · fast mode rates (the fastMode setting)
On a subscription the dollars are a yardstick, not your bill. How a cache read counts against the 5-hour and weekly limits is not documented.
Install
claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install cache-warm@kilimcininkoroglu-mods
Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.
To load it from a local checkout for one session:
claude --plugin-dir plugins/cache-warm
After installing
- Restart Claude Code.
- Disable every other keep-warm mod, for example
claude plugin disable cache-tax@claude-code-mods. Two keep-warm mods in one session send two pings per idle stretch. - Check once that a ping reads your cache, as "Prove it on your own session" below describes.
- To keep the cache warm with no end, run
/cache-warm alwaysonce. The switch is global: every later session of every project starts the loop by itself, and/cache-warm offends it for good. Without it, a window is armed only by/cache-warmor after a paid cold write.
What it can reach
Validated with claude plugin validate on Claude Code 2.1.288:
❯ ./register.ts hooks: session.start, classic.SessionStart, prompt.submit, command.run{command=cache-warm}, command.run{command=cache-status}, turn.step, turn.complete, session.compact
❯ ./register.ts calls: $.clock.after (via arm, scheduleKeepWarm), $.clock.every, $.clock.now, $.command.register (via registerCommands), $.command.run (via keepWarmAfterResume), $.env.get (via seedFromTranscript), $.fs.exists (via seedFromTranscript), $.fs.stat (via seedFromTranscript), $.model.fork (via forkPing), $.prompt.submit (via keepWarmAfterResume), $.session.id, $.session.model, $.session.root (via seedFromTranscript), $.session.usage, $.settings.read (via readFast), $.sidebar.set (via logEvent, toSidebar, toStream), $.store.delete (via prune, pruneRequests, startEndless, startWindow, stop), $.store.get, $.store.keys (via prune, pruneRequests), $.store.set (via afterTurn, keepLastRead, startWindow, warmCommand), $.ui.log (via logEvent, seedFromTranscript, toStream), $.ui.status (via showStatusAt)
❯ ./register.ts env writes: nothing
❯ ./register.ts env reads: CLAUDE_CONFIG_DIR, HOME
Reach L2: it drives Claude.
1. Reads: the time of each main-loop model request; the token counts and model id of each turn and of each ping; the live context size; the origin of each message, to arm a window again; the resume fields Claude Code computes for settings hooks; the session id and model; the last write time of the session's own transcript file, once when the module loads into a running conversation; the fastMode and fastModePerSessionOptIn settings, at the session's start and at each turn's end; its own $.store. It never reads a prompt's text, a file's content or a tool result.
2. Runs: one $.model.fork per idle stretch while a window or the always loop runs, 50 minutes after the last request unless the test setting is used (floor 1 minute); never while off; a ping that found the cache gone ends a window with an end, and under always the loop carries on; after a resume of an interactive session whose cache still holds, one keep-warm message through /cache-warm:send (a plugin prompt when the engine refuses the command), which is a real turn
3. Sends: the fork, an API request over the session's own transcript with a fixed one-line prompt, and after a resume the fixed keep-warm message as a turn of the conversation
4. Persists: in $.store, the window end, the ping period and the last main-loop request's time and the last ping or turn read (tokens, cost, time) under this session's id, and the global always switch, which the endless loop needs no window key beside; this session's ended window is deleted at stop and at its next start, another session's window one week after it ended, another session's request time and last read once they are an hour old; the cold-write tally lives in memory and ends with the session
5. Hostile input: the only text it parses is the argument of /cache-warm, matched against a duration pattern and three words; the fork's prompt is a constant, so nothing crafted can reach it
Prove it on your own session
The mock-clock tests prove the timer and the scoring, not that a fork reads the main cache. One ping proves that. In a warm session:
> Reply with one word: ready
> /cache-warm 1h every 1m
After a minute the status line should read last ping read <close to your context> $.... A stopped: the ping read ... line means the fork did not share the cache, and the mod has already stopped. /cache-warm off ends the test.
Limits
- The 50-minute ping assumes the 1-hour cache tier, which the main conversation uses.
- A warm ping only proves the cache was warm at that moment. A model or effort switch, an edited CLAUDE.md or a changed tool list breaks the prefix no matter the time, and your next message pays.
- A ping's output cannot be capped; a model at high effort may think before it answers. The status line prices what the ping really billed.
- The resume logic is covered by hook tests that raise
classic.SessionStartwith the resume fields, and by a live check of a resumed interactive session in tmux. Whether a resume keeps the whole prefix is out of the mod's hands: it re-wrote 42k of 116k tokens in one measured session and 4k in another. - The cold-write tally is per session and lives in memory.
/clearempties it. - Fast mode is read from the
fastModesetting, the saved preference, not from the request. The mod does not see Claude Code fall back to standard speed within a session (a fast mode rate-limit cooldown, usage credits that ran out, an organization that turned fast mode off). Those turns bill standard rates while the mod prices them as fast. - Whether a ping, a
$.model.fork, runs at fast speed while the session does has not been measured; the mod prices it at the session's rates.
Development
make install # eslint, typescript-eslint, typescript
make lint # complexity limit 10, the build fails above it
make typecheck # needs .claude/types/ from /plugin-types
make validate
make test # claude plugin test

