ClaudeMods
☰
ZH-CN
● 0 人在线 · 浏览 0 次
赞助提交作品
GitHub 仓库 · 发布者 KilimcininKorOglu

cache-warm

为您设定的时间段保持 1 小时提示缓存热状态,或在 always 模式下每个会话都保持(无时间限制),每个空闲阶段一个缓存共享分叉,恢复后一条保温消息,在付费冷写入后为您启用一个时间段,并显示缓存状态、冷价格和本会话的冷写入次数。

KilimcininKorOglu@KilimcininKorOglu

KilimcininKorOglu/claude-code-mods/tree/main/plugins/cache-warm

已翻译

关于这个 mod

cache-warm

Claude Code 将您的对话保存在提示缓存中 1 小时。如果您离开更长时间,下一条消息将以缓存写入率重写整个上下文:在 Claude Fable 5.1 上,这是 200k 令牌 $4.00,而从热缓存读取相同令牌仅需 $0.05。此 mod 为您选择的时间段保持缓存热状态,并显示冷缓存的成本。

它遵循 Karan Bansal (karanb192/claude-code-mods) 的 cache-tax mod 的行为,但没有其发送保护。代码是全新的。

它的功能

保持缓存热状态。 /cache-warm 启用一个 6 小时的时间段。在其中,在主循环最后一次模型请求后 50 分钟,mod 通过会话自己的记录发送一个无工具的 $.model.fork。服务器从缓存中回复,这会刷新 1 小时。每个新请求都会推迟 ping,因此您主动使用的会话根本不会发送 ping。会话的第一个请求,以及 /clear 或压缩后的第一个请求,也会设置 ping,所以即使工具调用在第一个回合内超过 1 小时,它也会在回合中间获得 ping(在 2.1.281 上测量:ping 在 sleep 110 运行时发出,回合正常结束)。

在 always 下无限运行。 /cache-warm always 不是一个时间段:ping 每 50 分钟发出一次,只要会话活跃,直到 /cache-warm off 结束。开关是 mod 自己 $.store 中的一个全局键,所以每个项目的每个后续会话在其开始和 /clear 后启动相同的循环。

  • 每个回合的结束在 $.store 的会话 ID 下保持最后一次请求的时间,每个 ping 也在那里保持自己的读取时间。因此,一个加载到正在运行的对话中的模块(/reload-plugins、更新)会从两者的较晚者准时 ping;仅由 ping 保温的会话的最后回合比其缓存更旧。当两者都超过 1 小时旧时,缓存已消失,循环等待下一个回合而不是为冷 ping 付费。
  • 会话的记录不用于此,因为 /reload-plugins 在其中写入自己的一行,文件的最后写入时间看起来像请求(测量:最后一个回合后 90 秒的重新加载会将 ping 设置为晚 90 秒)。只有没有保持时间的会话,即运行较旧版本的会话,才会读取一次其记录的最后写入(~/.claude/projects/<directory>/<session id>.jsonl,当设置时在 CLAUDE_CONFIG_DIR 下)。记录在会话开始的目录下查找,从不是 shell cd 移动到的目录,找不到的会在记录行中命名。
  • 找到缓存已消失的 ping 不会结束此循环。该 ping 支付的写入是新缓存:mod 在记录行中说明这一点,在会话的统计中计算写入并继续。热 ping 以读取率读取上下文,200k 令牌约 $0.05,所以空闲一天的 ping 成本约 $1.40。

缓存消失时停止。 这适用于有结束时间的时间段,不适用于 always。温 ping 读取上下文,仅写入自己的几个令牌。当 ping 什么都没读到,或写入读取内容的十分之一或更多时,缓存已经消失,ping 本身支付了写入,所以 mod 停止并显示原因。当 API 用错误回复分叉时也会停止(行命名其状态和类型),以及当分叉在回复前被切断时。当引擎没有分叉时,如在恢复流程的第一个回复前,时间段不会停止:它等待下一个回复,其回合重新启用 ping,行说明缓存的保持时间。没有文本的回复仍然读取了缓存,所以它计为 ping。在 always 下这样的失败仅在那个回合停止循环:下一个回合重新启动它,所以会话在运行时从不持有开关。

在付费冷写入后启用自己。 当一个回合重写大小超过 20k 令牌的上下文的至少一半时,mod 计算该冷写入并启用一个 6 小时的时间段,除非已经启用了更长的。在 always 下不启用 6 小时的时间段,因为无限循环已经保持该缓存。

显示状态。 /cache-status 打印模型(热或冷)、上下文大小、冷价格、时间段、平衡点和本会话的冷写入。

发送给冷缓存的消息永远不会被停止或延迟。缓存已过期的恢复会话会获得一行,显示其第一条消息的价格。Claude Code 从记录的最后回复对缓存进行日期计算,ping 从不写入,所以在 mod 为会话保持的最后 ping 在 1 小时内读取缓存时,该行被省略。

恢复后发送保温消息。 恢复流程无法在其自己的第一个回复前分叉:$.model.fork 回复 nothing-to-fork(使用无头 claude --resume 测量)。所以没有 ping 可以发出,如果您 30 分钟后关闭并重新打开会话,除非您写了什么,否则会在 1 小时处失去缓存。当交互式会话在时间段或 always 运行时被恢复,其上下文 50k 令牌或更多且其缓存仍持有时,mod 因此在恢复后 3 秒钟发送一条消息,通过其自己的 /cache-warm:send 命令:

/cache-warm:send 此消息由 cache-warm 插件发送,而非个人。会话已恢复,恢复的会话仅在回复后才能保持其提示缓存热状态。不要运行工具或继续任务。用单个单词回复:warm

这是一个真实的回合:它读取缓存,模型回答一个单词,两者保留在对话中,其结束再次启用 ping。当缓存已经消失时(您的下一条消息无论如何都会支付相同的写入)、在 -p 运行中或您在这三秒内发送消息时,不会发送任何东西。其他 mod 像任何其他回合一样看待它:task-poke 在任务打开时可能在其后发送继续提示,desk-notify 显示其回合结束通知。在 2.1.283 上测量在一个 116k 令牌的恢复交互式会话中:消息在恢复时发出,模型回答 warm,下一个 ping 分叉并读取 117k 令牌。恢复可以自己打破前缀的一部分(该会话重写了 116k 中的 42k;另一个,在最后一个 ping 后 40 分钟恢复,4k);保温回合在恢复时支付该写入,而不是您的第一条消息。三秒是从两个启动钩子的较晚者计算的,classic.SessionStart 和 session.start,因为它们没有固定顺序:加载了市场上每个 mod 时,session.start 在 classic.SessionStart 后四秒进行(在 2.1.285 上测量),仅由第一个启动的等待发现了一个尚未启动的会话并未发送任何东西。

命令

/cache-warm               保持温度 6 小时
/cache-warm 90m           保持温度您自己的时间段(也可以 2h30m)
/cache-warm always        保持缓存温度无限制,在每个项目的每个会话中
/cache-warm 6h every 2m   每两分钟 ping 一次;测试设置,下限 1 分钟,在此时间段后被遗忘
/cache-warm status        状态行文本
/cache-warm off           停止,忘记时间段,并关闭 always
/cache-status             卡片
/cache-warm:send <text>   mod 在恢复后发送的保温消息;其正文仅为文本

它的显示内容

时间段启用或停止后提示下的状态行:

cache-warm: 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
cache-warm: stopped: the ping read 0 and wrote 180k tokens ($3.60), the cache was already gone

last 部分命名最后读取缓存的请求,ping 或主循环回合,后者较晚:last ping read 200k $0.05 或 last turn read 250k $0.07,从不两者。回合的数字覆盖整个回合,所有请求和。括号中的时间是该答复何时到达,本地时间;来自较早日期的那一天带有其日期和月份,如 (22 Sep 23:10)。记录保存在 $.store 下的会话 ID 中,所以 /reload-plugins 或更新立即再次显示它。

当时间段运行时,交互式会话每分钟重绘该行,因此剩余时间和下一个 ping 的时间在回合和 ping 之间倒计时,last 部分保留在行上。ping 前的最后一分钟读取 ping now,因为分钟四舍五入且行每分钟绘制一次;当 ping 的分叉在外面时,它读取 pinging…。当引擎没有分叉时,行说明这个而不是倒计时:

cache-warm: always · no ping before the next reply · cache holds until 19:16

每个 ping 尝试一个流条目当边栏打开时。它也保存在边栏的日志文件中(~/.claude/sidebar/<project>-<date>.log),所以您以后可以回读 ping 是否发出以及它做了什么;当边栏关闭时相同的文本是记录行:

ping sent · read 901k · wrote 0 · $0.45
ping found the cache gone · read 0 · wrote 180k · $3.60
ping not sent: the conversation has no reply to fork yet; the ping waits for the next reply, and the cache holds until 19:16
ping failed: the ping failed, the API answered 529 (overloaded)
keep-warm message sent: the session was resumed and its cache holds until 19:16; a resumed session pings only after a reply

当 sidebar 打开时,状态行移到那里作为一个 cache window 部分,为会话停留并在每次更改时重写,状态行保持清晰。仅剩余时间(或 always)被着色:当时间段比一个 ping 周期更早结束时为黄色,保持时为绿色,等待第一个回合时为浅色。ping 详情在其后为浅色,停止的时间段显示其 stopped: 前面为红色,原因为默认颜色。没有边栏时,状态行如上所述绘制。

停止原因停留一个回合。在下一个回合处,部分显示空闲行:浅色,除了已支付的 N cold writes paid $X 为黄色。这样窗格显示现在的测量,而不是已结束时间段的最后一句。原因保留在记录中,当没有时间段运行时状态行为空:

cache window
off · 2 cold writes paid $6.30 · context 315k tokens

当时间段用尽时,只要已结束的时间和相同的 ping 周期,您的下一条消息会再次启用它,空闲行在等待时说明这一点:

cache window
off · 6h again at your next message · 1 cold write paid $4.01 · context 201k tokens

cache-warm: the 6h window ran out; this message arms another one. /cache-warm off stops it.

仅当时间段用尽时才回来。ping 停止的时间段不:缓存已在那里消失,下一条消息的冷写入启用其自己的 6 小时窗口。/cache-warm off 忘记等待回来的时间段。

在时间段下,部分保持第二条浅色行:最后的记录行、缩短的,具有黄色冷写入成本。时间段行说明缓存的保持时间;第二行说明 mod 最后做了什么:

cache window
6h left · ping in 50m
cold write 201k tokens paid ($4.01)

/cache-status 的卡片:

claude-fable-5-1
state       warm, 42m left
context     200,502 tokens
cold cost   $4.01 to re-write it (warm turn $0.05)
keep warm   on, 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
break-even  up to 80 pings at the read rate cost one cold write, about 2d 18h of idle at one ping per 50m
session     1 cold write paid, $4.01

一条记录行,未发送给模型,当冷写入启用时间段或恢复开始冷时。当观察器关闭(无时间段、always 关闭)时,行转到边栏的流,不向记录写入;关闭的边栏删除它。正在运行的观察器写入的行是不变的。

价格

表格 i

安装

请先查看作者 README 确认 marketplace 和插件名称;命令可能随仓库结构改变。

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install cache-warm
原文 / README

cache-warm

Claude Code keeps your conversation in a prompt cache for one hour. Step away for longer, and the next message re-writes the whole context at the cache-write rate: on Claude Fable 5.1 that is $4.00 for 200k tokens, against $0.05 to read the same tokens from a warm cache. This mod keeps the cache warm for a window you choose, and shows you what a cold cache would cost.

It follows the behaviour of the cache-tax mod by Karan Bansal (karanb192/claude-code-mods), without its send guard. The code is new.

What it does

Keeps the cache warm. /cache-warm arms a six-hour window. Inside it, 50 minutes after the main loop's last model request, the mod sends one tool-less $.model.fork over the session's own transcript. The server answers it from the cache, and that refreshes the hour. Every new request pushes the ping later, so a session you are actively using sends no ping at all. The first request of a session, and the first after /clear or a compaction, sets the ping as well, so even a tool call that runs past the hour inside the first turn gets its ping in the middle of the turn (measured on 2.1.281: the ping went out while sleep 110 ran, and the turn ended normally).

Runs with no end under always. /cache-warm always is not a window: the ping goes out every 50 minutes for as long as the session lives, and only /cache-warm off ends it. The switch is one global key in the mod's own $.store, so every later session of every project starts the same loop at its start and after /clear.

  • Each turn's end keeps the last request's time in $.store under the session's id, and each ping keeps its own read there too. A module loaded into a running conversation (/reload-plugins, an update) therefore pings on time from the later of the two; a session kept warm by pings alone has a last turn older than its cache. When both are more than an hour old, the cache is gone, and the loop waits for the next turn instead of paying for a cold ping.
  • The session's transcript is not used for this, because /reload-plugins writes a line of its own there and the file's last write would then look like a request (measured: a reload 90 seconds after the last turn would have set the ping 90 seconds late). Only a session with no time kept yet, one that ran an older version, reads the last write of its transcript (~/.claude/projects/<directory>/<session id>.jsonl, under CLAUDE_CONFIG_DIR when it is set) once. The transcript is looked up under the directory the session started in, never the one a shell cd moved to, and one it cannot find is named in a transcript line.
  • A ping that finds the cache gone does not end this loop. The write that ping paid for is the new cache: the mod says so in a transcript line, counts the write in the session's tally and keeps going. A warm ping reads the context at the read rate, about $0.05 for 200k tokens, so an idle day of pings costs about $1.40.

Stops when the cache is gone. This applies to a window with an end, not to always. A warm ping reads the context and writes only its own few tokens. When a ping reads nothing, or writes a tenth of what it read or more, the cache was already gone and the ping itself paid for the write, so the mod stops and shows why. It also stops when the API answers the fork with an error (the line names its status and kind), and when the fork is cut before it replies. When the engine has nothing to fork, as in a resumed process before its first reply, the window does not stop: it waits for the next reply, whose turn arms the ping again, and the line says until when the cache holds. A reply without text still read the cache, so it counts as a ping. Under always such a failure stops the loop for that turn only: the next turn starts it again, so the session never holds the switch while nothing runs.

Arms itself after a paid cold write. When a turn re-writes at least half of a context larger than 20k tokens, the mod counts that cold write and arms a six-hour window, unless a longer one is already armed. Under always no six-hour window is armed, because the endless loop already keeps that cache.

Shows the state. /cache-status prints the model, warm or cold, the context size, the cold price, the window, the break-even and this session's cold writes.

A message you send to a cold cache is never stopped or delayed. A resumed session whose cache has lapsed gets one line with the price of its first message. Claude Code dates the cache from the transcript's last reply, which a ping never writes, so that line is left out while the last ping the mod kept for the session read the cache within the hour.

Sends a keep-warm message after a resume. A resumed process cannot fork before its own first reply: $.model.fork answers nothing-to-fork (measured with a headless claude --resume). So no ping can go out, and a session you closed and opened again 30 minutes later would lose its cache at the hour unless you wrote something. When an interactive session is resumed while a window or always runs, its context is 50k tokens or more and its cache still holds, the mod therefore sends one message three seconds after the resume, through its own /cache-warm:send command:

/cache-warm:send This message was sent by the cache-warm plugin, not by the person. The session was resumed, and a resumed session can keep its prompt cache warm only after a reply. Do not run a tool or continue a task. Reply with the single word: warm

This is a real turn: it reads the cache, the model answers one word, the pair stays in the conversation, and its end arms the ping again. Nothing is sent when the cache is already gone (your next message pays the same write anyway), in a -p run, or when you sent a message within those three seconds. Other mods see it like any other turn: task-poke may send its continue prompt after it while tasks are open, and desk-notify shows its turn-end notification. Measured on 2.1.283 in a resumed interactive session of 116k tokens: the message went out at the resume, the model answered warm, and the next ping forked and read 117k tokens. A resume can break part of the prefix on its own (that session re-wrote 42k of the 116k; another, resumed 40 minutes after its last ping, 4k); the keep-warm turn pays that write at the resume instead of your first message. The three seconds count from the later of the two start hooks, classic.SessionStart and session.start, because they settle in no fixed order: with every mod of the marketplace loaded, session.start settled four seconds after classic.SessionStart (measured on 2.1.285), and a wait started by the first alone found a session that had not started and sent nothing.

Commands

/cache-warm               keep warm for six hours
/cache-warm 90m           keep warm for a window of your own (also 2h30m)
/cache-warm always        keep the cache warm with no end, in every session of every project
/cache-warm 6h every 2m   ping every two minutes; a test setting, floor 1m, forgotten after this window
/cache-warm status        the status line text
/cache-warm off           stop, forget the window, and turn always off
/cache-status             the card
/cache-warm:send <text>   the keep-warm message the mod sends after a resume; its body is the text alone

What it shows

A status line under the prompt while a window is armed or after a stop:

cache-warm: 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
cache-warm: stopped: the ping read 0 and wrote 180k tokens ($3.60), the cache was already gone

The last part names the last request that read the cache, a ping or a main-loop turn, whichever came later: last ping read 200k $0.05 or last turn read 250k $0.07, never both. A turn's figures cover the whole turn, every request summed. The time in brackets is when that answer came, in local time; one from an earlier day carries its day and month, as (22 Sep 23:10). The record is kept in $.store under the session's id, so /reload-plugins or an update shows it again at once.

While a window runs, an interactive session redraws the line every minute, so the time left and the time to the next ping count down between turns and pings, and the last part stays on the line. The last minute before a ping reads ping now, because minutes are rounded and the line is drawn once a minute; while the ping's fork is out it reads pinging…. When the engine has nothing to fork, the line says so instead of counting down:

cache-warm: always · no ping before the next reply · cache holds until 19:16

One stream entry per ping attempt while the sidebar is open. It is also kept in the sidebar's log file (~/.claude/sidebar/<project>-<date>.log), so you can read back later whether a ping went out and what it did; with the sidebar closed the same text is a transcript line:

ping sent · read 901k · wrote 0 · $0.45
ping found the cache gone · read 0 · wrote 180k · $3.60
ping not sent: the conversation has no reply to fork yet; the ping waits for the next reply, and the cache holds until 19:16
ping failed: the ping failed, the API answered 529 (overloaded)
keep-warm message sent: the session was resumed and its cache holds until 19:16; a resumed session pings only after a reply

With the sidebar open, the status line moves there as a cache window section that stays for the session and is rewritten at every change, and the status line stays clear. Only the time left (or always) is coloured: yellow when the window ends sooner than one ping period, green while it holds, faint while it waits for the first turn. The ping details after it are faint, and a stopped window shows its stopped: front in red with the reason in the default colour. Without the sidebar, the status line is drawn as above.

A stop reason stays for one turn. At the next turn the section shows the idle line instead: faint, except for a paid N cold writes paid $X, which is yellow. That way the pane shows a measurement of now, not the last sentence of a window that ended. The reason stays in the transcript, and the status line is empty while no window runs:

cache window
off · 2 cold writes paid $6.30 · context 315k tokens

A window that runs out of time is armed again by your next message, as long as the one that ended and with the same ping period, and the idle line says so while it waits:

cache window
off · 6h again at your next message · 1 cold write paid $4.01 · context 201k tokens

cache-warm: the 6h window ran out; this message arms another one. /cache-warm off stops it.

Only a window that ran out of time comes back. One that a ping stopped does not: the cache is already gone there, and the cold write of your next message arms its own 6h window. /cache-warm off forgets a window waiting to come back.

Under the window the section holds a second, faint line: the last transcript line, shortened, with the cost of a cold write in yellow. The window line says how long the cache is kept; the second line says what the mod did last:

cache window
6h left · ping in 50m
cold write 201k tokens paid ($4.01)

The card of /cache-status:

claude-fable-5-1
state       warm, 42m left
context     200,502 tokens
cold cost   $4.01 to re-write it (warm turn $0.05)
keep warm   on, 5h 10m left · ping in 37m · last ping read 200k $0.05 (05:42)
break-even  up to 80 pings at the read rate cost one cold write, about 2d 18h of idle at one ping per 50m
session     1 cold write paid, $4.01

One transcript line, not sent to the model, when a cold write arms the window or a resume starts cold. While the watcher is off (no window, always off), the line goes to the sidebar's stream instead and writes nothing to the transcript; a closed sidebar drops it. The line a running watcher writes is unchanged.

Prices

The table in hooks/pricing.ts holds the cache-read, 1-hour cache-write and output rates of every model on the Anthropic pricing page, read in September 2026. A model id takes the first family it contains, so claude-opus-4-1 is priced as Opus 4.1 ($1.50 / $30 / $75) and claude-opus-4-8 as Opus 4.8 ($0.50 / $10 / $25). claude-opus-5-5 also contains opus-5, so its own row comes first: Opus 5.5 is $0.20 / $8 / $20, cheaper than Opus 5. Sonnet 5.5 has its own row at the Sonnet 5 rates, $0.20 / $4 / $10. A ping is priced in full: the cache read, its cache write, its uncached input at the base rate (half the 1-hour write rate) and its output. An unknown model shows n/a.

Fast mode bills Opus 5.5, Opus 5 and Opus 4.8 at their own base rates ($8 and $10 input), with the cache multipliers applied on top. The mod prices Opus 5.5 at $0.40 / $16 / $40 and Opus 5 and 4.8 at $1 / $20 / $50 while the fastMode setting, which /fast writes, is on. It reads the settings at the session's start and at the end of each main-loop turn, so a /fast counts from the next turn. With fastModePerSessionOptIn set to true, every session starts with fast mode off, so the standard rates apply. Any other model keeps its standard rates, and the card mentions fast mode only when the rates changed:

claude-opus-5-5 · fast mode rates (the fastMode setting)

On a subscription the dollars are a yardstick, not your bill. How a cache read counts against the 5-hour and weekly limits is not documented.

Install

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install cache-warm@kilimcininkoroglu-mods

Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.

To load it from a local checkout for one session:

claude --plugin-dir plugins/cache-warm

After installing

  1. Restart Claude Code.
  2. Disable every other keep-warm mod, for example claude plugin disable cache-tax@claude-code-mods. Two keep-warm mods in one session send two pings per idle stretch.
  3. Check once that a ping reads your cache, as "Prove it on your own session" below describes.
  4. To keep the cache warm with no end, run /cache-warm always once. The switch is global: every later session of every project starts the loop by itself, and /cache-warm off ends it for good. Without it, a window is armed only by /cache-warm or after a paid cold write.

What it can reach

Validated with claude plugin validate on Claude Code 2.1.288:

❯ ./register.ts hooks: session.start, classic.SessionStart, prompt.submit, command.run{command=cache-warm}, command.run{command=cache-status}, turn.step, turn.complete, session.compact
❯ ./register.ts calls: $.clock.after (via arm, scheduleKeepWarm), $.clock.every, $.clock.now, $.command.register (via registerCommands), $.command.run (via keepWarmAfterResume), $.env.get (via seedFromTranscript), $.fs.exists (via seedFromTranscript), $.fs.stat (via seedFromTranscript), $.model.fork (via forkPing), $.prompt.submit (via keepWarmAfterResume), $.session.id, $.session.model, $.session.root (via seedFromTranscript), $.session.usage, $.settings.read (via readFast), $.sidebar.set (via logEvent, toSidebar, toStream), $.store.delete (via prune, pruneRequests, startEndless, startWindow, stop), $.store.get, $.store.keys (via prune, pruneRequests), $.store.set (via afterTurn, keepLastRead, startWindow, warmCommand), $.ui.log (via logEvent, seedFromTranscript, toStream), $.ui.status (via showStatusAt)
❯ ./register.ts env writes: nothing
❯ ./register.ts env reads: CLAUDE_CONFIG_DIR, HOME

Reach L2: it drives Claude.

1. Reads:    the time of each main-loop model request; the token counts and model id of each turn and of each ping; the live context size; the origin of each message, to arm a window again; the resume fields Claude Code computes for settings hooks; the session id and model; the last write time of the session's own transcript file, once when the module loads into a running conversation; the fastMode and fastModePerSessionOptIn settings, at the session's start and at each turn's end; its own $.store. It never reads a prompt's text, a file's content or a tool result.
2. Runs:     one $.model.fork per idle stretch while a window or the always loop runs, 50 minutes after the last request unless the test setting is used (floor 1 minute); never while off; a ping that found the cache gone ends a window with an end, and under always the loop carries on; after a resume of an interactive session whose cache still holds, one keep-warm message through /cache-warm:send (a plugin prompt when the engine refuses the command), which is a real turn
3. Sends:    the fork, an API request over the session's own transcript with a fixed one-line prompt, and after a resume the fixed keep-warm message as a turn of the conversation
4. Persists: in $.store, the window end, the ping period and the last main-loop request's time and the last ping or turn read (tokens, cost, time) under this session's id, and the global always switch, which the endless loop needs no window key beside; this session's ended window is deleted at stop and at its next start, another session's window one week after it ended, another session's request time and last read once they are an hour old; the cold-write tally lives in memory and ends with the session
5. Hostile input: the only text it parses is the argument of /cache-warm, matched against a duration pattern and three words; the fork's prompt is a constant, so nothing crafted can reach it

Prove it on your own session

The mock-clock tests prove the timer and the scoring, not that a fork reads the main cache. One ping proves that. In a warm session:

> Reply with one word: ready
> /cache-warm 1h every 1m

After a minute the status line should read last ping read <close to your context> $.... A stopped: the ping read ... line means the fork did not share the cache, and the mod has already stopped. /cache-warm off ends the test.

Limits

  • The 50-minute ping assumes the 1-hour cache tier, which the main conversation uses.
  • A warm ping only proves the cache was warm at that moment. A model or effort switch, an edited CLAUDE.md or a changed tool list breaks the prefix no matter the time, and your next message pays.
  • A ping's output cannot be capped; a model at high effort may think before it answers. The status line prices what the ping really billed.
  • The resume logic is covered by hook tests that raise classic.SessionStart with the resume fields, and by a live check of a resumed interactive session in tmux. Whether a resume keeps the whole prefix is out of the mod's hands: it re-wrote 42k of 116k tokens in one measured session and 4k in another.
  • The cold-write tally is per session and lives in memory. /clear empties it.
  • Fast mode is read from the fastMode setting, the saved preference, not from the request. The mod does not see Claude Code fall back to standard speed within a session (a fast mode rate-limit cooldown, usage credits that ran out, an organization that turned fast mode off). Those turns bill standard rates while the mod prices them as fast.
  • Whether a ping, a $.model.fork, runs at fast speed while the session does has not been measured; the mod prices it at the session's rates.

Development

make install     # eslint, typescript-eslint, typescript
make lint        # complexity limit 10, the build fails above it
make typecheck   # needs .claude/types/ from /plugin-types
make validate
make test        # claude plugin test

更多类似作品