
我做了 conPACT:任务完成后 Claude Code 自动压缩自己的上下文,并在提示缓存过期前提议压缩闲置的会话
conPACT 是一款 Claude Code 模组,任务完成后通过 MCP 工具自动压缩会话的上下文,并在提示缓存过期前提示压缩闲置的会话。它以进程内插件 hooks 模组的形式运行于 Claude Code 2.1.284 及以上版本,通过 Remote Control 时则改用 Stop hook 发送 /compact,也支持 Codex 与 ChatGPT Desktop。作者称约节省 4.7% 的花费,峰值上下文也更小,采用 MIT 许可证。

关于这个 mod
作者介绍 conPACT,它通过 MCP 工具 queue_compaction 排入自我压缩,让压缩在最终回答之后以真正的 /compact 执行,在总结实时上下文的同时保留聊天记录。
它支持焦点指示与最小大小设置。缓存到期前五分钟会弹出提示,让用户选择立即压缩,或为该会话开启自动压缩。提示栏上方还有一行状态。
作者用 Claude Code 开发,并从九月中旬起在自己的开发会话中使用。它以模组形式集成,使用 Claude Code 2.1.284 及以上版本新的进程内插件 hooks;通过 Remote Control 时,则由 Stop hook 发送 /compact。它也可以通过可选的 sidecar 支持 ChatGPT Desktop(Codex),并通过 Stop hook 支持 Codex CLI。
作者的重点心得:/compact 无法在 turn 进行中发送,所以压缩只会在 turn 结束后执行,且最多一次。时机比大小更重要。
以 187,978 次 API 调用为样本,conPACT 的 134 次压缩都在缓存仍有效(warm cache)时完成,手动压缩 121 次中只有 30 次在缓存有效时进行。按标价计算,每次压缩约为 $5.67 对 $0.20。峰值上下文的中位数从 501k 降到 337k tokens;冷启动恢复时重新读取的中位数从 493k 降到 255k tokens;冷启动在花费中的占比从 6.4% 降到 3.8%,而且没有返工的代价。
需求:Python 3.11 及以上,只使用标准库,可在 Windows、Linux、macOS 运行,采用 MIT 许可证。GitHub:https://github.com/st0nebridge/conPACT
安装
安装方法请查看原始来源。
原文 / README
Long Claude Code sessions cost you twice. The context keeps growing after the work that needed it is done, and if you come back after the prompt cache has expired, your next message re-reads all of it uncached. /compact fixes both, but only if you type it at the right moment. I usually didn't. What it does When a piece of work is finished, Claude queues a compaction of its own session through an MCP tool ( queue_compaction ). It runs after the final answer, as a real /compact : your chat history stays visible and only the live context is summarised. It can take a focus ("keep the plan and the open decisions") and a minimum size. When a big session sits idle, a toast appears five minutes before the cache expires and offers to compact it now, or always for that session. A row above the prompt follows the request: queued, compacting, then what it came to. How Claude Code was used I built it with Claude Code, and it has been compacting its own development sessions since mid-September. In Claude Code 2.1.284+ it ships as a mod (the new in-process plugin hooks), so the session compacts itself and draws the row above the prompt. Without the mod, a Stop hook sends /compact over Remote Control. It also works for ChatGPT Desktop (Codex) through an optional sidecar, and for the Codex CLI through a Stop hook. What I learned You can't send /compact mid-turn. A busy session receives it as plain text and it never runs. So the tool only records the request, and the compaction happens after the turn ends: always after the final answer, and at most once. When you compact matters more than how small. I measured it over 187,978 of my own API calls (12 days with conPACT, 80 before). All 134 conPACT compactions ran on a warm cache, against 30 of the 121 I'd typed by hand. A cold compaction re-reads the whole context at the cache-write price first, so that's roughly $5.67 against $0.20 per compaction at list price. Median peak context per session fell from 501k to 337k. The idle toast does its job: a cold restart now starts from a compacted context. The median re-read on a cold resume fell from 493k to 255k tokens, and cold resumes' share of spend fell from 6.4% to 3.8%. No rework penalty. Claude re-reads some files after a compaction, but the next prompt was a correction 5.5% of the time, against 6.1% in ordinary turns. Net: about 4.7% of spend saved at list price (1.8–13%, depending on what you assume I'd have done otherwise). It's modest, and the "after" period is only 12 active days, but every compaction saves more than it costs. Python 3.11+, standard library only. Windows, Linux and macOS (the test suite runs on all three; most of my live use is on Windows). MIT licensed. GitHub: https://github.com/st0nebridge/conPACT Feedback welcome, especially from anyone on macOS or Linux. submitted by /u/stonebrigade [link] [comments]

