ClaudeMods
☰
ZH-CN
● 0 人在线 · 浏览 0 次
赞助提交作品
GitHub 仓库 · 发布者 KilimcininKorOglu

gemini-compact

将压缩工作从 Claude 转交给 Gemini:在 summary 模式下,Gemini 总结对话,最新消息逐字保留;在 prune 模式下,所有消息都保留,由 Gemini 决定旧的工具调用是保留、裁剪还是删除。直到执行 /gemini-compact on 前都处于关闭状态;密钥、层级、模型和思考级别由 gemini-core 保存。

KilimcininKorOglu@KilimcininKorOglu

KilimcininKorOglu/claude-code-mods/tree/main/plugins/gemini-compact

已翻译

关于这个 mod

gemini-compact

当上下文填满时,Claude 会再发起一次 Claude 请求来压缩上下文:读取整个上下文并写出摘要,这次请求会计入 Claude 用量。此 Mod 把这项工作交给 Gemini。它有两种模式:

  • summary(默认):Gemini 总结最新 6 条消息之前的对话,这些消息会在摘要之后逐字保留。Claude 不写摘要;只有 Gemini 失败时才运行引擎内置的摘要。
  • prune:Gemini 针对每个较早的工具调用回答:保留调用及其输出、保留但裁剪输出,还是删除。每条用户和助手消息都逐字保留。不会生成摘要。

裁剪思路参考了 Tamara Tran 的 fast-jev-compaction(tamaratran/fast-jev-compaction),该项目会请求 TypeSafe Jev。这里的代码是新的,请求对象是 Gemini。

模式选择

两种模式都会用 Gemini 请求替代内置压缩中的 Claude 请求。

压缩之后,每次 Claude 请求都会读取剩余内容。在 summary 模式下,剩余内容是摘要和最新消息,接近内置摘要留下的内容。在 prune 模式下,剩余内容是对话的每条消息减去被删除的工具输出,体积更大,因此之后的每次请求都会读取更多内容。这些大小尚未在长会话中测量。

想最大限度节省 Claude 用量时选择 summary 模式。每条消息的准确措辞比体积更重要时选择 prune 模式。

Summary 模式

  1. 在 /compact、引擎自己的压缩,以及主循环某一回合以回答结束且上下文超过阈值之后,session.compact hook 会接管对话。
  2. 最新 6 条消息会保留。裁剪位置向前移到一条助手消息,因此不会在没有对应调用的情况下保留工具结果;保留部分会在摘要之后以一条助手消息开头。如果找不到助手消息作为裁剪点,就会总结全部内容。
  3. 一个 generateContent 请求会把裁剪点之前的所有内容发送给 Gemini:每条消息,以及每次调用的输入和完整输出。当文本超过 summaryMaxInputChars(默认 2,000,000 个字符)时,会把最长的输出裁到同一个长度,每个输出都保留开头和结尾;即使没有任何输出,超过限制的对话也会失败。
  4. Gemini 会用纯文本写出分为 9 个部分的摘要:请求和意图、技术概念、文件和代码、错误及修复、问题解决过程、每条用户消息的原文、待处理任务、当前工作和下一步。你在 /compact 后写的文字也会随之发送给 Gemini。
  5. 对话会变成一条用户消息(先是说明,再是摘要),后面接着保留的消息;这些消息会作为引擎自己的消息送回去。
  6. 在没有密钥、最新消息之前没有内容、Gemini 失败、摘要少于 200 个字符或因输出限制(32,768 个 token)被裁剪,或者结果没有比原对话更短时,内置摘要会运行,并用一行说明原因。

在两种模式中,gemini-core 都会用它为 gemini-compact 保存的密钥、模型和思考级别构建请求,并读取回答。HTTP 503(“high demand”)之后,Mod 会在 1 s、2 s 和 3 s 后重试,最多共 4 次,并且不会开始等待结束时间超过 60 s 的尝试。收到 429 或密钥错误后,如果 gemini-core 还有下一把密钥,会用它接手请求。

在 2.1.277 上的一次实测中,使用 gemini-3.5-flash-lite 执行 /compact 花了 2.6 秒,Claude 没有发送压缩请求;之后模型说出了一个单词和一个文件,而它们只出现在被总结的部分。

Prune 模式

  1. 同样的 3 个触发条件会到达 session.compact hook。
  2. 第一条消息或最新 6 条消息中的工具调用,或者结果位于其中一条消息里的工具调用,会完整保留。其他调用都会获得一个 ID(c1、c2、……)。如果没有这样的调用,就运行内置摘要。
  3. 一个 generateContent 请求会把对话发送给 Gemini:每条消息,以及每次调用的输入和完整输出。当超过 maxInputChars(默认 400,000 个字符)时,会用 summary 模式相同的方式裁剪最长输出。响应 schema 会对每个 ID 恰好给出一个回答:keep、truncate 或 drop。
  4. Mod 会检查回答(每个 ID 恰好一次,不能出现其他 ID),然后重建对话:
    • keep:保留调用和调用输出。
    • truncate:保留调用,输出保留前 300 个字符,再加上一行说明它已被裁剪。若输出只比这个长度多至多 120 个字符,则完整保留。
    • drop:删除调用及其输出。在最近的助手消息上添加说明,列出被移除的调用,例如 [gemini-compact removed 1 earlier tool call(s) and their output after a compaction; they ran: Bash(ls -la /usr/bin)]。没有这条说明时,模型会读到工作已经消失的回复,并说自己从未做过那项工作(在 2.1.277 上测得)。
    • 回答没有处理的消息会作为引擎自己的消息送回去。
  5. 当结果缩短不到 25%,或发生任何失败(没有密钥、HTTP 错误、回答破坏 schema)时,内置摘要会运行,并用一行说明原因。

在 2.1.277 上的一次实测中,使用 gemini-3.5-flash-lite 执行 /compact 花了 1.1 秒。Gemini 删除了一份 ls 列表,保留了用户准备编辑的文件的 cat 输出,对话缩短了 93%。

它显示的内容

transcript 中的一行,不会发送给模型;另外还有一条会保留 15 秒的提示:

gemini-compact: summary: 58 → 7 messages · 91% smaller · 312k in, 5k out
gemini-compact: kept 9/11 messages · 93% smaller · 1 dropped, 0 truncated · 5k in, 59 out
gemini-compact: built-in summary: under 25% smaller (kept 14/16 messages · 3% smaller · ...)
gemini-compact: built-in summary: Gemini HTTP 429: Resource has been exhausted

在免费层级中,提示会追加 · sent to Gemini free tier。引擎跳过或失败的自动压缩也会写一行(automatic compaction skipped: ...、automatic compaction failed: ...)。

命令

/gemini-compact              on or off, mode, the model and thinking level gemini-core holds, threshold, tier, whether a key is set, the last result
/gemini-compact on | off     off leaves every compaction to the built-in summary; on is refused while gemini-core has no key
/gemini-compact mode summary | mode prune
/gemini-compact at <1-99>    compact after a turn that ends with the context over this percentage
/gemini-compact at off       no automatic compaction; /compact and the engine's own compaction still ask Gemini
/gemini-compact reset        back to the plugin options, and off

安装后 Mod 处于关闭状态:每次压缩都使用内置压缩,不会在 Mod 的阈值上启动,也不会有任何内容发送给 Gemini,直到执行 /gemini-compact on。命令设置会跨会话保留并立即生效。一次压缩启动后,在某一回合结束且上下文低于阈值之前,Mod 不会启动另一次压缩;因此,如果上下文持续超标,也不会在每个回合后都压缩。

密钥、层级、模型(默认 gemini-3.5-flash-lite)和思考级别属于 gemini-core:

/gemini-core model compact gemini-3.5-flash
/gemini-core thinking compact low
/gemini-core paid

免费层级还是付费层级

对话包含你的 prompt、模型运行的命令以及它读取的文件内容。在免费层级中,Google 可能会使用这些内容,人工审查员也可能读取它们;gemini-core README 引用了 Gemini API Additional Terms。对于不会展示给 Google 的项目,请使用启用了计费的密钥,并设置 /gemini-core paid。没有 Mod 能判断密钥属于哪个层级;层级设置只会选择警告内容。

Google AI Studio 会显示每个模型的免费层级限制;文档不会显示。那些限制没有经过测量。长对话的摘要是一个大型请求,因此每分钟 token 限制可能会以 HTTP 429 拒绝请求;然后运行内置摘要。

安装

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install gemini-compact@kilimcininkoroglu-mods

它依赖 gemini-core,claude plugin install 会添加该依赖。函数 hooks 属于早期预览功能。Claude Code 2.1.288 及更高版本默认加载它们,因此不需要打开任何开关。

要在单个会话中从本地 checkout 加载它,请将 gemini-core 放在它旁边:

claude --plugin-dir plugins/gemini-core --plugin-dir plugins/gemini-compact

安装后

  1. 按照 gemini-core 的 After installing 章节设置 Gemini 密钥和层级,然后重启 Claude Code。
  2. 运行 /gemini-compact on。没有密钥时,它会回答 still off: gemini-core has no Gemini key 并保持关闭。
  3. 运行 /gemini-compact。第一行会显示 on · summary · <model> · thinking ... · automatic at 60% · <tier> tier · key set。
  4. 执行一次 /compact。transcript 行应以 gemini-compact: summary: 开头。以 built-in summary: 开头的行会说明没有使用 Gemini 的原因。

从 0.2.x 更新后:claude plugin update 不会添加 gemini-core(在 2.1.278 上测得),因此请运行一次 claude plugin install gemini-core@kilimcininkoroglu-mods。0.3.0 版将密钥、层级和模型移到了 gemini-core;之前保存的 apiKey、tier 和 model 选项,以及 /gemini-compact free|paid|model 设置,不再读取,因此请在 gemini-core 中重新设置。mode 和 at 设置会保留。0.4.0 版将 Mod 默认设为关闭:从更早版本更新后,除非之前运行过 /gemini-compact on,否则它会处于关闭状态,因此请运行一次 /gemini-compact on。

选项

| Option | Default | What it sets | |---|---|---| | mode | summary | summary 或 prune;/gemini-compact mode 会覆盖它 | | compactAtPercent | 60 | 自动阈值,0 到 99;0 会关闭它;/gemini-compact at 会覆盖它 | | keepRecent | 6 | 逐字保留的最新消息(summary),或永远不会送出进行判断的消息调用(prune);0 到 1000 | | minReduction | 0.25 | Prune 模式:低于此比例(0 到 1)时运行内置摘要 | | headChars | 300 | Prune 模式:裁剪输出时保留的字符数;0 到 100,000 | | maxInputChars | 400000 | Prune 模式:最多发送给 Gemini 的字符数;10,000 到 4,000,000 | | summaryMaxInputChars | 2000000 | Summary 模式:最多发送给 Gemini 的字符数;10,000 到 4,000,000 |

超出范围的值会回退到默认值。

它可以接触什么

在 Claude Code 2.1.283 上通过 claude plugin validate 验证:

❯ ./register.ts hooks: session.start, command.run{command=gemini-compact}, session.compact, turn.complete
❯ ./register.ts calls: $.clock.now (via askGemini), $.clock.sleep (via askGemini), $.command.register, $.gemini.enroll, $.gemini.read (via askGemini), $.gemini.request (via askGemini), $.gemini.settings (via compactWithGemini, runCommand, storePatch), $.http.fetch (via askGemini), $.session.compact (via maybeCompact), $.session.usage (via maybeCompact), $.store.delete (via runCommand), $.store.get (via loadConfig), $.store.set (via storePatch), $.ui.log, $.ui.toast (via report)

Reach L3,接触网络。

1. 读取:每次压缩时的对话(消息、工具输入和输出);每个主循环回合之后的上下文填充情况;自己的 $.store;以及来自 gemini-core、带有密钥的请求
2. 运行:不启动进程;在某回合结束且超过阈值后调用一次 $.session.compact,直到上下文再次低于阈值之前最多调用一次
3. 发送:对话(summary:除最新消息之外的全部内容;prune:全部内容),每次压缩一个请求(503 后最多 4 次,以及 429 或密钥错误后每多一把密钥再一次),发送到 gemini-core 构建的 URL(generativelanguage.googleapis.com),密钥放在 x-goog-api-key header 中,从不放入 URL
4. 持久化:在 $.store 中保存 3 个命令设置(enabled、mode、atPercent);上次结果保存在 i

安装

请先查看作者 README 确认 marketplace 和插件名称;命令可能随仓库结构改变。

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install gemini-compact
原文 / README

gemini-compact

When the context fills up, Claude compacts it with one more Claude request: it reads the whole context and writes a summary, and that request counts against your Claude usage. This mod hands that job to Gemini. It has two modes:

  • summary (the default): Gemini summarizes the conversation before the newest 6 messages, and those messages stay verbatim after the summary. Claude writes no summary; the engine's built-in summary runs only when Gemini fails.
  • prune: Gemini answers, for every older tool call, whether the call and its output stay, stay with a cut output, or go. Every user and assistant message stays verbatim. Nothing is summarized.

The prune idea follows fast-jev-compaction by Tamara Tran (tamaratran/fast-jev-compaction), which asks TypeSafe Jev. The code here is new and asks Gemini.

Which mode

Both modes replace the Claude request of the built-in compaction with a Gemini request.

After the compaction, every Claude request reads what is left. In summary mode that is the summary and the newest messages, close to what the built-in summary leaves. In prune mode it is every message of the conversation less the dropped tool output, which is larger, so each following request reads more. The sizes were not measured on a long session.

Pick summary mode to save the most Claude usage. Pick prune mode when the exact wording of every message matters more than the size.

Summary mode

  1. At /compact, at the engine's own compaction, and after a main-loop turn that ends with an answer and the context over the threshold, the session.compact hook takes the conversation.
  2. The newest 6 messages stay. The cut moves back to an assistant message, so a tool result is never kept without its call and the kept part opens with an assistant message after the summary. With no assistant message to cut at, everything is summarized.
  3. One generateContent request sends everything before the cut to Gemini: every message, and each call with its input and its full output. When the text is over summaryMaxInputChars (2,000,000 characters by default), the longest outputs are cut to one common length, each keeping its head and tail; a conversation over the limit even without any output fails.
  4. Gemini writes a plain-text summary in nine sections: the request and intent, technical concepts, files and code, errors and fixes, problem solving, every user message verbatim, pending tasks, the current work, and the next step. The text you write after /compact goes to Gemini with it.
  5. The conversation becomes one user message (a note, then the summary) followed by the kept messages, which go back as the engine's own messages.
  6. The built-in summary runs, and one line says why, when there is no key, nothing lies before the newest messages, Gemini fails, the summary is under 200 characters or cut at the output limit (32,768 tokens), or the result is not smaller than the conversation.

In both modes gemini-core builds the request with the key, the model and the thinking level it holds for gemini-compact, and reads the answer. After an HTTP 503 ("high demand") the mod asks again after 1 s, 2 s and 3 s, at most four attempts in all, and no attempt starts whose wait would end past 60 s. After a 429 or a key error, gemini-core hands over the request with its next key, when it holds one.

In a live check on 2.1.277, /compact took 2.6 seconds with gemini-3.5-flash-lite, Claude sent no compaction request, and after it the model named a word and a file that appeared only in the summarized part.

Prune mode

  1. The same three triggers reach the session.compact hook.
  2. A tool call in the first message or in the newest 6 messages, or whose result lies in one of them, is kept whole. Every other call gets an id (c1, c2, ...). With no such call, the built-in summary runs.
  3. One generateContent request sends the conversation to Gemini: every message, and each call with its input and its full output. Over maxInputChars (400,000 characters by default) the longest outputs are cut the same way as in summary mode. A response schema allows exactly one answer per id: keep, truncate or drop.
  4. The mod checks the answer (every id exactly once, no other id) and rebuilds the conversation:
    • keep: the call and its output stay.
    • truncate: the call stays, and the output keeps its first 300 characters and one line that says it was cut. An output at most 120 characters longer than that stays whole.
    • drop: the call and its output go. A note on the nearest assistant message names the removed calls, for example [gemini-compact removed 1 earlier tool call(s) and their output after a compaction; they ran: Bash(ls -la /usr/bin)]. Without the note, the model read a reply whose work was gone and said it had never done that work (measured on 2.1.277).
    • A message the answer does not touch goes back as the engine's own message.
  5. When the result is less than 25% smaller, or anything fails (no key, an HTTP error, an answer that breaks the schema), the engine's built-in summary runs and one line says why.

In a live check on 2.1.277, /compact took 1.1 seconds with gemini-3.5-flash-lite. Gemini dropped an ls listing and kept the cat output of a file the user was about to edit, and the conversation became 93% smaller.

What it shows

One line in the transcript, not sent to the model, and a toast that stays for 15 seconds:

gemini-compact: summary: 58 → 7 messages · 91% smaller · 312k in, 5k out
gemini-compact: kept 9/11 messages · 93% smaller · 1 dropped, 0 truncated · 5k in, 59 out
gemini-compact: built-in summary: under 25% smaller (kept 14/16 messages · 3% smaller · ...)
gemini-compact: built-in summary: Gemini HTTP 429: Resource has been exhausted

On the free tier the toast adds · sent to Gemini free tier. An automatic compaction the engine skips or that fails writes one line too (automatic compaction skipped: ..., automatic compaction failed: ...).

Command

/gemini-compact              on or off, mode, the model and thinking level gemini-core holds, threshold, tier, whether a key is set, the last result
/gemini-compact on | off     off leaves every compaction to the built-in summary; on is refused while gemini-core has no key
/gemini-compact mode summary | mode prune
/gemini-compact at <1-99>    compact after a turn that ends with the context over this percentage
/gemini-compact at off       no automatic compaction; /compact and the engine's own compaction still ask Gemini
/gemini-compact reset        back to the plugin options, and off

The mod is off after an install: every compaction is the built-in one, none starts on the mod's threshold, and nothing goes to Gemini until /gemini-compact on. The command settings are kept across sessions and take effect at once. After a compaction it started, the mod starts no other one until a turn ends with the context under the threshold, so a context that stays over it does not compact after every turn.

The key, the tier, the model (default gemini-3.5-flash-lite) and the thinking level belong to gemini-core:

/gemini-core model compact gemini-3.5-flash
/gemini-core thinking compact low
/gemini-core paid

Free tier or paid tier

The conversation holds your prompts, the commands the model ran and the contents of the files it read. On the free tier Google may use them and human reviewers may read them; the gemini-core README quotes the Gemini API Additional Terms. On a project you would not show to Google, use a key with billing enabled and set /gemini-core paid. No mod can tell which tier a key is on; the tier setting only chooses the warning.

Google AI Studio shows the free tier limits per model; the documentation does not. They were not measured. A summary of a long conversation is one large request, so a per-minute token limit can refuse it with HTTP 429; the built-in summary then runs.

Install

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install gemini-compact@kilimcininkoroglu-mods

It depends on gemini-core, which claude plugin install adds. Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.

To load it from a local checkout for one session, put gemini-core beside it:

claude --plugin-dir plugins/gemini-core --plugin-dir plugins/gemini-compact

After installing

  1. Set the Gemini key and the tier in gemini-core, as its After installing section says, then restart Claude Code.
  2. Run /gemini-compact on. Without a key it answers still off: gemini-core has no Gemini key and stays off.
  3. Run /gemini-compact. The first line reads on · summary · <model> · thinking ... · automatic at 60% · <tier> tier · key set.
  4. Run /compact once. The transcript line should start with gemini-compact: summary:. A line that starts with built-in summary: names why Gemini was not used.

After an update from 0.2.x: claude plugin update does not add gemini-core (measured on 2.1.278), so run claude plugin install gemini-core@kilimcininkoroglu-mods once. Version 0.3.0 moved the key, tier and model to gemini-core; the apiKey, tier and model options and the settings /gemini-compact free|paid|model stored before are no longer read, so set them again in gemini-core. The mode and at settings stay. Version 0.4.0 made the mod off by default: after an update from an earlier version it is off unless you ran /gemini-compact on before, so run /gemini-compact on once.

Options

| Option | Default | What it sets | |---|---|---| | mode | summary | summary or prune; /gemini-compact mode overrides it | | compactAtPercent | 60 | The automatic threshold, 0 to 99; 0 turns it off; /gemini-compact at overrides it | | keepRecent | 6 | Newest messages kept verbatim (summary) or whose calls are never sent for a decision (prune); 0 to 1000 | | minReduction | 0.25 | Prune mode: below this fraction (0 to 1) the built-in summary runs | | headChars | 300 | Prune mode: characters kept of a truncated output; 0 to 100,000 | | maxInputChars | 400000 | Prune mode: characters sent to Gemini at most; 10,000 to 4,000,000 | | summaryMaxInputChars | 2000000 | Summary mode: characters sent to Gemini at most; 10,000 to 4,000,000 |

A value outside its range falls back to the default.

What it can reach

Validated with claude plugin validate on Claude Code 2.1.283:

❯ ./register.ts hooks: session.start, command.run{command=gemini-compact}, session.compact, turn.complete
❯ ./register.ts calls: $.clock.now (via askGemini), $.clock.sleep (via askGemini), $.command.register, $.gemini.enroll, $.gemini.read (via askGemini), $.gemini.request (via askGemini), $.gemini.settings (via compactWithGemini, runCommand, storePatch), $.http.fetch (via askGemini), $.session.compact (via maybeCompact), $.session.usage (via maybeCompact), $.store.delete (via runCommand), $.store.get (via loadConfig), $.store.set (via storePatch), $.ui.log, $.ui.toast (via report)

Reach L3, reaches the network.

1. Reads:    the conversation at each compaction (messages, tool inputs and outputs); the context fill after each main-loop turn; its own $.store; from gemini-core, the request with the key
2. Runs:     no process; one $.session.compact after a turn that ends over the threshold, at most once until the context was under it again
3. Sends:    the conversation (summary: all but the newest messages; prune: all of it), one request per compaction (up to four after a 503, and once more per extra key after a 429 or a key error), to the URL gemini-core builds (generativelanguage.googleapis.com) with the key in the x-goog-api-key header, never in the URL
4. Persists: in $.store, the three command settings (enabled, mode, atPercent); the last result lives in memory
5. Hostile input: the Gemini answer is untrusted: a prune answer is applied only in the schema shape with every candidate id once; a summary becomes the text of one user message the model reads, so a hostile summary can steer the model, as text in a file it reads can; anything malformed falls back to the built-in summary

Limits

  • A summary and a drop are a model's judgment. A summary loses detail the newest messages do not repeat. The prune note tells the model which calls ran, so it can run a tool again.
  • After the compaction the context is written to the cache again. In prune mode it stays larger than a built-in summary, so the next message pays a larger cache write.
  • A summary of a long conversation takes Gemini longer, and the compaction waits for it. Only short conversations were timed.
  • A subagent's own compaction is left to the engine.
  • A compaction the engine precomputes (precompute) also asks Gemini. Whether the engine reuses that result for the compaction that follows was not measured.
  • The test engine of claude plugin test passes no trigger to a $.session.compact() call. The plugin trigger and the hook that answers it were measured in a live session.

Development

make install     # eslint, typescript-eslint, typescript
make lint        # complexity limit 10, the build fails above it
make typecheck   # needs .claude/types/ from /plugin-types
make validate
make test        # claude plugin test

更多类似作品