ClaudeMods
☰
ZH-TW
● 0 人在線上 · 瀏覽 0 次
贊助提交作品
GitHub 儲存庫 · 發布者 KilimcininKorOglu

gemini-compact

把壓縮工作從 Claude 交給 Gemini:在 summary 模式中,Gemini 摘要對話,最新訊息逐字保留;在 prune 模式中,所有訊息都保留,由 Gemini 決定較舊的工具呼叫要保留、裁切或刪除。直到執行 /gemini-compact on 前都保持關閉;金鑰、層級、模型與思考層級由 gemini-core 保存。

KilimcininKorOglu@KilimcininKorOglu

KilimcininKorOglu/claude-code-mods/tree/main/plugins/gemini-compact

已翻譯

關於這個 mod

gemini-compact

當上下文填滿時,Claude 會再送出一個 Claude 請求來壓縮上下文:讀取整個上下文並寫出摘要,而這個請求會計入 Claude 用量。這個 Mod 將這項工作交給 Gemini。它有兩種模式:

  • summary(預設):Gemini 摘要最新 6 則訊息之前的對話,而這些訊息會在摘要後逐字保留。Claude 不會寫摘要;只有 Gemini 失敗時才會執行階段內建的摘要。
  • prune:Gemini 會針對每一個較舊的工具呼叫回答:保留呼叫及其輸出、保留但裁切輸出,或刪除。每一則使用者與助理訊息都逐字保留。不會進行摘要。

裁切概念沿用了 Tamara Tran 的 fast-jev-compaction(tamaratran/fast-jev-compaction),它會詢問 TypeSafe Jev。這裡的程式碼是新的,詢問的是 Gemini。

選擇模式

兩種模式都會用 Gemini 請求取代內建壓縮中的 Claude 請求。

壓縮後,每個 Claude 請求都會讀取剩下的內容。在 summary 模式中,剩下的是摘要與最新訊息,接近內建摘要留下的內容。在 prune 模式中,剩下的是對話中的每一則訊息,扣除被刪除的工具輸出;內容較大,因此後續每個請求都會讀取更多內容。這些大小尚未在長工作階段中測量。

想最大幅度節省 Claude 用量時選擇 summary 模式。每則訊息的精確措辭比大小更重要時,選擇 prune 模式。

Summary 模式

  1. 在 /compact、執行階段本身的壓縮,以及主要迴圈某一回合以回答結束且上下文超過門檻後,session.compact hook 會接手對話。
  2. 最新 6 則訊息會保留。裁切位置會往前移到助理訊息,因此不會在沒有對應呼叫時保留工具結果;保留的部分會在摘要後以助理訊息開頭。沒有助理訊息可作為裁切點時,所有內容都會摘要。
  3. 一個 generateContent 請求會把裁切點之前的所有內容傳給 Gemini:每一則訊息,以及每個呼叫的輸入和完整輸出。當文字超過 summaryMaxInputChars(預設為 2,000,000 個字元)時,最長的輸出會裁切到同一個長度,每個輸出都保留開頭和結尾;即使沒有任何輸出,超過限制的對話也會失敗。
  4. Gemini 會以純文字寫出 9 個章節的摘要:請求與意圖、技術概念、檔案與程式碼、錯誤與修正、問題解決過程、每一則使用者訊息的原文、待辦事項、目前的工作,以及下一步。你在 /compact 後寫的文字也會一起傳給 Gemini。
  5. 對話會變成一則使用者訊息(先是備註,再是摘要),後面接著保留的訊息;這些訊息會以執行階段本身的訊息送回去。
  6. 沒有金鑰、最新訊息之前沒有內容、Gemini 失敗、摘要少於 200 個字元或因輸出限制(32,768 個 token)而被裁切,或結果沒有比對話更短時,內建摘要會執行,並以一行說明原因。

在兩種模式中,gemini-core 都會使用它為 gemini-compact 保存的金鑰、模型與思考層級建構請求,並讀取回答。HTTP 503(「high demand」)後,Mod 會在 1 s、2 s 與 3 s 後再次詢問,最多總共 4 次,而且不會開始等待結束時間超過 60 s 的嘗試。收到 429 或金鑰錯誤後,如果 gemini-core 還有下一把金鑰,就會用它接手請求。

在 2.1.277 上的一次實測中,使用 gemini-3.5-flash-lite 執行 /compact 花了 2.6 秒,Claude 沒有送出壓縮請求;之後模型說出了一個只出現在摘要部分的單字和檔案。

Prune 模式

  1. 同樣的 3 個觸發條件會到達 session.compact hook。
  2. 第一則訊息或最新 6 則訊息中的工具呼叫,或其結果位於其中一則訊息的工具呼叫,會完整保留。其他呼叫都會取得 ID(c1、c2、……)。沒有這類呼叫時,會執行內建摘要。
  3. 一個 generateContent 請求會把對話傳給 Gemini:每一則訊息,以及每個呼叫的輸入和完整輸出。超過 maxInputChars(預設 400,000 個字元)時,會用 summary 模式相同的方式裁切最長的輸出。回應 schema 會對每個 ID 恰好給出一個答案:keep、truncate 或 drop。
  4. Mod 會檢查答案(每個 ID 恰好一次,不得有其他 ID),然後重建對話:
    • keep:保留呼叫及其輸出。
    • truncate:保留呼叫,輸出保留前 300 個字元,並加上一行說明它已被裁切。輸出若只比這個長度多至多 120 個字元,就會完整保留。
    • drop:移除呼叫及其輸出。在最近的助理訊息上加註被移除的呼叫,例如 [gemini-compact removed 1 earlier tool call(s) and their output after a compaction; they ran: Bash(ls -la /usr/bin)]。沒有這則註記時,模型會讀到工作已經消失的回覆,並說自己從未做過那項工作(在 2.1.277 上測得)。
    • 答案沒有處理的訊息會以執行階段本身的訊息送回去。
  5. 結果縮短不到 25%,或發生任何失敗(沒有金鑰、HTTP 錯誤、答案破壞 schema)時,內建摘要會執行,並以一行說明原因。

在 2.1.277 上的一次實測中,使用 gemini-3.5-flash-lite 執行 /compact 花了 1.1 秒。Gemini 刪除了一份 ls 清單,保留了使用者準備編輯的檔案之 cat 輸出,對話縮短了 93%。

顯示內容

逐字稿中的一行,不會傳給模型;另外還有一則會停留 15 秒的提示訊息:

gemini-compact: summary: 58 → 7 messages · 91% smaller · 312k in, 5k out
gemini-compact: kept 9/11 messages · 93% smaller · 1 dropped, 0 truncated · 5k in, 59 out
gemini-compact: built-in summary: under 25% smaller (kept 14/16 messages · 3% smaller · ...)
gemini-compact: built-in summary: Gemini HTTP 429: Resource has been exhausted

在免費層級中,提示訊息會加上 · sent to Gemini free tier。執行階段跳過或失敗的自動壓縮也會寫一行(automatic compaction skipped: ...、automatic compaction failed: ...)。

指令

/gemini-compact              on or off, mode, the model and thinking level gemini-core holds, threshold, tier, whether a key is set, the last result
/gemini-compact on | off     off leaves every compaction to the built-in summary; on is refused while gemini-core has no key
/gemini-compact mode summary | mode prune
/gemini-compact at <1-99>    compact after a turn that ends with the context over this percentage
/gemini-compact at off       no automatic compaction; /compact and the engine's own compaction still ask Gemini
/gemini-compact reset        back to the plugin options, and off

安裝後 Mod 會關閉:每次壓縮都是內建的,不會在 Mod 的門檻啟動,在執行 /gemini-compact on 前也不會有任何內容傳給 Gemini。指令設定會跨工作階段保留並立即生效。一次壓縮啟動後,直到某個回合結束且上下文低於門檻前,Mod 不會再啟動其他壓縮;因此上下文持續超過門檻時,不會在每個回合後都壓縮。

金鑰、層級、模型(預設 gemini-3.5-flash-lite)與思考層級屬於 gemini-core:

/gemini-core model compact gemini-3.5-flash
/gemini-core thinking compact low
/gemini-core paid

免費層級或付費層級

對話包含你的 prompt、模型執行的指令,以及它讀取的檔案內容。在免費層級中,Google 可能會使用這些內容,人工檢閱者也可能讀取;gemini-core README 引用了 Gemini API Additional Terms。對於不會給 Google 看的專案,請使用已啟用計費的金鑰,並設定 /gemini-core paid。沒有任何 Mod 能判斷金鑰屬於哪個層級;層級設定只會決定警告內容。

Google AI Studio 顯示每個模型的免費層級限制;文件沒有顯示。這些限制尚未測量。長對話的摘要是一個大型請求,因此每分鐘 token 限制可能會以 HTTP 429 拒絕它;接著會執行內建摘要。

安裝

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install gemini-compact@kilimcininkoroglu-mods

它相依於 gemini-core,claude plugin install 會加入它。函式 hook 屬於搶先體驗功能。Claude Code 2.1.288 及更新版本預設會載入它們,因此不必切換任何設定。

若要在單一工作階段從本機 checkout 載入,請把 gemini-core 放在它旁邊:

claude --plugin-dir plugins/gemini-core --plugin-dir plugins/gemini-compact

安裝後

  1. 按照 gemini-core 的 After installing 章節設定 Gemini 金鑰和層級,然後重新啟動 Claude Code。
  2. 執行 /gemini-compact on。沒有金鑰時,它會回答 still off: gemini-core has no Gemini key 並保持關閉。
  3. 執行 /gemini-compact。第一行會顯示 on · summary · <model> · thinking ... · automatic at 60% · <tier> tier · key set。
  4. 執行一次 /compact。逐字稿那一行應以 gemini-compact: summary: 開頭。以 built-in summary: 開頭的行會說明沒有使用 Gemini 的原因。

從 0.2.x 更新後:claude plugin update 不會加入 gemini-core(在 2.1.278 上測得),因此請執行一次 claude plugin install gemini-core@kilimcininkoroglu-mods。0.3.0 版將金鑰、層級和模型移到 gemini-core;之前儲存的 apiKey、tier、model 選項,以及 /gemini-compact free|paid|model 設定都不會再讀取,因此請在 gemini-core 中重新設定。mode 和 at 設定會保留。0.4.0 版預設關閉 Mod:從較早版本更新後,除非之前執行過 /gemini-compact on,否則它會關閉,因此請執行一次 /gemini-compact on。

選項

| Option | Default | What it sets | |---|---|---| | mode | summary | summary 或 prune;/gemini-compact mode 會覆寫它 | | compactAtPercent | 60 | 自動門檻,0 到 99;0 會關閉;/gemini-compact at 會覆寫它 | | keepRecent | 6 | 逐字保留的最新訊息(summary),或永遠不送出判斷的呼叫(prune);0 到 1000 | | minReduction | 0.25 | Prune 模式:低於這個比例(0 到 1)時執行內建摘要 | | headChars | 300 | Prune 模式:裁切輸出時保留的字元數;0 到 100,000 | | maxInputChars | 400000 | Prune 模式:最多送給 Gemini 的字元數;10,000 到 4,000,000 | | summaryMaxInputChars | 2000000 | Summary 模式:最多送給 Gemini 的字元數;10,000 到 4,000,000 |

超出範圍的值會回復為預設值。

可以接觸的內容

在 Claude Code 2.1.283 上使用 claude plugin validate 驗證:

❯ ./register.ts hooks: session.start, command.run{command=gemini-compact}, session.compact, turn.complete
❯ ./register.ts calls: $.clock.now (via askGemini), $.clock.sleep (via askGemini), $.command.register, $.gemini.enroll, $.gemini.read (via askGemini), $.gemini.request (via askGemini), $.gemini.settings (via compactWithGemini, runCommand, storePatch), $.http.fetch (via askGemini), $.session.compact (via maybeCompact), $.session.usage (via maybeCompact), $.store.delete (via runCommand), $.store.get (via loadConfig), $.store.set (via storePatch), $.ui.log, $.ui.toast (via report)

Reach L3,會接觸網路。

1. 讀取:每次壓縮時的對話(訊息、工具輸入與輸出);每個主要迴圈回合後的上下文填充量;自己的 $.store;以及來自 gemini-core、帶有金鑰的請求
2. 執行:不啟動處理程序;在回合結束且超過門檻後呼叫一次 $.session.compact,直到上下文再次低於門檻前最多一次
3. 傳送:對話(summary:最新訊息以外的全部內容;prune:全部內容),每次壓縮一個請求(503 後最多 4 次,429 或金鑰錯誤後每多一把金鑰再一次),傳送到 gemini-core 建立的 URL(generativelanguage.googleapis.com),金鑰放在 x-goog-api-key header,不放進 URL
4. 持續儲存:在 $.store 中保存 3 個指令設定(enabled、mode、atPercent);上次結果會保存在 i

安裝

請先查看作者 README,確認 marketplace 與外掛名稱;指令可能隨儲存庫結構而變動。

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install gemini-compact
原文 / README

gemini-compact

When the context fills up, Claude compacts it with one more Claude request: it reads the whole context and writes a summary, and that request counts against your Claude usage. This mod hands that job to Gemini. It has two modes:

  • summary (the default): Gemini summarizes the conversation before the newest 6 messages, and those messages stay verbatim after the summary. Claude writes no summary; the engine's built-in summary runs only when Gemini fails.
  • prune: Gemini answers, for every older tool call, whether the call and its output stay, stay with a cut output, or go. Every user and assistant message stays verbatim. Nothing is summarized.

The prune idea follows fast-jev-compaction by Tamara Tran (tamaratran/fast-jev-compaction), which asks TypeSafe Jev. The code here is new and asks Gemini.

Which mode

Both modes replace the Claude request of the built-in compaction with a Gemini request.

After the compaction, every Claude request reads what is left. In summary mode that is the summary and the newest messages, close to what the built-in summary leaves. In prune mode it is every message of the conversation less the dropped tool output, which is larger, so each following request reads more. The sizes were not measured on a long session.

Pick summary mode to save the most Claude usage. Pick prune mode when the exact wording of every message matters more than the size.

Summary mode

  1. At /compact, at the engine's own compaction, and after a main-loop turn that ends with an answer and the context over the threshold, the session.compact hook takes the conversation.
  2. The newest 6 messages stay. The cut moves back to an assistant message, so a tool result is never kept without its call and the kept part opens with an assistant message after the summary. With no assistant message to cut at, everything is summarized.
  3. One generateContent request sends everything before the cut to Gemini: every message, and each call with its input and its full output. When the text is over summaryMaxInputChars (2,000,000 characters by default), the longest outputs are cut to one common length, each keeping its head and tail; a conversation over the limit even without any output fails.
  4. Gemini writes a plain-text summary in nine sections: the request and intent, technical concepts, files and code, errors and fixes, problem solving, every user message verbatim, pending tasks, the current work, and the next step. The text you write after /compact goes to Gemini with it.
  5. The conversation becomes one user message (a note, then the summary) followed by the kept messages, which go back as the engine's own messages.
  6. The built-in summary runs, and one line says why, when there is no key, nothing lies before the newest messages, Gemini fails, the summary is under 200 characters or cut at the output limit (32,768 tokens), or the result is not smaller than the conversation.

In both modes gemini-core builds the request with the key, the model and the thinking level it holds for gemini-compact, and reads the answer. After an HTTP 503 ("high demand") the mod asks again after 1 s, 2 s and 3 s, at most four attempts in all, and no attempt starts whose wait would end past 60 s. After a 429 or a key error, gemini-core hands over the request with its next key, when it holds one.

In a live check on 2.1.277, /compact took 2.6 seconds with gemini-3.5-flash-lite, Claude sent no compaction request, and after it the model named a word and a file that appeared only in the summarized part.

Prune mode

  1. The same three triggers reach the session.compact hook.
  2. A tool call in the first message or in the newest 6 messages, or whose result lies in one of them, is kept whole. Every other call gets an id (c1, c2, ...). With no such call, the built-in summary runs.
  3. One generateContent request sends the conversation to Gemini: every message, and each call with its input and its full output. Over maxInputChars (400,000 characters by default) the longest outputs are cut the same way as in summary mode. A response schema allows exactly one answer per id: keep, truncate or drop.
  4. The mod checks the answer (every id exactly once, no other id) and rebuilds the conversation:
    • keep: the call and its output stay.
    • truncate: the call stays, and the output keeps its first 300 characters and one line that says it was cut. An output at most 120 characters longer than that stays whole.
    • drop: the call and its output go. A note on the nearest assistant message names the removed calls, for example [gemini-compact removed 1 earlier tool call(s) and their output after a compaction; they ran: Bash(ls -la /usr/bin)]. Without the note, the model read a reply whose work was gone and said it had never done that work (measured on 2.1.277).
    • A message the answer does not touch goes back as the engine's own message.
  5. When the result is less than 25% smaller, or anything fails (no key, an HTTP error, an answer that breaks the schema), the engine's built-in summary runs and one line says why.

In a live check on 2.1.277, /compact took 1.1 seconds with gemini-3.5-flash-lite. Gemini dropped an ls listing and kept the cat output of a file the user was about to edit, and the conversation became 93% smaller.

What it shows

One line in the transcript, not sent to the model, and a toast that stays for 15 seconds:

gemini-compact: summary: 58 → 7 messages · 91% smaller · 312k in, 5k out
gemini-compact: kept 9/11 messages · 93% smaller · 1 dropped, 0 truncated · 5k in, 59 out
gemini-compact: built-in summary: under 25% smaller (kept 14/16 messages · 3% smaller · ...)
gemini-compact: built-in summary: Gemini HTTP 429: Resource has been exhausted

On the free tier the toast adds · sent to Gemini free tier. An automatic compaction the engine skips or that fails writes one line too (automatic compaction skipped: ..., automatic compaction failed: ...).

Command

/gemini-compact              on or off, mode, the model and thinking level gemini-core holds, threshold, tier, whether a key is set, the last result
/gemini-compact on | off     off leaves every compaction to the built-in summary; on is refused while gemini-core has no key
/gemini-compact mode summary | mode prune
/gemini-compact at <1-99>    compact after a turn that ends with the context over this percentage
/gemini-compact at off       no automatic compaction; /compact and the engine's own compaction still ask Gemini
/gemini-compact reset        back to the plugin options, and off

The mod is off after an install: every compaction is the built-in one, none starts on the mod's threshold, and nothing goes to Gemini until /gemini-compact on. The command settings are kept across sessions and take effect at once. After a compaction it started, the mod starts no other one until a turn ends with the context under the threshold, so a context that stays over it does not compact after every turn.

The key, the tier, the model (default gemini-3.5-flash-lite) and the thinking level belong to gemini-core:

/gemini-core model compact gemini-3.5-flash
/gemini-core thinking compact low
/gemini-core paid

Free tier or paid tier

The conversation holds your prompts, the commands the model ran and the contents of the files it read. On the free tier Google may use them and human reviewers may read them; the gemini-core README quotes the Gemini API Additional Terms. On a project you would not show to Google, use a key with billing enabled and set /gemini-core paid. No mod can tell which tier a key is on; the tier setting only chooses the warning.

Google AI Studio shows the free tier limits per model; the documentation does not. They were not measured. A summary of a long conversation is one large request, so a per-minute token limit can refuse it with HTTP 429; the built-in summary then runs.

Install

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install gemini-compact@kilimcininkoroglu-mods

It depends on gemini-core, which claude plugin install adds. Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.

To load it from a local checkout for one session, put gemini-core beside it:

claude --plugin-dir plugins/gemini-core --plugin-dir plugins/gemini-compact

After installing

  1. Set the Gemini key and the tier in gemini-core, as its After installing section says, then restart Claude Code.
  2. Run /gemini-compact on. Without a key it answers still off: gemini-core has no Gemini key and stays off.
  3. Run /gemini-compact. The first line reads on · summary · <model> · thinking ... · automatic at 60% · <tier> tier · key set.
  4. Run /compact once. The transcript line should start with gemini-compact: summary:. A line that starts with built-in summary: names why Gemini was not used.

After an update from 0.2.x: claude plugin update does not add gemini-core (measured on 2.1.278), so run claude plugin install gemini-core@kilimcininkoroglu-mods once. Version 0.3.0 moved the key, tier and model to gemini-core; the apiKey, tier and model options and the settings /gemini-compact free|paid|model stored before are no longer read, so set them again in gemini-core. The mode and at settings stay. Version 0.4.0 made the mod off by default: after an update from an earlier version it is off unless you ran /gemini-compact on before, so run /gemini-compact on once.

Options

| Option | Default | What it sets | |---|---|---| | mode | summary | summary or prune; /gemini-compact mode overrides it | | compactAtPercent | 60 | The automatic threshold, 0 to 99; 0 turns it off; /gemini-compact at overrides it | | keepRecent | 6 | Newest messages kept verbatim (summary) or whose calls are never sent for a decision (prune); 0 to 1000 | | minReduction | 0.25 | Prune mode: below this fraction (0 to 1) the built-in summary runs | | headChars | 300 | Prune mode: characters kept of a truncated output; 0 to 100,000 | | maxInputChars | 400000 | Prune mode: characters sent to Gemini at most; 10,000 to 4,000,000 | | summaryMaxInputChars | 2000000 | Summary mode: characters sent to Gemini at most; 10,000 to 4,000,000 |

A value outside its range falls back to the default.

What it can reach

Validated with claude plugin validate on Claude Code 2.1.283:

❯ ./register.ts hooks: session.start, command.run{command=gemini-compact}, session.compact, turn.complete
❯ ./register.ts calls: $.clock.now (via askGemini), $.clock.sleep (via askGemini), $.command.register, $.gemini.enroll, $.gemini.read (via askGemini), $.gemini.request (via askGemini), $.gemini.settings (via compactWithGemini, runCommand, storePatch), $.http.fetch (via askGemini), $.session.compact (via maybeCompact), $.session.usage (via maybeCompact), $.store.delete (via runCommand), $.store.get (via loadConfig), $.store.set (via storePatch), $.ui.log, $.ui.toast (via report)

Reach L3, reaches the network.

1. Reads:    the conversation at each compaction (messages, tool inputs and outputs); the context fill after each main-loop turn; its own $.store; from gemini-core, the request with the key
2. Runs:     no process; one $.session.compact after a turn that ends over the threshold, at most once until the context was under it again
3. Sends:    the conversation (summary: all but the newest messages; prune: all of it), one request per compaction (up to four after a 503, and once more per extra key after a 429 or a key error), to the URL gemini-core builds (generativelanguage.googleapis.com) with the key in the x-goog-api-key header, never in the URL
4. Persists: in $.store, the three command settings (enabled, mode, atPercent); the last result lives in memory
5. Hostile input: the Gemini answer is untrusted: a prune answer is applied only in the schema shape with every candidate id once; a summary becomes the text of one user message the model reads, so a hostile summary can steer the model, as text in a file it reads can; anything malformed falls back to the built-in summary

Limits

  • A summary and a drop are a model's judgment. A summary loses detail the newest messages do not repeat. The prune note tells the model which calls ran, so it can run a tool again.
  • After the compaction the context is written to the cache again. In prune mode it stays larger than a built-in summary, so the next message pays a larger cache write.
  • A summary of a long conversation takes Gemini longer, and the compaction waits for it. Only short conversations were timed.
  • A subagent's own compaction is left to the engine.
  • A compaction the engine precomputes (precompute) also asks Gemini. Whether the engine reuses that result for the compaction that follows was not measured.
  • The test engine of claude plugin test passes no trigger to a $.session.compact() call. The plugin trigger and the hook that answers it were measured in a live session.

Development

make install     # eslint, typescript-eslint, typescript
make lint        # complexity limit 10, the build fails above it
make typecheck   # needs .claude/types/ from /plugin-types
make validate
make test        # claude plugin test

更多類似作品