ClaudeMods
☰
JA
● 0 人がオンライン ・閲覧 0 回
スポンサー作品を投稿
GitHub リポジトリ · 投稿者 KilimcininKorOglu

gemini-compact

コンパクションを Claude から Gemini に移します。summary モードでは Gemini が会話を要約し、最新のメッセージは逐語的に残ります。prune モードではすべてのメッセージを残し、Gemini が古いツール呼び出しを保持、カット、削除します。/gemini-compact on までオフで、キー、ティア、モデル、思考レベルは gemini-core が保持します。

KilimcininKorOglu@KilimcininKorOglu

KilimcininKorOglu/claude-code-mods/tree/main/plugins/gemini-compact

翻訳済み

この mod について

gemini-compact

コンテキストがいっぱいになると、Claude はもう1回 Claude リクエストを行ってコンパクションします。コンテキスト全体を読み、要約を書き、そのリクエストも Claude の使用量に計上されます。この Mod はその仕事を Gemini に渡します。モードは2つあります。

  • summary(デフォルト):Gemini は最新 6 メッセージより前の会話を要約し、それらのメッセージは要約の後に逐語的に残ります。Claude は要約を書きません。Gemini が失敗したときだけ、エンジン内蔵の要約が実行されます。
  • prune:Gemini は古い各ツール呼び出しについて、呼び出しと出力を残すか、出力をカットして残すか、削除するかを答えます。ユーザーとアシスタントのすべてのメッセージは逐語的に残ります。要約は行いません。

prune の考え方は Tamara Tran の fast-jev-compaction(tamaratran/fast-jev-compaction)に従っています。これは TypeSafe Jev に問い合わせます。ここでのコードは新しく、Gemini に問い合わせます。

モードの選択

どちらのモードも、内蔵コンパクションの Claude リクエストを Gemini リクエストに置き換えます。

コンパクション後は、すべての Claude リクエストが残った内容を読みます。summary モードでは要約と最新メッセージで、内蔵要約が残すものに近い内容です。prune モードでは、削除されたツール出力を除く会話の全メッセージなので、より大きくなり、その後の各リクエストが多く読みます。長いセッションでのサイズは測定していません。

Claude の使用量を最も節約したいなら summary モードを選びます。各メッセージの正確な文言がサイズより重要なら prune モードを選びます。

Summary モード

  1. /compact、エンジン自身のコンパクション、そしてメインループのターンが回答で終わりコンテキストがしきい値を超えた後に、session.compact hook が会話を受け取ります。
  2. 最新 6 メッセージが残ります。カット位置はアシスタントメッセージまで戻るため、ツール結果だけが呼び出しなしで残ることはなく、残った部分は要約の後のアシスタントメッセージから始まります。カットできるアシスタントメッセージがなければ、すべてを要約します。
  3. 1つの generateContent リクエストがカット位置より前のすべてを Gemini に送ります。すべてのメッセージと、各呼び出しの入力および完全な出力です。テキストが summaryMaxInputChars(デフォルト 2,000,000 文字)を超えると、最長の出力を共通の長さまでカットし、それぞれの先頭と末尾を残します。出力がなくても会話が制限を超えていれば失敗します。
  4. Gemini はプレーンテキストの要約を9セクションで書きます。リクエストと意図、技術概念、ファイルとコード、エラーと修正、問題解決、すべてのユーザーメッセージの逐語的な内容、保留中のタスク、現在の作業、次のステップです。/compact の後に書いたテキストも一緒に Gemini へ送られます。
  5. 会話は1つのユーザーメッセージ(注記、続いて要約)と、その後の保持メッセージになります。保持メッセージはエンジン自身のメッセージとして戻されます。
  6. キーがない、最新メッセージより前に何もない、Gemini が失敗した、要約が 200 文字未満、出力上限(32,768 トークン)でカットされた、または結果が会話より短くない場合、内蔵要約が実行され、理由が1行で示されます。

どちらのモードでも、gemini-core は gemini-compact 用に保持しているキー、モデル、思考レベルでリクエストを作り、回答を読みます。HTTP 503(「high demand」)の後、Mod は 1 s、2 s、3 s 後に再試行し、最大 4 回です。待機が 60 s を過ぎて終わる試行は開始しません。429 またはキーエラーの後は、gemini-core に次のキーがあれば、そのキーでリクエストを引き継ぎます。

2.1.277 のライブチェックでは、gemini-3.5-flash-lite で /compact に 2.6 秒かかり、Claude はコンパクションリクエストを送りませんでした。その後モデルは、要約された部分にだけ現れた単語とファイルを挙げました。

Prune モード

  1. 同じ3つのトリガーが session.compact hook に到達します。
  2. 最初のメッセージまたは最新 6 メッセージにあるツール呼び出し、またはその結果がそれらのどれかにある呼び出しは完全に残ります。それ以外の呼び出しには ID(c1、c2、…)が付与されます。そのような呼び出しがなければ、内蔵要約が実行されます。
  3. 1つの generateContent リクエストが会話を Gemini に送り、すべてのメッセージと各呼び出しの入力および完全な出力を含めます。maxInputChars(デフォルト 400,000 文字)を超えると、Summary モードと同じ方法で最長の出力をカットします。レスポンススキーマは各 ID に対して keep、truncate、drop のいずれか1つだけを許可します。
  4. Mod は回答を検査し(すべての ID が1回ずつで、他の ID はなし)、会話を再構築します。
    • keep:呼び出しと出力を残します。
    • truncate:呼び出しを残し、出力の最初の 300 文字と、カットされたことを示す1行を残します。その長さより 120 文字以内しか長くない出力は全体を残します。
    • drop:呼び出しと出力を削除します。最も近いアシスタントメッセージに、削除した呼び出しを示す注記を付けます。例:[gemini-compact removed 1 earlier tool call(s) and their output after a compaction; they ran: Bash(ls -la /usr/bin)]。この注記がないと、モデルは作業が消えた返信を読み、その作業をしたことがないと言います(2.1.277 で測定)。
    • 回答が触れなかったメッセージはエンジン自身のメッセージとして戻ります。
  5. 結果が 25% 未満しか小さくならない場合、またはキーなし、HTTP エラー、スキーマを破る回答など何かが失敗した場合、内蔵要約が実行され、理由が1行で示されます。

2.1.277 のライブチェックでは、gemini-3.5-flash-lite で /compact に 1.1 秒かかりました。Gemini は ls の一覧を削除し、ユーザーが編集しようとしていたファイルの cat 出力を残し、会話は 93% 小さくなりました。

表示されるもの

モデルには送られないトランスクリプトの1行と、15秒間残るトーストです。

gemini-compact: summary: 58 → 7 messages · 91% smaller · 312k in, 5k out
gemini-compact: kept 9/11 messages · 93% smaller · 1 dropped, 0 truncated · 5k in, 59 out
gemini-compact: built-in summary: under 25% smaller (kept 14/16 messages · 3% smaller · ...)
gemini-compact: built-in summary: Gemini HTTP 429: Resource has been exhausted

無料ティアではトーストに · sent to Gemini free tier が追加されます。エンジンがスキップまたは失敗した自動コンパクションも1行を書きます(automatic compaction skipped: ...、automatic compaction failed: ...)。

コマンド

/gemini-compact              on or off, mode, the model and thinking level gemini-core holds, threshold, tier, whether a key is set, the last result
/gemini-compact on | off     off leaves every compaction to the built-in summary; on is refused while gemini-core has no key
/gemini-compact mode summary | mode prune
/gemini-compact at <1-99>    compact after a turn that ends with the context over this percentage
/gemini-compact at off       no automatic compaction; /compact and the engine's own compaction still ask Gemini
/gemini-compact reset        back to the plugin options, and off

インストール後、Mod はオフです。すべてのコンパクションは内蔵のものになり、Mod のしきい値では開始せず、/gemini-compact on まで Gemini に何も送られません。コマンド設定はセッションをまたいで保持され、すぐに反映されます。コンパクションを開始した後は、コンテキストがしきい値未満になるターンが終わるまで別のものを開始しません。そのため、コンテキストが超過したままでも各ターン後にコンパクトすることはありません。

キー、ティア、モデル(デフォルト gemini-3.5-flash-lite)、思考レベルは gemini-core に属します。

/gemini-core model compact gemini-3.5-flash
/gemini-core thinking compact low
/gemini-core paid

無料ティアまたは有料ティア

会話にはプロンプト、モデルが実行したコマンド、読んだファイルの内容が含まれます。無料ティアでは Google がそれらを使用する可能性があり、人間のレビュアーが読む可能性もあります。gemini-core README は Gemini API Additional Terms を引用しています。Google に見せないプロジェクトでは、請求を有効にしたキーを使い、/gemini-core paid を設定してください。キーがどのティアにあるかを Mod が判断することはできず、ティア設定は警告だけを選びます。

Google AI Studio はモデルごとの無料ティア制限を表示しますが、ドキュメントにはありません。測定はしていません。長い会話の要約は1つの大きなリクエストなので、毎分のトークン制限が HTTP 429 で拒否することがあり、その場合は内蔵要約が実行されます。

インストール

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install gemini-compact@kilimcininkoroglu-mods

gemini-core に依存し、claude plugin install が追加します。関数フックはアーリーアクセスです。Claude Code 2.1.288 以降はデフォルトで読み込むので、切り替えるものはありません。

1回のセッションでローカル checkout から読み込むには、gemini-core を隣に置きます。

claude --plugin-dir plugins/gemini-core --plugin-dir plugins/gemini-compact

インストール後

  1. gemini-core の After installing セクションに従って Gemini キーとティアを設定し、Claude Code を再起動します。
  2. /gemini-compact on を実行します。キーがなければ still off: gemini-core has no Gemini key と答えてオフのままです。
  3. /gemini-compact を実行します。最初の行は on · summary · <model> · thinking ... · automatic at 60% · <tier> tier · key set です。
  4. /compact を1回実行します。トランスクリプトの行は gemini-compact: summary: で始まるはずです。built-in summary: で始まる行には Gemini を使わなかった理由が示されます。

0.2.x から更新する場合:claude plugin update は gemini-core を追加しません(2.1.278 で測定)。そのため claude plugin install gemini-core@kilimcininkoroglu-mods を1回実行してください。0.3.0 ではキー、ティア、モデルが gemini-core に移りました。以前保存した apiKey、tier、model オプションと /gemini-compact free|paid|model 設定はもう読まれないので、gemini-core で再設定します。mode と at の設定は残ります。0.4.0 では Mod がデフォルトでオフになりました。以前のバージョンから更新した場合、前に /gemini-compact on を実行していない限りオフなので、1回実行してください。

オプション

| Option | Default | What it sets | |---|---|---| | mode | summary | summary または prune。/gemini-compact mode が上書きします | | compactAtPercent | 60 | 自動しきい値、0〜99。0 でオフ、/gemini-compact at が上書きします | | keepRecent | 6 | 逐語的に保持する最新メッセージ(summary)、または判断のために送らない呼び出し(prune)。0〜1000 | | minReduction | 0.25 | Prune モード:この割合(0〜1)未満なら内蔵要約を実行 | | headChars | 300 | Prune モード:切り詰めた出力から保持する文字数。0〜100,000 | | maxInputChars | 400000 | Prune モード:Gemini に送る文字数の上限。10,000〜4,000,000 | | summaryMaxInputChars | 2000000 | Summary モード:Gemini に送る文字数の上限。10,000〜4,000,000 |

範囲外の値はデフォルトに戻ります。

到達範囲

Claude Code 2.1.283 で claude plugin validate を実行して検証済みです。

❯ ./register.ts hooks: session.start, command.run{command=gemini-compact}, session.compact, turn.complete
❯ ./register.ts calls: $.clock.now (via askGemini), $.clock.sleep (via askGemini), $.command.register, $.gemini.enroll, $.gemini.read (via askGemini), $.gemini.request (via askGemini), $.gemini.settings (via compactWithGemini, runCommand, storePatch), $.http.fetch (via askGemini), $.session.compact (via maybeCompact), $.session.usage (via maybeCompact), $.store.delete (via runCommand), $.store.get (via loadConfig), $.store.set (via storePatch), $.ui.log, $.ui.toast (via report)

Reach L3、ネットワークに到達します。

1. 読み取り:各コンパクションの会話(メッセージ、ツールの入力と出力)、各メインループターン後のコンテキストの埋まり具合、自分の $.store、gemini-core からキー付きのリクエスト
2. 実行:プロセスなし。しきい値を超えて終わったターンの後に $.session.compact を1回、コンテキストが再びしきい値未満になるまで最大1回
3. 送信:会話(summary:最新メッセージ以外、prune:すべて)をコンパクションごとに1リクエスト(503 後は最大4回、429 またはキーエラーの後は追加のキーごとにもう1回)。gemini-core が作る URL(generativelanguage.googleapis.com)へ、キーを x-goog-api-key ヘッダーに入れて送り、URL には決して入れません
4. 永続化:$.store に3つのコマンド設定(enabled、mode、atPercent)を保存し、最後の結果は i

インストール

まず作者の README で marketplace とプラグイン名を確認してください。コマンドはリポジトリの構成によって変わる場合があります。

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install gemini-compact
原文 / README

gemini-compact

When the context fills up, Claude compacts it with one more Claude request: it reads the whole context and writes a summary, and that request counts against your Claude usage. This mod hands that job to Gemini. It has two modes:

  • summary (the default): Gemini summarizes the conversation before the newest 6 messages, and those messages stay verbatim after the summary. Claude writes no summary; the engine's built-in summary runs only when Gemini fails.
  • prune: Gemini answers, for every older tool call, whether the call and its output stay, stay with a cut output, or go. Every user and assistant message stays verbatim. Nothing is summarized.

The prune idea follows fast-jev-compaction by Tamara Tran (tamaratran/fast-jev-compaction), which asks TypeSafe Jev. The code here is new and asks Gemini.

Which mode

Both modes replace the Claude request of the built-in compaction with a Gemini request.

After the compaction, every Claude request reads what is left. In summary mode that is the summary and the newest messages, close to what the built-in summary leaves. In prune mode it is every message of the conversation less the dropped tool output, which is larger, so each following request reads more. The sizes were not measured on a long session.

Pick summary mode to save the most Claude usage. Pick prune mode when the exact wording of every message matters more than the size.

Summary mode

  1. At /compact, at the engine's own compaction, and after a main-loop turn that ends with an answer and the context over the threshold, the session.compact hook takes the conversation.
  2. The newest 6 messages stay. The cut moves back to an assistant message, so a tool result is never kept without its call and the kept part opens with an assistant message after the summary. With no assistant message to cut at, everything is summarized.
  3. One generateContent request sends everything before the cut to Gemini: every message, and each call with its input and its full output. When the text is over summaryMaxInputChars (2,000,000 characters by default), the longest outputs are cut to one common length, each keeping its head and tail; a conversation over the limit even without any output fails.
  4. Gemini writes a plain-text summary in nine sections: the request and intent, technical concepts, files and code, errors and fixes, problem solving, every user message verbatim, pending tasks, the current work, and the next step. The text you write after /compact goes to Gemini with it.
  5. The conversation becomes one user message (a note, then the summary) followed by the kept messages, which go back as the engine's own messages.
  6. The built-in summary runs, and one line says why, when there is no key, nothing lies before the newest messages, Gemini fails, the summary is under 200 characters or cut at the output limit (32,768 tokens), or the result is not smaller than the conversation.

In both modes gemini-core builds the request with the key, the model and the thinking level it holds for gemini-compact, and reads the answer. After an HTTP 503 ("high demand") the mod asks again after 1 s, 2 s and 3 s, at most four attempts in all, and no attempt starts whose wait would end past 60 s. After a 429 or a key error, gemini-core hands over the request with its next key, when it holds one.

In a live check on 2.1.277, /compact took 2.6 seconds with gemini-3.5-flash-lite, Claude sent no compaction request, and after it the model named a word and a file that appeared only in the summarized part.

Prune mode

  1. The same three triggers reach the session.compact hook.
  2. A tool call in the first message or in the newest 6 messages, or whose result lies in one of them, is kept whole. Every other call gets an id (c1, c2, ...). With no such call, the built-in summary runs.
  3. One generateContent request sends the conversation to Gemini: every message, and each call with its input and its full output. Over maxInputChars (400,000 characters by default) the longest outputs are cut the same way as in summary mode. A response schema allows exactly one answer per id: keep, truncate or drop.
  4. The mod checks the answer (every id exactly once, no other id) and rebuilds the conversation:
    • keep: the call and its output stay.
    • truncate: the call stays, and the output keeps its first 300 characters and one line that says it was cut. An output at most 120 characters longer than that stays whole.
    • drop: the call and its output go. A note on the nearest assistant message names the removed calls, for example [gemini-compact removed 1 earlier tool call(s) and their output after a compaction; they ran: Bash(ls -la /usr/bin)]. Without the note, the model read a reply whose work was gone and said it had never done that work (measured on 2.1.277).
    • A message the answer does not touch goes back as the engine's own message.
  5. When the result is less than 25% smaller, or anything fails (no key, an HTTP error, an answer that breaks the schema), the engine's built-in summary runs and one line says why.

In a live check on 2.1.277, /compact took 1.1 seconds with gemini-3.5-flash-lite. Gemini dropped an ls listing and kept the cat output of a file the user was about to edit, and the conversation became 93% smaller.

What it shows

One line in the transcript, not sent to the model, and a toast that stays for 15 seconds:

gemini-compact: summary: 58 → 7 messages · 91% smaller · 312k in, 5k out
gemini-compact: kept 9/11 messages · 93% smaller · 1 dropped, 0 truncated · 5k in, 59 out
gemini-compact: built-in summary: under 25% smaller (kept 14/16 messages · 3% smaller · ...)
gemini-compact: built-in summary: Gemini HTTP 429: Resource has been exhausted

On the free tier the toast adds · sent to Gemini free tier. An automatic compaction the engine skips or that fails writes one line too (automatic compaction skipped: ..., automatic compaction failed: ...).

Command

/gemini-compact              on or off, mode, the model and thinking level gemini-core holds, threshold, tier, whether a key is set, the last result
/gemini-compact on | off     off leaves every compaction to the built-in summary; on is refused while gemini-core has no key
/gemini-compact mode summary | mode prune
/gemini-compact at <1-99>    compact after a turn that ends with the context over this percentage
/gemini-compact at off       no automatic compaction; /compact and the engine's own compaction still ask Gemini
/gemini-compact reset        back to the plugin options, and off

The mod is off after an install: every compaction is the built-in one, none starts on the mod's threshold, and nothing goes to Gemini until /gemini-compact on. The command settings are kept across sessions and take effect at once. After a compaction it started, the mod starts no other one until a turn ends with the context under the threshold, so a context that stays over it does not compact after every turn.

The key, the tier, the model (default gemini-3.5-flash-lite) and the thinking level belong to gemini-core:

/gemini-core model compact gemini-3.5-flash
/gemini-core thinking compact low
/gemini-core paid

Free tier or paid tier

The conversation holds your prompts, the commands the model ran and the contents of the files it read. On the free tier Google may use them and human reviewers may read them; the gemini-core README quotes the Gemini API Additional Terms. On a project you would not show to Google, use a key with billing enabled and set /gemini-core paid. No mod can tell which tier a key is on; the tier setting only chooses the warning.

Google AI Studio shows the free tier limits per model; the documentation does not. They were not measured. A summary of a long conversation is one large request, so a per-minute token limit can refuse it with HTTP 429; the built-in summary then runs.

Install

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install gemini-compact@kilimcininkoroglu-mods

It depends on gemini-core, which claude plugin install adds. Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.

To load it from a local checkout for one session, put gemini-core beside it:

claude --plugin-dir plugins/gemini-core --plugin-dir plugins/gemini-compact

After installing

  1. Set the Gemini key and the tier in gemini-core, as its After installing section says, then restart Claude Code.
  2. Run /gemini-compact on. Without a key it answers still off: gemini-core has no Gemini key and stays off.
  3. Run /gemini-compact. The first line reads on · summary · <model> · thinking ... · automatic at 60% · <tier> tier · key set.
  4. Run /compact once. The transcript line should start with gemini-compact: summary:. A line that starts with built-in summary: names why Gemini was not used.

After an update from 0.2.x: claude plugin update does not add gemini-core (measured on 2.1.278), so run claude plugin install gemini-core@kilimcininkoroglu-mods once. Version 0.3.0 moved the key, tier and model to gemini-core; the apiKey, tier and model options and the settings /gemini-compact free|paid|model stored before are no longer read, so set them again in gemini-core. The mode and at settings stay. Version 0.4.0 made the mod off by default: after an update from an earlier version it is off unless you ran /gemini-compact on before, so run /gemini-compact on once.

Options

| Option | Default | What it sets | |---|---|---| | mode | summary | summary or prune; /gemini-compact mode overrides it | | compactAtPercent | 60 | The automatic threshold, 0 to 99; 0 turns it off; /gemini-compact at overrides it | | keepRecent | 6 | Newest messages kept verbatim (summary) or whose calls are never sent for a decision (prune); 0 to 1000 | | minReduction | 0.25 | Prune mode: below this fraction (0 to 1) the built-in summary runs | | headChars | 300 | Prune mode: characters kept of a truncated output; 0 to 100,000 | | maxInputChars | 400000 | Prune mode: characters sent to Gemini at most; 10,000 to 4,000,000 | | summaryMaxInputChars | 2000000 | Summary mode: characters sent to Gemini at most; 10,000 to 4,000,000 |

A value outside its range falls back to the default.

What it can reach

Validated with claude plugin validate on Claude Code 2.1.283:

❯ ./register.ts hooks: session.start, command.run{command=gemini-compact}, session.compact, turn.complete
❯ ./register.ts calls: $.clock.now (via askGemini), $.clock.sleep (via askGemini), $.command.register, $.gemini.enroll, $.gemini.read (via askGemini), $.gemini.request (via askGemini), $.gemini.settings (via compactWithGemini, runCommand, storePatch), $.http.fetch (via askGemini), $.session.compact (via maybeCompact), $.session.usage (via maybeCompact), $.store.delete (via runCommand), $.store.get (via loadConfig), $.store.set (via storePatch), $.ui.log, $.ui.toast (via report)

Reach L3, reaches the network.

1. Reads:    the conversation at each compaction (messages, tool inputs and outputs); the context fill after each main-loop turn; its own $.store; from gemini-core, the request with the key
2. Runs:     no process; one $.session.compact after a turn that ends over the threshold, at most once until the context was under it again
3. Sends:    the conversation (summary: all but the newest messages; prune: all of it), one request per compaction (up to four after a 503, and once more per extra key after a 429 or a key error), to the URL gemini-core builds (generativelanguage.googleapis.com) with the key in the x-goog-api-key header, never in the URL
4. Persists: in $.store, the three command settings (enabled, mode, atPercent); the last result lives in memory
5. Hostile input: the Gemini answer is untrusted: a prune answer is applied only in the schema shape with every candidate id once; a summary becomes the text of one user message the model reads, so a hostile summary can steer the model, as text in a file it reads can; anything malformed falls back to the built-in summary

Limits

  • A summary and a drop are a model's judgment. A summary loses detail the newest messages do not repeat. The prune note tells the model which calls ran, so it can run a tool again.
  • After the compaction the context is written to the cache again. In prune mode it stays larger than a built-in summary, so the next message pays a larger cache write.
  • A summary of a long conversation takes Gemini longer, and the compaction waits for it. Only short conversations were timed.
  • A subagent's own compaction is left to the engine.
  • A compaction the engine precomputes (precompute) also asks Gemini. Whether the engine reuses that result for the compaction that follows was not measured.
  • The test engine of claude plugin test passes no trigger to a $.session.compact() call. The plugin trigger and the hook that answers it were measured in a live session.

Development

make install     # eslint, typescript-eslint, typescript
make lint        # complexity limit 10, the build fails above it
make typecheck   # needs .claude/types/ from /plugin-types
make validate
make test        # claude plugin test

関連作品