ClaudeMods
☰
KO
● 0 명 접속 중 · 조회 0 회
후원프로젝트 제출
GitHub 저장소 · 작성자 KilimcininKorOglu

gemini-compact

압축을 Claude에서 Gemini로 옮깁니다. summary 모드에서는 Gemini가 대화를 요약하고 최신 메시지는 그대로 남기며, prune 모드에서는 모든 메시지를 남기고 Gemini가 오래된 도구 호출을 유지하거나 자르거나 삭제합니다. /gemini-compact on 전까지 꺼져 있으며 키, 티어, 모델과 사고 수준은 gemini-core가 보관합니다.

KilimcininKorOglu@KilimcininKorOglu

KilimcininKorOglu/claude-code-mods/tree/main/plugins/gemini-compact

번역 완료

이 mod 소개

gemini-compact

컨텍스트가 가득 차면 Claude는 Claude 요청을 한 번 더 보내 압축합니다. 전체 컨텍스트를 읽고 요약을 작성하며 이 요청도 Claude 사용량에 포함됩니다. 이 Mod는 그 작업을 Gemini에 넘깁니다. 모드는 2가지입니다.

  • summary(기본값): Gemini가 최신 6개 메시지 이전의 대화를 요약하고, 해당 메시지는 요약 뒤에 그대로 남습니다. Claude는 요약을 작성하지 않으며 Gemini가 실패할 때만 엔진 내장 요약이 실행됩니다.
  • prune: Gemini가 오래된 각 도구 호출에 대해 호출과 출력이 남는지, 잘린 출력과 함께 남는지, 삭제되는지를 답합니다. 모든 사용자와 어시스턴트 메시지는 그대로 남습니다. 요약하지 않습니다.

prune 아이디어는 Tamara Tran의 fast-jev-compaction(tamaratran/fast-jev-compaction)을 따릅니다. 이 프로젝트는 TypeSafe Jev에 질의합니다. 여기의 코드는 새로 작성되었고 Gemini에 질의합니다.

모드 선택

두 모드 모두 내장 압축의 Claude 요청을 Gemini 요청으로 바꿉니다.

압축 후 모든 Claude 요청은 남은 내용을 읽습니다. summary 모드에서는 요약과 최신 메시지이며, 내장 요약이 남기는 내용과 비슷합니다. prune 모드에서는 삭제된 도구 출력만 제외한 대화의 모든 메시지이므로 더 크고, 이후 각 요청이 더 많이 읽습니다. 긴 세션에서 크기를 측정하지는 않았습니다.

Claude 사용량을 가장 많이 아끼려면 summary 모드를 선택하세요. 각 메시지의 정확한 표현이 크기보다 중요하면 prune 모드를 선택하세요.

Summary 모드

  1. /compact, 엔진 자체의 압축, 그리고 메인 루프 턴이 답변으로 끝나면서 컨텍스트가 임계값을 넘은 뒤에 session.compact hook이 대화를 받습니다.
  2. 최신 6개 메시지가 남습니다. 자르는 지점은 어시스턴트 메시지로 돌아가므로 도구 호출 없이 도구 결과만 남지 않으며, 남은 부분은 요약 뒤의 어시스턴트 메시지로 시작합니다. 자를 어시스턴트 메시지가 없으면 모두 요약합니다.
  3. 하나의 generateContent 요청이 자르는 지점 이전의 모든 내용을 Gemini에 보냅니다. 모든 메시지와 각 호출의 입력 및 전체 출력입니다. 텍스트가 summaryMaxInputChars(기본 2,000,000자)를 넘으면 가장 긴 출력을 같은 길이로 자르고 각각 앞과 뒤를 남깁니다. 출력이 없어도 대화가 제한을 넘으면 실패합니다.
  4. Gemini는 일반 텍스트 요약을 9개 섹션으로 작성합니다. 요청과 의도, 기술 개념, 파일과 코드, 오류와 수정, 문제 해결, 모든 사용자 메시지의 원문, 보류 중인 작업, 현재 작업, 다음 단계입니다. /compact 뒤에 작성한 텍스트도 함께 Gemini로 전송됩니다.
  5. 대화는 사용자 메시지 하나(메모, 이어서 요약)와 보존된 메시지가 됩니다. 보존된 메시지는 엔진 자체의 메시지로 돌아갑니다.
  6. 키가 없거나, 최신 메시지 앞에 아무것도 없거나, Gemini가 실패하거나, 요약이 200자 미만이거나 출력 한도(32,768토큰)에서 잘리거나, 결과가 대화보다 작지 않으면 내장 요약이 실행되고 이유를 한 줄로 표시합니다.

두 모드 모두 gemini-core가 gemini-compact용으로 보관한 키, 모델, 사고 수준으로 요청을 만들고 답변을 읽습니다. HTTP 503(“high demand”) 후 Mod는 1 s, 2 s, 3 s 뒤에 다시 요청하며 총 4번까지 시도합니다. 기다림이 60 s 이후에 끝나는 시도는 시작하지 않습니다. 429 또는 키 오류 후 gemini-core에 다음 키가 있으면 그 키로 요청을 넘깁니다.

2.1.277에서 실시간으로 확인했을 때 gemini-3.5-flash-lite를 사용한 /compact는 2.6초가 걸렸고 Claude는 압축 요청을 보내지 않았습니다. 이후 모델은 요약된 부분에만 있던 단어와 파일을 말했습니다.

Prune 모드

  1. 같은 3개 트리거가 session.compact hook에 도달합니다.
  2. 첫 메시지 또는 최신 6개 메시지에 있는 도구 호출, 혹은 그 결과가 그중 하나에 있는 호출은 통째로 남습니다. 나머지 호출은 ID(c1, c2, …)를 받습니다. 그런 호출이 없으면 내장 요약이 실행됩니다.
  3. 하나의 generateContent 요청이 대화를 Gemini에 보냅니다. 모든 메시지와 각 호출의 입력 및 전체 출력입니다. maxInputChars(기본 400,000자)를 넘으면 summary 모드와 같은 방식으로 가장 긴 출력을 자릅니다. 응답 스키마는 각 ID에 정확히 하나의 답인 keep, truncate 또는 drop만 허용합니다.
  4. Mod는 답을 확인하고(각 ID가 정확히 한 번씩이며 다른 ID는 없음) 대화를 다시 만듭니다.
    • keep: 호출과 출력을 유지합니다.
    • truncate: 호출을 유지하고 출력의 첫 300자와 잘렸음을 알리는 한 줄을 남깁니다. 이 길이보다 최대 120자만 더 긴 출력은 그대로 둡니다.
    • drop: 호출과 출력을 삭제합니다. 가장 가까운 어시스턴트 메시지에 제거된 호출을 적은 메모를 붙입니다. 예: [gemini-compact removed 1 earlier tool call(s) and their output after a compaction; they ran: Bash(ls -la /usr/bin)]. 이 메모가 없으면 모델은 작업이 사라진 응답을 읽고 자신은 그 작업을 한 적이 없다고 말합니다(2.1.277에서 측정).
    • 답이 건드리지 않은 메시지는 엔진 자체의 메시지로 돌아갑니다.
  5. 결과가 25% 미만으로만 작아지거나 키 없음, HTTP 오류, 스키마를 깨는 답 등 무엇이든 실패하면 내장 요약이 실행되고 이유를 한 줄로 표시합니다.

2.1.277에서 실시간으로 확인했을 때 gemini-3.5-flash-lite를 사용한 /compact는 1.1초가 걸렸습니다. Gemini는 ls 목록을 삭제하고 사용자가 편집하려던 파일의 cat 출력을 남겼으며 대화가 93% 작아졌습니다.

표시 내용

모델에 전송되지 않는 transcript의 한 줄과 15초 동안 남는 토스트가 표시됩니다.

gemini-compact: summary: 58 → 7 messages · 91% smaller · 312k in, 5k out
gemini-compact: kept 9/11 messages · 93% smaller · 1 dropped, 0 truncated · 5k in, 59 out
gemini-compact: built-in summary: under 25% smaller (kept 14/16 messages · 3% smaller · ...)
gemini-compact: built-in summary: Gemini HTTP 429: Resource has been exhausted

무료 티어에서는 토스트에 · sent to Gemini free tier가 추가됩니다. 엔진이 건너뛰거나 실패한 자동 압축도 한 줄을 기록합니다(automatic compaction skipped: ..., automatic compaction failed: ...).

명령

/gemini-compact              on or off, mode, the model and thinking level gemini-core holds, threshold, tier, whether a key is set, the last result
/gemini-compact on | off     off leaves every compaction to the built-in summary; on is refused while gemini-core has no key
/gemini-compact mode summary | mode prune
/gemini-compact at <1-99>    compact after a turn that ends with the context over this percentage
/gemini-compact at off       no automatic compaction; /compact and the engine's own compaction still ask Gemini
/gemini-compact reset        back to the plugin options, and off

설치 후 Mod는 꺼져 있습니다. 모든 압축은 내장 방식이고 Mod의 임계값에서는 시작하지 않으며 /gemini-compact on 전에는 아무것도 Gemini로 보내지 않습니다. 명령 설정은 세션 간 유지되고 즉시 적용됩니다. 압축이 시작된 뒤에는 컨텍스트가 임계값 아래로 내려간 턴이 끝날 때까지 다른 압축을 시작하지 않으므로 컨텍스트가 계속 초과해도 매 턴 후 압축하지 않습니다.

키, 티어, 모델(기본 gemini-3.5-flash-lite), 사고 수준은 gemini-core에 속합니다.

/gemini-core model compact gemini-3.5-flash
/gemini-core thinking compact low
/gemini-core paid

무료 티어 또는 유료 티어

대화에는 프롬프트, 모델이 실행한 명령, 읽은 파일의 내용이 들어 있습니다. 무료 티어에서는 Google이 이를 사용할 수 있고 사람 검토자가 읽을 수도 있습니다. gemini-core README는 Gemini API Additional Terms를 인용합니다. Google에 보여 주지 않을 프로젝트에서는 결제가 활성화된 키를 사용하고 /gemini-core paid를 설정하세요. 어떤 Mod도 키가 어느 티어에 있는지 알 수 없으며, 티어 설정은 경고만 선택합니다.

Google AI Studio에는 모델별 무료 티어 한도가 표시되지만 문서에는 없습니다. 측정하지 않았습니다. 긴 대화의 요약은 큰 요청 하나이므로 분당 토큰 한도가 HTTP 429로 거부할 수 있고, 그러면 내장 요약이 실행됩니다.

설치

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install gemini-compact@kilimcininkoroglu-mods

gemini-core에 의존하며 claude plugin install이 추가합니다. 함수 hook은 얼리 액세스입니다. Claude Code 2.1.288 이상은 기본으로 로드하므로 켤 것이 없습니다.

한 세션 동안 로컬 checkout에서 로드하려면 gemini-core를 옆에 둡니다.

claude --plugin-dir plugins/gemini-core --plugin-dir plugins/gemini-compact

설치 후

  1. gemini-core의 After installing 섹션에 따라 Gemini 키와 티어를 설정한 다음 Claude Code를 다시 시작합니다.
  2. /gemini-compact on을 실행합니다. 키가 없으면 still off: gemini-core has no Gemini key라고 답하고 꺼진 상태로 유지합니다.
  3. /gemini-compact를 실행합니다. 첫 줄은 on · summary · <model> · thinking ... · automatic at 60% · <tier> tier · key set으로 표시됩니다.
  4. /compact를 한 번 실행합니다. transcript 줄은 gemini-compact: summary:로 시작해야 합니다. built-in summary:로 시작하는 줄에는 Gemini를 사용하지 않은 이유가 표시됩니다.

0.2.x에서 업데이트한 경우: claude plugin update는 gemini-core를 추가하지 않습니다(2.1.278에서 측정). 따라서 claude plugin install gemini-core@kilimcininkoroglu-mods를 한 번 실행하세요. 0.3.0에서 키, 티어, 모델이 gemini-core로 이동했습니다. 이전에 저장한 apiKey, tier, model 옵션과 /gemini-compact free|paid|model 설정은 더 이상 읽히지 않으므로 gemini-core에서 다시 설정합니다. mode와 at 설정은 유지됩니다. 0.4.0에서 Mod가 기본적으로 꺼졌습니다. 이전 버전에서 업데이트하면 전에 /gemini-compact on을 실행하지 않은 한 꺼져 있으므로 한 번 실행하세요.

옵션

| Option | Default | What it sets | |---|---|---| | mode | summary | summary 또는 prune; /gemini-compact mode가 덮어씀 | | compactAtPercent | 60 | 자동 임계값 0~99; 0이면 꺼지고 /gemini-compact at이 덮어씀 | | keepRecent | 6 | 그대로 유지할 최신 메시지(summary) 또는 결정을 위해 절대 보내지 않을 호출(prune); 0~1000 | | minReduction | 0.25 | Prune 모드: 이 비율(0~1) 아래면 내장 요약 실행 | | headChars | 300 | Prune 모드: 잘린 출력에서 보존할 문자 수; 0~100,000 | | maxInputChars | 400000 | Prune 모드: Gemini에 보낼 최대 문자 수; 10,000~4,000,000 | | summaryMaxInputChars | 2000000 | Summary 모드: Gemini에 보낼 최대 문자 수; 10,000~4,000,000 |

범위를 벗어난 값은 기본값으로 돌아갑니다.

접근 범위

Claude Code 2.1.283에서 claude plugin validate로 검증했습니다.

❯ ./register.ts hooks: session.start, command.run{command=gemini-compact}, session.compact, turn.complete
❯ ./register.ts calls: $.clock.now (via askGemini), $.clock.sleep (via askGemini), $.command.register, $.gemini.enroll, $.gemini.read (via askGemini), $.gemini.request (via askGemini), $.gemini.settings (via compactWithGemini, runCommand, storePatch), $.http.fetch (via askGemini), $.session.compact (via maybeCompact), $.session.usage (via maybeCompact), $.store.delete (via runCommand), $.store.get (via loadConfig), $.store.set (via storePatch), $.ui.log, $.ui.toast (via report)

Reach L3, 네트워크에 접근합니다.

1. 읽음: 각 압축의 대화(메시지, 도구 입력과 출력), 각 메인 루프 턴 후 컨텍스트 채움 정도, 자체 $.store, gemini-core에서 키가 포함된 요청
2. 실행: 프로세스 없음. 임계값을 넘겨 끝난 턴 뒤 $.session.compact 한 번, 컨텍스트가 다시 임계값 아래가 될 때까지 최대 한 번
3. 전송: 대화(summary: 최신 메시지 제외 전부, prune: 전부)를 압축마다 한 요청(503 후 최대 4번, 429 또는 키 오류 후 추가 키마다 한 번 더)으로 gemini-core가 만든 URL(generativelanguage.googleapis.com)에 보냅니다. 키는 x-goog-api-key 헤더에 넣고 URL에는 절대 넣지 않습니다.
4. 저장: $.store에 세 명령 설정(enabled, mode, atPercent)을 저장하며 마지막 결과는 i

설치

먼저 작성자의 README에서 marketplace와 플러그인 이름을 확인하세요. 저장소 구조에 따라 명령어가 달라질 수 있습니다.

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install gemini-compact
원문 / README

gemini-compact

When the context fills up, Claude compacts it with one more Claude request: it reads the whole context and writes a summary, and that request counts against your Claude usage. This mod hands that job to Gemini. It has two modes:

  • summary (the default): Gemini summarizes the conversation before the newest 6 messages, and those messages stay verbatim after the summary. Claude writes no summary; the engine's built-in summary runs only when Gemini fails.
  • prune: Gemini answers, for every older tool call, whether the call and its output stay, stay with a cut output, or go. Every user and assistant message stays verbatim. Nothing is summarized.

The prune idea follows fast-jev-compaction by Tamara Tran (tamaratran/fast-jev-compaction), which asks TypeSafe Jev. The code here is new and asks Gemini.

Which mode

Both modes replace the Claude request of the built-in compaction with a Gemini request.

After the compaction, every Claude request reads what is left. In summary mode that is the summary and the newest messages, close to what the built-in summary leaves. In prune mode it is every message of the conversation less the dropped tool output, which is larger, so each following request reads more. The sizes were not measured on a long session.

Pick summary mode to save the most Claude usage. Pick prune mode when the exact wording of every message matters more than the size.

Summary mode

  1. At /compact, at the engine's own compaction, and after a main-loop turn that ends with an answer and the context over the threshold, the session.compact hook takes the conversation.
  2. The newest 6 messages stay. The cut moves back to an assistant message, so a tool result is never kept without its call and the kept part opens with an assistant message after the summary. With no assistant message to cut at, everything is summarized.
  3. One generateContent request sends everything before the cut to Gemini: every message, and each call with its input and its full output. When the text is over summaryMaxInputChars (2,000,000 characters by default), the longest outputs are cut to one common length, each keeping its head and tail; a conversation over the limit even without any output fails.
  4. Gemini writes a plain-text summary in nine sections: the request and intent, technical concepts, files and code, errors and fixes, problem solving, every user message verbatim, pending tasks, the current work, and the next step. The text you write after /compact goes to Gemini with it.
  5. The conversation becomes one user message (a note, then the summary) followed by the kept messages, which go back as the engine's own messages.
  6. The built-in summary runs, and one line says why, when there is no key, nothing lies before the newest messages, Gemini fails, the summary is under 200 characters or cut at the output limit (32,768 tokens), or the result is not smaller than the conversation.

In both modes gemini-core builds the request with the key, the model and the thinking level it holds for gemini-compact, and reads the answer. After an HTTP 503 ("high demand") the mod asks again after 1 s, 2 s and 3 s, at most four attempts in all, and no attempt starts whose wait would end past 60 s. After a 429 or a key error, gemini-core hands over the request with its next key, when it holds one.

In a live check on 2.1.277, /compact took 2.6 seconds with gemini-3.5-flash-lite, Claude sent no compaction request, and after it the model named a word and a file that appeared only in the summarized part.

Prune mode

  1. The same three triggers reach the session.compact hook.
  2. A tool call in the first message or in the newest 6 messages, or whose result lies in one of them, is kept whole. Every other call gets an id (c1, c2, ...). With no such call, the built-in summary runs.
  3. One generateContent request sends the conversation to Gemini: every message, and each call with its input and its full output. Over maxInputChars (400,000 characters by default) the longest outputs are cut the same way as in summary mode. A response schema allows exactly one answer per id: keep, truncate or drop.
  4. The mod checks the answer (every id exactly once, no other id) and rebuilds the conversation:
    • keep: the call and its output stay.
    • truncate: the call stays, and the output keeps its first 300 characters and one line that says it was cut. An output at most 120 characters longer than that stays whole.
    • drop: the call and its output go. A note on the nearest assistant message names the removed calls, for example [gemini-compact removed 1 earlier tool call(s) and their output after a compaction; they ran: Bash(ls -la /usr/bin)]. Without the note, the model read a reply whose work was gone and said it had never done that work (measured on 2.1.277).
    • A message the answer does not touch goes back as the engine's own message.
  5. When the result is less than 25% smaller, or anything fails (no key, an HTTP error, an answer that breaks the schema), the engine's built-in summary runs and one line says why.

In a live check on 2.1.277, /compact took 1.1 seconds with gemini-3.5-flash-lite. Gemini dropped an ls listing and kept the cat output of a file the user was about to edit, and the conversation became 93% smaller.

What it shows

One line in the transcript, not sent to the model, and a toast that stays for 15 seconds:

gemini-compact: summary: 58 → 7 messages · 91% smaller · 312k in, 5k out
gemini-compact: kept 9/11 messages · 93% smaller · 1 dropped, 0 truncated · 5k in, 59 out
gemini-compact: built-in summary: under 25% smaller (kept 14/16 messages · 3% smaller · ...)
gemini-compact: built-in summary: Gemini HTTP 429: Resource has been exhausted

On the free tier the toast adds · sent to Gemini free tier. An automatic compaction the engine skips or that fails writes one line too (automatic compaction skipped: ..., automatic compaction failed: ...).

Command

/gemini-compact              on or off, mode, the model and thinking level gemini-core holds, threshold, tier, whether a key is set, the last result
/gemini-compact on | off     off leaves every compaction to the built-in summary; on is refused while gemini-core has no key
/gemini-compact mode summary | mode prune
/gemini-compact at <1-99>    compact after a turn that ends with the context over this percentage
/gemini-compact at off       no automatic compaction; /compact and the engine's own compaction still ask Gemini
/gemini-compact reset        back to the plugin options, and off

The mod is off after an install: every compaction is the built-in one, none starts on the mod's threshold, and nothing goes to Gemini until /gemini-compact on. The command settings are kept across sessions and take effect at once. After a compaction it started, the mod starts no other one until a turn ends with the context under the threshold, so a context that stays over it does not compact after every turn.

The key, the tier, the model (default gemini-3.5-flash-lite) and the thinking level belong to gemini-core:

/gemini-core model compact gemini-3.5-flash
/gemini-core thinking compact low
/gemini-core paid

Free tier or paid tier

The conversation holds your prompts, the commands the model ran and the contents of the files it read. On the free tier Google may use them and human reviewers may read them; the gemini-core README quotes the Gemini API Additional Terms. On a project you would not show to Google, use a key with billing enabled and set /gemini-core paid. No mod can tell which tier a key is on; the tier setting only chooses the warning.

Google AI Studio shows the free tier limits per model; the documentation does not. They were not measured. A summary of a long conversation is one large request, so a per-minute token limit can refuse it with HTTP 429; the built-in summary then runs.

Install

claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install gemini-compact@kilimcininkoroglu-mods

It depends on gemini-core, which claude plugin install adds. Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.

To load it from a local checkout for one session, put gemini-core beside it:

claude --plugin-dir plugins/gemini-core --plugin-dir plugins/gemini-compact

After installing

  1. Set the Gemini key and the tier in gemini-core, as its After installing section says, then restart Claude Code.
  2. Run /gemini-compact on. Without a key it answers still off: gemini-core has no Gemini key and stays off.
  3. Run /gemini-compact. The first line reads on · summary · <model> · thinking ... · automatic at 60% · <tier> tier · key set.
  4. Run /compact once. The transcript line should start with gemini-compact: summary:. A line that starts with built-in summary: names why Gemini was not used.

After an update from 0.2.x: claude plugin update does not add gemini-core (measured on 2.1.278), so run claude plugin install gemini-core@kilimcininkoroglu-mods once. Version 0.3.0 moved the key, tier and model to gemini-core; the apiKey, tier and model options and the settings /gemini-compact free|paid|model stored before are no longer read, so set them again in gemini-core. The mode and at settings stay. Version 0.4.0 made the mod off by default: after an update from an earlier version it is off unless you ran /gemini-compact on before, so run /gemini-compact on once.

Options

| Option | Default | What it sets | |---|---|---| | mode | summary | summary or prune; /gemini-compact mode overrides it | | compactAtPercent | 60 | The automatic threshold, 0 to 99; 0 turns it off; /gemini-compact at overrides it | | keepRecent | 6 | Newest messages kept verbatim (summary) or whose calls are never sent for a decision (prune); 0 to 1000 | | minReduction | 0.25 | Prune mode: below this fraction (0 to 1) the built-in summary runs | | headChars | 300 | Prune mode: characters kept of a truncated output; 0 to 100,000 | | maxInputChars | 400000 | Prune mode: characters sent to Gemini at most; 10,000 to 4,000,000 | | summaryMaxInputChars | 2000000 | Summary mode: characters sent to Gemini at most; 10,000 to 4,000,000 |

A value outside its range falls back to the default.

What it can reach

Validated with claude plugin validate on Claude Code 2.1.283:

❯ ./register.ts hooks: session.start, command.run{command=gemini-compact}, session.compact, turn.complete
❯ ./register.ts calls: $.clock.now (via askGemini), $.clock.sleep (via askGemini), $.command.register, $.gemini.enroll, $.gemini.read (via askGemini), $.gemini.request (via askGemini), $.gemini.settings (via compactWithGemini, runCommand, storePatch), $.http.fetch (via askGemini), $.session.compact (via maybeCompact), $.session.usage (via maybeCompact), $.store.delete (via runCommand), $.store.get (via loadConfig), $.store.set (via storePatch), $.ui.log, $.ui.toast (via report)

Reach L3, reaches the network.

1. Reads:    the conversation at each compaction (messages, tool inputs and outputs); the context fill after each main-loop turn; its own $.store; from gemini-core, the request with the key
2. Runs:     no process; one $.session.compact after a turn that ends over the threshold, at most once until the context was under it again
3. Sends:    the conversation (summary: all but the newest messages; prune: all of it), one request per compaction (up to four after a 503, and once more per extra key after a 429 or a key error), to the URL gemini-core builds (generativelanguage.googleapis.com) with the key in the x-goog-api-key header, never in the URL
4. Persists: in $.store, the three command settings (enabled, mode, atPercent); the last result lives in memory
5. Hostile input: the Gemini answer is untrusted: a prune answer is applied only in the schema shape with every candidate id once; a summary becomes the text of one user message the model reads, so a hostile summary can steer the model, as text in a file it reads can; anything malformed falls back to the built-in summary

Limits

  • A summary and a drop are a model's judgment. A summary loses detail the newest messages do not repeat. The prune note tells the model which calls ran, so it can run a tool again.
  • After the compaction the context is written to the cache again. In prune mode it stays larger than a built-in summary, so the next message pays a larger cache write.
  • A summary of a long conversation takes Gemini longer, and the compaction waits for it. Only short conversations were timed.
  • A subagent's own compaction is left to the engine.
  • A compaction the engine precomputes (precompute) also asks Gemini. Whether the engine reuses that result for the compaction that follows was not measured.
  • The test engine of claude plugin test passes no trigger to a $.session.compact() call. The plugin trigger and the hook that answers it were measured in a live session.

Development

make install     # eslint, typescript-eslint, typescript
make lint        # complexity limit 10, the build fails above it
make typecheck   # needs .claude/types/ from /plugin-types
make validate
make test        # claude plugin test

비슷한 프로젝트