mrzzmrzz/claude-code-mods/tree/main/plugins/omni-token
omni-token
프롬프트 위 띠에 실시간 컨텍스트 창 예측, tok/s, 프롬프트 캐시 적중률, TTL을 보여 주고 선택적으로 캐시 자동 워밍과 /omni-token 제어를 제공하는 Claude Code 플러그인입니다.
이 mod 소개
omni-token은 claude-code-mods 마켓플레이스의 Claude Code 플러그인으로, 매 턴 뒤 갱신되는 컨텍스트 창의 실시간 예측을 프롬프트 위 띠에 그립니다. 화면에는 날씨 스타일의 채움 표시기(Clear에서 Compact soon까지), 절대값과 백분율 컨텍스트 사용량, 12턴 sparkline, 마지막 턴의 초당 출력 token, 프롬프트 캐시 적중률, 프롬프트 캐시 만료까지의 카운트다운이 표시됩니다. 5분 아래에서는 노란색으로 바뀌고 만료 뒤에는 expired로 표시됩니다. /plugin marketplace add mrzzmrzz/claude-code-mods 및 /plugin install omni-token@claude-code-mods로 설치합니다. 옵션은 cacheTtl(5m 또는 1h), autoWarm, warmHours입니다. auto-warm을 켜면 유휴 세션이 캐시가 만료되기 직전에 작은 fork 요청 하나를 보내 TTL을 다시 시작합니다. 이 요청은 트랜스크립트에 들어가지 않으며 warmHours에 도달하거나 캐시가 이미 식었으면 중지됩니다. /omni-token 명령으로 상태를 보고 세션별 워밍을 끄거나 켜거나 초기화할 수 있습니다.
설치
먼저 작성자의 README에서 marketplace와 플러그인 이름을 확인하세요. 저장소 구조에 따라 명령어가 달라질 수 있습니다.
claude plugin marketplace add mrzzmrzz/claude-code-mods claude plugin install omni-token
원문 / README
claude-code-mods
Mods for Claude Code, packaged as a plugin marketplace.
Install
In Claude Code:
/plugin marketplace add mrzzmrzz/claude-code-mods
/plugin install omni-token@claude-code-mods
Mods
omni-token
A live forecast of your context window, shown in the band above the prompt and updated after every turn:
Cloudy 67% 134.4k / 200k ▂▃▄▅▆█ ▲ +98.3k last turn 95 tok/s cache 96% 58:12
| Fill | Forecast (color) | | --- | --- | | < 25% | Clear (yellow) | | 25–49% | Cloudy (cyan) | | 50–74% | Showers (blue) | | 75–89% | Storm (magenta) | | ≥ 90% | Compact soon (red) |
- The sparkline covers the last 12 turns, scaled to the highest of them.
tok/sis the last turn's output tokens (thinking included) over the time from each request's start to its response's end.cache 96%is the last turn's prompt-cache hit rate: cache reads over all input tokens.58:12counts down to when the prompt cache expires: the last main-loop request's start plus the cache TTL. It turns yellow under 5 minutes and readsexpiredafter.
Options
| Option | Values | Default |
| --- | --- | --- |
| cacheTtl | 5m, 1h | 1h |
| autoWarm | true, false | false |
| warmHours | hours | 24 |
Auto-warm
With autoWarm on, while the session is idle the mod sends one tiny forked request over the main thread's own prefix shortly before the cache expires (3 minutes before on 1h, 1 minute on 5m). The cache read restarts the TTL, so the next real message reads the cache instead of rewriting the whole context. The fork never enters the transcript.
- Each warm bills a cache read of the whole context (on Claude Opus 5.5, context × $0.20/MTok) plus a few hundred output tokens.
- It stops
warmHoursafter the last turn started, and never warms an expired cache (that would pay the full write it exists to avoid) or a context under 50k tokens. warmed 3xshows how many warms ran since the last turn;warm missedmeans a warm found the cache already cold, and warming stops until the next turn.
/omni-token
| Command | Effect |
| --- | --- |
| /omni-token | Status: auto-warm on or off and why, cache TTL and time left, warms since the last turn, option names |
| /omni-token warm off | Stop auto-warming in this session, at once; other sessions and the autoWarm option are unchanged |
| /omni-token warm on | Auto-warm this session even with autoWarm off |
| /omni-token warm reset | Follow the autoWarm option again |
On a 1-hour TTL, keeping a cache warm costs 0.2/7.8 of a cold restart per hour, so it pays off for gaps under about 39 hours, if you come back.
The API's usage figures don't say which TTL a session runs on, so set it to match yours: 1h on most Claude subscriptions, 5m on the API default or in usage overage.
Developing
Load a mod straight from this checkout, with hot reload on save:
claude --plugin-dir ./plugins/omni-token
Check one before committing:
claude plugin validate ./plugins/omni-token
claude plugin test ./plugins/omni-token
