KilimcininKorOglu/claude-code-mods/tree/main/plugins/flaky-memory
flaky-memory
작업 트리의 지문을 채취하고 동일한 코드에서 어떤 테스트가 통과하고 실패했는지 기억하여 모델이 작업 코드를 변경하는 대신 불안정한 테스트를 다시 실행하는 Claude Code 모드입니다.
이 mod 소개
flaky-memory records 테스트는 작업 트리 지문(git rev-parse HEAD, git diff HEAD 및 64 비트 FNV-1a로 해시된 추적되지 않은 파일 이름)에 대해 실행됩니다. 테스트가 실패했지만 지난 7일 내에 동일한 지문에서 통과 및 실패가 모두 발생한 경우, 모드는 코드를 변경하기 전에 모델에 테스트를 다시 실행하라는 메모를 Bash 결과에 추가하고 one 줄 요약을 사용자에게 표시합니다(또는 사이드바에 스트리밍합니다).
go test, pytest, jest, vitest, bun, Cargo test/nextest, PHPUnit, npm/pnpm/yarn test, deno test, rspec, make test, mvn test, gradle test 및 dotnet test에 대한 명령을 인식하고 pass/fail 출력을 구문 분석합니다. 해당 기간 내에 실패한 테스트만 저장됩니다(테스트당 최대 50 실행). 명령에는 /flaky-memory, /flaky-memory 재설정 [id] 및 /flaky-memory on|off가 포함됩니다. 상태는 $.store를 통해 git 공통 디렉토리별로 유지되며 머신에는 아무것도 남지 않습니다. Reach는 L2입니다. 각 테스트 명령 전에 argv로 읽기 전용 git 명령을 실행합니다. claude 플러그인 마켓플레이스로 설치하고 KilimcininKorOglu/claude-code-mods를 추가하고 claude 플러그인을 설치하여 flaky-memory@kilimcininkoroglu-mods를 설치합니다.
설치
먼저 작성자의 README에서 marketplace와 플러그인 이름을 확인하세요. 저장소 구조에 따라 명령어가 달라질 수 있습니다.
claude plugin marketplace add KilimcininKorOglu/claude-code-mods claude plugin install flaky-memory
원문 / README
flaky-memory
A test fails, the model assumes its last change broke it and starts "fixing" code that was fine. Sometimes the test is simply flaky: on the very same code it passed an hour ago. This mod remembers which tests failed on which code. When a test fails that both passed and failed on the same code in the last 7 days, it adds a note to that Bash result, so the model runs the test again instead of changing code for a flaky test.
What it does
Before a Bash command that runs tests, the mod takes a fingerprint of the working tree: git rev-parse HEAD, git diff HEAD and the names of the untracked files, hashed with 64-bit FNV-1a. After the command, it reads which tests the output names as passed or failed and stores one run per test with that fingerprint.
A test is flaky when one fingerprint holds both a pass and a failure of it. A failure after a code change is not flaky, because the fingerprint is different.
After the Bash result of a failed run, the model reads this note:
flaky-memory: go:TestFlip failed 2 of 3 runs in the last 7 days and both passed and failed on the same code once. It may be flaky rather than broken by this change: run it again before you change code for it.
At the same moment you get one line, so you see what the model was told. It holds the finding alone, without the instruction:
flaky-memory: go:TestFlip failed 2 of 3 runs in the last 7 days and both passed and failed on the same code once
With the sidebar open, that line goes into its stream instead, with failed 2 of 3 in red and the explanation faint, and the transcript stays clean. When the window no longer holds a pass and a failure of that test on one tree, the entry goes away and a new one says so, with is no longer flaky in green:
flaky-memory: no longer flaky
go:TestFlip is no longer flaky: nothing in the last 7 days has it passing and failing on the same code
/flaky-memory reset takes the entry down without a closing line, because you asked for it. Without the sidebar, the line lands in the transcript as above.
Test commands
A Bash command counts as a test command when it contains one of these: go test, pytest, python -m pytest (also python3), jest, vitest, bun test, cargo test, cargo nextest, phpunit (also vendor/bin/phpunit), npm test, pnpm test, yarn test (also with run), bun run test, deno test, rspec, make test, mvn test, gradle test (also ./gradlew test), dotnet test. Other commands pass through untouched, and no git command runs for them.
What the output must show
| Runner | Failed | Passed |
|---|---|---|
| go test | --- FAIL: TestX | --- PASS: TestX (with -v) |
| pytest | FAILED path::test, ERROR path::test | path::test PASSED (-v), PASSED path::test (-rA) |
| jest, vitest, bun | ✕, ×, ✗, (fail) lines | ✓, √, (pass) lines |
| cargo test | test x ... FAILED | test x ... ok |
| PHPUnit | 1) Class::method | none |
| deno test | name ... FAILED | name ... ok |
| dotnet test | Failed Name [12 ms] | Passed Name [1 ms] |
| rspec | the rspec path:line # name rerun list | none |
| Maven surefire | name(Class) Time elapsed … <<< FAILURE! | none |
| Gradle | Class > test FAILED | none |
A run that names no passing test still counts as a pass for the tests the same command failed in its last failing run, as long as it exits 0. That way go test ./... without -v, PHPUnit, rspec, Maven and Gradle work too: their failures are read, and their next run that exits 0 counts those tests as passed.
Only tests that failed within the window are stored, so a suite of thousands of passing tests stores nothing. Each test keeps at most 50 runs from the last 7 days.
Command
/flaky-memory the flaky tests of this repository, the most failing first
/flaky-memory reset forget the runs of this repository
/flaky-memory reset <test id> forget the runs of one test, for example go:TestFlip
/flaky-memory on | off record test runs or not (on by default); off keeps the stored runs
The repository is its git common directory, so the worktrees of one repository share their runs.
Install
claude plugin marketplace add KilimcininKorOglu/claude-code-mods
claude plugin install flaky-memory@kilimcininkoroglu-mods
Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.
After installing
- Restart Claude Code. The mod needs no key and no setting.
- It starts recording at the first test run inside a git repository.
What it can reach
Validated with claude plugin validate on Claude Code 2.1.283:
❯ ./register.ts hooks: session.start, command.run{command=flaky-memory}, tool.call{tool=Bash}
❯ ./register.ts calls: $.clock.now (via learn, runCommand), $.command.register, $.process.run (via git), $.session.cwd, $.sidebar.clear (via dropEntry), $.sidebar.set (via toPerson), $.store.delete (via forget), $.store.get (via isEnabled, loadHistory), $.store.set (via forget, learn, runCommand), $.ui.log
Reach L2: it runs git.
1. Reads: the output of each Bash test command; the working tree through git
2. Runs: git rev-parse, git diff HEAD and git ls-files --others, read-only, by argv, before each test command
3. Sends: a note to the model after a failed run of a flaky test, and one line to you; nothing leaves the machine
4. Persists: per repository, in $.store: each failed test's runs of the last 7 days (time, fingerprint, passed) and the tests each command failed last
5. Hostile input: test output is untrusted text; it is matched against fixed line patterns, and a test name is only stored and echoed back, never run
Limits
- Outside a git repository nothing is recorded.
- A diff over 4 MiB gets no fingerprint, and that run is not recorded.
- Only the names of untracked files go into the fingerprint, not their content. A change inside an untracked file does not change the fingerprint.
- State outside the tree (a database, a cache, a file under
/tmp) is not in the fingerprint. A test that depends on it can look flaky. - An interrupted run and a run sent to the background are not recorded, because their output is partial.
- The runner prefixes of the ids (
go:,pytest:,js:,cargo:,phpunit:,deno:,dotnet:,rspec:,maven:,gradle:) keep two runners' test names apart. The go id carries no package name, so two packages with the same test name share one id.
Development
make install # eslint, typescript-eslint, typescript
make lint # complexity limit 10, the build fails above it
make typecheck # needs .claude/types/ from /plugin-types
make validate
make test # claude plugin test