AdamAwan/LubbDubb/tree/main/pr-assistant
pr-assistant
풀 리퀘스트를 지원합니다. /pr-assistant:pr map은 전체를 한 페이지로 그리고, /pr-assistant:pr walk는 그 맵을 그린 뒤 정류 지점·diff·메모 패널을 옆에 띄워 함께 살펴봅니다.
이 mod 소개
<p align="center"> <picture> <source media="(prefers-color-scheme: dark)" srcset="docs/brand/logo-dark.svg"> <img src="docs/brand/logo.svg" alt="LubbDubb" width="120"> </picture> </p>LubbDubb
한 소프트웨어 엔지니어의 작업을 위한 셀프 호스팅 상시 실행 오케스트레이션 하네스입니다. 이 _콕핏_은 이슈, PR, CI, 리뷰 댓글 같은 입력을 지켜보고, 하트비트마다 할 일을 결정한 다음 Claude Code 에이전트에게 맡깁니다. 실제 판단이 필요한 일만 사용자에게 에스컬레이션합니다.
이름은 하트비트에서 왔습니다. 서버의 핵심은 모든 것을 움직이는 주기적 펄스입니다.
무엇을 어디서 읽는가.
docs/mission.md는 이 도구가 왜 존재하는지, 작업을 어떻게 바꾸는지 설명합니다. 이 파일은 하네스의 기능, 작업 흐름, 첫날부터 중요한 설정을 다루는 개요입니다. 다른 워크플로가 들어갈 자리까지 포함한 전체 흐름은docs/workflow.md에 있습니다.docs/spec/는 애플리케이션 각 부분의 현재 동작을 사실로 기록한 사양입니다. 모든 설정 키, 디스패치 규칙, 에이전트 런타임, API, 콕핏이 포함됩니다. 아래 내용의 상세가 필요하면 그곳을 보세요.docs/feature-timeline.md는 지금의 형태가 되기까지를 기록합니다.
펄스
하트비트가 구동하는 반복 사이클입니다(플릿이 바쁠 때 heartbeatIntervalMs 기본값은 30s이고, 유휴 상태일 때 idleHeartbeatIntervalMs는 5분). 요청에 따라 직접 실행할 수도 있습니다.
snapshot the world → diff against the last snapshot → reconcile plans
→ decide (dispatcher) → execute (executor) → audit
모든 단계를 기록합니다. 디스패처의 판단 근거, 내보낸 모든 작업, 각 작업의 결과를 저장하므로 유휴 사이클도 바쁜 사이클만큼 설명할 수 있습니다. 에이전트는 풀링된 git worktree(코드 작업)나 임시 디렉터리(데스크 작업)에서 실행되고, 타입이 지정된 도구 채널로 결과를 보냅니다.
하는 일
루프에서 차지하는 위치에 따라 묶었습니다. 각 행은 담당 사양으로 연결됩니다.
작업 받기
| 기능 | 설명 |
| --- | --- |
| 선택적 감시 | ${labelPrefix}-watch 태그 하나가 이슈와 풀 리퀘스트에서 처리할 대상을 결정합니다. 태그 밖의 내용은 건드리지 않습니다. → [06][s06] |
| 두 제공자 | GitHub와 Azure DevOps를 기능별 경계 뒤에 두고, 전체 테스트와 데모는 fake 제공자에서도 실행됩니다. → [15][s15] |
| 트래커 상태 | 제공자에 워크플로 상태가 있으면 상태를 충족한 항목만 가져오고, 하네스가 진행 중과 리뷰 중으로 옮깁니다. → [06][s06] |
| 우선순위와 순서 | 우선순위 라벨이 가져올 순서에 가중치를 주고, 순위가 매겨진 계획을 다음 작업으로 보여 줍니다. 현재 여력에서 선을 긋고 사용자가 다시 정렬할 수 있습니다. → [05][s05] |
| 티켓 보드 | 열린 항목과 닫힌 항목을 표로 보거나 상태별 열이 있는 보드에서 카드를 옮길 수 있습니다. 카드를 놓기 전에 투입 비용을 알려 줍니다. → [17][s17] |
| 첨부 파일 | 브리프에 첨부한 이미지는 해당 이슈를 처리하는 에이전트까지 따라가며 모든 worktree 밖에 저장됩니다. → [12][s12] |
결정하기
| 기능 | 설명 | | --- | --- | | 규칙 파이프라인 | 이름이 있는 스무 개 남짓한 규칙을 선언된 순서대로 거치며, 각 규칙이 세계 상태에서 작업을 제안합니다. 순서는 데이터이지 댓글에 적힌 숫자가 아닙니다. → [05][s05] | | 제한된 어휘 | 디스패처가 요청할 수 있는 것은 검증된 열한 가지 작업뿐입니다. 잘못된 형식은 거부하고 감사 기록을 남기며 실행하지 않습니다. → [05][s05] | | 검사별 CI 정책 | 어떤 검사가 빨간색이 되었는지에 따라 수정, 안내를 곁들인 수정, 우리 소유가 아니므로 보류, 한 번만 에스컬레이션 중 하나를 고릅니다. → [02][s02] | | 결정 로그 | 실행, 지연, 거부, 건너뛰기를 모두 이유와 해당 규칙과 함께 기록하며, 그 규칙이 왜 존재하는지도 펼쳐 볼 수 있습니다. → [18][s18] | | 쿨다운과 상한 | 출처별 시도 상한과 쿨다운을 두어 계속 실패하는 디스패치가 반복 대신 에스컬레이션으로 넘어가게 합니다. → [05][s05] |
작업하기
| 기능 | 설명 | | --- | --- | | worktree의 에이전트 | 제한된 체크아웃 풀을 브랜치에 빌려 주고 다시 만들지 않고 전환하므로, 돌아온 브랜치는 준비된 상태로 시작합니다. → [09][s09] | | 타입 지정 도구 채널 | 에이전트가 호출하는 MCP 서버입니다. 세계를 읽고, 배운 내용을 올리고, 풀 리퀘스트를 열고, 검사를 보고합니다. → [11][s11] | | 권한 안전망 | 허용 목록 밖의 명령을 만난 에이전트는 멈추지 않고 사용자에게 묻습니다. → [11][s11] | | 규칙별 모델 | 디스패치 규칙마다 모델과 실행 깊이를 담은 이름 있는 프로필을 배정하므로 충돌 수정과 계획의 가격이 같지 않습니다. → [02][s02] | | 작업과 일정 | 콕핏에서 임시 프롬프트를 큐에 넣거나 cron 식으로 사용자를 위해 예약할 수 있습니다. 둘 다 다른 작업처럼 슬롯을 기다립니다. → [13][s13] | | 충돌 복구 | 재시작으로 고아가 된 에이전트는 보류되고 펄스도 멈춥니다. 사용자가 복구·재큐잉·제거할 때까지 그대로입니다. → [10][s10] | | 실시간 트랜스크립트 | 에이전트를 클릭해 하는 일을 읽고, 입력하고, 실행 중간에 만든 결과를 볼 수 있습니다. → [10][s10], [12][s12] |
풀 리퀘스트
| 기능 | 설명 | | --- | --- | | 건강 조건 | 실패한 CI, 기반보다 뒤처지거나 충돌하는 상태, 처리하지 않은 리뷰 스레드, 병합 준비 상태를 확인합니다. 브랜치마다 에이전트 하나를 두고 가장 큰 우려부터 처리합니다. → [07][s07] | | 플릿 리뷰 | 기본값은 꺼짐입니다. 사람이 요청받기 전에 하네스가 풀 리퀘스트를 직접 읽고 얼마나 깊게 볼지 분류합니다. → [07][s07] | | 리뷰 답변 | 처리되지 않은 각 스레드를 에이전트 하나에 보내고 하네스를 통해 답합니다. 제공자 자체의 스레드에서 서명하고 기록하며 해결합니다. → [07][s07] | | 스택 | 한 부분은 의존하는 부분을 기반으로 합니다. 빨간 기반은 그것을 소유한 PR에 귀속하고 아래에서 위로 병합합니다. → [07][s07] | | 사용자 대기 | 누군가 사용자에게 배정한 풀 리퀘스트, 요청한 사람, 그리고 얼마나 오래 놓여 있는지를 보여 줍니다. → [17][s17] |
목표별 퍼널
| 기능 | 설명 |
| --- | --- |
| 목표 평가 | 여기서 작업할 목표가 있습니까? 거부할 때 티켓에 부족한 내용을 적고, 목표 텍스트가 바뀌면 해제합니다. → [08][s08] |
| 계획 | 하나의 풀 리퀘스트이거나, 각 부분에 자체 브랜치와 범위가 있는 의존성 연결 분해입니다. 또는 「이미 끝남」일 수도 있습니다. → [08][s08] |
| 계획 승인 | 항상 필요합니다. 계획에는 위험, 범위 제외, 성공을 확인할 방법이 담기며 대화형 플래너와 논의할 수 있습니다. → [08][s08] |
| 평가 | 에이전트의 자신감이 아니라 전달된 결과를 평가합니다. no 쪽은 재계획, 부분 추가, 에스컬레이션으로 이어집니다. → [08][s08] |
| 검증 | 전달된 것을 실행해야만 답할 수 있는 검사를 사람에게 맞게 작성하고, 필요하면 자신의 Claude Code에 넘깁니다. → [20][s20] |
| 회고 | 전달된 목표마다 하나씩, 공유 스크래치패드와 하네스가 기록한 비용으로 작성합니다. → [13][s13] |
| 마무리 | 하네스는 티켓을 닫지 않습니다. 사용자 이름으로 계속 남는 의무를 등록하고, 트래커가 더 이상 열림으로 표시하지 않을 때 정리합니다. → [13][s13] |
도착한 뒤
| 기능 | 설명 |
| --- | --- |
| 환경 | 기본값은 꺼짐입니다. 각 PR이 도착한 커밋과 각 환경에 그것이 있는지를 사용자 명령으로 묻고, 결과는 세 값 중 하나입니다. → [24][s24] |
| 도착 | 어딘가에 도착하면 전달된 목표가 사용자에게 남긴 의무를 열거나 티켓에 한 줄을 추가할 수 있습니다. 환경별 선택 사항입니다. → [24][s24] |
| 배포 후 감시 | 목표가 실행 중인 시스템이 보여야 할 것을 선언하고, 도착이 창을 열면 일정에 따라 텔레메트리를 묻습니다. 모델은 관여하지 않습니다. → [29][s29] |
| 장애물 | 플릿을 가로막는 것을 문장으로 맞추지 않고 키로 다룹니다. 독립된 두 목소리, 소유자 하나, 사람일 수 없는 종료 조건입니다. → [27][s27] |
| 플릿 간 풀 | fleet 위의 거리입니다. 공유 저장소에 플릿마다 네임스페이스를 하나씩 두고, 상호 확인 모델과 다이제스트를 제공합니다. → [28][s28] |
하네스 자체 감시
| 기능 | 설명 | | --- | --- | | 사용자 필요 | 에스컬레이션, 계획 승인, 외부 제안, 권한 요청, 설정 상태, 전달이 남긴 의무를 한 레일에 모읍니다. → [17][s17] | | 인사이트 | 실행 비용과 결과를 보여 주고, 버킷 중앙값의 몇 배에 이르는 실행을 찾는 실시간 소모 감시도 제공합니다. → [18][s18] | | 활주로 | 플릿에 남은 작업이 있는지 플릿 시간으로 측정하고, 남지 않은 이유가 사용자인지도 보여 줍니다. → [25][s25] | | 설정 상태 | 플릿을 조용히 멈출 수 있는 설정을 한 줄씩 보여 주며, 각 줄은 조언이 아니라 실제 세계에 대한 검사로 끝납니다. → [26][s26] | | 오류 로그 | 포착된 모든 실패가 모이는 경로입니다. 저장하고 stderr에 미러링한 뒤 콕핏으로 스트리밍합니다. → [18][s18] | | 자체 업데이트 | 하네스가 자신의 빌드를 감시하고 비운 뒤, 자신을 교체할 슈퍼바이저에게 넘깁니다. → [21][s21] | | 로컬 실행 | 콕핏에서 시작하고 중지하는 머신의 유일한 개발 환경이며, 출력은 플릿이 이미……
설치
먼저 작성자의 README에서 marketplace와 플러그인 이름을 확인하세요. 저장소 구조에 따라 명령어가 달라질 수 있습니다.
claude plugin marketplace add AdamAwan/LubbDubb claude plugin install pr-assistant
원문 / README
LubbDubb
A self-hosted, always-running orchestration harness for one software engineer's work — a cockpit that watches your inputs (issues, PRs, CI, review comments), decides what to do on a heartbeat, and dispatches Claude Code agents to do it, escalating to you only what genuinely needs judgment.
The name is the heartbeat: the server's core is a periodic pulse that drives everything.
Where to read what.
docs/mission.mdis why it exists and what it changes about the job. This file is the overview: what the harness does, how work flows through it, and the configuration that matters on day one.docs/workflow.mdis the workflow in full, including where a different one slots in.docs/spec/is the specification of how every part of the application behaves today, written as fact — every config key, every dispatch rule, the agent runtimes, the API, the cockpit. If you want the detail behind anything below, it is in there.docs/feature-timeline.mdis how it got this way.
The pulse
One repeating cycle, driven by a heartbeat (heartbeatIntervalMs, default 30s while the fleet is
busy; idleHeartbeatIntervalMs, 5 minutes, while it is not) and also
triggerable on demand:
snapshot the world → diff against the last snapshot → reconcile plans
→ decide (dispatcher) → execute (executor) → audit
Every step is recorded. The dispatcher's rationale, every action it emitted, and the outcome of each action are persisted — so an idle cycle is as explainable as a busy one. Agents run in pooled git worktrees (code work) or scratch dirs (desk work), and report back over a typed tool channel.
What it does
Grouped by where in the loop it sits. Each line links to the spec that owns it.
Taking work in
| Feature | What it is |
| ---------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| Opt-in watching | One ${labelPrefix}-watch tag decides what is acted on, on issues and pull requests alike. Nothing outside it is touched. → 06 |
| Two providers | GitHub and Azure DevOps behind per-capability seams, plus a fake provider the whole suite and the demo run on. → 15 |
| Tracker states | Where the provider has workflow states, pickup is gated on them and the harness moves the item to in-progress and in-review. → 06 |
| Priority and order | Priority labels weight pickup; the ranked plan ships as Up next with a cut-line at current headroom, re-orderable by you. → 05 |
| Tickets board | Every open and closed item, as a table or a column-per-state board you drag cards across — the drop's cost said before it lands. → 17 |
| Attachments | Images attached to a brief follow the issue to whichever agent works it, stored outside every worktree. → 12 |
Deciding
| Feature | What it is | | ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------- | | A rule pipeline | Two dozen named rules walked in a declared order, each proposing work from the world. The order is data, not numbers on a comment. → 05 | | A bounded vocabulary | The dispatcher can only ever ask for one of eleven validated actions; anything malformed is rejected and audited, never executed. → 05 | | Per-check CI policy | What to do about which check went red — fix it, fix it with guidance, hold it because it is not ours, or escalate once. → 02 | | The decision log | Executed, deferred, rejected and skipped alike, each with its reason and the rule that produced it, expandable to why that rule exists. → 18 | | Cooldowns and caps | Per-origin attempt caps and cooldowns, so a dispatch that keeps failing escalates instead of looping. → 05 |
Doing the work
| Feature | What it is | | ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | | Agents in worktrees | A bounded pool of checkouts leased to branches and switched rather than recreated, so a branch that comes back starts warm. → 09 | | A typed tool channel | An MCP server agents call back on — read the world, raise what they learned, open a pull request, report a check. → 11 | | A permission backstop | An agent hitting a command outside the allow-list asks you rather than hanging. → 11 | | Models per rule | Named profiles — a model and the depth it runs at — assigned per dispatch rule, so a conflict fix and a plan are not priced alike. → 02 | | Jobs and schedules | An ad-hoc prompt queued from the cockpit, or queued for you on a cron expression. Both wait for a slot like everything else. → 13 | | Crash recovery | Agents orphaned by a restart are parked, and the pulse is held, until you restore, requeue or remove each one. → 10 | | Live transcripts | Click an agent and read what it is doing, type into it, and see what it produced mid-run. → 10, 12 |
Pull requests
| Feature | What it is | | ---------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | | Health predicates | Failing CI, behind or conflicting with its base, unhandled review threads, ready to merge — one agent per branch, top concern first. → 07 | | The fleet review | Off by default: the harness reads a pull request of its own before a person is asked, with a triage that picks how thoroughly. → 07 | | Answering a review | Every unhandled thread goes to one agent, replied to through the harness — signed, recorded, and resolved on the provider's own threads. → 07 | | Stacks | A part is based on the part it depends on; a red base is attributed to the PR that owns it, and merging is bottom-up. → 07 | | Waiting on you | A pull request somebody assigned you, who asked, and how long it has been sitting there. → 17 |
The funnel, per goal
| Feature | What it is |
| ------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Goal appraisal | Is there a goal here to work from? A refusal says what is missing, on the ticket, and lifts when the goal text changes. → 08 |
| Planning | One pull request, or a dependency-chained decomposition into parts each with its own branch and scope — or "this is already done". → 08 |
| Plan approval | Always. A plan carries risks, scope-outs and how anyone will know it worked, and can be discussed with a conversational planner. → 08 |
| Assessment | Asked of what was delivered, not of the agent's confidence. Its no arm replans, adds a part, or escalates. → 08 |
| Validation | Checks that can only be answered by running the delivered thing, written for a person — or handed to your own Claude Code. → 20 |
| Retrospective | One write-up per delivered goal, from the shared scratchpad and the harness's own record of what it cost. → 13 |
| Close-out | The harness never closes a ticket. It files a standing obligation with your name on it, which settles once the tracker stops listing it open. → 13 |
After it lands
| Feature | What it is |
| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| Environments | Off by default. The commit each PR landed as, and whether each environment has it yet — asked with your own command, three-valued. → 24 |
| Arrivals | Arriving somewhere can be what opens what a delivered goal owes you, and what puts a line on the ticket. Both opt-in, per environment. → 24 |
| Post-deploy watch | A goal declares what a running system must show; an arrival opens a window, and your telemetry is asked on a schedule. No model in it. → 29 |
| Obstacles | What is in the fleet's way, keyed rather than matched on prose: two independent voices, an owner, and an exit that is never a person. → 27 |
| The cross-fleet pool | The distance above fleet: one namespace per fleet in a shared repository, with a corroboration model and a digest. → 28 |
Watching the harness itself
| Feature | What it is | | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | | Needs you | One rail: escalations, plan approvals, outbound proposals, permission requests, config health, and the obligations a delivery leaves. → 17 | | Insights | What runs cost, what they yielded and what came out — plus a live burn watch that surfaces a run several times its bucket's median. → 18 | | The runway | Whether there is work left for the fleet, measured in fleet time — and whether the reason there is not is you. → 25 | | Config health | One row per setting that can stop the fleet silently, each ending in a check against the real world rather than advice. → 26 | | The error log | The one path every caught failure funnels through: persisted, mirrored to stderr, streamed to the cockpit. → 18 | | Self-update | The harness watches its own build, drains, and hands off to a supervisor that replaces it. → 21 | | Local runs | The machine's one dev environment, brought up and taken down from the cockpit, with its output in the pane the fleet already has. → 23 | | Pets | A vivarium at the foot of the rail. Your actions drop eggs; the eggs hatch. It gates nothing. → 22 |
The flow of work
Two entry points, one path. A prompt states a goal and a ticket is found or created for it; a ticket states its own. Everything downstream keys on the ticket, so work started from a prompt is as recoverable, reviewable and reportable as work started from the tracker.
flowchart TD
P([Start with a prompt]) --> G[Goal is stated]
T([Start with a ticket]) --> TK
G -- find or create --> TK[Ticket]
TK --> V{Enough information<br/>to proceed?}
V -- no --> AL[Say what is missing,<br/>on the ticket]
AL --> UW([Stop working it])
UW -. the goal text changes .-> TK
V -- yes --> PL[Plan the work]
PL --> AP{Plan accepted?}
AP -- no, revise --> PL
AP -- yes --> WK[/Do the work/]
WK --> QG{Quality gates<br/>review, CI, a person}
QG -- not satisfied --> WK
QG -- satisfied --> M[Merge]
M --> GC{Goal achieved?}
GC -- no --> PL
GC -- yes --> DV[Deliver: validate, write up,<br/>hand back what is yours]
DV --> AR{Configured<br/>environments?}
AR -- yes --> EN[Watch it arrive,<br/>then watch it behave]
AR -- no --> D
EN --> D([Done])
The standard steps
| Step | What happens |
| ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Intake | Every open issue/PR is fetched and shown. What is acted on is decided by the watch tag, plus tracker workflow states where the provider has them. |
| Enough information? | One agent reads the ticket against the repository and says whether there is a goal here to work from. Only an explicit unclear holds anything. |
| Plan the work | A planning agent reads the repo and returns either one PR will do, a decomposition into dependency-chained parts, or this goal is already met. |
| Plan accepted? | Every plan verdict appears in Needs you. Accepting releases the parts to be scheduled; rejecting falls the issue back to the single-PR path, and it can be held. |
| Do the work | An agent per part (or one for the whole issue), each in its own worktree. Code is the most common arm, not the only one — a part may finish with a report. |
| Quality gates | The fleet's own review of the diff (opt-in), then tests, static analysis, pipeline health and human review. Each failing check is classified per check. |
| Merge | A green, approved, mergeable, comment-clear PR is merged — bottom-up for a stack, and never while it is based on another in-flight branch. |
| Goal achieved? | Asked of what was actually delivered, not of the agent's confidence. A no returns to planning, because what is missing may be a different decomposition. |
| Deliver | A validation sheet of checks only a person or a running system can answer; a retrospective written from the record; the tracker state moved and a status comment. |
| Close out | Nothing closes the ticket. A standing obligation with your name on it is filed, and settles itself once the tracker stops listing the item open. |
| Arrival | Where environments are configured: the commit each PR landed as, whether each environment has it yet, and — where declared — whether it is behaving now it is there. |
The gates that carry the loop
Each is a decision something has to make, not a step that always passes.
- Enough information to proceed rejects a goal nothing can act on, before an agent spends itself discovering that. Refusal is not silent — it says what is missing, on the ticket — and it is not permanent: the hold ends when the goal text changes, or when anyone comments.
- Plan accepted is where you see the shape of the work before it happens. There is no switch for it: a plan that started itself can only be undone by a replan, which is strictly worse.
- Quality gates are a set, not a list — tests, static analysis, pipeline health, human review, and the fleet's own review where it is on. The classification is per check, and the third reading is the one that matters: red, but not ours holds and says why, rather than sending an agent at a wall.
- Goal achieved is asked of the delivered work. Its
noarm proposes a replan, a follow-up part, or escalates — depending on what the assessment says fell short.
Priorities when headroom is scarce
The dispatcher ranks every candidate, then applies the concurrency cut. Roughly, highest first:
- An operator queued a job from the cockpit — takes the next free slot.
- A PR with problems: failing CI, a stale or conflicting base, an unhandled review comment.
- A PR that is ready to merge.
- Planning, approval, appraisal and assessment for issues.
- Plan parts, then fresh issue pickup.
PR work runs before new issue pickup, so a PR in trouble is always worked ahead of starting new tickets. The full ordered plan ships to the cockpit as Up next, with a cut-line at the current headroom; you can re-order it, and the override persists.
Where you are in the loop
The harness owns the loop; you own the verdicts. You are asked when — and only when — a decision is genuinely yours:
- Needs you collects escalations, plan approvals, outbound-act proposals, permission requests from agents that hit a command outside the allow-list, the config checks, and what a delivered goal still owes you — a validation sheet to run, a ticket to close.
- Nothing side-effectful leaves without a human, save two standing authorities that are yours: a
stack landing you clicked over named pull requests, and
sendPrRepliesWithoutApproval, on by default, which sends a drafted review reply straight to the thread. A rejection stands until the world gives a reason to ask again — a push, a CI result, an approval, a comment — and the reason you typed is handed to the next agent that works that item. - A restart never decides for you. Agents orphaned by a crash or shutdown are parked, and the heartbeat is held until you restore, requeue or remove each one.
Two things are deliberately fixed: every act reaching the outside world is authorized, and an agent declares that it finished — silence never reads as success.
Getting started
Node >=22.12 and git. A fresh clone installs with npm ci — node-pty is a
native build, so it is not instant (on npm 12+ it builds because package.json
lists it in allowScripts).
npm ci # native dep: node-pty
cp lubbdubb.config.example.json lubbdubb.config.json # your local config (gitignored)
npm start # builds the cockpit, serves on 127.0.0.1:4300
Every key in the config is optional and the harness boots with no file at all — but the defaults
select the real Claude Code runtime, while the shipped example selects the mock one (agentMode: "raw"
against the built-in fake providers). So copy it for a first run and the whole loop turns with no
model or provider credentials. From then on the cockpit's Config page edits that same file
directly, key by key, leaving its comments and ordering alone — hand-editing and the form are two ways
at one file.
npm start builds the cockpit bundle and then runs the server. Two variants matter:
npm run start:server skips the build and serves whatever web/dist already holds, and
npm run serve runs the server under the supervisor that can replace it — which is what
self-update needs, and the way to run it under systemd, NSSM or a long-lived terminal.
→ docs/spec/21
Boot prints what it decided, and the link to open:
[lubbdubb] cockpit listening on 127.0.0.1:4300
[lubbdubb] open the cockpit: http://127.0.0.1:4300/#t=<token>
[lubbdubb] token minted at .lubbdubb/cockpit-token (0600) — reused on the next start
[lubbdubb] heartbeat=300000ms cap=3
[lubbdubb] agent tools: on
Open it once per browser and the cockpit remembers the token. The harness binds loopback only and
every route needs that token, because the cockpit can queue a job — and a job spawns a real agent with
write access to your repo. The token file is gitignored, along with the rest of .lubbdubb/ (the
SQLite database, worktrees, desk scratch dirs and attachments all live under it).
One click: the Claude Code plugin
The harness ships a Claude Code plugin for your own Claude Code: the /lubbdubb:… skills every
Open in Claude Code link in the cockpit calls (/lubbdubb:check 284:C, /lubbdubb:plan 284,
/lubbdubb:ask 284, …), the desktop tool channel those skills talk to, and a notice board: a status
line above your prompt and a panel beside the transcript with what the harness is waiting on you for,
feature progress, pull request states, the fleet and what is up next. Boot writes it to ~/.lubbdubb/plugin:
[lubbdubb] Claude Code plugin 1.0.0-… written to ~/.lubbdubb/plugin — install or update it from the cockpit's MCP tab
Until it is installed the cockpit shows a banner under the top bar; Install it opens the MCP tab,
whose button runs claude plugin marketplace add and claude plugin install for you at user scope, and
removes the hand-registered lubbdubb MCP server and the old /lubbdubb skill if you had them. Restart
open Claude Code sessions to pick it up. Skip it and nothing breaks: every check simply falls to the fleet.
→ docs/spec/11
Then: use Inject event to simulate the world moving (a CI failure, a review comment) and watch the
harness react; click an agent to see its live transcript and type into it; answer items in Needs
you; use New job to launch an ad-hoc prompt, or New schedule to have one queued on a cron
expression (0 9 * * 1-5 — weekdays at nine, read in the harness's own timezone). The Decision log
shows what was decided each cycle and which rule produced it; Activity shows how the world itself
changed.
Needs you also carries the configuration checks — one row per setting that can stop the fleet silently, each ending in a check against the real world rather than a sentence of advice. On a fresh install that rail is the shortest route from the mock loop to a working deployment. → docs/spec/26
Configuration
Every key is optional, and every key, its default and its precedence is in
docs/spec/02-configuration.md. These are the ones that decide
whether a deployment works at all:
| Key | Default | Why it matters |
| ------------------------------ | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| repoRoot | the directory you launch in | The git repository worktrees are cut from. Left alone, the harness works on its own checkout. |
| integrations | all fake | Which provider serves each capability — sourceControl, issues, pool. fake is the mock world the demo and the suite run on. |
| github / azureDevOps | unset | The provider's own target: owner and repo, or organization, project and repository. |
| userId | unset | Who you are to the provider. Pickup reads label authorship, so without it nothing is ever picked up and nothing says why. → 06 |
| ownWorkOnly | true | Whether the world arrives filtered to you: your watch tags, your pull requests. A team decision, so it belongs in the project layer. |
| agentMode | stream | stream runs a real model. raw is the mock agent — argv over a terminal — and what the example config and the tests use. |
| maxConcurrentAgents | 3 | The fleet's one size knob: the worktree pool is this plus a slack of two, read live, so raising it raises the pool with it. |
| defaultBranch | "main" | The integration branch. Not auto-detected, and a PR targeting anything else is treated as stacked. |
| heartbeatIntervalMs | 30000 | The gap between timer-driven cycles while the fleet is busy; idleHeartbeatIntervalMs (5 min) while it is not. Everything is also triggerable on demand. |
| labelPrefix | "lubbdubb" | Derives the one -watch tag. Everything is opt-in; an empty prefix turns the gate off entirely. |
| ci.checks | [] | Per-check policy: dispatch, dispatch with guidance, ignore, or escalate. Empty means every red check gets an agent. → 02 |
| agentModels | unset | Named profiles — a model and the depth it runs at — assigned per dispatch rule. Omitted, no launch carries --model. |
| agentAllowedTools | toolchain + read/skill/web | What an unattended agent may run without asking. Anything else is routed to you rather than hanging. Never set this via claudeArgs. |
| sendPrRepliesWithoutApproval | true | Send a drafted review reply straight to the thread. false is the stricter setting: every draft waits in the inbox. |
| review.enabled | false | The fleet reads a pull request of its own before a person is asked. With two or more review.modes, a triage picks how thoroughly. |
| environments | [] | Where landed work travels: a command per environment printing the commit it is at, optionally what an arrival opens and what to watch for. |
| host / port / auth | 127.0.0.1 / 4300 / on | Loopback and a bearer token. A host reachable off this machine with auth.enabled: false is refused at load. |
A first real deployment is about six lines:
{
"repoRoot": "/path/to/your/repo",
"integrations": { "sourceControl": "github", "issues": "github" },
"github": { "owner": "acme", "repo": "app" },
"userId": "your-github-login",
"agentMode": "stream",
"maxConcurrentAgents": 3,
"defaultBranch": "main"
}
Secrets are never config keys. GITHUB_TOKEN for GitHub, AZURE_DEVOPS_PAT (or a logged-in az
CLI) for Azure DevOps, LUBBDUBB_TOKEN for the cockpit — all from the environment, so
lubbdubb.config.json stays safe to paste. Agents inherit your shell's model credentials; a stray
ANTHROPIC_API_KEY silently moves the whole fleet onto API billing.
Several switches no longer exist. Planning, plan approval, the goal appraisal, the assessment, the retrospective, validation, the tool channel and the permission backstop are all unconditional. A config still naming one of them is warned about and ignored, and the warning says what replaced it. → docs/spec/02
Sharing a config with your team
lubbdubb.config.json is yours and is gitignored. A lubbdubb.project.json committed at the root of
the repository the harness works on is the team's: everyone pointed at that repo picks it up, and
each person's own file wins over it key by key. So the CI routing, the environments, the integration
branch and the tracker's state names are written once in the project, while who you are (userId),
which models you dispatch on (agentModels) and how many agents your machine runs stay local. Every
key is legal in it except repoRoot — the file is found through repoRoot, so it cannot be the
thing that sets it. → docs/spec/02
Development
npm run dev # server with reload
npm run web:dev # cockpit with HMR (proxies /api + /ws)
npm test # unit + integration tests (node:test)
npm run smoke # full end-to-end with real node-pty + a git worktree
npm run check # format, lint, typecheck ×2, knip, test — concurrently, in one shot
npm run check is the gate CI enforces. It runs its stages in parallel and reports every failure
rather than stopping at the first; a warm run costs about as long as the test suite alone.
See docs/spec/19-development.md for the test seams, the coverage and
security workflows, and the hosted GitHub Pages demo build.
Documentation
| Where | What it is |
| ------------------------------------------------------ | ------------------------------------------------------------------------------ |
| docs/operating.md | How to operate the harness: what changes about the job, and what stays yours |
| docs/operating.html | The same guide as a page to skim — open it in a browser |
| docs/workflow.md | The end-to-end workflow, its variation points, and what is narrower than drawn |
| docs/spec/ | The specification of every subsystem as it behaves today |
| docs/feature-timeline.md | What landed when, from the walking skeleton onwards |
| docs/prompt-templates/ | The rule dispatcher's built-in prompt bodies, ready to override |
| CLAUDE.md | Operating notes for agents working in this repo — the sharp edges |
License
MIT. Copyright (c) 2026 Adam Awan.

