yodem/agentic-label-engineering
ale
Agentic Label Engineering (ALE): typed task labels, an append-only event board, and a zero-LLM watchdog for multi-agent coding work, shipped as an `ale` CLI plus a Claude Code plugin with executor hooks, skills and an `/ale-board` status Mod.
About this mod
ALE moves completion judgments out of coding agents and into code: an orchestrator writes one label per task (owner, allowed files, 2–5 acceptance commands), every claim/heartbeat/submission goes into an append-only event log, and only ale verify can mark a task accepted after running the acceptance commands itself. A model-free watchdog flags stale claims, stuck tasks and overruns.
The package bundles the whole Claude Code plugin (agent catalog, hooks, launchers); it adds the /label-layer and /ale:board skills and the /ale-board status Mod, which needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1. Install via pip install git+https://github.com/yodem/agentic-label-engineering.git or uv tool install, then /plugin marketplace add yodem/agentic-label-engineering and /plugin install ale@agentic-label-engineering. Status is pre-1.0; interfaces may change.
Installation
Check the author's README for the marketplace and plugin name first. Commands may change as the repository evolves.
claude plugin marketplace add yodem/agentic-label-engineering claude plugin install ale
Original text / README
Agentic Label Engineering (ALE)
Typed task labels, an append-only board, and a zero-LLM watchdog for multi-agent coding work.
When several coding agents work from one plan, the weak point is the word "done". An agent can say it finished while the tests fail, drift outside the files it was meant to touch, or stall without anyone noticing. ALE moves those judgments out of the agents and into code:
- An orchestrator writes one label per task: who should do it, which files it may change, and the 2 to 5 shell commands that prove it is finished.
- Every claim, heartbeat and submission goes into an append-only event log. State is computed from the log, never stored.
- Only
ale verifycan mark a task accepted, and only after it runs the acceptance commands itself. No agent can mark its own work done. - A watchdog with no model in it flags stale claims, stuck tasks and overruns.
A visual explainer is published at https://yodem.github.io/agentic-label-engineering/
(source: docs/index.html).
PLAN.md ──bake──▶ labels ──init-run──▶ board (events.jsonl)
│
dispatch ─▶ claim ─▶ heartbeat ─▶ submit ─▶ ale verify ─┬─▶ accepted ─▶ integrate
└─▶ rejected ─▶ fix task
Status: pre-1.0. The executor protocol works with Claude Code, Pi and Codex. Interfaces may still change.
Install
Requirements: Python 3.9 or newer, git, and a POSIX system (macOS or Linux).
Install the ale CLI from GitHub, into a virtualenv or as a uv tool:
python3 -m venv ~/.venvs/ale && . ~/.venvs/ale/bin/activate
pip install git+https://github.com/yodem/agentic-label-engineering.git
ale --help
uv tool install git+https://github.com/yodem/agentic-label-engineering.git # alternative
ale --version prints the installed version.
The package bundles the whole plugin (agent catalog, hooks, launchers), so it works outside a checkout,
and workers it starts load the same hooks as the Claude Code plugin.
To hack on ALE itself, install from a clone with pip install -e . (see CONTRIBUTING.md).
Claude Code plugin
The plugin adds the executor hooks, the /label-layer and /ale:board skills, and the /ale-board
status Mod. It calls the ale CLI, so install that first.
/plugin marketplace add yodem/agentic-label-engineering
/plugin install ale@agentic-label-engineering
The /ale-board Mod also needs CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1; without it the rest of the
plugin works and the Mod stays silent. See mod/README.md.
Pi and Codex
- Pi: load
adapters/pi/ale.tsas an extension. - Codex: copy
adapters/codex/hooks.jsoninto your Codex hook config.
Quick start: one task by hand, no agent
This five-minute walkthrough plays both roles, orchestrator and executor, so you can watch the
protocol refuse a false "done". Run it with ale on your PATH (see Install).
1. A throwaway project with a two-task plan. .ale/ holds run state and stays out of git.
mkdir /tmp/ale-demo && cd /tmp/ale-demo
git init -q && printf '.ale/\n' > .gitignore && git add .gitignore && git commit -qm init
ale setup
cat > PLAN.md <<'EOF'
# Plan
## Task 1: Add a greeting module
**Files:** Create: `src/greet.py`
Write `greet(name)` returning `Hello, <name>!`.
Run: `test -f src/greet.py`
Run: `python3 -c "import sys; sys.path.insert(0, 'src'); from greet import greet; assert greet('Ada') == 'Hello, Ada!'"`
## Task 2: Document the greeting
**Files:** Modify: `README.md`
Depends on Task 1.
Run: `test -f README.md`
Run: `grep -q greet README.md`
EOF
2. Bake labels. ALE reads the plan literally: Files: lines become allowed paths, Run: lines
become acceptance checks, and Depends on Task N becomes a dependency. It writes an ale-label
block under each task heading and a PLAN.md.ale-provenance.json record, then exits 1 and lists
the gaps. That's expected.
ale plan bake PLAN.md --write # exit 1: gaps listed
T1: lane
T1: lane_reason
T2: lane
T2: lane_reason
Lane (inline, workflow or pane) and the reason for it are always the planner's call; no
classifier guesses them. Fill them in, check the plan, and commit it:
sed -i.bak 's/"lane":null/"lane":"inline"/; s/"lane_reason": null/"lane_reason": "Small and watched, so inline."/' PLAN.md && rm PLAN.md.bak
ale plan bake PLAN.md
git add -A && git commit -qm plan # PLAN.md and its .ale-provenance.json
3. Start the run and dispatch. init-run freezes the labels. ready lists tasks whose
dependencies are met. dispatch --no-exec creates a git worktree for T1 without launching an agent.
It prints a spawn request (the prompt an agent would get), saved here to a file.
ale init-run --plan PLAN.md --set-current
ale ready
ale dispatch --no-exec > .ale/dispatch.jsonl
AGENT=$(python3 -c "import json; print(json.loads(open('.ale/dispatch.jsonl').readline())['agent_id'])")
WT=.ale/runs/PLAN/wt/T1
4. Act as the executor, and claim done without doing the work.
ale claim --task T1 --agent "$AGENT"
ale heartbeat --task T1 --agent "$AGENT" --step "writing greet.py"
ale submit --task T1 --agent "$AGENT" --summary "done"
ale verify --task T1 --cwd "$WT" # acceptance failed: A1, A2 (exit 1)
ale status # T1 rejected
5. Do the work, then verify again. reopen is the lead's decision to retry verification.
mkdir -p "$WT/src" && printf 'def greet(name):\n return "Hello, %%s!" %% name\n' > "$WT/src/greet.py"
ale reopen --task T1 --reason "greet.py written"
ale verify --task T1 --cwd "$WT" # exit 0
ale integrate --task T1 # commits the allowed changes and merges branch ale/PLAN/T1
ale status # T1 accepted, integrated=yes; T2 ready
ale watchdog # [] - no breaches
With an agent, you write the plan, fill the gaps, and the runner does steps 3 to 5 for every task:
ale run PLAN.md # or, in Claude Code: /label-layer PLAN.md
ale run dispatches through the executors in your roster, verifies submissions, integrates accepted
work, and opens up to two focused fix tasks per rejection. Watch it with ale status,
ale timeline, ale meta, or the web board (ale board --open, or /ale:board in Claude Code).
Concepts
- Label:
labels(closed vocabulary fromroster.json: role, model_tier, lane, risk, effort, locality),context(spec, pointers, allowed paths, dependencies),acceptance(2 to 5 commands),watchthresholds. - Roster: your vocabulary plus the
(role, model_tier) -> executor, modeltable. A new model is a one-line change here.ale setupwrites one to.ale/roster.json. - Board:
events.jsonl, append only. State is computed from it. Handoff files underhandoff/are rendered views for humans and successor agents. - Lease: a claim lives while heartbeats arrive. The watchdog releases dead claims.
- Completion barrier: only
ale verifywritesaccepted, after running the acceptance commands itself. - Stacked tasks: a task with one dependency can start from that dependency's accepted commit. See docs/label-layer.md.
- Event authorship: system events (
verified,accepted,rejected,failed,canceled,lease_expired,released,input_answered) are applied only when written with no agent id. Executor events (claim,heartbeat,submit,input-required,note) must come from the task's current owner. Usage is recorded by the orchestrator or an adapter, not the executor, and is therefore not owner-guarded.
Executors
A label's executor names a harness: claude, codex, pi, or any harness the roster
declares under harnesses. Code, not a model, then decides how and where it runs
(see docs/label-layer.md):
| Mode | When | Runs |
| --- | --- | --- |
| in-session | inline or workflow lane, a harness that can run in the lead's session (claude) and an Anthropic model | an in-session subagent the lead starts; dispatch creates and records its worktree |
| headless | any other inline or workflow task on the local host | the harness's headless argv (for example codex exec --json --skip-git-repo-check …) wrapped by bin/ale-exec |
| pane | a pane lane, or any non-in-session task on the roster's remote_host | a herdr pane of the harness's kind, through the launcher in ALE_HERDR_EXEC |
Work that does not run in-session goes to the roster's remote_host unless the label's locality
is local; it needs a provisioned remote worktree (ALE_REMOTE_WORKTREE_<TASK>), otherwise
dispatch releases the task instead of running it locally. A roster-declared harness looks like:
"harnesses": {"gemini": {"headless": ["gemini", "-m", "{model}", "-p", "{prompt}"], "herdr_kind": "gemini"}}
The older ids claude-headless, codex-exec, pi-print, claude-subagent (and claude_code) and
herdr-pane still load: each names a harness and fixes its mode. ale plan route <plan> --task T --json prints the harness, model, mode and host code picks for a task.
Setup
ale setup creates .ale/roster.json. With --answers FILE, interactively on a terminal, or through
the /ale:setup skill (which asks you each question), it also configures the harnesses it finds on
PATH (with their --version), the model per harness and tier, the CandleKeep handbook refs file,
the remote host from herdr-exec.toml, the judge mode (off by default for private repositories) and
free-text project notes that every prompt carries. ale setup --questions --json lists the questions;
ale setup --check reports what is configured and what is broken. Setup never stores credentials:
it prints the login commands for you to run.
Optional judge
ALE can collect second-opinion votes on labels from an external judge command. It is off by default, runs in shadow mode (it never changes a decision), and any executable that speaks a small JSON contract works. See docs/judge.md.
Exit codes
0 ok, 1 check failed, 2 usage, 3 claim lost, 4 lease lost, 5 needs sign-off, 6 breaches found.
Eval loop
ALE scores its own runs. ale init-run indexes each run under ~/.ale/ (or $ALE_HOME/.ale/).
ale analyze grades every indexed run against bars committed in
ale/schema/analyze_thresholds.json, writes a dated report and findings.json, and exits 1 on a
breached check. ale eval cases --ci replays evalcases/cases.jsonl, a regression suite where
each case is a real run failure, and exits 1 on a failed or regressed case. Both append to one
eval ledger. See docs/analyze.md.
Watchdog
ale watchdog scans open tasks for breaches: a stale lease (no heartbeat within
heartbeat_timeout_s), a claim stuck without progress past stuck_after_s, a run past
max_duration_s, a task left in submitted longer than heartbeat_timeout_s (breach type
unverified: nobody ran ale verify on it in time), and attempts past max_attempts. Run it from
a loop or scheduler; exit 6 means it found at least one breach.
Harness support
The status in each cell describes the shipped integration, not a claim about what the underlying harness could support in a future adapter.
| Rule | Claude Code hooks | Codex hooks | Pi extension | ale-exec wrapper |
| --- | --- | --- | --- | --- |
| Path guard | Enforced: PreToolUse edit denial | Enforced: PreToolUse denial | Enforced: edit and write events | Not possible: verify only |
| Auto heartbeat | Enforced: PostToolUse, 60 s throttle | Not possible: current adapter has no PostToolUse entry | Enforced: post-tool hook | Enforced: timer heartbeat |
| Stop/submit gate | Enforced: Stop acceptance gate | Not possible: current adapter has no Stop entry | Advisory: settlement runs the gate but print mode cannot block | Enforced: exit-time check then submit or input-required |
| Usage capture | Enforced: transcript IDs are deduplicated | Advisory: --usage-from codex-json on wrapper | Enforced: assistant message usage | Enforced: printed JSON usage when configured |
| Session context | Enforced: SessionStart stdout | Not possible: current adapter has no SessionStart entry | Enforced: session start injection | Not possible: wrapper has no context injection |
The Claude Code and Codex hook contracts, Pi event limits, and transcript
fields are recorded in docs/harness-facts.md. Shell commands can write
anywhere, so ale verify --base remains the containment backstop.
Limits and trust boundary
- POSIX only. Local filesystems only: the append guarantee does not hold on network mounts.
- Acceptance commands run with
shell=True. Labels are code. Only run labels you or your orchestrator wrote. See SECURITY.md. verify --baseexpects one task per working tree. Use a git worktree per executor.laneis never chosen by a classifier. The planner answers three questions and recordslane_reason.agent_idis self-asserted. The log guards against accidents and honest mistakes, not against a malicious local process that forges events.- Changed paths are normalised before containment is checked; absolute paths and paths that escape the project are always violations.
ale integrateneeds no uncommitted changes to tracked files (untracked files are ignored). Keep.ale/and virtualenvs in.gitignore.cost_gate.max_concurrentis part of the roster schema but is not enforced yet.
Documentation
| Read | For |
| --- | --- |
| Explainer page | A visual walkthrough of the whole idea |
| docs/label-layer.md | Plan format, label fields, worktrees, fix tasks, the run loop |
| docs/analyze.md | The eval loop: ale analyze, ale eval cases, the ledger, fix records |
| docs/agents.md | Agent taxonomy, catalog lookup, rule enforcement |
| docs/hooks.md | What the hooks enforce, per harness |
| docs/labeling.md | The label cascade and judge shadow decisions |
| docs/judge.md | The optional judge command and its contract |
| docs/harness-facts.md | Verified hook facts for Claude Code, Pi and Codex |
| EXECUTOR.md | The protocol an executor agent follows |
| AGENTS.md, CLAUDE.md | Instructions for coding agents that work on this repository |
| CHANGELOG.md | Release history |
Contributing
See CONTRIBUTING.md. Tests: uvx --python 3.9 pytest -q and cd mod && bun test.
License
MIT. See LICENSE. Parts of agents/ and catalog/refs/ derive from
OrchestKit under MIT; see NOTICE.
