TheSmokeDev/taskchad-os/tree/master/.claude/plugins/persona-cognition

TheSmokeDev/taskchad-os/tree/master/.claude/plugins/persona-cognition

#TaskChad OS
永続メモリ、リアルタイム音声、multi-agent オーケストレーション、ブラウザ制御、Telegram/Discord、および operator-controlled ファクトリーを備えた self-hosted コグニティブ エージェント OS およびパーソナル AI アシスタント。
音声メモではなく、あなたの声を生で聞きます。 ChatGPT Voice が Codex アプリで行うのと同じように、エージェントに実際の仕事について大声で話します。ただし、それが自分自身の第 2 の脳であり、オープンソースである点が異なります。実行中のエージェント mid-flight を音声で操作すると、電話を切った後も会話全体を記憶します。 2 つのドア: ダッシュボード、または Discord 音声チャネルの /talk join。
独自の音声をインストールすることもできます。あなたの第 2 の脳は、talk-mode-setup スキルを実行し、キーをチェックし、サイドカーを配線し、デフォルトで Codex subscription を使用するため、per-minute メーターはありません。
音声の下には、チャット ラッパーではなく、実際のコグニティブ OS があります。9-layer 認知スタック、dependency-tracked コンボイ グラフ上の multi-agent オーケストレーション、デフォルトで同意するのではなく信念を形成し、矛盾にフラグを立てるメモリです。単に喜ばせるだけではなく、反発するように作られています。これは、Claude、Codex、Gemini、または OpenAI-compatible バックエンドで実行され、Telegram、Slack、WhatsApp、Web、および CLI を通じてアクセスされます。回帰スイートは、エージェント システムが通常失敗するステートフル境界に集中しています。 証明: テスト + 演算子ループ を参照してください。
TaskChad OS は、分離されたメモリ、範囲指定されたツール、 スケジュールされた作業、チームルーム、承認キュー、永続的な領収書。それができます エージェント作業用の dark-factory ツールキット: 反復的な機械は、 オペレータがスイッチ、承認、証拠を保持したまま移動します。
「闇の工場」とは、無制限の自律性を意味するわけではありません。外部書き込みは残る default-denied、コーディングディスパッチはoperator-controlledのままで、生成された アーティファクトは、その成果物がなければ、プッシュ、デプロイ、公開、または証明されたものとして扱われることはありません。 対応する領収書。
|工場の表面 |組み立てるもの |ここから始めましょう | |---|---|---| |ペルソナファクトリー |アイデンティティ、記憶、ツール、準備、学習 | ペルソナ ブループリント | |リポジトリファクトリー |有界 issue-to-worktree コーディング ディスパッチ | Archon リポジトリのディスパッチ | |ビジュアルファクトリー |根拠のあるイメージのコンセプトと検証済みのプロンプト パック | イメージ ノード ファクトリ | |権威工場 | Evidence-gated SEO/GEO ページとリリース ウェーブ | TokenMax 権限スタック | | Client-site工場 |ブランド化されたテスト可能な service-business Web サイト | クライアント サイト ファクトリ |
TaskChad OS は、このフレームワークのオリジナルのパブリック エクスポートであり、によって維持されています。 TaskChad OS の貢献者。コールと TheHomie から進化しました コミュニティの Claude Code Second Brain ワークショップ、その後 identity-first に成長 独自のメモリを備えたエージェント OS、オーケストレーション、multi-channel イングレス、オペレーティング 部屋とデスクトップの表面。
OpenClaw、Hermes Agent、OpenSouls、および ClaudeClaw がエコシステムとしてクレジットされています 影響を及ぼします。 TaskChad OS は独立したプロジェクトであり、 それらのプロジェクトによって後援または承認されています。 NOTICE.md および 著者.md。

この 45-second alpha-era 製品ツアーでは、ダッシュボード、デスクトップ スタック コントロール、モバイル アクセス、 ブラウザ ビューア、作業キュー、コンボイ、手術室、クリーン シャットダウンの証明。 これは、現在の full-product ウォークスルーではなく、歴史的証拠として保存されています。 リアルタイム音声、quick-agent ステアリング、走行、およびその後の工場出荷時の路面いる 最新リリース に記載されています およびオペレーターマニュアル。
Full-quality MP4 は v0.1.0-alpha.1 リリース。
このスイートは .claude/scripts/tests/ に住んでいます。カウントを公表する代わりに、
生成されたエクスポートが変更されるたびにドリフトし、TaskChad OS は実際のエクスポートを指します
カバレッジ サーフェスとそれを再現するコマンド:
|サブシステム |代表的な取材 |
|----------|--------------------------|
|オーケストレーション | test_orchestration_api.py、test_executor_boundary.py、チームおよびメールボックス スイート |
|認知 + 記憶 | living-self 行為、test_living_memory.py、test_episodes.py、test_session_brief.py、リコールおよび信念スイート |
|トークモード音声 |セッション、実行、ステアリング、フラッシュ、Discord 受信、および報告スイート |
|ランタイム + レーン ルーティング |選択、レーン、プロバイダー アダプター、診断、および quiet-JSON スイート |
|メモリパイプライン |リフレクション、毎週の合成、夢、インデックス作成、およびリコールスイート |
|可観測性 | Langfuse trace-shape および failure-visibility スイート |
ユニットの対象範囲に加えて、フレームワークは operator-loop を通じて実行され、 happy-path アサーションだけでなくスモーク テスト:
/mission、/chat、/mobile、/browser、
/work、/convoy、および /teams。変更したサーフェスに対してターゲット スイートを実行し、広範なローカル回帰に uv run pytest tests/ -q を使用します。証明の境界(まだまだではないもの)
主張されている)は、現在の証明境界 にリストされています。
# Linux/macOS/WSL
curl -sSL https://raw.githubusercontent.com/TheSmokeDev/taskchad-os/master/install.sh | bash
# Windows PowerShell
irm https://raw.githubusercontent.com/TheSmokeDev/taskchad-os/master/install.ps1 | iex
手動パス:
git clone https://github.com/TheSmokeDev/taskchad-os.git
cd taskchad-os/.claude/scripts
uv sync
cp .env.example .env
uv run python setup_wizard.py
uv run thehomie chat
## はじめる
thehomie chat # Start a conversation
thehomie setup # Configure providers and integrations
thehomie setup --check # Verify setup without changing anything
thehomie status --json # Machine-readable health report
thehomie doctor # Diagnostics with fix hints
thehomie desktop --shell # Launch the Desktop dashboard app
thehomie team list # Inspect team sessions
まず作者の README で marketplace とプラグイン名を確認してください。コマンドはリポジトリの構成によって変わる場合があります。
claude plugin marketplace add TheSmokeDev/taskchad-os claude plugin install homie-persona-cognition
A self-hosted cognitive agent OS and personal AI assistant with persistent memory, realtime voice, multi-agent orchestration, browser control, Telegram/Discord, and operator-controlled factories.
It hears you live, not voice notes. Talk your agents through real work out loud, the way ChatGPT Voice does on the Codex app, except it's your own second brain and it's open source. Steer a running agent mid-flight by voice, and it remembers the whole conversation after you hang up. Two doors: the dashboard, or /talk join in your Discord voice channel.
It can even install its own voice. Your second brain runs the talk-mode-setup skill, checks your keys, wires the sidecar, and defaults to your Codex subscription so there is no per-minute meter.
Under the voice is a real cognitive OS, not a chat wrapper: a 9-layer cognition stack, multi-agent orchestration over a dependency-tracked convoy graph, and a memory that forms beliefs and flags contradictions instead of agreeing by default. It is built to push back, not just please. It runs on Claude, Codex, Gemini, or an OpenAI-compatible backend and reaches you through Telegram, Slack, WhatsApp, the web, and the CLI. The regression suite is concentrated on the stateful boundaries where agent systems usually fail; see Proof: Tests + Operator Loops.
TaskChad OS can assemble specialized personas with isolated memory, scoped tools, scheduled work, team rooms, approval queues, and durable receipts. That makes it a dark-factory toolkit for agent work: the repetitive machinery can keep moving while the operator retains the switches, approvals, and evidence.
"Dark factory" does not mean unbounded autonomy. External writes stay default-denied, coding dispatch stays operator-controlled, and a generated artifact is never treated as pushed, deployed, published, or proven without its corresponding receipt.
| Factory surface | What it assembles | Start here | |---|---|---| | Persona factory | Identity, memory, tools, readiness, and learning | Persona Blueprints | | Repository factory | Bounded issue-to-worktree coding dispatch | Archon Repo Dispatch | | Visual factory | Grounded image concepts and validated prompt packs | Image Node Factory | | Authority factory | Evidence-gated SEO/GEO pages and release waves | TokenMax Authority Stack | | Client-site factory | Branded, testable service-business websites | Client Site Factory |
TaskChad OS is the original public export of this framework, maintained by the TaskChad OS contributors. It evolved from Cole and the TheHomie Community's Claude Code Second Brain workshop, then grew into an identity-first agent OS with its own memory, orchestration, multi-channel ingress, Operating Room, and desktop surfaces.
OpenClaw, Hermes Agent, OpenSouls, and ClaudeClaw are credited as ecosystem influences. TaskChad OS is an independent project and is not affiliated with, sponsored by, or endorsed by those projects. See NOTICE.md and AUTHORS.md.

This 45-second alpha-era product tour shows dashboard, Desktop Stack controls, Mobile Access, Browser Viewer, Work Queue, Convoy, Operating Room, and clean shutdown proof. It is preserved as historical evidence, not a current full-product walkthrough. Realtime voice, quick-agent steering, Runs, and the later factory surfaces are documented in the latest release and the operator manual.
Full-quality MP4 is attached on the v0.1.0-alpha.1 release.
The suite lives in .claude/scripts/tests/. Instead of publishing a count that
drifts whenever a generated export changes, TaskChad OS points to the actual
coverage surfaces and the commands that reproduce them:
| Subsystem | Representative coverage |
|-----------|-------------------------|
| Orchestration | test_orchestration_api.py, test_executor_boundary.py, team and mailbox suites |
| Cognition + memory | living-self acts, test_living_memory.py, test_episodes.py, test_session_brief.py, recall and belief suites |
| Talk Mode voice | session, runs, steering, flush, Discord receive, and debrief suites |
| Runtime + lane routing | selection, lane, provider adapter, diagnostics, and quiet-JSON suites |
| Memory pipelines | reflection, weekly synthesis, dream, indexing, and recall suites |
| Observability | Langfuse trace-shape and failure-visibility suites |
On top of unit coverage, the framework is exercised through operator-loop and smoke testing, not just happy-path assertions:
/mission, /chat, /mobile, /browser,
/work, /convoy, and /teams.Run the targeted suite for the surface you changed and use uv run pytest tests/ -q for the broad local regression. Proof boundaries (what is not yet
claimed) are listed under Current Proof Boundaries.
# Linux/macOS/WSL
curl -sSL https://raw.githubusercontent.com/TheSmokeDev/taskchad-os/master/install.sh | bash
# Windows PowerShell
irm https://raw.githubusercontent.com/TheSmokeDev/taskchad-os/master/install.ps1 | iex
Manual path:
git clone https://github.com/TheSmokeDev/taskchad-os.git
cd taskchad-os/.claude/scripts
uv sync
cp .env.example .env
uv run python setup_wizard.py
uv run thehomie chat
thehomie chat # Start a conversation
thehomie setup # Configure providers and integrations
thehomie setup --check # Verify setup without changing anything
thehomie status --json # Machine-readable health report
thehomie doctor # Diagnostics with fix hints
thehomie desktop --shell # Launch the Desktop dashboard app
thehomie team list # Inspect team sessions
| Start here | What it covers |
|---|---|
| Install Guide | Prerequisites, setup wizard, channel credentials, Docker, systemd, vault setup |
| Talk Mode Showcase | Give your AI a real-time voice co-founder — architecture, the battle-tested receive pipeline, build-your-own guide |
| Operator Manual | Public feature map, source-of-truth files, operator entry points, tests, proof boundaries |
| Persona Harness Learning | Daily learning operations, evidence, defaults, pause/resume, rollback, and troubleshooting |
| Learning Developer Guide | Existing lifecycle hooks, domain evidence integration, and isolated executable examples |
| Desktop v0 | Dashboard-first Electron app, portable/package smoke proof, Desktop/Hono/Python lifecycle |
| Multi-Channel Adapters | Telegram attachments, grouped documents, quick-turn batching, Queue/Steer controls |
| Runtime Status And Model Control | /provider, /model, lane-first runtime behavior, quiet JSON contract |
| FRAMEWORK.md | Compact development guide generated during public framework export |
/talk, OpenAI Realtime
WebRTC) and in Discord voice channels, with tool calling, real execution, and
the session-end vault debrief on both surfaces — every slice adversarially
reviewed and live-canary proven. The Discord voice receive audio bug
(mid-word zero-splices from a mic-pump timing race) was root-caused from a
live PCM capture and fixed with a paced jitter buffer — two adversarial
review rounds plus an independent design gate, locked by a deterministic
virtual-clock test harness. See the
Talk Mode Showcase for the full story.It's 6:30am. You open a session.
Instead of "Hi, how can I help you today?" — you get:
"Morning. While you were out — your business had 3 new leads overnight, the loan you flagged is 5 days from maturity, and there's an inbound email from a backlink partner worth reviewing. Yesterday you were mid-decision on the routing refactor. Pick that up, or hit the leads first?"
You didn't set up a notification. You didn't write a morning brief. TaskChad OS was watching. Its memory isn't a static file you load — it's a living record tended between sessions. Its identity isn't a document you edit — it's a self that amends when the evidence is strong enough.
The load-bearing walls are up. The "while you were out" brief is a shipped feature — the Session Opening Brief composes fresh heartbeat observations, new threads, episodes written while you were away, and applied memory amendments into the first turn after an absence, with zero extra LLM calls (cognition/proactive_brief.py, 51 tests in test_session_brief.py). Vault, tiered recall, daily reflection, weekly synthesis, dream consolidation, WorkingMemory-owned prompt state, and the self-evolution replay loop all ship today. Ambient monitoring runs on the heartbeat; durable identity amendments only apply after clearing the default-deny evidence + policy gate described below.
The vault is where TaskChad OS's mind actually lives. Not a notes folder it writes to — the substrate it thinks on. Every recall, every reflection, every promotion reads and writes here. When you edit SOUL.md, you're editing the agent's personality. When concepts/YourBusiness.md accumulates a new section, the agent learned something.
| Layer | What's in it |
|-------|--------------|
| Identity | SOUL.md (personality, values, tone), SELF.md (self-model — capabilities, failure modes), USER.md (you — projects, accounts, preferences) |
| Memory | MEMORY.md (long-term decisions/lessons), GOALS.md (objectives + metrics), daily/YYYY-MM-DD.md, weekly/YYYY-WNN.md, WORKING.md (cross-session scratchpad) |
| Knowledge graph | concepts/ (auto-compiled entity pages), connections/ (cross-domain insight articles), qa/ (filed Q&A from /file), raw/ (immutable original sources) |
| Indexes & log | INDEX.md (whole-wiki catalog, auto-refreshed), concepts/INDEX.md (concept drill-down), LOG.md (append-only compilation timeline) |
| Structure | wikilinks ([[YourBusiness]]), backlinks, MOCs, dashboards, Dataview queries, canvases, graph view |
| Tooling | vault_lint.py (8 health checks, zero LLM cost), entity_extractor.py (extract / compile / contradictions / backfill / sweep / index / preserve-raw / archive), automatic raw-source preservation |
| Pipelines | daily reflection (8 AM), weekly synthesis (Sunday 8 PM), dream consolidation (nightly ~3 AM + post-weekly + on-demand) |
| Sync state | _state/ — memory candidates, self-model inferences, sync manifest. Optional Obsidian Sync via _state/ exclusion patterns. |
Is Obsidian required? No. The vault is plain Markdown — edit it with anything. Obsidian is the recommended editor because the wikilinks, backlinks, graph view, Dataview, and canvas all light up natively. TaskChad OS itself only needs the files.
Where does the vault live? Default vault/memory/, override with HOMIE_VAULT_DIR=/path/to/your/vault (env var honored across runtime, bootstrap, heartbeat, team memory, finance, sanitizer).
TaskChad OS is provider-agnostic. Claude SDK, Codex, Gemini, OpenRouter, OpenAI-compatible — interchangeable batteries. The framework runs the same regardless. Editor adapters (Claude Code project instructions, hooks, MCP bridges) are integration surfaces layered on top of the framework, not part of it. When the heartbeat runs through Codex or Gemini fallback, those editor instructions are not touched.
# Chat
thehomie chat # Interactive REPL
thehomie chat -q "hello" # Single query, stdout response
thehomie chat -q "hello" -Q # JSON output (machine/API contract)
thehomie chat --resume <id> # Resume session by ID
thehomie chat -c # Resume most recent session
thehomie chat -m claude # Force a specific provider/lane
# In-chat commands (any channel)
/working # Show open threads / hypotheses / questions
/working add "text" # Append to scratchpad
/working resolve <N> # Move item N to archived
/file # File the last bot answer as a vault note (with entity cascade)
# Budget (personal finance, optional)
/budget # Snapshot — balances, bills, loans, allocations
/budget transactions # Last 20 bank transactions
/budget spending # Spending by category (current month)
/budget accounts # Connected bank accounts
/budget connect # Connect new bank (Teller / Plaid)
/forecast # Forecast cash flow + bill timing
# System
thehomie setup # Interactive onboarding wizard
thehomie setup --check # Verify all integrations without changing anything
thehomie status # System health overview
thehomie status --json # JSON health report
thehomie doctor # Deep diagnostics with actionable fix hints
# Multi-agent convoy
thehomie convoy create ... # Create convoy with subtasks + deps
thehomie convoy list # List convoys (optional: --status active)
thehomie convoy show <id> # Convoy detail + subtask status
thehomie convoy dispatch <sid> # Dispatch a subtask via executor
thehomie convoy complete <sid> # Mark subtask complete
thehomie convoy fail <sid> # Mark subtask failed
thehomie convoy cancel <id> # Cancel convoy
thehomie convoy add-task <id> # Add subtask to existing convoy
# Mailbox
thehomie mailbox send ... # Send typed inter-agent message
thehomie mailbox inbox <agent> # Check agent inbox
thehomie mailbox claim <agent> # Claim deliveries
thehomie mailbox ack <did> # Acknowledge delivery
# Team sessions
thehomie team list # List active team sessions
thehomie team status <id> # Team detail + members + mailbox backlog
thehomie team members <id> # Member list with roles
thehomie team shutdown <id> # Request graceful shutdown
thehomie team ping <id> # Bump activity timestamp
thehomie team close <id> # Force-close team session
CHANNELS COGNITIVE ENGINE RUNTIME (lane-first)
────────── ──────────────── ─────────────────────
Telegram ─┐ ChatRouter._handle_inner() selection.py
Slack ────┤ │ lane_router.py
Discord ──┤ IncomingMessage ConversationEngine │
WhatsApp ─┤ ──────────────→ ├─ Tier Gate (rules, no LLM) ├─ Claude SDK (Max sub)
Web/MC ───┤ ├─ Recall (dual search+graph) ├─ Codex CLI (ChatGPT sub)
CLI ──────┘ ├─ Region Assembly (frozen) └─ openai-compatible
│ identity · self · user · durable (Gemini · OpenRouter ·
│ working · recent_conversation OpenAI · local)
│ + dynamic regions
├─ Mental Process Detection
├─ Runtime dispatch Health-aware fallback,
└─ Post-response learning manual /provider control,
cost tracking, retry
MEMORY SUBSTRATE (the vault) BACKGROUND PIPELINES ORCHESTRATION
──────────────── ──────────────────── ─────────────
Obsidian-compatible Markdown Heartbeat ───── every 30 min Convoy DAGs
SOUL · SELF · USER · MEMORY Reflection ──── 8 AM daily Typed mailbox
GOALS · WORKING · HEARTBEAT Weekly ──────── Sunday 8 PM Team sessions
daily/ · weekly/ Dream ───────── post-weekly + Backend fallback
concepts/ · connections/ on-demand (auto → paperclip
qa/ · raw/ · _state/ All via recall_service.recall() → workflow → local)
INDEX.md · LOG.md · MOCs (sole entrypoint, Invariant I-3) Local API :4322
Hybrid search (FTS5 + 768-dim
BGE vector + LLM re-rank) COMPILATION ENGINE
Memory graph (1-hop + hub boost) ──────────────────
entity_extractor.py (pure Python heuristic)
Ingest → extract → compile → connect
→ contradict → reindex → log → archive
Fires automatically on ingest, /file,
daily reflection, weekly synthesis (8 entry points)
L9 SELF-EVOLUTION Belief + contradiction engine (operator_beliefs.py,
belief_conflicts.py); identity-file amendments behind a
default-deny evidence + policy gate (amendments.py,
evidence_gate.py); Evolve replay-veto harness
L8 CONTINUITY Session persistence, full cognition on resume (no skip),
recent_conversation region (600 tok), compaction flush,
open-loop tracking
L7 THINKING Immutable WorkingMemory + gated cognitive pass that never
enters the transcript (working_memory.py, cognitive_pass.py)
L6 LEARNING Auto-capture → staging → promotion → skills → inference
L5 RECALL 3-tier gate + dual (keyword+vector) search + 1-hop graph
traversal + hub-score boost + Tier-1 haiku re-rank (recall.py)
L4 MEMORY MEMORY.md + daily/weekly logs + hybrid search index
L3 UNDERSTANDING USER.md + Theory of Mind (inference tracker, confidence)
L2 SELF-AWARENESS SELF.md — capabilities, patterns, failure modes
L1 IDENTITY SOUL.md — personality, values, boundaries, tone
L0 FOUNDATION Obsidian vault graph + MOCs + autolink
62 cognition modules in .claude/chat/cognition/, covered by 479 tests across
13 core files (recall, beliefs, episodes, working memory, session briefs). Every L5–L9
claim above resolves to a named module and test file — see the
Proof table. PageRank and Brandes betweenness are
implemented in graph.py, but the live recall path boosts by a simpler
link-centrality (hub) score; treat the heavier centrality measures as available,
not as what currently drives ranking. Full breakdown: docs/architecture.md.
L0-L9 is the engineering view. The product story has five dimensions; the operator-facing public map lives in docs/manual/README.md. Private PRDs, PRPs, and vault notes stay outside the public framework export.
| Dimension | The question it answers | Status | |-----------|-------------------------|--------| | 1. Identity | Who am I, and how do I know? | ✅ SOUL/SELF/USER injected every turn + session-opening briefing engine | | 2. Memory | What do I know, and how do I find it cheaply? | ✅ Vault + FTS5 + 768-dim BGE vector + graph + Tier-1 LLM re-rank + briefing compression | | 3. Continuity | Do I remember yesterday, and can I pick up mid-thought? | ✅ WORKING.md scratchpad + full cognition on resume + dream consolidation | | 4. Ambient Awareness | Am I watching when you're not here? | 🔄 Heartbeat live; reliability hardening + ambient monitor tasks in flight | | 5. Self-Evolution | Can I grow without manual edits? | ✅ Belief/contradiction engine + identity amendments behind a default-deny evidence + policy gate; 🔄 broader auto-apply scope still expanding |
| | Invariant | Rule |
|---|---|---|
| I-1 | Canonical Ingress | All 6 channels enter ChatRouter._handle_inner(). No bypasses. |
| I-2 | Durable Session Identity | session_key (conversation) separated from request_id (transport). |
| I-3 | One Recall Service | recall_service.recall() is the sole entrypoint — chat, heartbeat, reflection, weekly. |
| I-4 | UI Through APIs | The TaskChad OS Dashboard calls framework APIs, not raw DB. |
| I-5 | Runtime Contract | Provider invocation only through runtime/. No leaky provider hints. |
npm install -g @anthropic-ai/claude-code (handles auth + model access)# 1. Dependencies
cd .claude/scripts && uv sync
# 2. Configure
cp .env.example .env # Add TELEGRAM_BOT_TOKEN, OWNER_NAME, provider keys
# 3. Integrations (Google OAuth, Asana, Slack)
uv run python setup_auth.py # Walk through each integration
uv run python setup_auth.py --check # Verify everything is connected
# 4. Build the memory search index
uv run python memory_index.py --rebuild # ~80MB ONNX model, one-time download
# 5. Start the agent
uv run python ../chat/main.py # Foreground
bash ../chat/run_chat.sh # Background (writes bot.log, bot.pid)
# 6. Schedule background jobs (Windows — Task Scheduler)
# Creates: heartbeat (30 min), daily reflection (8 AM),
# weekly synthesis (Sun 8 PM), dream consolidation (post-weekly + on-demand)
powershell -ExecutionPolicy Bypass -File .claude/scripts/setup_scheduler.ps1 # Run as Admin
Your agent's persistent memory lives in vault/memory/ by default — override with HOMIE_VAULT_DIR=/path/to/your/vault. Auto-loaded at session start (provider-agnostic — works for Claude SDK, Codex, Gemini, OpenRouter).
| File | What It Holds |
|------|---------------|
| SOUL.md | Personality, values, communication style, behavioral rules |
| SELF.md | Self-model — capabilities, patterns, failure modes |
| USER.md | Your profile — projects, accounts, integrations, preferences |
| MEMORY.md | Long-term memory — decisions, lessons, important facts |
| GOALS.md | Quarterly objectives, key metrics, active projects |
| HEARTBEAT.md | What to check and surface each heartbeat run |
| WORKING.md | Cross-session scratchpad — open threads, hypotheses, unresolved questions |
| daily/YYYY-MM-DD.md | Session logs, heartbeat entries, daily context |
| weekly/YYYY-WNN.md | Weekly summaries — patterns, progress, decisions |
| concepts/, connections/, qa/, raw/ | Auto-compiled knowledge graph (see Knowledge Compilation) |
Runtime selection is lane-first: /model claude, /model codex, /model gemini, /model openrouter, /model openai, and /model auto choose where the next request runs. Provider-specific model pins use provider:model, but Codex also accepts short GPT-style aliases:
uv run thehomie chat -q "/model codex:default" -Q # Codex plan default; no --model flag passed
uv run thehomie chat -q "/model codex:gpt-5.5" -Q # Pin a concrete Codex model
uv run thehomie chat -q "/model gpt5.5" -Q # Same pin, easier shorthand
uv run thehomie chat -q "/model codex 5.5" -Q # Same pin, provider + version shorthand
uv run thehomie chat -m codex:gpt-5.5 -q "Reply OK" -Q
codex:default, codex latest, and gpt latest clear the Codex model pin and leave the Codex CLI/ChatGPT plan to choose its hidden backend model. Pinned values such as codex:gpt-5.5, gpt5.5, gpt 5.5, gbt 5.5, codex 5.5, and codec 5.5 are normalized to gpt-5.5.
/provider, /diagnostics, and thehomie status --json report the configured model. When Codex is set to chatgpt-plan-default, the CLI/ChatGPT plan chooses the concrete backend model and The Homie reports that backend as unobserved.
.env)| Variable | Description |
|----------|-------------|
| OWNER_NAME | Your name — used in heartbeat prompts and memory |
| HOMIE_VAULT_DIR | Absolute path to your vault (default vault/memory/) |
| TELEGRAM_BOT_TOKEN | Bot token from @BotFather |
| TELEGRAM_ALLOWED_USER_IDS | Comma-separated user IDs allowed to chat |
| HEARTBEAT_TIMEZONE | IANA timezone (e.g. America/Chicago) |
| LANGFUSE_SECRET_KEY | Langfuse API key for observability (optional) |
| ORCHESTRATION_API_TOKEN | Bearer token for the local orchestration API (optional) |
Full reference: INSTALL.md
cd .claude/scripts
uv run python memory_search.py "query" # Hybrid (recommended)
uv run python memory_search.py "query" --mode keyword # Fast, exact
uv run python memory_search.py "query" --mode semantic # Conceptual match
uv run python memory_search.py "topic" --path-prefix daily/
uv run python memory_index.py --stats # Index stats
uv run python memory_index.py --rebuild # Force full reindex
Ported from Karpathy's LLM Wiki pattern: when a document is ingested, the compilation engine extracts entities, creates concept pages, detects connections, and flags contradictions. The vault compounds automatically.
cd .claude/scripts
# Extract entities from any document (prints JSON)
uv run python entity_extractor.py extract "path/to/doc.md"
# Compile: extract + create/update concept pages + connections
uv run python entity_extractor.py compile "path/to/doc.md" --vault-dir "vault/memory"
# Bootstrap: compile ALL existing vault notes (one-time)
uv run python entity_extractor.py backfill --vault-dir "vault/memory" --dry-run
uv run python entity_extractor.py backfill --vault-dir "vault/memory"
# Sweep: compile only notes without concept coverage
uv run python entity_extractor.py sweep --vault-dir "vault/memory"
# Check contradictions on a concept page
uv run python entity_extractor.py contradictions "vault/memory/concepts/LANGFUSE.md"
# Generate/regenerate concepts/INDEX.md (grouped by entity type)
uv run python entity_extractor.py index --vault-dir "vault/memory"
# Generate/regenerate root INDEX.md (whole-wiki catalog: identity + MOCs + concepts + dirs)
uv run python entity_extractor.py index-root --vault-dir "vault/memory"
# Preserve a source into raw/ as an immutable archive (Karpathy raw/ pattern)
uv run python entity_extractor.py preserve-raw "path/to/source.md" --vault-dir "vault/memory"
# Archive stale orphan concept pages
uv run python entity_extractor.py archive --vault-dir "vault/memory" --dry-run
uv run python entity_extractor.py archive --vault-dir "vault/memory" --page "SOME-SLUG"
# Vault health lint (8 checks, zero LLM cost)
uv run python vault_lint.py --vault-dir "vault/memory"
uv run python vault_lint.py --vault-dir "vault/memory" --check broken_wikilinks
uv run python vault_lint.py --vault-dir "vault/memory" --format json
Knowledge graph structure:
| Folder | Contents | Created By |
|--------|----------|-----------|
| concepts/ | Auto-compiled entity pages — accumulate claims from multiple sources | Compilation cascade |
| connections/ | Cross-cutting insight articles linking 2+ related concepts | Compilation cascade |
| qa/ | Filed Q&A answers from /file bot command | /file command |
| raw/ | Immutable original sources (never modified) | Vault ingest workflow |
| BUILD-LOG.md | Chronological record of every compilation run | Compilation cascade |
When compilation fires automatically:
/file — Instant filing of bot answers with entity cascade/file nudge — Auto-suggested after long analytical responses (>800 chars)Vault health: vault_lint.py runs 8 checks (orphans, broken wikilinks, frontmatter, tag audit against SCHEMA.md, stale content, page size, index completeness, contradiction scan). Zero LLM cost — pure Python. Wired into daily reflection as an automatic post-step.
Provider-agnostic: entity extraction is pure Python (heuristic — headings, bold, wikilinks, frontmatter). No API calls needed. Heading numbers (1. , 3- ) auto-stripped from slugs. The ingest workflow can enhance extraction when running in an LLM context.
The local API (port 4322) exposes convoy, mailbox, and team endpoints. The public operator map starts in docs/manual/README.md; private agent instructions are not part of the public export.
Team dispatch uses a BackendSelector with auto → paperclip → workflow → local fallback. Team memory is stored per team-id in the vault with secret guardrails (8 credential patterns rejected before write).
Langfuse self-hosted or cloud — every message produces a single nested trace:
chat_message (ROOT)
├─ session_lookup
├─ process_detection
├─ recall (classify_tier + recall_pipeline)
├─ region_assembly
├─ runtime execution ← model/provider/cost tracked where the active runtime exposes it
└─ post_response
Set LANGFUSE_ENABLED=true in .env and point LANGFUSE_BASE_URL at your instance. The cognitive-loop smoke was validated locally with trace c14af2029d3188b8a6f7526cda68946d, which captured root chat_message plus session_lookup, process_detection, region_assembly, recall, recall_pipeline, classify_tier, and post_response. With SENTRY_DSN configured, the SDK also returned an event id for an isolated controlled exception.
Identity files (SOUL.md, SELF.md, USER.md, MEMORY.md) are not edited blind. The self-evolution loop captures behavior corrections, accumulates evidence, and proposes amendments to an append-only ledger. A proposed amendment only applies if it clears a default-deny gate — a confidence floor, a vault-confined evidence read that bounds every cited path, secret rejection, and a deterministic regression floor — and writes a rollback snapshot before it touches the target file. Candidate identity/config deltas are additionally replayed against a stratified golden corpus and hard-vetoed on regression. The gate is the source of truth here, not a tagline: an empty-evidence, high-confidence "I read the doc" amendment is rejected on a real falsifiable check, leaving SELF.md byte-unchanged (see tests/test_living_self_act4.py).
| Component | Where | What it does |
|-----------|-------|--------------|
| InferenceTracker | cognition/self_model.py | Captures confirmed observations, decays old inferences, surfaces high-confidence beliefs |
| Skills conflict guard | skills.py | Prevents duplicate skill registration during auto-generation |
| evolve subsystem | .claude/scripts/evolve/ | Replay engine for proposed identity / config deltas |
| └─ replay.py | | Runs candidate overrides against golden_queries.json |
| └─ replay_tracing.py | | Tags replay runs into a dedicated evolve-replay Langfuse namespace (opt-in via --trace, isolated from production cost data) |
| └─ regression.py | | Bootstrap CIs + hard-veto on regression against regression_queries.json |
| └─ goldens.py | | Stratified golden-query management |
| └─ veto.py | | Configurable veto rules (schema in veto_rules.schema.json) |
| └─ compare.py / statistics.py | | Side-by-side replay comparison with confidence intervals |
What's shipped:
EVOLVE_TRACE_REPLAYS env var), isolated from production cost dataTwo-phase ship rhythm: every Evolve increment goes through ship → adversarial Codex review → harden → claim done. That review caught a recurring class of bug across five PRs (tunable config bound in default args, derived-cache trusted as source of truth, optional-provider calls bypassing the enabled-flag helper) that unit tests missed — the three anti-patterns are now written up as enforced review rules with grep checks.
| | OpenClaw | Hermes Agent | TaskChad OS |
|---|---|---|---|
| Thesis | Channel breadth - 25+ adapters | Self-improving skills loop | A real partner - identity + memory + proactive judgment + the nerve to push back |
| Interface | Many chat channels | TUI, CLI, gateway, and desktop workbench | CLI, Telegram/Slack/Discord/WhatsApp/web relay, dashboard, and Desktop v0 shell |
| Runtime | Adapter-first routing | Broad provider/model support plus terminal backends | Lane-first runtime with /provider, /model, status/doctor, and quiet JSON contract |
| Learning loop | Notes and commands | Skills from experience, skill improvement, memory nudges, session search | Belief/contradiction engine, evidence-gated identity amendments, staged memory promotion, replay-veto safety |
| Memory | Plain-text notes | MEMORY.md, user modeling, FTS session search | 9-layer vault: identity, graph traversal + hub boost, dual search, daily/weekly synthesis, staged promotion |
| Knowledge graph | No | Not the focus | Entity compilation engine: concept pages, connections, contradictions, Q&A filing, Tier-1 LLM re-ranking |
| Operator surface | Bot-style access | Gateway and terminal workbench | Operating Room, Capability Gateway, Team Room, Desktop v0, public manual surfaces |
| Multi-agent | No | Subagents and parallel workstreams | Convoy DAGs with dependency-edge parallel release + exactly-once executor callbacks, typed mailbox, team sessions, backend fallback |
cd .claude/scripts
uv run pytest tests/ -v # full active suite
uv run ruff check . # Lint
uv run ruff format . # Format
uv run thehomie --help # Verify CLI
See CONTRIBUTING.md for the full contributor guide.
cp .claude/scripts/.env.example .claude/scripts/.env
docker compose config
docker compose up # bot + scheduler (heartbeat · reflection · weekly synthesis)
MIT. Built by the TaskChad OS contributors.