ClaudeModsClaude Code 模組目錄
☰
● 0 人在線上 · 瀏覽 0 次
+ 提交作品
← 返回作品集
Reddit 貼文 · 其他

live-vibe:功能齊備的 Claude Code 全雙工語音 Mod。基於 Claude Code Mods,支援 CUDA 與 Apple silicon,兩行指令安裝,採用 MIT 授權。 — Claude Code 插件/外掛 · ClaudeMods

回報這個作品這是你的作品?認領

在 ClaudeMods 了解這款 Claude Code 插件/外掛,查看原始來源與所需權限。 我打造 live-vibe,是因為我想像使用 ChatGPT Codex 的語音模式一樣,與 Claude Code 對話。我希望麥克風一直開著,在我說完一個想法時回答,並在我插話時停止說話。Claude Code 原本沒有這個功能,但新的 Mods 能力允許外掛在代理的執行流程中執行程式碼、在終端機內繪圖,足以把整套功能做成外掛。它採用 MIT 授權,放在我稱為 ottomation 的市集中,兩行指令就能安裝。 /live 提供與 Claude 的全雙工語音對話。你插話時,它在你說出第一個字就會停下來。但在它說話時回應一聲 "mm-hm" 不算打斷;它會從暫停的音訊取樣位置繼續播放,所以你可以像跟人聊天一樣,不時發出回應聲。/vibe 是導演模式,Claude 負責閱讀並指揮工作子代理,本身絕不編輯檔案。/livevibe 同時啟用兩者。一個小巧快速的語音模型負責與我交談,把真正的工作交給擔任導演的 Claude。Claude 完成後,語音會用一兩句話告訴我結果;期間我也能詢問 Claude 在做什麼,Claude 繼續工作時仍能得到回答。 所需的一切

PPale-Soil-2524@Pale-Soil-2524
已翻譯

關於這個 mod

我打造 live-vibe,是因為我想像使用 ChatGPT Codex 的語音模式一樣,與 Claude Code 對話。我希望麥克風一直開著,在我說完一個想法時回答,並在我插話時停止說話。Claude Code 原本沒有這個功能,但新的 Mods 能力允許外掛在代理的執行流程中執行程式碼、在終端機內繪圖,足以把整套功能做成外掛。它採用 MIT 授權,放在我稱為 ottomation 的市集中,兩行指令就能安裝。

/live 提供與 Claude 的全雙工語音對話。你插話時,它在你說出第一個字就會停下來。但在它說話時回應一聲 "mm-hm" 不算打斷;它會從暫停的音訊取樣位置繼續播放,所以你可以像跟人聊天一樣,不時發出回應聲。/vibe 是導演模式,Claude 負責閱讀並指揮工作子代理,本身絕不編輯檔案。/livevibe 同時啟用兩者。一個小巧快速的語音模型負責與我交談,把真正的工作交給擔任導演的 Claude。Claude 完成後,語音會用一兩句話告訴我結果;期間我也能詢問 Claude 在做什麼,Claude 繼續工作時仍能得到回答。

所需元件都隨外掛提供。/live setup 會下載並檢查 Kyutai 串流 STT、Kokoro TTS、WebRTC 回音消除與語音前端,每台機器只需一次。支援 CUDA,也透過 MLX 支援 Apple silicon。對話層使用輕量的本機模型;若不想執行本機模型,則改用 Anthropic API。在 WSL2 上,語音透過原生 Windows 播放器播放,因為 WSLg 的 RDP 音訊會有爆音,我不想整天聽那個聲音。

我花最多時間處理的是輪流發言的判斷,因為這決定你能否邊想邊說。外掛讀取 Kyutai 的暫停預測頭,預測你是否即將繼續說話。多數發言會在最後一個字後約半秒結束;結尾的 "and" 或逗號會多留一些時間,讓你把話說完。0.7.0 也包含一條實驗路徑,是我讀了 OpenAI 關於 GPT-Live 的文章後做的。它依模型信心程度調整等待時間,在你還沒完全說完前開始草擬回答,並可在你啟用時提供簡短的聆聽回應。這條路徑預設開啟,若覺得煩,一個設定就能關閉。

你需要 Claude Code 2.1.287 或更新版本、uv、麥克風與喇叭。先執行 /plugin marketplace add potto007/ottomation,再執行 /plugin install live-vibe@ottomation,最後執行 /live setup。儲存庫:https://github.com/potto007/ottomation

如果你試用了,請告訴我你的機器上輪流發言的感受。由 /u/Pale-Soil-2524 投稿 [link] [comments]

安裝

安裝方式請查看原始來源。

原文 / README

I built live-vibe because I want to talk to Claude Code the way I can talk to ChatGPT Codex in Voice Mode. I want the mic open the whole time, and I want it to answer when I have finished the thought and shut up when I talk over it. Claude Code did not have that. It does have the new Mods capability, which lets a plugin run code in the agent's path and draw into the terminal, and that turned out to be enough to build the whole thing as a plugin. It is MIT licensed, in a marketplace I'm calling ottomation, and it installs in two lines. /live is full-duplex voice with Claude. Talk over it and it stops on your first word. An "mm-hm" while it is speaking does not count, and it carries on from the sample it paused on, so you can grunt along the way you would with a person. /vibe is director mode, where Claude reads and directs worker subagents and never edits a file itself. /livevibe is both at once. A small, fast voice model holds the conversation with me and hands the real work to Claude as director. When Claude finishes, the voice tells me what happened in a sentence or two, and in the meantime I can ask it what Claude is up to and get an answer while Claude keeps working. Everything you need comes with the plugin. Kyutai streaming STT, Kokoro TTS, WebRTC echo cancellation and the voice front are all pulled down and checked by /live setup, once per machine. CUDA works and so does Apple silicon through MLX. The conversation layer is a lightweight local model, and if you would rather not run one it falls back to the Anthropic API. If you are on WSL2, speech plays through a native Windows player, because WSLg's RDP audio crackles and I was not willing to listen to that all day. Turn-taking is where most of my time went, because it decides whether you can think out loud. The plugin reads Kyutai's pause-forecast heads, which predict whether you are about to keep talking. Most turns close about half a second after your last word, and a trailing "and" or a comma buys you time to finish. 0.7.0 also ships an experimental path I put together after reading OpenAI's write-up on GPT-Live. It scales the wait on the model's confidence, starts drafting a reply before you have quite finished, and can backchannel if you turn that on. It is on by default and one setting turns it off if it annoys you. You need Claude Code 2.1.287 or newer, uv, a mic and a speaker. Install with /plugin marketplace add potto007/ottomation , then /plugin install live-vibe@ottomation , then /live setup . Repo: https://github.com/potto007/ottomation If you try it, tell me how the turn-taking feels on your machine. submitted by /u/Pale-Soil-2524 [link] [comments]

查看原始來源 ↗

更多類似作品