ClaudeModsClaude Code mod ディレクトリ
☰
● 0 人がオンライン ・閲覧 0 回
+ 作品を投稿
← 作品一覧へ
Reddit の投稿 · その他

live-vibe:必要な機能を備えた Claude Code 用の全二重音声 Mod。Claude Code Mods を基盤に CUDA と Apple silicon に対応、2 行でインストール、MIT ライセンス。 — Claude Code プラグイン · ClaudeMods

この作品を報告あなたの作品ですか?所有者確認へ

ClaudeMods でこの Claude Code プラグインのソースと必要な権限を確認できます。 live-vibe を作ったのは、ChatGPT Codex の音声モードのように Claude Code と話したかったからです。マイクをずっと開いたままにして、考えを言い終えたら返答し、私が割り込んだら黙ってほしい。Claude Code にはその機能がありませんでした。でも、新しい Mods 機能はエージェントの実行経路でプラグインのコードを動かし、ターミナルに描画できます。それだけで全体をプラグインとして作れました。MIT ライセンスで、ottomation と名付けたマーケットプレイスにあり、2 行でインストールできます。 /live は Claude との全二重音声会話です。割り込むと最初の一言で止まります。ただし話している間の "mm-hm" は割り込みと見なさず、一時停止した音声サンプルから続けます。人と話すように相づちを打てます。/vibe はディレクターモードで、Claude は読み取りと作業用サブエージェントへの指示を行い、自分ではファイルを一切編集しません。/livevibe は両方を同時に使います。小型で高速な音声モデルが私と会話し、実作業はディレクターの Claude に渡します。Claude が終わると音声が結果を一、二文で知らせます。その間にも Claude が何をしているか尋ねられ、Claude は作業を続けたまま返答が得られます。 必要なものはすべて

PPale-Soil-2524@Pale-Soil-2524
翻訳済み

この mod について

live-vibe を作ったのは、ChatGPT Codex の音声モードのように Claude Code と話したかったからです。マイクをずっと開いたままにして、考えを言い終えたら返答し、私が割り込んだら黙ってほしい。Claude Code にはその機能がありませんでした。でも、新しい Mods 機能はエージェントの実行経路でプラグインのコードを動かし、ターミナルに描画できます。それだけで全体をプラグインとして作れました。MIT ライセンスで、ottomation と名付けたマーケットプレイスにあり、2 行でインストールできます。

/live は Claude との全二重音声会話です。割り込むと最初の一言で止まります。ただし話している間の "mm-hm" は割り込みと見なさず、一時停止した音声サンプルから続けます。人と話すように相づちを打てます。/vibe はディレクターモードで、Claude は読み取りと作業用サブエージェントへの指示を行い、自分ではファイルを一切編集しません。/livevibe は両方を同時に使います。小型で高速な音声モデルが私と会話し、実作業はディレクターの Claude に渡します。Claude が終わると音声が結果を一、二文で知らせます。その間にも Claude が何をしているか尋ねられ、Claude は作業を続けたまま返答が得られます。

必要なものはすべてプラグインに含まれます。Kyutai のストリーミング STT、Kokoro TTS、WebRTC のエコーキャンセル、音声フロントエンドは /live setup がダウンロードして確認します。各マシンで一度だけです。CUDA に対応し、Apple silicon も MLX 経由で使えます。会話層は軽量のローカルモデルで、動かしたくなければ Anthropic API にフォールバックします。WSL2 ではネイティブの Windows プレーヤーで音声を再生します。WSLg の RDP 音声はノイズが出て、一日中それを聞きたくなかったからです。

一番時間をかけたのは発話交代です。考えながら話せるかどうかが決まるからです。プラグインは Kyutai のポーズ予測ヘッドを読み、まだ話し続けそうかを予測します。ほとんどの発話は最後の言葉から約半秒で閉じますが、末尾の "and" やカンマがあれば言い終える時間を確保します。0.7.0 には、OpenAI の GPT-Live の解説を読んで作った実験的な経路も含まれます。モデルの確信度で待ち時間を調整し、話し終える少し前から返答を下書きし、有効にすれば相づちも打ちます。デフォルトで有効ですが、気になる場合は一つの設定で無効にできます。

Claude Code 2.1.287 以降、uv、マイク、スピーカーが必要です。/plugin marketplace add potto007/ottomation、次に /plugin install live-vibe@ottomation、最後に /live setup を実行してください。Repo: https://github.com/potto007/ottomation

試したら、あなたのマシンでの発話交代の感触を教えてください。投稿者 /u/Pale-Soil-2524 [link] [comments]

インストール

インストール方法は元のソースをご確認ください。

原文 / README

I built live-vibe because I want to talk to Claude Code the way I can talk to ChatGPT Codex in Voice Mode. I want the mic open the whole time, and I want it to answer when I have finished the thought and shut up when I talk over it. Claude Code did not have that. It does have the new Mods capability, which lets a plugin run code in the agent's path and draw into the terminal, and that turned out to be enough to build the whole thing as a plugin. It is MIT licensed, in a marketplace I'm calling ottomation, and it installs in two lines. /live is full-duplex voice with Claude. Talk over it and it stops on your first word. An "mm-hm" while it is speaking does not count, and it carries on from the sample it paused on, so you can grunt along the way you would with a person. /vibe is director mode, where Claude reads and directs worker subagents and never edits a file itself. /livevibe is both at once. A small, fast voice model holds the conversation with me and hands the real work to Claude as director. When Claude finishes, the voice tells me what happened in a sentence or two, and in the meantime I can ask it what Claude is up to and get an answer while Claude keeps working. Everything you need comes with the plugin. Kyutai streaming STT, Kokoro TTS, WebRTC echo cancellation and the voice front are all pulled down and checked by /live setup, once per machine. CUDA works and so does Apple silicon through MLX. The conversation layer is a lightweight local model, and if you would rather not run one it falls back to the Anthropic API. If you are on WSL2, speech plays through a native Windows player, because WSLg's RDP audio crackles and I was not willing to listen to that all day. Turn-taking is where most of my time went, because it decides whether you can think out loud. The plugin reads Kyutai's pause-forecast heads, which predict whether you are about to keep talking. Most turns close about half a second after your last word, and a trailing "and" or a comma buys you time to finish. 0.7.0 also ships an experimental path I put together after reading OpenAI's write-up on GPT-Live. It scales the wait on the model's confidence, starts drafting a reply before you have quite finished, and can backchannel if you turn that on. It is on by default and one setting turns it off if it annoys you. You need Claude Code 2.1.287 or newer, uv, a mic and a speaker. Install with /plugin marketplace add potto007/ottomation , then /plugin install live-vibe@ottomation , then /live setup . Repo: https://github.com/potto007/ottomation If you try it, tell me how the turn-taking feels on your machine. submitted by /u/Pale-Soil-2524 [link] [comments]

元のソースを見る ↗

関連作品