ClaudeModsClaude Code 模组目录
☰
● 0 人在线 · 浏览 0 次
+ 提交作品
← 返回作品集
Reddit 帖子 · 其他

live-vibe:功能齐全的 Claude Code 全双工语音 Mod。基于 Claude Code Mods,支持 CUDA 和 Apple silicon,两行命令安装,采用 MIT 许可证。 — Claude Code 插件 · ClaudeMods

回报这个作品这是你的作品?认领

在 ClaudeMods 了解这款 Claude Code 插件,查看原始来源与所需权限。 我开发 live-vibe,是因为我想像使用 ChatGPT Codex 的语音模式一样,与 Claude Code 对话。我希望麦克风一直开启,在我说完一个想法时回答,并在我插话时停止说话。Claude Code 原本没有这个功能,但新的 Mods 能力允许插件在智能体的执行流程中运行代码、在终端内绘图,足以把整套功能做成插件。它采用 MIT 许可证,放在我称为 ottomation 的市场中,两行命令就能安装。 /live 提供与 Claude 的全双工语音对话。你插话时,它在你说出第一个字就会停下来。但在它说话时回应一声 "mm-hm" 不算打断;它会从暂停的音频采样位置继续播放,所以你可以像跟人聊天一样,不时发出回应声。/vibe 是导演模式,Claude 负责阅读并指挥工作子智能体,本身绝不编辑文件。/livevibe 同时启用两者。一个小巧快速的语音模型负责与我交谈,把真正的工作交给担任导演的 Claude。Claude 完成后,语音会用一两句话告诉我结果;期间我也能询问 Claude 在做什么,Claude 继续工作时仍能得到回答。 所需的一切

PPale-Soil-2524@Pale-Soil-2524
已翻译

关于这个 mod

我开发 live-vibe,是因为我想像使用 ChatGPT Codex 的语音模式一样,与 Claude Code 对话。我希望麦克风一直开启,在我说完一个想法时回答,并在我插话时停止说话。Claude Code 原本没有这个功能,但新的 Mods 能力允许插件在智能体的执行流程中运行代码、在终端内绘图,足以把整套功能做成插件。它采用 MIT 许可证,放在我称为 ottomation 的市场中,两行命令就能安装。

/live 提供与 Claude 的全双工语音对话。你插话时,它在你说出第一个字就会停下来。但在它说话时回应一声 "mm-hm" 不算打断;它会从暂停的音频采样位置继续播放,所以你可以像跟人聊天一样,不时发出回应声。/vibe 是导演模式,Claude 负责阅读并指挥工作子智能体,本身绝不编辑文件。/livevibe 同时启用两者。一个小巧快速的语音模型负责与我交谈,把真正的工作交给担任导演的 Claude。Claude 完成后,语音会用一两句话告诉我结果;期间我也能询问 Claude 在做什么,Claude 继续工作时仍能得到回答。

所需组件都随插件提供。/live setup 会下载并检查 Kyutai 流式 STT、Kokoro TTS、WebRTC 回声消除和语音前端,每台机器只需一次。支持 CUDA,也通过 MLX 支持 Apple silicon。对话层使用轻量的本地模型;若不想运行本地模型,则回退到 Anthropic API。在 WSL2 上,语音通过原生 Windows 播放器播放,因为 WSLg 的 RDP 音频会有爆音,我不想整天听那个声音。

我花最多时间处理的是轮流发言的判断,因为这决定你能否边想边说。插件读取 Kyutai 的暂停预测头,预测你是否即将继续说话。多数发言会在最后一个字后约半秒结束;结尾的 "and" 或逗号会多留一些时间,让你把话说完。0.7.0 也包含一条实验路径,是我读了 OpenAI 关于 GPT-Live 的文章后做的。它按模型置信度调整等待时间,在你还没完全说完前开始起草回答,并可在你启用时提供简短的倾听回应。这条路径默认开启,若觉得烦,一个设置就能关闭。

你需要 Claude Code 2.1.287 或更新版本、uv、麦克风和扬声器。先运行 /plugin marketplace add potto007/ottomation,再运行 /plugin install live-vibe@ottomation,最后运行 /live setup。仓库:https://github.com/potto007/ottomation

如果你试用了,请告诉我你的机器上轮流发言的感受。由 /u/Pale-Soil-2524 投稿 [link] [comments]

安装

安装方法请查看原始来源。

原文 / README

I built live-vibe because I want to talk to Claude Code the way I can talk to ChatGPT Codex in Voice Mode. I want the mic open the whole time, and I want it to answer when I have finished the thought and shut up when I talk over it. Claude Code did not have that. It does have the new Mods capability, which lets a plugin run code in the agent's path and draw into the terminal, and that turned out to be enough to build the whole thing as a plugin. It is MIT licensed, in a marketplace I'm calling ottomation, and it installs in two lines. /live is full-duplex voice with Claude. Talk over it and it stops on your first word. An "mm-hm" while it is speaking does not count, and it carries on from the sample it paused on, so you can grunt along the way you would with a person. /vibe is director mode, where Claude reads and directs worker subagents and never edits a file itself. /livevibe is both at once. A small, fast voice model holds the conversation with me and hands the real work to Claude as director. When Claude finishes, the voice tells me what happened in a sentence or two, and in the meantime I can ask it what Claude is up to and get an answer while Claude keeps working. Everything you need comes with the plugin. Kyutai streaming STT, Kokoro TTS, WebRTC echo cancellation and the voice front are all pulled down and checked by /live setup, once per machine. CUDA works and so does Apple silicon through MLX. The conversation layer is a lightweight local model, and if you would rather not run one it falls back to the Anthropic API. If you are on WSL2, speech plays through a native Windows player, because WSLg's RDP audio crackles and I was not willing to listen to that all day. Turn-taking is where most of my time went, because it decides whether you can think out loud. The plugin reads Kyutai's pause-forecast heads, which predict whether you are about to keep talking. Most turns close about half a second after your last word, and a trailing "and" or a comma buys you time to finish. 0.7.0 also ships an experimental path I put together after reading OpenAI's write-up on GPT-Live. It scales the wait on the model's confidence, starts drafting a reply before you have quite finished, and can backchannel if you turn that on. It is on by default and one setting turns it off if it annoys you. You need Claude Code 2.1.287 or newer, uv, a mic and a speaker. Install with /plugin marketplace add potto007/ottomation , then /plugin install live-vibe@ottomation , then /live setup . Repo: https://github.com/potto007/ottomation If you try it, tell me how the turn-taking feels on your machine. submitted by /u/Pale-Soil-2524 [link] [comments]

查看原始来源 ↗

更多类似作品