dsh-plugin-speech
メンテナンス: 活発nakamuraia/dsh-plugin-speech
DeepSeek Harness のアシスタント返信を読み上げる音声合成プラグイン。音声合成プロバイダーとストリーミング再生を利用できます。
$ dsh plugin add dsh-plugin-speech4
スター
0
フォーク
TypeScript
言語
MIT
ライセンス
2026-09-13
作成日
2026-09-14
最終プッシュ
README
dsh-plugin-speech
English | 中文

Read assistant replies aloud in DeepSeek Harness, with the audio streamed while it is still being generated.
What it does
Each finished assistant message gets a play control in its action row. Press it and the reply is spoken through a text-to-speech service; press it again and it stops. A settings row in General owns the service, the voice, and the tuning.
The reply is sent in one request and the audio is played as it arrives, so speech starts on the first frames instead of waiting for the whole file. The equivalent of watching an answer stream in, for audio.
Providers
| Service | Key needed | Voices |
|---|---|---|
| Microsoft Edge (default) | no | the service publishes its full catalog, every language it serves |
Adding a provider means adding one folder under src/providers/ and one line in src/providers/registry.ts — the settings screen builds its fields from what the provider declares.
Requirements
- A DeepSeek Harness Web surface (
dsh web), version0.1.5-rc.2or newer. - Network access to the speech service. Nothing else: there is no model to download, and the Edge provider needs no account.
Install
-
Add the package to your harness checkout:
pnpm add -w @nakamuraia/dsh-plugin-speechInstalling from GitHub works the same way, if you would rather track the source:
pnpm add -w github:NakamuraIA/dsh-plugin-speech -
Copy
speech.patch.ymlnext to the checkout and start the Web surface with it:dsh web --patch ./speech.patch.yml
The patch adds one Loader row. The row name is the package name and also the browser module id the built bundle registers under, so the two stay in step — do not rename one without the other.
Settings
Settings opens on General, where a Read aloud row owns:
- Service — which provider speaks.
- Language and Voice — the catalog entry to use.
- Speed, Volume, Pitch — sent to the service with every request.
- Skip code — leave fenced code and inline code out of the reading.
- Test voice — speaks a sample through the settings on screen, saved or not.
Code blocks, tables, and pasted spreadsheets are never read; links keep their label and lose their target; emoji, quotes, and standalone symbols are dropped. Sentence punctuation is kept, because the voice uses it for pauses.
How it works
browser host service
─────── ──── ───────
click ─ ▶ POST /speech/speak { text }
select provider from
the durable settings
────────────────────────── ▶ synthesize
◀ ─ audio bytes stream back ─ ─ ─ ─ ─ ─ ─ ─ ─ ┘
play as they arrive
Synthesis lives on the host because a browser cannot hold an API key and cannot reach most services across CORS. The browser only posts text and plays bytes.
Development
The source is developed inside a DeepSeek Harness checkout, where the type packages and the client build preset live; this repository carries the source and the built artifacts. To rebuild, copy src/ into packages/client/ui-speech/ of a checkout, rename the package there, and run its build.
License
MIT. The Edge provider drives Microsoft's read-aloud endpoint through msedge-tts (MIT); that endpoint is not a documented, supported Microsoft API, so treat availability as best effort.
音声 の他のプラグイン
vox-director
by alisa0808
1 トピックから Vox 風ペーパーコラージュ解説・広告動画を Atlas Cloud + ffmpeg で E2E 自動生成するエージェントスキル
dsh-voice-ai-girlfriend
by beiyege-01
Whisper 音声入力と Qwen3-TTS 音声合成による会話型コンパニオンプラグイン。文単位のストリーミング読み上げ、割り込み/キュー両モード対応
dsh-voice-scribe
by pensivefei
DSH音声入力プラグイン。Altを押して話すと入力欄に文字起こし。Web Speech標準(設定不要)・任意でOpenAI互換ASR・DSH LLMで推敲。
dsh-plugin-tts
by 1624318455
Edge TTS 音声プラグイン:返信読み上げ、自動読取トグル、無料・キー不要
