TypeScript DSH plugin (dsh-plugin) that gives DeepSeek Harness a voice: mic-in speech-to-text, cloned-voice TTS read-aloud, interrupt / queue modes, a streaming DUIX digital-human window, and QQ push — all running in the DSH Web GUI with a local voice_bridge.
DSH integration
Native runtime
Author-claimed
Safety audit
Unaudited
Last verified
2026-09-02
License
NOASSERTION
01What can it help you accomplish?
Talk to DeepSeek Harness as a speaking voice companion through the mic
Sentence-by-sentence TTS read-aloud of each reply, plus auto-pushed text / voice / image replies to your QQ when you step away from the desk
DSH users who want a chatty, voice-first AI girlfriend experience inside the DeepSeek Harness desktop GUI
Turn each DeepSeek Harness reply into a streaming talking-head digital-human video
Real-time DUIX lip-synced talking-head clips (≤10s segments, TTS + video synthesized in parallel), togglable via the digital-human switch, up to 200 clips retained
Creators and DSH users who want replies rendered as a live, animated digital-human window instead of plain text
02How to install into DeepSeek Harness
Not specified by the author
03DSH integration and capability boundaries
Native DSH plugin (dsh-plugin) that runs inside the DSH Web GUI (:3080) with a local voice_bridge (:8765) handling STT / TTS / digital-human / QQ — built for DeepSeek Harness
Mic speech-to-text (FunASR)
microphone audio from the DSH Web GUI→Chinese ASR text fed to the agent reply
Voice cloning & TTS (OmniVoice)
reply text plus optional reference audio or voice parameters→cloned-voice speech, voice designable across 600+ languages
Streaming digital human (DUIX)
reply text→lip-synced talking-head video segments, idle/active states in the girlfriend window
QQ push bridge (OneBot)
reply text / voice / image→messages sent to your QQ account
sends text / voice / image messages to the configured QQ account over OneBot HTTP+WS
04Who is it for? When not to use it?
Good for
- DSH users who want a chatty, voice-first AI girlfriend experience inside the DeepSeek Harness desktop GUI
- Creators and DSH users who want replies rendered as a live, animated digital-human window instead of plain text
Not for
- The voice companion opens local ports — the DSH Web GUI on :3080 and the voice_bridge on :8765 — and bridges replies out to QQ over the network, so reply text / audio is pushed to an external QQ account and the local ports must be reachable.
05Compatibility, maintenance and safety notes
- The QQ push feature pushes text / voice / image replies to a QQ account through the plugin's QQ bridge over OneBot HTTP+WS, so it needs a QQ account and a configured OneBot bridge; without them the QQ push stays inactive.
- TTS voice cloning (OmniVoice) is described as running with WSL2 + FlashInfer acceleration on the voice_bridge, so a Windows / WSL2 host with accelerator setup affects voice-clone quality and latency; non-WSL2 hosts may not get that acceleration path.
- The voice companion opens local ports — the DSH Web GUI on :3080 and the voice_bridge on :8765 — and bridges replies out to QQ over the network, so reply text / audio is pushed to an external QQ account and the local ports must be reachable.
TypeScript · license not declared (NOASSERTION) · 60 stars · last push 2026-08-30
06Frequently asked questions
What is dsh-voice-ai-girlfriend and how does it connect to DeepSeek Harness?
It is a DSH plugin (dsh-plugin) titled "DSH 语音 AI 女友" that runs inside the DSH Web GUI (:3080). A local voice_bridge (:8765) exposes /api/stt, /api/tts, /api/dh/* and /api/qq/* for speech-to-text, TTS, the digital human, and the QQ bridge.
How do the spoken replies work?
Click the mic and it listens; the agent replies and TTS reads the answer aloud sentence by sentence. The README says she opens her mouth within 0.5s — FunASR Chinese ASR in ~150ms and near-instant TTS.
Can I interrupt her while she is talking?
Yes. As the README puts it, chip in and she stops to listen; toggle the switch to return to queue mode if you want her to finish speaking.
What is the streaming digital human (DUIX)?
Replies are turned into real-time talking-head clips: long replies split into ≤10s segments, video and the next segment's TTS synthesized in parallel and lip-synced, played back continuously. There is a digital-human switch and up to 200 clips are kept.
How does the QQ push work and what does it need?
Replies auto-push to your QQ (text + voice + image) through the plugin's QQ bridge using OneBot HTTP+WS. It needs a QQ account and a configured OneBot bridge; without them the push stays inactive.
07Related DSH workflows
vox-director
by alisa0808
Turn one topic into a finished Vox-style paper-collage explainer/ad video — automated end to end on Atlas Cloud + ffmpeg. An agent skill.
dsh-voice-scribe
by pensivefei
DSH voice input plugin: tap Alt to talk, get text in composer. Web Speech default (zero config), optional OpenAI-compatible ASR, polish via DSH LLM.
dsh-plugin-tts
by 1624318455
Edge TTS voice plugin for DeepSeek Harness: read assistant replies aloud, auto-read toggle, voice settings panel (free, no API key)
dsh-voice
by goodandready
Voice input for DeepSeek Harness: dictation and voice messages with multi-provider STT fallback chains
08Data and sources
# DSH 语音 AI 女友(Voice AI Girlfriend)
浏览器(DSH Web GUI :3080)
This page is generated from the project’s public documentation, repository metadata and a structured parse of DSH Plugins; last verified on 2026-09-02. Found an error? Submit a correction.
Best DeepSeek Harness Plugins
Twelve plugins worth installing first — picked from the whole catalog, across every category.
