Native Rapid-MLX provider for DeepSeek Harness: dsh auto-discovers served models and manages them in-session, with compaction timed to the Mac's unified-memory ceiling.
DSH integration
Native runtime
Author-claimed
Safety audit
Unaudited
Last verified
2026-08-30
License
Apache-2.0
01What can it help you accomplish?
Run local Rapid-MLX models inside DeepSeek Harness without hand-maintaining per-model facts
`dsh` auto-discovers served models from `/v1/models` — context window, reasoning/tool parsers, MoE/hybrid, modalities — with no hand-written settings.yaml
DeepSeek Harness users on Apple Silicon running Rapid-MLX who want accurate, server-driven model facts
See, pull, remove and health-check Rapid-MLX models without leaving the dsh session
Five agent tools — `rapid_mlx_serving`, `rapid_mlx_cached`, `rapid_mlx_pull`, `rapid_mlx_remove`, `rapid_mlx_health` — plus a `/rapid-mlx` command
Agents and developers who want to manage served/downloadable models and check health in-session
Get truthful reasoning controls and machine-fitted context compaction
Reasoning selector only appears for models that actually have a reasoning parser; compaction timed to the server's `max_model_len` (unified-memory ceiling), not a hand-written number
Developers hit by silent failures from stale context windows or dead reasoning selectors when switching Rapid-MLX models
02How to install into DeepSeek Harness
Prerequisites
- Node ≥ 22.15 (dsh imports Node's Zstd stream API without declaring it)
- a running Rapid-MLX server
Installation steps
- 01
$ dsh plugin --profile web add @raullenchai/dsh-provider
- 02
$ dsh plugin --profile web add github:raullenchai/rapid-mlx-dsh-provider
- 03
export RAPID_MLX_BASE_URL=http://localhost:8000/v1 # optional; this is the default
- 04
$ dsh web
- 05
# $DSH_HOME/settings.yaml agent-default-model: provider: rapid-mlx model: qwen3.6-35b-8bit
Verify the integration
- Installs and activates as a profile layer; the entry shows up in `dsh --profile headless --dump-config` with no "declares no dsh.bundle" warning.
- Registers the `rapid-mlx` route with `ctx.llm` and serves real queries.
03DSH integration and capability boundaries
Native dsh LLM adapter: installed via `dsh plugin --profile web add`, registers the `rapid-mlx` route and reads model facts from the Rapid-MLX server's `/v1/models`.
Server-driven model discovery
Rapid-MLX `/v1/models` HTTP endpoint→served models with context window, reasoning/tool parsers, MoE/hybrid, modalities — deduped
In-session model management tools
the active dsh agent session→five tools (`rapid_mlx_serving`, `rapid_mlx_cached`, `rapid_mlx_pull`, `rapid_mlx_remove`, `rapid_mlx_health`) plus a `/rapid-mlx` command
`rapid_mlx_pull` and `rapid_mlx_remove` change on-disk cached models (non-interactive, forced `-y`)Memory-fitted context compaction
dsh-compaction-basic compaction request→compacts at `thresholdRatio × capacity` (0.8 default) using server `max_model_len` when available, else `context_window`
Conformant LLM adapter (cookbook contract)
dsh LLM calls — chat, tool calls, streaming, abort→streaming responses meeting the official adapter protocol obligations
04Who is it for? When not to use it?
Good for
- DeepSeek Harness users on Apple Silicon running Rapid-MLX who want accurate, server-driven model facts
- Agents and developers who want to manage served/downloadable models and check health in-session
- Developers hit by silent failures from stale context windows or dead reasoning selectors when switching Rapid-MLX models
Not for
- The route is registered as `rapid-mlx`. If your settings.yaml also declares a `rapid-mlx` provider under `llm-pi-ai`, the two compete for one route name (`registerAdapter` owns provider exclusivity). Use one or rename ours.
- Several server-reported facts are read but not yet acted on — `recommended_sampling`, `tool_call_parser` (no fast-fail on models that can't emit tool_calls), and `is_hybrid`/`is_moe`/`capabilities`; images are refused with `UNSUPPORTED` rather than carried.
05Compatibility, maintenance and safety notes
- The route is registered as `rapid-mlx`. If your settings.yaml also declares a `rapid-mlx` provider under `llm-pi-ai`, the two compete for one route name (`registerAdapter` owns provider exclusivity). Use one or rename ours.
- Several server-reported facts are read but not yet acted on — `recommended_sampling`, `tool_call_parser` (no fast-fail on models that can't emit tool_calls), and `is_hybrid`/`is_moe`/`capabilities`; images are refused with `UNSUPPORTED` rather than carried.
- DSH is still a developer preview that moves fast; dsh 0.1.0-rc.8 is API-compatible with rc.7 but the author treats compatibility as tracking a moving target, not a frozen promise.
Apache-2.0 · published to npm as @raullenchai/dsh-provider (latest release v0.2.0, 2026-08-19)
06Frequently asked questions
How do I install dsh-provider for DeepSeek Harness?
Run `dsh plugin --profile web add @raullenchai/dsh-provider` (or `dsh plugin --profile web add github:raullenchai/rapid-mlx-dsh-provider` for source). You need Node ≥ 22.15 and a running Rapid-MLX server, then `dsh web`.
What are the prerequisites?
Node ≥ 22.15 (dsh imports Node's Zstd stream API without declaring it) and a running Rapid-MLX server. Set `RAPID_MLX_BASE_URL` only if your server isn't at the default http://localhost:8000/v1.
How does it connect to DeepSeek Harness?
It registers a native LLM adapter as the `rapid-mlx` route; dsh talks to the local Rapid-MLX server's OpenAI-compatible `/v1` endpoint and reads model facts from `/v1/models`. Point `agent-default-model` in settings.yaml at `provider: rapid-mlx`.
How is this different from dsh's generic openai-completions route?
The generic route makes you hand-maintain per-model facts in settings.yaml and knows nothing beyond what you typed. This adapter reads `/v1/models`, so switching models needs no re-setup, and the reasoning selector only appears for models that actually have a reasoning parser.
Troubleshooting: my settings.yaml already has a rapid-mlx provider — what happens?
Both register the `rapid-mlx` route and compete for one route name (`registerAdapter` owns provider exclusivity). Use only one, or rename this plugin's route. Note dsh 0.1.0-rc.8 is API-compatible with rc.7 but DSH is a fast-moving developer preview.
07Related DSH workflows
dsh-routing-suite
by yjh051108
dsh-routing-suite — injector + router-standard kit: install the runtime injector first, then the task-aware reasoning-mode router preset (measured P1-P23).
brooks-lint
by hyhmrright
AI code reviews grounded in 12 classic engineering books — decay risk diagnostics with book citations, severity labels, and 6 analysis modes including full-sweep auto-fix
dsh-plugin-shop
by livxue
The most comprehensive DeepSeek Harness plugin market — refreshed daily, sourced across the Internet, reviewed before publishing.
dsh-our-free-model
by zouyuxuan122
Adds selectable remote models, availability checks, a local token dashboard, and an optional authenticated local forwarding endpoint.
08Data and sources
A native [Rapid-MLX](https://github.com/raullenchai/Rapid-MLX) provider for
Registers the `rapid-mlx` route with `ctx.llm` and serves real queries.
This page is generated from the project’s public documentation, repository metadata and a structured parse of DSH Plugins; last verified on 2026-08-30. Found an error? Submit a correction.
Best DeepSeek Harness Plugins
Twelve plugins worth installing first — picked from the whole catalog, across every category.
