dsh-clawrouter
キュレーション掲載メンテナンス: 活発blockrunai/dsh-clawrouter
危険なツール呼び出しを実行前により強いモデルがレビューするセーフティゲート。67 モデルを x402 の従量課金で利用可能
$ dsh plugin add dsh-clawrouter21
スター
3
フォーク
TypeScript
言語
MIT
ライセンス
2026-08-14
作成日
2026-09-05
最終プッシュ
README
A second brain for your DeepSeek Harness agent
DeepSeek is fast and cheap — keep it for the loop.
This adds what it cannot do: a stronger model reviews the dangerous command before it runs.
75 models from one credential — an API key from user.blockrun.ai, or a Solana or Base wallet if you would rather not sign up.
English | 中文
dsh-clawrouter is a DeepSeek Harness plugin that puts a stronger model in front of your agent's dangerous actions. When the agent proposes
rm -rf ~, a reviewer model reads it and answers allow / deny / ask — enforced by the real tool executor, not by a prompt. It also registers a BlockRun provider route, so the reviewer (and any of 75 models) is reachable from one credential: a BlockRun API key from user.blockrun.ai, billed at exact token usage — or, if you would rather not sign up for anything, a Solana or Base wallet paying per request in USDC over x402. MIT licensed.
dsh plugin --profile web add dsh-clawrouter
Why this exists
Two things people keep asking for in the Harness discussions:
「是否有类似 Codex 或者 CC 的审查模式?即额外调用模型审查指令,以解放双手?Full Access 还是太让人担心了。」 — #421 Is there a review mode like Codex or Claude Code — call an extra model to review the command, to free up my hands? Full Access is too worrying.
「使用 Full Access 模式创建并测试插件时误删了我的整个家目录」 — #461 Testing a plugin in Full Access mode, it deleted my entire home directory.
Full Access is all-or-nothing: approve every command by hand, or approve nothing and hope. This adds a third option.
How it compares
| Approve everything | Full Access | Permission rules | dsh-clawrouter | |
|---|---|---|---|---|
| Hands-free | No | Yes | Yes | Yes |
Catches rm -rf ~ |
Only if you notice | No | Only if you wrote the rule | Yes |
| Understands intent | You do | Nothing does | No — literal match | Yes, a model reads it |
| Enforced where | UI prompt | — | Executor | Executor |
| Fails | — | open | closed | to a human, never open |
| Reviews ordinary work | Everything | Nothing | Nothing | Nothing |
What it does
1. Review gate
When the agent proposes something destructive, a strong model (default anthropic/claude-opus-5) reads it and answers:
| Verdict | What happens |
|---|---|
| safe | proceeds to the normal permission chain, untouched |
| dangerous | denied, with a reason the agent can act on |
| uncertain | escalated to you — the normal approval prompt |
It only ever narrows. A call the reviewer clears still faces every sandbox, permission, and approval gate you already have — and an escalation defers to them too: if a stricter policy would have denied the call, you get that denial rather than an approval prompt. This does not replace your permission system; it sits in front of it.
Enable it in your profile's cordis.patch.yml:
- id: blockrun-review
config:
enabled: true
reviewerModel: anthropic/claude-opus-5
What gets reviewed. Deliberately narrow — a gate that fires on ordinary work gets switched off, and then it protects nobody. Reads, edits and builds are never reviewed. The shipped rules flag recursive deletes, raw disk writes, fork bombs, curl … | sh, force-pushes and hard resets, chmod 777, sudo, and anything touching ~/.ssh, ~/.aws, or /etc/passwd — plus destruction that isn't spelled rm: git clean -fdx, find … -delete, git checkout -- ., terraform destroy, and npm publish (a registry will not let you take a release back).
Mentioning a command is not running one — grep -rn "rm -rf" docs/ is not flagged — and neither is writing one: a Makefile containing rm -rf build, a cleanup script, or a README quoting git reset --hard are all ordinary work. File-body arguments (content, new_string, diff, …) are treated as data, because what a file eventually does happens when something executes it, and that execution is a separate call this gate still reads. Add your own rules:
extraRules:
- name: no-prod-deploy
pattern: "deploy\\s+--env[= ]prod"
If you mistype reviewerModel, every flagged command escalates or is denied — which looks exactly like the gate working cautiously. The failure now carries the cause, so a denial reads "BlockRun does not serve model … Did you mean …?" rather than a bare timeout, and a warning is logged wherever a log exporter is composed.
When the reviewer is unreachable, the gate escalates to you (onReviewerFailure: ask, the default). It never silently allows — a safety gate that fails open is worse than none — and never hard-blocks on a network blip. Unattended automation can set deny.
What it costs to leave on
Measured, because this is the question that decides whether you keep it enabled:
| Fires on ordinary work | never — 0 of 59, including commands that merely mention a destructive one (grep -rn "rm -rf" docs/, echo "DROP TABLE" >> notes.md) |
| Misses dangerous work | none of 39, across git, containers, clusters, cloud storage, databases, and host state |
| Catches files that execute later | git hooks, CI workflows, shell startup files, launch agents, .gitconfig, .env, npm postinstall, sandbox escalation — 10 of 10, 0 false positives across 15 ordinary file edits |
| Survives evasion | \rm -rf /, command rm, env rm, eval "rm -rf $DIR", bash -c "…", | xargs rm, and heredocs piped into a shell |
| Cost when it does fire | $0.0057 on claude-opus-5, at the 512-token reviewer cap — $0.0249 without it |
| Latency when it does fire | ~3s |
| What the reviewer sees | ~356 tokens — the flagged call, not your conversation |
That figure depends on the cap. This gateway quotes from the request — input size plus the max_tokens asked for — and settles that amount whichever way the model answers, so a review that asks for room it never uses pays for it every time the gate fires. reviewerMaxTokens (512) is what keeps a two-field JSON verdict priced like one. Before 0.10.0 the reviewer inherited claude-opus-5's advertised 128,000-token output and cost $0.28–0.33 per review; if you are on an earlier version, upgrade rather than switching to a weaker reviewer.
So during normal work it is invisible: no latency, no cost, no prompts. It bills roughly half a cent on the rare command that deserves a second opinion. Both corpora are tests, so a rule that starts flagging npm test — or stops flagging kubectl delete namespace — fails CI rather than your session.
Not every dangerous action is a shell command. Writing .git/hooks/pre-commit, .github/workflows/ci.yml, or an npm postinstall runs code later — on the next commit, the next CI run with your secrets, the next npm install on someone else's machine. These are quieter than rm -rf, and worse for it: a user watching for destruction sees nothing happen at all. Measured before those rules existed, 2 of 10 were flagged.
Recall is the ceiling on everything above: a command the matcher never flags is a command the reviewer never sees. An earlier version of this table claimed nothing was missed, measured against the six commands the rules had been written for. Against the 39 above, those same rules caught one. The corpus exists so that number can never again be taken on faith.
2. /spend
/spend
What this route has cost since the process started — total, per model, tokens and fees.
How the figure is computed depends on which credential you are on, because the two are billed on genuinely different arithmetic. /spend says which one it used, in the sentence under the total.
On an API key: priced from the tokens the provider reported, at the catalog's published per-million rates. That is the same basis your account is invoiced on — no per-call fee, no minimum — so the figure lines up with /dashboard/activity rather than approximating it. Measured against the live account host: openai/gpt-5.5 on 16 input / 17 output tokens reports $0.000590, which is 16/1M × $5 + 17/1M × $30. A model the catalog publishes no rate for is counted as unpriced and named as such, never silently as $0.
On a wallet: the flat quote the gateway gave, per call. The two chains are quoted differently for the same request — measured 2026-09-05, 2000 µUSDC on Base against 1000 on Solana — so /spend uses the figure for the chain the call actually settled on (requestFeeUsd, solanaRequestFeeUsd). A session that used both reports mixed and explains both.
Everything below this line describes the wallet path, whose numbers are quotes rather than usage. Before 0.11.0 those numbers were reported for API-key deployments too, which overstated a session of small calls by orders of magnitude.
You pay for what you request, not what you get. The gateway quotes from the request — input size plus the max_tokens you ask for — and settles that quoted amount whichever way the model answers. Measured against production:
max_tokens requested |
claude-opus-5 |
deepseek-chat |
|---|---|---|
| 16 | $0.0020 | $0.0020 |
| 1,000 | $0.0036 | $0.0020 |
| 8,000 | $0.0211 | $0.0020 |
| 60,000 | $0.1511 | $0.0027 |
Two things follow, and the second one costs real money.
There is a floor of $0.002 — a $0.001 minimum payment plus a flat $0.001 transaction fee. Below it everything quotes the same, which is why deepseek-chat barely moves in that table: it is cheap enough that even 8,000 output tokens stays under the floor. An earlier version of this section concluded from exactly that observation that billing was per request rather than per token. It was measured only on deepseek-chat, the cheapest model on the route, where the floor hides the rate entirely.
A large max_tokens is billed even when the reply is short. This is why defaultMaxTokens is capped at maxOutputCeiling (8,192) rather than taken from a model's advertised max_output. Left uncapped, claude-opus-5 advertises 128,000, and a request carrying that default quotes $0.3211 — against $0.0216 with no cap at all and $0.0036 capped at 1,000. Eighty-nine times the cost, decided by a field the caller never set. Raise maxOutputCeiling when a workload genuinely needs long replies; you are then paying for them deliberately.
Input size drives the other half of the quote. The same request at growing prompt sizes, max_tokens held small:
| Model | small | ~22K in | ~112K in |
|---|---|---|---|
openai/gpt-4.1-nano |
$0.002 | $0.005 | $0.023 |
deepseek/deepseek-chat |
$0.002 | $0.007 | $0.031 |
google/gemini-3.5-flash |
$0.002 | $0.066 | $0.325 |
anthropic/claude-opus-5 |
$0.002 | $0.217 | $1.081 |
Everything starts at the same floor and then diverges by more than thirty-fold. A coding agent holding a 100K-token context pays roughly fifteen times the floor per call on DeepSeek — and five hundred times on Opus. /spend says so whenever your average call carries a large context, and points you at your own model's rate rather than one number. It is also blind to a request that failed after paying. Your wallet balance is the authority.
A free model is $0, and /spend says so. The 7 models the catalog bills as free never reach the x402 handshake — the gateway answers them with 200 and no 402, so no quote is signed and nothing settles. Their rows read $0 N calls (free tier — no payment was signed), and the large-context warning below is computed from the paid calls alone, since a call that was never quoted cannot be under-quoted. Before 0.10.3 every call was charged the flat request price regardless, so a session spent trying the free tier reported a cost that was entirely invented.
It also says when a different model answered. The gateway substitutes silently behind the free tier — a request for nvidia/nemotron-3-ultra-550b is answered by a 30B about one time in three — and the substitute is named on the chunks it streams. /spend prints it under the row for the model you asked for:
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning $0 3 calls (free tier — no payment was signed)
answered by nvidia/nemotron-3-nano-30b on 1 of 3 — the gateway substituted a different model
answered by nvidia/nemotron-3-super-120b on 2 of 3 — the gateway substituted a different model
The call still counts against the id you requested rather than the one that answered, because this meter prices every call at the same flat figure — which row it lands on cannot move the total, and the id you chose is the one you would recognize.
Reading a 402 quote is free, so every figure above is reproducible without spending anything.
The default requestFeeUsd is 0.002 because that is what the gateway quotes: a 402 for a ~17-token request returns {"amount":"0.002000"}. BlockRun's published pricing page currently says $0.001.
3. /review
/review <paste a diff, a plan, or the agent's conclusion>
Runs the same strong model over material you choose. For the case one user reported: the agent read the right evidence, drew the wrong conclusion, and only a direct challenge surfaced the real bug.
4. /gate — check the net is actually up
/gate # is the gate armed, and with what?
/gate drill # put a dangerous command through the live reviewer
A safety feature that is quietly off is worse than one never installed, because you stopped watching. This gate can be off while everything a user can see looks right: enabled defaults to false, a patch layer replaces a row's whole config rather than merging keys, and /review registers either way — so a working /review tells you the plugin loaded and nothing about whether tool calls are being inspected.
/gate is therefore registered whether or not the gate is armed, and says which. /gate drill sends rm -rf / --no-preserve-root through the risk matcher and the real reviewer — never to a tool — and reports each stage separately, because they fail for unrelated reasons: a rule that stopped matching is a policy problem, an unreachable reviewer is a wallet or model problem. At runtime those both collapse into "escalate", which is indistinguishable from the gate working. The drill is what tells them apart. It costs one reviewer call.
5. Vision — give your agent eyes it does not have
DeepSeek serves no vision model, so this is capability rather than savings. Attach an image and a vision model reads it:
- id: blockrun-llm
config:
visionModels: [google/gemini-3.5-flash] # narrows the measured default; widen as you verify
The gateway's vision tag is not sufficient, so this plugin does not trust it. Every tagged chat model is sent a solid PNG and asked its colour, in three different colours — one is guessable — and only the models that answer all three correctly are offered image input. Re-measure it yourself with npm run probe:vision, which prints the list to paste back. Measured 2026-08-31, 34 of 40 answered:
| Result | Models |
|---|---|
| answered correctly | every tagged Anthropic, Google, Moonshot, xAI and Z.ai model; OpenAI's non-pro models plus gpt-5.6-sol-pro, gpt-5.6-terra-pro and now gpt-5.6-luna-pro; deepseek/deepseek-v4-flash-vision-exp and xiaomi/mimo-v2.5 |
refuses the image — INVALID_REQUEST on all three |
openai/gpt-5.2-pro, gpt-5.4-pro, gpt-5.5-pro |
| answers that no image was sent, because the gateway never sent it | nvidia/llama-3.2-11b-vision |
| never measured — a different model answered | nvidia/nemotron-3-nano-omni-30b-a3b-reasoning (served by nemotron-3-nano-30b), qwen/qwen3.8-flash (served by qwen3.7-flash) |
The llama-3.2-11b-vision row is not the model's fault: the gateway strips every image part on both its NVIDIA paths before dispatch, unconditionally, under a comment claiming NVIDIA models have no vision. Probed directly against NVIDIA, that model and nemotron-3-nano-omni both name the colour correctly. It stays out of the list because this list is what works through this gateway, but it is a gateway bug and it is being fixed.
The last row is the one to understand before widening the list yourself. Those two are not failing; they are not being reached. nemotron-3-nano-omni passes on roughly half of probe runs and is answered by a model with no vision on the rest, so admitting it would buy an image path that works when the gateway's cascade feels like it — and when it does not, the reply is a confident wrong answer about an image that was never seen. qwen3.8-flash is a paid model substituted by a cheaper paid one, so this is not a free-tier quirk. Reported as BlockRunAI/blockrun#450.
gpt-5.6-luna-pro moved the other way, though not for the reason it looks like. It was not an image failure at all: the catalog priced it below what OpenAI charges, the gateway derives its upstream cost ceiling from that price, so every request matched no endpoint and — with no fallback declared for it — failed outright. Repricing it fixed the model, and the image happened to work all along. The separate fix for the three pro entries above is still open, which is why they still refuse.
The first measurement (2026-08-16) was far worse: OpenAI returned HTTP 400 after payment, xAI 503, and Anthropic streamed [Error: 400 {"message":"Could not process image"}] as assistant text, so the harness saw an ordinary successful turn and the agent acted on the error string as though the model wrote it. The gateway has since fixed all three, and this plugin still detects that exact shape — the whole message being nothing but a relayed error — and finishes the request as a failure with the status mapped as if it had arrived as one. An answer that merely mentions an error, or a turn that also called a tool, is left alone. So a model is offered image input only when the gateway tags it vision and it appears in visionModels, which defaults to the measured set. Both signals must agree — the tag alone over-claims, and the list alone would keep claiming vision for a model the gateway has since retagged.
Widen it yourself as you verify others; that is a config change, not a release here.
6. Reasoning effort
Reasoning models get high and max, declared per model from the catalog's reasoning tag.
max is DeepSeek's vocabulary, which the harness adopts. OpenAI's is low | medium | high, and it returns HTTP 400 after taking payment for anything else — so max is translated to each vendor's nearest value rather than refused. Asking for the most thinking available should not fail over a spelling.
Asking a model that does not reason at all is a different case, and is refused locally, before paying: openai/gpt-4o charges and then rejects reasoning_effort outright. The catalog says which models qualify, so that costs nothing to discover.
7. 75 models from one credential
Registers a blockrun provider route reaching every model BlockRun serves. That matters most for the ones DeepSeek does not — Claude, GPT, Gemini, Grok — which is exactly what a reviewer needs.
There are three ways to pay for it, and you pick one. They are not three doors onto the same thing: different host, different billing, different failure when the money runs out.
| API key (recommended) | Solana wallet | Base wallet | |
|---|---|---|---|
| Credential | brk_live_… from user.blockrun.ai |
a bs58 Solana secret key | an EVM private key |
| Host | api.blockrun.ai |
sol.blockrun.ai/api |
blockrun.ai/api |
| Billed | exact token usage against the published price sheet — no per-call fee, no minimum | a flat quote per request, signed and settled on chain over x402 | the same, on Base |
| Top up | card or wire, in the portal | send SPL USDC to your own address | send Base USDC to your own address |
| Where the spend shows up | your account ledger at user.blockrun.ai/dashboard |
the wallet balance itself | the wallet balance itself |
| Signup | Google sign-in | none at all | none at all |
| Config key | apiKeyEnv |
solanaWalletKeyEnv |
walletKeyEnv |
They are checked in that order. The API key wins because it is the credential you chose on purpose — paying from a wallet you merely happen to have exported would spend money on a call you meant to put on the account. Solana is checked before Base: a deployment holding both wallets has said which chains it can pay on, not which it prefers, and this route picks Solana. Both gateways serve the same catalog, verified id for id.
Adding a credential does not disturb an existing deployment: set none of the three and nothing about your setup changes.
Paying on Solana needs @solana/web3.js and @solana/spl-token. They are optional peer dependencies, so an API-key or Base-only deployment does not carry them:
npm install @solana/web3.js @solana/spl-token
セキュリティ・認証 の他のプラグイン
cc-safety-net
by kenryu42
AIコーディングエージェント向けの実行前セキュリティガード。Gitやファイル操作の破壊的コマンド、機密ファイルへのアクセスをツール実行前に検知してブロックします。
tencentmeeting-cli
by tencentcloud
テンセント会議(Tencent Meeting)の CLI ツール。オープンプラットフォームの OAuth2 認証を利用し、会議管理、録画管理、参加者レポートなどに対応
dsh-auto-review
by perrylink
DeepSeek Harness 承認リクエスト向け第 2 モデル AI 自動レビュー:読み取り専用レビューサブエージェントが構造化 allow/deny 判定を返す。デフォルトフェイルクローズ、セッションログで完全監査可能
dsh-permission-rules
by perrylink
Claude Code 風宣言的 permission rules:allow/deny/ask、glob/regex、監査ログ
