MIT-licensed Python eval harness for agent skills, distributed as a dsh-plugin via the skills marketplace. Run your real agent (Claude Code, Codex, Pi, Hermes) against a spec, track a success rate, ablate and compare runs, and see token/time cost per attempt.
DSH integration
Ecosystem-related
Author-claimed
Safety audit
Unaudited
Last verified
2026-09-01
License
MIT
01What can it help you accomplish?
Measure whether an agent skill actually works before you ship it
A tracked **success rate** you can watch over time, plus a saved per-run transcript to inspect
Developers and teams maintaining skills for Claude Code, Codex, Pi, or Hermes who want proof a skill helps — and what it costs in tokens
Prove a skill (not the bare agent) is doing the work, and catch activation drift
An ablated-vs-full diff via `caliper compare`, plus a per-skill activation scoreboard showing when a skill fires on a neighbour's prompt
Skill authors who need to know a skill still fires correctly and stays in its lane after a model update or a one-line prompt edit
02How to install into DeepSeek Harness
Prerequisites
- Python 3.10+ — the CLI install step notes `pipx install caliper-eval # requires Python 3.10+`
- A target agent CLI installed and authenticated — claude-code, codex, pi, or hermes (Caliper runs skills only through CLI agents; there is no direct-API backend)
Installation steps
- 01
Install the skills (agentic path): `npx skills@latest add edonadei/caliper`
$ npx skills@latest add edonadei/caliper
- 02
Or install the CLI (Python path): `pipx install caliper-eval # requires Python 3.10+`
- 03
Write a spec, then run it: `caliper run my-skill.eval.yaml --k 3 # --ablate <skill> for a run to diff against`
- 04
Diff two runs: `caliper compare .caliper/results/commit-commands/<evaluation-run>.json .caliper/results/commit-commands/<ablated-run>.json`
Verify the integration
Not specified by the author
03DSH integration and capability boundaries
Distributed as a dsh-plugin through the skills marketplace (npx skills add); evaluates agent skills for Claude Code, Codex, Pi and Hermes, the agent backends DeepSeek Harness runs
Spec-driven evaluation harness
a `.eval.yaml` spec describing the skills, judge, and tasks to run→a tracked success rate and a saved transcript per run
Ablation and run comparison
a spec plus an `--ablate <skill>` run and a full run→`caliper compare` diffs the two runs task by task
Bundled agent skills: evaluate-skill and grill-skill
your skill's `SKILL.md` (grill-skill) or an `.eval.yaml` (evaluate-skill)→creates, runs, validates and manages evals from inside the agent; grill-skill generates a 3-task spec
Token and time usage reporting
any run→per-attempt token volume and wall-clock time, rolled up per run
04Who is it for? When not to use it?
Good for
- Developers and teams maintaining skills for Claude Code, Codex, Pi, or Hermes who want proof a skill helps — and what it costs in tokens
- Skill authors who need to know a skill still fires correctly and stays in its lane after a model update or a one-line prompt edit
Not for
- Caliper runs skills only through CLI agents, so every backend can actually load and run a skill. There is no direct-API backend: to run against API-priced billing you configure one of these CLIs with an API key rather than selecting a separate backend.
- Running a spec that declares `mcp:` on a backend that can't honor it is a hard error rather than a silent no-op. `pi` does not and will not honor `mcp:` natively; expose the capability as a CLI tool the skill drives, or run the eval on claude-code / hermes / codex instead.
05Compatibility, maintenance and safety notes
- Caliper runs skills only through CLI agents, so every backend can actually load and run a skill. There is no direct-API backend: to run against API-priced billing you configure one of these CLIs with an API key rather than selecting a separate backend.
- Running a spec that declares `mcp:` on a backend that can't honor it is a hard error rather than a silent no-op. `pi` does not and will not honor `mcp:` natively; expose the capability as a CLI tool the skill drives, or run the eval on claude-code / hermes / codex instead.
- Caliper never pastes your skill into the prompt. It installs it where the agent looks for skills and lets the agent decide, so a run measures the description (does it fire?) and the body (does it work?) together.
MIT · latest release v0.11.0 (2026-08-29); repo last pushed 2026-08-30
06Frequently asked questions
How do I install Caliper for DeepSeek Harness?
Add it as a skill through the skills marketplace with `npx skills@latest add edonadei/caliper`, or install the CLI directly with `pipx install caliper-eval` (requires Python 3.10+). The `evaluate-skill` skill installs Caliper automatically if it's missing.
Which agents can Caliper evaluate skills against?
Claude Code, Codex, Pi, and Hermes. Caliper runs skills only through CLI agents — there is no direct-API backend — so you configure one of those CLIs with an API key (e.g. `ANTHROPIC_API_KEY` / `OPENAI_API_KEY`) for API-priced runs.
How do I prove a skill is actually helping and not the bare agent?
Run the spec once with `--ablate <skill>` (the skill removed) and keep that run, then `caliper compare` it against the full run. The diff shows per-task pass-rate and token deltas, isolating the skill's real contribution.
Does Caliper track cost, not just pass rate?
Yes. It records token volume and wall-clock time per attempt and rolls them up per run, so two runs with the same score can be compared on what they cost. Dollar cost is deliberately not tracked — tokens are the volume signal.
Can a spec use MCP servers?
Yes, via the optional `mcp:` block — but `pi` does not honor `mcp:` natively and declaring `mcp:` on a backend that can't honor it is a hard error. For pi, expose the capability as a CLI tool the skill drives instead.
07Related DSH workflows
deepseek-reasonix
by esengine
DeepSeek-native AI coding agent for your terminal. Engineered around prefix-cache stability — leave it running.
mnemon
by mnemon-dev
LLM-supervised persistent memory for AI agents — graph-based recall, cross-session knowledge, single binary. Works with DeepSeek Harness, Claude Code, OpenClaw, and any agent runtime.
phi
by pulseaiclub
a coding agent, rpc plugin, sub-agents, hashline edits, and mcp
sivtr
by ariestar
A unified memory workspace for agents and people, making terminal output and AI session context searchable and reusable across local workspaces.
08Data and sources
npx skills@latest add edonadei/caliper
This page is generated from the project’s public documentation, repository metadata and a structured parse of DSH Plugins; last verified on 2026-09-01. Found an error? Submit a correction.
Best DeepSeek Harness Plugins
Twelve plugins worth installing first — picked from the whole catalog, across every category.
