Back to directory

caliper

Maintenance: Active

edonadei/caliper

Run your real agent with and without your skills, MCPs, and rules. See which ones actually help, and what they cost in tokens. Supports Claude Code, Codex, Pi, and Hermes.

View on GitHubHomepage
$ npx skills@latest add edonadei/caliper

195

stars

19

forks

Python

Language

MIT

License

2026-05-13

Created

2026-09-28

Last push

MIT-licensed Python eval harness for agent skills, distributed as a dsh-plugin via the skills marketplace. Run your real agent (Claude Code, Codex, Pi, Hermes) against a spec, track a success rate, ablate and compare runs, and see token/time cost per attempt.

DSH integration

Ecosystem-related

Author-claimed

Safety audit

Unaudited

Last verified

2026-09-01

License

MIT

01What can it help you accomplish?

  • Measure whether an agent skill actually works before you ship it

    A tracked **success rate** you can watch over time, plus a saved per-run transcript to inspect

    Developers and teams maintaining skills for Claude Code, Codex, Pi, or Hermes who want proof a skill helps — and what it costs in tokens

  • Prove a skill (not the bare agent) is doing the work, and catch activation drift

    An ablated-vs-full diff via `caliper compare`, plus a per-skill activation scoreboard showing when a skill fires on a neighbour's prompt

    Skill authors who need to know a skill still fires correctly and stays in its lane after a model update or a one-line prompt edit

02How to install into DeepSeek Harness

Prerequisites

  • Python 3.10+ — the CLI install step notes `pipx install caliper-eval # requires Python 3.10+`
  • A target agent CLI installed and authenticated — claude-code, codex, pi, or hermes (Caliper runs skills only through CLI agents; there is no direct-API backend)

Installation steps

  1. 01

    Install the skills (agentic path): `npx skills@latest add edonadei/caliper`

    $ npx skills@latest add edonadei/caliper

  2. 02

    Or install the CLI (Python path): `pipx install caliper-eval # requires Python 3.10+`

  3. 03

    Write a spec, then run it: `caliper run my-skill.eval.yaml --k 3 # --ablate <skill> for a run to diff against`

  4. 04

    Diff two runs: `caliper compare .caliper/results/commit-commands/<evaluation-run>.json .caliper/results/commit-commands/<ablated-run>.json`

Verify the integration

Not specified by the author

03DSH integration and capability boundaries

DSH integrationEcosystem-related

Distributed as a dsh-plugin through the skills marketplace (npx skills add); evaluates agent skills for Claude Code, Codex, Pi and Hermes, the agent backends DeepSeek Harness runs

  • Spec-driven evaluation harness

    a `.eval.yaml` spec describing the skills, judge, and tasks to run→a tracked success rate and a saved transcript per run

  • Ablation and run comparison

    a spec plus an `--ablate <skill>` run and a full run→`caliper compare` diffs the two runs task by task

  • Bundled agent skills: evaluate-skill and grill-skill

    your skill's `SKILL.md` (grill-skill) or an `.eval.yaml` (evaluate-skill)→creates, runs, validates and manages evals from inside the agent; grill-skill generates a 3-task spec

  • Token and time usage reporting

    any run→per-attempt token volume and wall-clock time, rolled up per run

04Who is it for? When not to use it?

Good for

  • Developers and teams maintaining skills for Claude Code, Codex, Pi, or Hermes who want proof a skill helps — and what it costs in tokens
  • Skill authors who need to know a skill still fires correctly and stays in its lane after a model update or a one-line prompt edit

Not for

  • Caliper runs skills only through CLI agents, so every backend can actually load and run a skill. There is no direct-API backend: to run against API-priced billing you configure one of these CLIs with an API key rather than selecting a separate backend.
  • Running a spec that declares `mcp:` on a backend that can't honor it is a hard error rather than a silent no-op. `pi` does not and will not honor `mcp:` natively; expose the capability as a CLI tool the skill drives, or run the eval on claude-code / hermes / codex instead.

05Compatibility, maintenance and safety notes

  • Caliper runs skills only through CLI agents, so every backend can actually load and run a skill. There is no direct-API backend: to run against API-priced billing you configure one of these CLIs with an API key rather than selecting a separate backend.
  • Running a spec that declares `mcp:` on a backend that can't honor it is a hard error rather than a silent no-op. `pi` does not and will not honor `mcp:` natively; expose the capability as a CLI tool the skill drives, or run the eval on claude-code / hermes / codex instead.
  • Caliper never pastes your skill into the prompt. It installs it where the agent looks for skills and lets the agent decide, so a run measures the description (does it fire?) and the body (does it work?) together.
2026-05-132026-08-30v0.11.0

MIT · latest release v0.11.0 (2026-08-29); repo last pushed 2026-08-30

06Frequently asked questions

How do I install Caliper for DeepSeek Harness?

Add it as a skill through the skills marketplace with `npx skills@latest add edonadei/caliper`, or install the CLI directly with `pipx install caliper-eval` (requires Python 3.10+). The `evaluate-skill` skill installs Caliper automatically if it's missing.

Which agents can Caliper evaluate skills against?

Claude Code, Codex, Pi, and Hermes. Caliper runs skills only through CLI agents — there is no direct-API backend — so you configure one of those CLIs with an API key (e.g. `ANTHROPIC_API_KEY` / `OPENAI_API_KEY`) for API-priced runs.

How do I prove a skill is actually helping and not the bare agent?

Run the spec once with `--ablate <skill>` (the skill removed) and keep that run, then `caliper compare` it against the full run. The diff shows per-task pass-rate and token deltas, isolating the skill's real contribution.

Does Caliper track cost, not just pass rate?

Yes. It records token volume and wall-clock time per attempt and rolls them up per run, so two runs with the same score can be compared on what they cost. Dollar cost is deliberately not tracked — tokens are the volume signal.

Can a spec use MCP servers?

Yes, via the optional `mcp:` block — but `pi` does not honor `mcp:` natively and declaring `mcp:` on a backend that can't honor it is a hard error. For pi, expose the capability as a CLI tool the skill drives instead.

08Data and sources

  • Author-claimedgithub.com91590cd87976…

    npx skills@latest add edonadei/caliper

This page is generated from the project’s public documentation, repository metadata and a structured parse of DSH Plugins; last verified on 2026-09-01. Found an error? Submit a correction.

🏆

Best DeepSeek Harness Plugins

Twelve plugins worth installing first — picked from the whole catalog, across every category.

DSH Plugins is an independent community directory of DeepSeek Harness plugins. Not affiliated with or endorsed by DeepSeek. Third-party plugins are not security-audited — review the source before installing.

New DeepSeek Harness plugins, weekly. No spam.