Skip to main content
New

claude plugin eval: finally measure what your plugin brings

Since v2.1.269 (September 2026), Claude Code can test a plugin on a suite of cases and compare it against a session without the plugin. Step-by-step tutorial, costs and pitfalls.

  • Tutorial
  • Tooling
Published

Quick answer

You wrote a Claude Code plugin with a skill or two. It seems to work. But does it work better than Claude on its own? Until now, the only possible answer was "I think so".

The claude plugin eval command, which shipped in v2.1.269 during the week of September 7 to 11, 2026, changes that. It runs your plugin against a suite of test cases, scores the results and, by default, reruns each case without the plugin to show what it really contributes.

The closest analogy is a clinical trial with a control group. You don't just measure that patients get better, you compare with those who didn't take the treatment.

The idea in three concepts

A case: a realistic prompt, the way a user of your plugin would type it, plus one or more graders.

A grader: a pass/fail check on what Claude produced. For example a regex over the reply, whether a specific tool was called, or a rubric a second model judges the reply against.

The no-plugin baseline: each case runs with the plugin (WITH column) and without it (W/OUT). The difference, Δ, is what your plugin contributes. If a case scores 1.0 in both columns, your plugin isn't what makes it pass.

Each case runs three times by default, because a single run of a non-deterministic agent says little. A case's score is the mean across its runs.

Requirements

  • Claude Code v2.1.269 or later (claude --version, then claude update if needed).
  • Git 2.31 or later, if git is installed.
  • A plugin directory with a plugin.json or .claude-plugin/plugin.json manifest.

Your first suite, step by step

1

Have Claude draft the cases

From your plugin's root, run claude plugin eval init. An interactive session opens. Claude reads your plugin, asks what a good result looks like, proposes prompts that should and shouldn't trigger the plugin, designs graders, tries them once and writes one directory per case under evals/. Exit the session with /exit when Claude says the suite is ready.

2

Run the suite

Still at the plugin root, run claude plugin eval . to run every case.

3

Read the table

A summary table shows the WITH, W/OUT and Δ columns, the run count and the estimated cost. A detailed HTML report is written under evals/results/.

4

Iterate

Adjust your plugin, rerun, compare.

# From the plugin root
claude plugin eval init
claude plugin eval .

Sample output from the docs:

CASE WITH W/OUT Δ RUNS COST NOTES
first-case 1.00 0.33 +0.67 6 $0.41
1 case(s) · mean Δ +0.67 · 74s · $0.41

Here, the plugin takes the case from 0.33 to 1.00: it clearly adds something.

The most common diagnosis

The docs say it themselves: the most common first finding is a Δ near zero with the grader that checks the skill call failing. Translation: Claude isn't choosing your skill when users phrase their request naturally.

The fix almost always goes through the skill's description field. That's the text Claude reads to decide whether to use it. Our skills guide covers how to write it.

To iterate quickly and cheaply on one case, run a single run without the baseline, then confirm with the default three runs before drawing conclusions:

claude plugin eval . --case <case-name> --runs 1 --ablation none

Picking good graders

There are six grader types. Four are computed from the transcript and files and cost nothing: regex, tool_used, tool_order and file_exists. The other two, llm and baseline, call a judge model and add to the cost. There are no custom-code graders.

The docs' advice for stable scores:

  • For long output, like a generated file, prefer a regex over the file. Keep llm graders for short outputs, with concrete PASS and FAIL conditions.
  • Give each case one grader on the result and one on the path taken (tool_used or tool_order). One tells you if the answer is right, the other if your plugin produced it.
  • If the skill grader passes but Δ is negative, suspect the judge first. Rerun with --judge-model sonnet for a stronger judge.

An example grader that checks your skill was called:

---
type: tool_used
tool: Skill
input_match: '"skill"\s*:\s*"(?:[\w-]+:)?your-skill-name"'
---

Security: what a run can do

Runs never stop to ask for permission. By default, only the read-only tools listed in the case are allowed. Bash, Write, Edit, WebFetch and WebSearch are removed unless you grant them with --allow-tools.

If you grant Bash, every command runs in Claude Code's OS-level sandbox: writes confined to the run's workspace, home directory unreadable, network limited to granted domains. Native Windows has no sandbox backend, so suites that grant a shell must run under WSL2.

Your plugin's real MCP servers don't start during a run unless you ask. You can replace them with Markdown mocks under evals/mocks/.

In CI

claude plugin eval can act as a gate in a pipeline. The docs recommend pinning the model under test with --model, so a default model rollout isn't mistaken for a plugin regression. Also add evals/results/ to your .gitignore.

Next steps