Eval design for AI product managers

Know what good
looks like before you ship.

Evals Coach turns a PRD or feature idea into a small, testable eval your team can run, critique and improve. No eval experience or repository access required.

Install the Claude plugin View on GitHub ↗

Best as a Claude plugin. Also a skill in Codex and other assistants. Then just say use Evals Coach.

“The AI should be helpful and accurate.” That isn’t an eval.

Most AI features are shipped on vibes because turning fuzzy product intent into something measurable is genuinely hard. These are the traps Evals Coach is built to catch.

Vague quality words you can’t test

“Helpful”, “accurate”, “on-brand”. Until they become observable must / must-not behaviours, no grader can score them and no one can agree whether a run passed.

Evals that pass while the product fails

A green dashboard measuring the wrong thing is worse than none. Evals Coach hunts for what could score well while real users still get hurt.

Failures with no regression net

The same production incident keeps recurring because nothing turned it into a test case. Feed in the failures; get back regression cases, without inventing evidence.

An LLM judge nobody calibrated

An uncalibrated judge quietly gating your release is a coin flip in a lab coat. Get a path to check it agrees with human labels before it holds the gate.

Built for PMs, not eval engineers.

No eval experience

You bring product judgement. It brings the eval-design method: modes, cases, graders, thresholds.

No repository access

Start from a PRD, feature idea, workflow, traces, feedback, or an existing eval. No codebase required.

Your stack, not ours

The output isn’t tied to any eval platform, so it imports into whatever your team already runs.

Ask in your own words. It picks the right mode.

Four things it does, whichever way you describe the task.

Create

From intent to eval

Turn a PRD, capability, user job or product idea into a practical eval plan and an importable test set.

Critique

Stress-test what you have

Review an existing eval for weak criteria, thin coverage, unreliable graders, or misleading release thresholds.

Expand

Failures into regressions

Turn production failures, traces, support cases or user feedback into regression cases that lock the fix in.

Calibrate

Trust your graders

Improve agreement between human judgement and an automated grader before it decides what ships.

A real eval, not generic advice.

Every time, it follows the same path: from the decision you need to make to a handoff your team can run.

1

The product decision the eval must support

Ship or hold? For which unit of work? Everything downstream is anchored to a real call you’re about to make.

2

Observable must / must-not behaviours

Vague quality words are replaced with behaviours you can actually see in a transcript and point to.

3

A minimum viable test set

The smallest set that can inform the decision: normal, edge, adversarial and critical cases, with the costliest failure covered first.

4

The right grader for each criterion

Deterministic, trace, LLM-as-judge or human, matched to the behaviour, with judge prompts written for you.

5

Calibration guidance and explicit release gates

What agreement to require, and the threshold that actually decides ship or hold, stated out loud, not implied.

6

A handoff engineering can run

An eval plan, a validated test-cases.csv, judge prompts, and a path from manual review to CI and production feedback.

Add it to Claude, then ask for an eval.

Recommended as a Claude plugin, which brings the workbench with it. It also runs as a skill in Codex and other assistants. Set it up once.

In the Claude app

Recommended · adds the workbench

Desktop, web and Cowork. No terminal needed.

  1. 1
    Open Settings → Customize → Plugins
  2. 2
    Click Add, then Add marketplace
  3. 3
    Enter justshipai/evals-coach and confirm
  4. 4
    Click Install on Evals Coach, then enable it

Then ask for an eval in your own words, or ask for the eval workbench.

Claude Code (CLI)

Terminal

Two commands in your terminal.

/plugin marketplace add justshipai/evals-coach /plugin install evals-coach@justshipai

Then ask for an eval, or ask for the eval workbench.

Codex & other assistants

Clone as a skill

Runs anywhere. Clone into your assistant’s skills directory.

git clone https://github.com/justshipai/evals-coach.git ~/.agents/skills/evals-coach

Then, for example: $evals-coach Turn this PRD into the smallest eval that can inform a ship decision.

The Eval Workbench

A seven-step wizard, published to your own account as a private Artifact. Describe your feature; it drafts the evaluation question, criteria, test cases, graders, judge prompts and release gate, each one editable, and hands back an eval plan, a test-cases.csv and judge prompts.

Claude-only: it asks Claude from inside the page, on the account of whoever opens it. Ask for “the eval workbench” once the Claude plugin is installed. Opened without drafting, it degrades gracefully to a worksheet with the same structure and starters.

1 Describe your feature
2 Confirm the evaluation question
3 Set good & bad
4 Edit the test cases
5 Approve the graders
6 Choose the release gate
7 Take the handoff away

Does it actually work? We measured.

A six-task blinded comparison using GPT-5.6 Sol at medium reasoning, Evals Coach versus a no-skill baseline.

100/65

Rubric score, Evals Coach vs baseline (%)

6/6

Tasks won, out of six

0/2

Critical failures vs baseline

The same run exposed a weakness (early skill outputs ran long) that we’ve since tightened, adding PM usability to the rubric. Because the rubric was developed alongside the skill, with one run per task and no independent replication yet, treat this as promising early evidence rather than a universal claim. Read the full method, scores and limitations →

Define good. Then measure it.

Install once and turn your next feature idea into an eval that can actually inform a ship decision.

Install the Claude plugin View on GitHub ↗