Evals Coach turns a PRD or feature idea into a small, testable eval your team can run, critique and improve. No eval experience or repository access required.
Best as a Claude plugin. Also a skill in Codex and other assistants. Then just say use Evals Coach.
Most AI features are shipped on vibes because turning fuzzy product intent into something measurable is genuinely hard. These are the traps Evals Coach is built to catch.
“Helpful”, “accurate”, “on-brand”. Until they become observable must / must-not behaviours, no grader can score them and no one can agree whether a run passed.
A green dashboard measuring the wrong thing is worse than none. Evals Coach hunts for what could score well while real users still get hurt.
The same production incident keeps recurring because nothing turned it into a test case. Feed in the failures; get back regression cases, without inventing evidence.
An uncalibrated judge quietly gating your release is a coin flip in a lab coat. Get a path to check it agrees with human labels before it holds the gate.
You bring product judgement. It brings the eval-design method: modes, cases, graders, thresholds.
Start from a PRD, feature idea, workflow, traces, feedback, or an existing eval. No codebase required.
The output isn’t tied to any eval platform, so it imports into whatever your team already runs.
Four things it does, whichever way you describe the task.
Turn a PRD, capability, user job or product idea into a practical eval plan and an importable test set.
Review an existing eval for weak criteria, thin coverage, unreliable graders, or misleading release thresholds.
Turn production failures, traces, support cases or user feedback into regression cases that lock the fix in.
Improve agreement between human judgement and an automated grader before it decides what ships.
Every time, it follows the same path: from the decision you need to make to a handoff your team can run.
Ship or hold? For which unit of work? Everything downstream is anchored to a real call you’re about to make.
Vague quality words are replaced with behaviours you can actually see in a transcript and point to.
The smallest set that can inform the decision: normal, edge, adversarial and critical cases, with the costliest failure covered first.
Deterministic, trace, LLM-as-judge or human, matched to the behaviour, with judge prompts written for you.
What agreement to require, and the threshold that actually decides ship or hold, stated out loud, not implied.
An eval plan, a validated test-cases.csv, judge prompts, and a path from manual review to CI and production feedback.
Recommended as a Claude plugin, which brings the workbench with it. It also runs as a skill in Codex and other assistants. Set it up once.
Desktop, web and Cowork. No terminal needed.
justshipai/evals-coach and confirmThen ask for an eval in your own words, or ask for the eval workbench.
Two commands in your terminal.
Then ask for an eval, or ask for the eval workbench.
Runs anywhere. Clone into your assistant’s skills directory.
Then, for example: $evals-coach Turn this PRD into the smallest eval that can inform a ship decision.
A seven-step wizard, published to your own account as a private Artifact. Describe your feature; it drafts the evaluation question, criteria, test cases, graders, judge prompts and release gate, each one editable, and hands back an eval plan, a test-cases.csv and judge prompts.
Claude-only: it asks Claude from inside the page, on the account of whoever opens it. Ask for “the eval workbench” once the Claude plugin is installed. Opened without drafting, it degrades gracefully to a worksheet with the same structure and starters.
A six-task blinded comparison using GPT-5.6 Sol at medium reasoning, Evals Coach versus a no-skill baseline.
Rubric score, Evals Coach vs baseline (%)
Tasks won, out of six
Critical failures vs baseline
The same run exposed a weakness (early skill outputs ran long) that we’ve since tightened, adding PM usability to the rubric. Because the rubric was developed alongside the skill, with one run per task and no independent replication yet, treat this as promising early evidence rather than a universal claim. Read the full method, scores and limitations →
Install once and turn your next feature idea into an eval that can actually inform a ship decision.