01Run it
Replace the <placeholders> with your own values, or use the builder below.
p202 eval run --agent-cmd <agent_cmd> --cases <cases>02What you get
--json, --csv or --ndjson for scripts, or -q for ids only. An AI agent gets compact JSON without asking. All output formats03Build your command
Pick values and the command line writes itself, quoted and ready to paste.
p202 eval runSet flags below; the command updates as you type.
04Flags
7 flags, plus the global flags every command takes.
| Flag | What it does |
|---|---|
| --agent-cmdstringrequired | Shell command that runs the agent under test (required) |
| --casesstringrequired | Case file or directory of *.json case files (required) |
| --judge-cmdstring | Shell command that grades rubrics (stdin: JSON; stdout: PASS/FAIL verdict) |
| --onlystring | Comma-separated case ids to run |
| --p202-binstring | Real p202 binary the capture shim execs (default: this binary) |
| --prioritystring | Priorities to run: critical, high, medium, low; comma-separated for severalcriticalhighmediumlow |
| --timeoutint · default 300 | Per-command timeout in seconds (agent, setup, checks) |
05How it works
Run behavioral snapshot evals: for each case the runner performs the setup
commands, hands the ask to --agent-cmd (stdin and $P202_EVAL_ASK), captures
every p202 invocation the agent makes (a PATH shim — no adapter needed),
re-reads instance state, and grades the case's expectations. Cases follow
the shape in .claude/skills/p202-agent-evals/SKILL.md; a starter file ships
in tests/fixtures/agent-eval/cases/.
The agent command must read the ask, drive p202 (the one on PATH — that
is the capture shim), print the agent's reply to stdout, and exit 0.
The shim records each command's exit status: runs_one_of credits only a
command that exited 0 (or the case's runs_one_of_exit), and names the
status of a matching command that failed; never_runs counts every
attempt. Each result's commands_failed counts what did not exit 0.
Rubric grading is optional: --judge-cmd receives {id, ask, rubric, reply,
commands, runs} as JSON on stdin (runs: each command with its exit_code)
and must print a line starting with PASS or FAIL; without it, rubric
cases report needs_judge instead of pass.
Exit codes: 0 when nothing failed; 5 (partial_failure) when cases failed or errored — results are still on stdout.
06For agents
Running this from an agent
- Read the same facts as JSON:
p202 commands eval run --json. - With
AI_AGENT,CLAUDECODEor another agent variable set, output is compact JSON and errors arrive on stderr as a JSON envelope with ahint. - Exit codes: 0 ok, 1 bad input, 2 auth, 3 network, 4 server error, 5 partial failure.
{
"path": "p202 eval run",
"use": "run",
"short": "Run eval cases against a pluggable agent command",
"long": "Run behavioral snapshot evals: for each case the runner performs the setup\ncommands, hands the ask to --agent-cmd (stdin and $P202_EVAL_ASK), captures\nevery `p202` invocation the agent makes (a PATH shim — no adapter needed),\nre-reads instance state, and grades the case's expectations. Cases follow\nthe shape in .claude/skills/p202-agent-evals/SKILL.md; a starter file ships\nin tests/fixtures/agent-eval/cases/.\n\nThe agent command must read the ask, drive `p202` (the one on PATH — that\nis the capture shim), print the agent's reply to stdout, and exit 0.\nThe shim records each command's exit status: runs_one_of credits only a\ncommand that exited 0 (or the case's runs_one_of_exit), and names the\nstatus of a matching command that failed; never_runs counts every\nattempt. Each result's commands_failed counts what did not exit 0.\nRubric grading is optional: --judge-cmd receives {id, ask, rubric, reply,\ncommands, runs} as JSON on stdin (runs: each command with its exit_code)\nand must print a line starting with PASS or FAIL; without it, rubric\ncases report needs_judge instead of pass.\n\nExit codes: 0 when nothing failed; 5 (partial_failure) when cases failed\nor errored — results are still on stdout.",
"runnable": true,
"flags": [
{
"name": "agent-cmd",
"type": "string",
"default": "",
"usage": "Shell command that runs the agent under test (required)",
"required": true
},
{
"name": "cases",
"type": "string",
"default": "",
"usage": "Case file or directory of *.json case files (required)",
"required": true
},
{
"name": "judge-cmd",
"type": "string",
"default": "",
"usage": "Shell command that grades rubrics (stdin: JSON; stdout: PASS/FAIL verdict)",
"required": false
},
{
"name": "only",
"type": "string",
"default": "",
"usage": "Comma-separated case ids to run",
"required": false
},
{
"name": "p202-bin",
"type": "string",
"default": "",
"usage": "Real p202 binary the capture shim execs (default: this binary)",
"required": false
},
{
"name": "priority",
"type": "string",
"default": "",
"usage": "Priorities to run: critical, high, medium, low; comma-separated for several",
"allowed_values": [
"critical",
"high",
"medium",
"low"
],
"value_list": true,
"required": false
},
{
"name": "timeout",
"type": "int",
"default": "300",
"usage": "Per-command timeout in seconds (agent, setup, checks)",
"required": false
}
]
}