The composer defines a prompt, expert path, regex bank, rubric weighting, and optional resource ceiling. The candidate gets the same real terminal used by published scenarios. The evaluator returns bounded evidence bands and the server computes the same five-dimension score. Authoring remains available only to configured design partners: not generally available on Recruit seats.
not generally available · design partners · illustrative chrome below
When a candidate submits, the regex bank you wrote runs first, same as a public scenario. Claude runs second against the rubric weighting you set. If your weighting changes after launch we re-grade; old certs version to the rubric they were graded against. You do not have to think about the grading code; you just write the scenario.
Do not write a 6-step exam. Pick one ticket, one expected fix path. Multi-step puzzles drift into gotcha grading.
Hard cap the scenario clock at 12 minutes. Anything longer reads as a take-home assignment; candidates resent it.
Give the candidate room to back out: the methodology dim rewards reversibility. Do not write scenarios where the first wrong move ends it.
Soft-launch to your team. Read 5 transcripts before publishing. If the wrong-but-clever solutions outnumber the right ones, the prompt is broken, not the candidates.
authoring is design-partner only · contact to confirm availability