Skip to content

Learn · Rulebooks

Test a rulebook

Replay a category corpus through the pack and read the variance it produces.

Checked against the product on · written for people publishing rule packs

A pack is measured on steadiness: given the same event twice, does it apply the same labels. The harness replays a corpus of sessions through the pack several times and counts the sessions whose runs disagree.

The measure is the set of labels a pass proposes. A session counts as unsteady when any repeat produces a different set.

How the number becomes a grade.
GradeSessions that disagreeReads as
SUnder 0.1%The pack answers the same way every time.
AUnder 1%Steady across a large corpus.
BUnder 5%Steady on the common cases.
CUnder 15%Steady on the clear cases.
D15% or moreThe pack answers differently on the same input.

Run the harness

  1. Pick the category your pack belongs to.

    A category holds the templates and the replay corpus the harness uses.

  2. Pick a template and a repeat count.

  3. Start the run.

    The replay happens in a sandbox. Nothing is written to any work area.

  4. Read the variance and the sessions that disagreed.

    Each one names the labels that moved between runs.

Note. Two labels whose boundary is unstated is the most common cause of a session disagreeing with itself. Tightening the two fragments usually moves the grade a full letter.

The report card

A grade is set at publication. A report card is how the pack is doing in production afterwards, across every work area that installed it. The report card is visible to buyers and leaves the grade alone.