Skip to main content

Purpose

A pilot exists to confirm that scoring lands correctly for your organization before you lean on it for high-stakes decisions. The most important design choice is running enough volume to trust the calibration.

The calibration window

Scoring is calibrated to your organization during Foundation, but calibration is only confirmed by running it. Treat your first phase, up to about 100 evaluation sessions, as a calibration-validation window.
Below roughly 100 sessions, treat scores as directional and keep human review especially close. Reserve high-stakes reliance for after scoring is confirmed at that volume. This is a hard gate on the Production Checklist.

Designing the pilot

1

Pick one role with real volume

Choose a role that will actually generate around 100 sessions in the pilot window, so calibration has something to validate against.
2

Keep a human close on every decision

During the window, review evidence on every evaluation, not a sample. This is how you catch calibration issues early.
3

Track agreement

Compare reviewer judgment with the score. Persistent disagreement on a dimension is a calibration signal, not noise to ignore.
4

Confirm, then expand

Once scoring is confirmed at volume, expand to more roles and departments. See Enterprise Rollout.

What good looks like

Enough volume

Around 100 sessions on the pilot role before high-stakes reliance.

Close review

Human review on every evaluation during the window.

Documented agreement

A record of where reviewers and scores agreed and diverged.

Deliberate expansion

Expansion only after calibration is confirmed, not before.

Calibration

How scoring stays reproducible as evaluations run.

Go-Live Readiness

Clearing setup before the first live session.

Production Checklist

The gate where the volume threshold lives.

Enterprise Rollout

Expanding after the pilot.