Purpose
A pilot exists to confirm that scoring lands correctly for your organization before you lean on it for high-stakes decisions. The most important design choice is running enough volume to trust the calibration.The calibration window
Scoring is calibrated to your organization during Foundation, but calibration is only confirmed by running it. Treat your first phase, up to about 100 evaluation sessions, as a calibration-validation window.Below roughly 100 sessions, treat scores as directional and keep human review
especially close. Reserve high-stakes reliance for after scoring is confirmed
at that volume. This is a hard gate on the Production
Checklist.
Designing the pilot
1
Pick one role with real volume
Choose a role that will actually generate around 100 sessions in the pilot
window, so calibration has something to validate against.
2
Keep a human close on every decision
During the window, review evidence on every evaluation, not a sample. This
is how you catch calibration issues early.
3
Track agreement
Compare reviewer judgment with the score. Persistent disagreement on a
dimension is a calibration signal, not noise to ignore.
4
Confirm, then expand
Once scoring is confirmed at volume, expand to more roles and departments.
See Enterprise Rollout.
What good looks like
Enough volume
Around 100 sessions on the pilot role before high-stakes reliance.
Close review
Human review on every evaluation during the window.
Documented agreement
A record of where reviewers and scores agreed and diverged.
Deliberate expansion
Expansion only after calibration is confirmed, not before.
Related
Calibration
How scoring stays reproducible as evaluations run.
Go-Live Readiness
Clearing setup before the first live session.
Production Checklist
The gate where the volume threshold lives.
Enterprise Rollout
Expanding after the pilot.

