The shape of a score
A score is built in two levels, not one.Answers become a competency score
Competencies become a round score
Why two levels matter
Because the number of questions asked in each competency is not a decision you made — it is an artefact of how that particular round ran. If questions were weighted individually, a panel that spent 18 of 20 questions on Technical would silently turn your 50% Technical weighting into 94%. Averaging inside the competency first means your configured mix is the mix that gets applied, whether a competency received two questions or twenty. Question count still matters — but it affects confidence, not weight. See What a thin round can conclude.Worked example
Your competency mix: Technical 50% · Communication 30% · Culture 20%. The panel asks 18 technical questions (candidate averages 8/10) and 2 communication questions (averages 3/10). No culture questions come up.What counts as evidence
Three states, kept strictly separate:Questions asked outside your configured list
Interviewers improvise, and that is intended — a panel brief gives the interviewer a goal for the round, not a script. Off-script questions are real evidence and they count. Each one is classified against your competencies:Mapped to one of your competencies
Mapped to one of your competencies
Classified with low confidence
Classified with low confidence
Outside your competency matrix
Outside your competency matrix
Touching a protected characteristic
Touching a protected characteristic
Coverage is reported, not deducted — where the gap was not the candidate’s
The guiding rule: a candidate is scored only on what they controlled.What a thin round can conclude
A round’s completeness is measured by how much of your competency matrix was assessed — not how many questions were asked. Those are very different things:- 20% of your list asked, but conversation covered every competency → complete assessment
- 60% of your list asked, all of it inside one competency → incomplete assessment
Score bands
The guarantees
Reproducible
Recomputable
Version-pinned
Never model-declared
What this does not claim
Being straight about the limits is part of being trustworthy about the method.- This makes aggregation fair. It does not make the underlying judgement perfect. The 0–10 score on each answer is still an AI assessment of that answer.
- Scoring depends on the transcript being correct. If a transcript misattributes who said what, the scores will be confidently wrong.
- A score is a recommendation, never a decision. Every outcome is subject to human review, and no adverse decision is issued automatically from an incomplete round.
FAQs
Why did a candidate score lower than the individual answers suggest?
Why did a candidate score lower than the individual answers suggest?
Can I recompute a score myself?
Can I recompute a score myself?
Does changing my weights change past scores?
Does changing my weights change past scores?
Does a candidate lose points for questions the interviewer skipped?
Does a candidate lose points for questions the interviewer skipped?
What can we tell a candidate who asks why they were scored the way they were?
What can we tell a candidate who asks why they were scored the way they were?
Related
- Report Glossary — every report type and score dimension defined
- How to Set Scoring Weights — configuring your competency mix
- Scoring Quality — getting better signal out of rounds
- Calibration — keeping scores aligned to your outcomes

