A form full of numbers is not automatically a good assessment. If a 3 means one thing to one manager and something else to the colleague next to them, the issue is often the method rather than the person being assessed. A useful assessment should help two people talk about the same facts, understand what is working and choose where to invest energy to improve.
Results and skills are not the same thing
Results describe what was achieved: goals, projects, timing, quality, service levels. Skills describe how the person approaches the work: technical knowledge, interpersonal skills, planning, problem solving, leadership and other dimensions relevant to the role.
Keeping the two separate makes the conversation easier to read. An important result can be achieved in unsustainable ways; likewise, a person can show solid skills during a period when some results also depend on external factors.
A skill has to become observable
Leadership, accuracy or collaboration are labels that are too broad if everyone can fill them with a different meaning. To use them in an assessment you need observable behaviors.
For example, planning can mean: sets explicit priorities, flags risks and dependencies in time, updates the plan when conditions change. It is not a perfect formula, but at least it gives assessor and employee something concrete to discuss.
- Use observable verbs: communicates, checks, plans, documents, involves.
- Avoid turning personality traits into skills.
- Define the expected level relative to the role and seniority.
- Don't include a skill if nobody will actually be able to observe it.
The scale needs a shared meaning
A scale from 1 to 5 can work, but only if the numbers have a meaning. Without descriptors, a 3 risks being on track for one person and barely sufficient for another.
It helps to define at least three clear anchors: below the role's expectations, in line with expectations, above expectations. The intermediate levels can represent transitional situations. The scale measures the distance from the role's expectations, not the value of the person.
Examples come before judgment
A sentence like I find you not very collaborative almost always opens a discussion about the label. A concrete episode, instead, opens a discussion about what actually happened.
Before the conversation, manager and employee can gather a few examples from the period: a successful project, a difficulty handled, feedback received, a mistake that led to a change. This also helps reduce the excessive weight of the last few weeks in memory.
Bias and calibration: making assessments more comparable
Even with clear criteria, whoever assesses remains exposed to cognitive shortcuts. Recency bias leads to giving too much weight to the last few weeks; the halo effect can stretch a very positive or very negative result to dimensions that were not really observed. To reduce these risks you need examples spread across the period, not impressions gathered at the end.
When several managers assess comparable people or roles, a short calibration discussion helps check that words and levels have the same meaning. There is no need to artificially standardize the scores: the point is to ask whether similar evidence is getting similar readings, and to make the differences explicit when the role context justifies them.
- Gather evidence throughout the period, not just before the conversation.
- If a rating is very high or very low, ask which concrete episodes support it.
- Compare similar cases across managers to check the meaning of the levels.
- Use calibration to make criteria consistent, not to force a distribution of scores.
Self-assessment and the manager's assessment must be able to meet
Self-assessment is not about seeing whether the person guesses the manager's score. It is about surfacing different perceptions before the conversation. A gap between the two readings is not necessarily a problem: it can be the most interesting point of the conversation.
The discussion works better when you start from the examples and only then reach the summary. Asking what leads you to this assessment? is far more useful than immediately debating whether the right number is 3 or 4.
Close with a small but verifiable development step
If an assessment ends with seven areas for improvement, chances are that after a few weeks none of them is actually being worked on. Better to pick one or two priorities and turn them into concrete actions.
It can be shadowing, a project, training, leading a meeting, periodic feedback or a new responsibility. What matters is also agreeing on how to tell whether the action is working.
Checklist finale
- Results and skills are kept separate
- Each skill has observable behaviors
- The scale has shared descriptors
- The period being assessed is clear
- The conversation starts from examples, not labels
- Bias and calibration are considered before the final summary
- Differences between self-assessment and manager are explored
- Development actions are few, concrete and verifiable