Project Manager Assessment: Validity, Fairness, and What Really Predicts Performance
Hiring? Send candidates a scored, cheat-resistant assessment.Start free
The question hiring leaders should ask
You've built a PM assessment. Scenario problem, prioritization, risk assessment, behavioral interview. Candidates who score 4+ do well on the job. Candidates who score 2.5 or below fail. But have you verified that? And is the assessment fair?
This post walks through what validity means for PM assessments, how to measure it, and what fairness looks like in practice.
What validity means
An assessment is valid if it predicts the job outcome you care about. For PM hiring, that's: "Does this person ship projects on time, manage risk well, and build team trust?"
There are three types:
1. Predictive validity
Does the assessment score predict future job performance?
How to measure it:
- Hire 10+ PMs using your assessment.
- After 6 months, rate them on job performance (360 review, manager feedback, project delivery metrics).
- Compare assessment score to performance rating.
- If high scorers perform well and low scorers struggle, you have predictive validity.
What good looks like:
- Correlation of 0.6+ between assessment score and performance rating (strong).
- Correlation of 0.4-0.6 (moderate, still useful).
- Correlation below 0.3 (low, the assessment isn't predictive).
Real data point: Teams using scenario-based PM assessments typically see 0.5-0.7 correlation. Teams using unstructured behavioral interviews see 0.2-0.3. The difference is real.
2. Construct validity
Does the assessment actually measure what it claims to measure?
For PM assessment, you claim to measure:
- Decision-making under constraint
- Prioritization judgment
- Risk awareness
- Stakeholder influence
How to verify: Do candidates who score high on "decision-making" actually demonstrate decision-making on the job? Or are they just good at taking the test?
Red flag: A candidate scores 4.5 on the scenario (decision-making) but on the job tends to hedge and seek consensus. The assessment didn't measure what matters.
How to prevent it: After hiring, have the hiring manager rate the candidate on each of the four dimensions independently (at 3 months and 6 months). Compare their rating to the assessment score. If there's a big gap, your assessment is measuring the wrong thing.
3. Content validity
Does the assessment include realistic problems that candidates will actually face?
Examples of high content validity:
- "You have a customer threatening to leave unless you ship by October 1" (real PM problem).
- "Rank these features given these constraints" (real PM problem).
- "Three teams are in parallel but one is a dependency; identify the risks" (real PM problem).
Examples of low content validity:
- "Write a 10-page project plan from scratch" (PMs don't do this in day-to-day work).
- "Explain Agile vs. Waterfall" (tests knowledge, not judgment).
- "Tell me about a time you managed a team" (behavioral, not work-sample).
How to measure it: Show your assessment to three PMs currently in role. Ask: "Do these problems look like what you actually face?" If they say no, you're testing something other than job performance.
Validity is not automatically there
Many organizations assume: "If the assessment looks good to us, it must be predictive." Not true.
Common assessment patterns that look rigorous but aren't predictive:
Pattern 1: Detailed Gantt chart assignment. Looks: professional, organized, technical. Actually measures: ability to use project management software, not PM judgment. Predictive validity: low (0.2-0.3).
Pattern 2: Unstructured behavioral interview. Looks: thorough, gets to know the person. Actually measures: interview confidence and storytelling skill. Predictive validity: low (0.2-0.3).
Pattern 3: Case study with no live debrief. Looks: candidates think deeply about a problem. Actually measures: consulting-style writing and analysis. Predictive validity: medium (0.4-0.5).
Pattern 4: Scenario problem + live prioritization + risk assessment. Looks: rigorous and expensive. Actually measures: decision-making, judgment, and systems thinking. Predictive validity: high (0.6-0.7).
How to verify your own assessment's validity
Step 1: Define what "good performance" means on the job
Before you even check if the assessment predicts it, define the outcome:
- Timeline: PMs ship milestones on the committed date or provide early warning.
- Scope: PMs ship the scope they committed to or explicitly rescope with stakeholder agreement.
- Risk: PMs surface dependency risks proactively, not after they blow up.
- Team: PMs maintain team engagement and psychological safety through change.
Make these behavioral, not fuzzy. "Ships on time" is behavioral. "Is a good leader" is fuzzy.
Step 2: Hire using your assessment and track outcomes
Hire 10-15 PMs over 6 months. Track their performance at 3, 6, and 12 months using the behavioral definition above.
How to measure:
- 360 review (manager, skip-level, peer) anchored to the four behaviors.
- Project delivery metrics (on-time delivery rate, scope changes, team retention).
- Skip-level conversations: "How is this PM's communication? Do you get surprised by risk?"
Step 3: Compare assessment scores to outcomes
Create a simple spreadsheet:
| Candidate | Assessment Score | Job Performance Rating (at 6 mo) | Match? |
|---|---|---|---|
| Alice | 4.2 | 4.1 | Yes |
| Bob | 3.5 | 3.4 | Yes |
| Carol | 3.0 | 2.8 | Yes |
| Dan | 4.8 | 3.2 | No (overpredict) |
| Eva | 2.8 | 2.1 | Yes |
If most rows match, you have validity. If several rows show mismatches, your assessment is not predictive.
Step 4: Fix mismatches
If a high-scorer (4.5 on assessment) performs poorly (2.5 on job):
- They may have gotten help on the scenario.
- The assessment may be measuring something other than job performance (e.g., you're good at taking tests but not at stakeholder communication).
- They may have landed in a role or environment that doesn't suit them (hired as a PM for a Scrum Master role).
If a low-scorer (2.8 on assessment) performs well (4.0 on job):
- Your assessment may be too harsh or is measuring the wrong thing.
- They may have transferred in from another role and learned on the job.
Either way, investigate and adjust your assessment.
Fairness: Is the assessment biased?
Validity is about prediction. Fairness is about equal opportunity.
An assessment can be valid (predicts performance) but unfair (biases against certain groups). Example: a scenario written in business jargon familiar to Ivy League candidates but not to community college candidates. Both groups can PM well, but one group is filtered out unfairly.
Common fairness problems in PM assessments
Problem 1: Assuming a specific industry background. Scenario assumes knowledge of SaaS metrics. Candidates from manufacturing, healthcare, or government are disadvantaged. Fix: Don't assume domain knowledge. Test PM thinking, not domain facts.
Problem 2: Timed scenarios that advantage people without caregiving responsibilities. "30-minute response, due by 5pm." Candidates juggling childcare or elder care are disadvantaged. Fix: Async assessments with flexible deadlines. 24 hours to respond is reasonable.
Problem 3: Language/jargon barriers. Scenario uses specific PM terminology (WIP, burn-down, etc.) without defining it. Non-native English speakers are disadvantaged. Fix: Assume no PM background. Define terms. Test thinking, not vocabulary.
Problem 4: Live verbal component that favors extroverts. Prioritization problem is done verbally in real time. Introverts who think best in writing are disadvantaged. Fix: Offer written or verbal option for prioritization. Both are valid.
Problem 5: Scenarios that assume a specific culture fit. Scenario assumes a startup mentality: "We're scrappy and ship fast." Candidates from risk-averse industries see this as irresponsible and score lower. Fix: Make scenarios industry-agnostic. Test PM thinking, not cultural values.
How to audit for fairness
After you've run your assessment on 20+ candidates:
- Group candidates by demographics (if you track: gender, race, education background, etc.).
- Compare average assessment scores across groups.
- If one group scores systematically lower, investigate:
- Is the group really lower-performing on the job? (Check against actual performance data.)
- Or is the assessment measuring something other than job readiness? (Ask that group: "Did the assessment feel fair?")
What you're looking for: Equal average scores across groups, or if there's a gap, that gap should match the job performance gap (not be larger).
Example:
- Group A scores 3.8 on assessment, performs at 3.7 on the job. ✓ Fair.
- Group B scores 3.2 on assessment, performs at 3.5 on the job. ✗ Assessment under-predicted; something's wrong with the assessment, not the group.
Red flags for invalidity or unfairness
Invalidity:
- Your high-scorers (4+) don't consistently perform well on the job.
- You can't articulate what the assessment is measuring (if you can't say, you probably don't know).
- You haven't measured job performance empirically (you're just guessing).
Unfairness:
- Certain groups score systematically lower, and you haven't verified they underperform on the job.
- You're using language or scenarios that assume a specific background or culture.
- Candidates from non-traditional PM backgrounds (bootcamp, internal promotions) are filtered out at the assessment stage.
Building valid and fair assessment
The best PM assessments:
- Use work samples (scenario + prioritization) to test actual judgment, not knowledge.
- Are industry-agnostic or test across multiple industries so no background is assumed.
- Are async when possible to accommodate different working styles and responsibilities.
- Define what success looks like (the rubric) and then verify that rubric predicts job performance.
- Are audited for fairness — run the numbers every 6-12 months.
An assessment that's valid and fair doesn't guarantee a PM will succeed. But it dramatically improves your odds.
How to validate your PM assessment
If you're using a standardized PM assessment, ask the provider: "What's the predictive validity of this assessment?" Real vendors have run the studies. If they haven't, that's a red flag.
If you've built your own assessment, run the simple four-step validation above (define success, hire and track, compare scores to outcomes, fix mismatches). It takes 6 months but pays for itself in hiring accuracy.