chevron-left
Go back
Case study

Nine Ways to Ensure Your Psychometric Assessments Are Legal, Accurate and Reliable

June 1, 2026

The difference between an authentic psychometric assessment and personality quizzes that circulate online is consequential. The two are easily confused only because they borrow the same vocabulary: one is a measurement instrument, built and validated by occupational psychologists to predict something specific about performance. The other is often just entertainment. 

But the distinction is important, because the entire point of using an assessment is to make better and faster hiring decisions, and an instrument that isn’t accurate is a waste of time and money. A defensible process is not complicated, but it is exacting. 

Here are nine things that separate a psychometric assessment worth trusting from one that isn’t.

Use Instruments Built By Occupational Psychologists

The first question about any assessment is who built it and on what evidence. A genuine psychometric assessment is designed by qualified specialists using the established principles of test theory, tested rigorously, and documented with published validity and reliability data, rather than assembled from questions that merely sound insightful. 

Decades of academic research point in a consistent direction: personality and cognitive ability are among the strongest predictors of job performance, because they measure stable psychological constructs that the work genuinely requires, while education, experience and unstructured interviews correlate far more weakly. An instrument built by occupational psychologists to measure those constructs - like ours at Thrive - is doing something fundamentally different from a quiz that scores your answers based on averages or vague concepts unsupported by data. The practical test here is a simple one: if a provider cannot explain the science behind their assessment and show you how it was validated, it’s not worth using.

Match The Assessment To The Specific Role

Validity isn’t a fixed property that a test either has or lacks; it has to be relative to what you’re trying to predict. This means an assessment that works well for one role can be close to meaningless for another. The psychological constructs that make a strong salesperson are not the ones that make a strong analyst, and a serious process begins by defining what a particular role really demands and then measuring that, rather than reaching for a generic profile and hoping it fits. When the assessment is matched to the competencies a role actually requires, the resulting data means something. When it’s not, you end up measuring the wrong things precisely, which is no improvement on measuring nothing at all.

Check That It Predicts Performance

An assessment earns its place only if its scores relate to how people go on to perform in the job, which is what criterion and predictive validity mean. Validity, in plain terms, is the degree to which an assessment measures what it claims to measure, and predictive validity specifically asks whether today's scores forecast tomorrow's performance. This is not a detail to gloss over. Decades of selection research show that structured, job-relevant methods predict performance far better than intuition or loosely related traits, and the strength of that link is the single most important thing to establish before you rely on any tool. A provider should be able to tell you not just that their assessment is valid, but which kinds of validity they have evidenced and how, because validity is measured on a spectrum rather than claimed as a flat yes or no.

Confirm That The Test Is Reliable

Where validity asks whether an assessment measures the right thing, reliability asks whether it measures consistently. A reliable assessment gives the same person broadly the same result regardless of the day, the mood or the surroundings, and an instrument that produces a different answer each time is simply measuring noise. The two qualities are related but distinct: an assessment can be reliable without being valid, like a marksman who hits the same wrong spot every time, but it cannot be valid without first being reliable. Psychologists check this in several ways, from test-retest reliability, where the same person is assessed twice, to internal consistency, the preferred modern method, which examines whether the individual items hang together and point to the same underlying trait. No amount of polished presentation can rescue a score that will not hold still, so a reliability figure, and the method behind it, is something to ask for by name.

Standardise How The Test Is Administered And Scored

Comparisons are only fair when everyone is assessed under the same conditions against the same criteria. That means consistent instructions, consistent timing where it is relevant, and a scoring system that does not drift with the mood of whoever is marking. 

Psychologists distinguish between systematic errors, which are flaws baked into the design of a test, and unsystematic ones, which come from the circumstances of a particular sitting, such as a noisy room or a distracted candidate, and standardisation is how you suppress both. The same discipline should extend to the human parts of hiring. An assessment works best as one of several inputs feeding a structured and standardised interview, so that every candidate is judged on the same evidence in the same way, rather than on the accidents of who interviewed them and how the conversation happened to flow. The moment conditions vary between candidates, results stop being comparable and bias finds its way back in.

Monitor For Adverse Impact

An assessment can look perfectly neutral and still disadvantage a particular group, which under the Equality Act 2010 can amount to indirect discrimination unless the practice is a proportionate way of meeting a genuine need. For example, language can penalise non-native speakers, culturally specific scenarios can disadvantage those from different backgrounds, and, most importantly, an assessment is only as fair as the norm group used to score it: if the sample the results are benchmarked against was narrow, the assessment may simply not apply to everyone who takes it. Good design fights this deliberately, by building assessments on large and heterogeneous samples, offering high-quality translations, and re-testing and re-norming over time. The legal test is demanding, because if an equally valid but less discriminatory alternative exists, the original can be unlawful, so checking your own results for disproportionate effects is both an ethical and legal obligation.

Make Reasonable Adjustments

The law requires employers to adjust assessments for disabled candidates who would otherwise be placed at a substantial disadvantage, and this is an area where standard tests routinely fall short. Most psychometric assessments are normed on neurotypical people, so a neurodivergent candidate - someone with autism, ADHD or dyslexia - may score lower not because they are less able but because the format does not accommodate how they think. In one well-known case, a multiple-choice test that a candidate's autism made unfairly difficult was found to have discriminated against her, precisely because the format, rather than her ability, was the barrier. Sensible adjustments, such as extra time or the option of open-ended answers, remove that barrier without lowering the bar. Building this flexibility into how an assessment is taken is not a courtesy extended to a few candidates, it is a legal requirement and a condition of measuring everyone fairly.

Keep A Human In The Loop And Be Transparent

Where assessment is automated, the law increasingly expects a person to remain meaningfully in control and expects candidates to be told that a system is being used. Under the new emerging European rules on artificial intelligence, tools used to screen or evaluate candidates are treated as high-risk and carry specific obligations for human oversight and transparency. The principle is sound regardless of jurisdiction. An assessment score is powerful because it looks objective, and that authority is exactly why it should inform a human decision rather than replace one outright. Candidates deserve to know when a system is assessing them and, within reason, how it works, and a responsible process keeps a qualified person accountable for the final judgement rather than deferring to a number. Transparency is not a constraint on good assessment. It is part of what makes assessment trustworthy in the first place.

Document Everything And Review It Regularly

A defensible process is a documented one: what assessment you used, why you chose it, how it performed, and what you found when you checked it for bias. It’s important not to see this as bureaucracy for its own sake, but the evidence you would need if a decision were ever challenged, whether by a candidate, a tribunal or your own leadership. It also reflects the fact that an assessment is not a fixed asset but a living one. Roles evolve, workforces change, and the best providers treat validation as continuous, using the data they gather to refine their assessments rather than assuming yesterday's evidence still holds. The final discipline, then, is to revisit rather than to trust indefinitely: to keep records, monitor outcomes, and re-validate as the world the assessment operates in moves on. An assessment you can account for is one you can defend, and one you can keep improving.

Final Thoughts

The reward for all of this discipline is considerable. An assessment process built this way, supported by a [hiring evaluation platform] like ours, produces decisions that are accurate, fair and explicable if anyone ever asks. It also produces something valuable long after the hire, which is reliable behavioural data about employees. This data can then inform how they are developed and managed, which is one of the more practical routes to improving employee engagement - and retention - once they’re through the door. 

Done casually, psychometric assessment is a liability dressed as rigour. Done properly, with valid and reliable instruments, careful attention to fairness, and a human being accountable at the end of it, it’s one of the few tools that makes hiring measurably better and keeps it firmly on the right side of the law.

Stay in the know

Subscribe and get Thrive's latest updates and articles straight to your inbox.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Website-Resoures-Signup