chevron-left
Go back
Case study

How to Build Hiring Assessments That AI Cannot Game

June 11, 2026

For most of the history of hiring, the assessment was a test the candidate faced alone. That assumption has now collapsed. Today a candidate can sit an online assessment with a large language model open in another window, feeding it the questions and relaying its answers. 

What’s interesting is that both sides of the table now hold the same tool: employers use AI to screen candidates, and candidates use AI to get past the screen, which leaves the assessment caught in the middle, frequently measuring the quality of someone's software rather than the quality of their thinking. Building assessments that resist this is now a core part of hiring well. 

The Reality Of The Gamed Assessment

In one recent survey, around four in ten candidates admitted to using AI during a job application process, and among UK graduates the figure is higher still, with roughly two-thirds using tools like ChatGPT on their applications and around half admitting to overstating their skills

The problem is not simply that candidates use AI, but that standard assessments cannot tell when they have. Controlled research on employment interviews has found that AI-generated answers can raise a candidate's score while slipping past the scoring methods meant to judge them, which means the tool does not just help candidates compete, it actively distorts the result. An assessment that can be outsourced without anyone noticing is not measuring the person in front of you. It is measuring their access to technology, which is precisely the thing a fair process is supposed to see past.

Why Off-The-Shelf Tests Fail First

The assessments most vulnerable to this are the generic, widely used ones: large language models are trained on the public internet, so they tend to be exceptionally good at anything that has been asked and answered many times before. 

When researchers tested this directly, an AI could solve something like three-quarters of standard, off-the-shelf assessment questions but only about a quarter of genuinely novel ones it had never encountered before in training. That gap is the whole game. A recycled logic puzzle or a familiar competency question is trivial for a model that has effectively seen the answer key. A problem written specifically for the role, one that does not exist anywhere online, forces a candidate to actually think, and thinking is the one thing the machine cannot do on their behalf without them understanding it well enough to relay it convincingly. The defence against AI, in other words, begins with originality: assessments built for the specific role rather than bought off the shelf.

Why The Format Matters

Originality is the start, but the format of an assessment matters just as much, and some formats are far harder to game than others. Forced-choice questions, which ask a candidate to choose between equally attractive options rather than agree or disagree with a flattering statement, are much more resistant to manipulation, because there is no obviously correct answer for a model to select. Cognitive assessments that are timed and delivered live, measuring not just whether someone reaches an answer but how quickly and how they get there, are harder to outsource than a leisurely questionnaire. 

A spoken conversation, in which a candidate has to reason aloud and respond to unscripted follow-ups, is harder still, because relaying a chatbot's output in real time without hesitation is far more difficult than pasting it into a text box. Underpinning all of these is human oversight: the point of a well-built assessment is not to wage an arms race of surveillance, but to design the task so that genuine ability is the easiest way through it, and to keep a person in the loop who can notice when something doesn’t add up.

From Defensible Assessment To Trustworthy Data

There is a larger reason to get this right, beyond keeping individual hires honest. An assessment a candidate cannot game produces data you can actually trust afterwards. When the signal from your hiring process is real rather than borrowed from a language model, everything downstream improves: the picture of who your strong performers are, the patterns that reveal what actually predicts success in your organisation, the whole body of HR analytics and workforce insights that only means something if the underlying measurements were honest. 

This is the real payoff of a gaming-resistant [candidate screening solution]: it stops the wrong people slipping through, and it keeps the data clean enough to learn from.

Final Thoughts

The arrival of capable, ubiquitous AI has not made assessment pointless, but it has raised the bar for doing it well. The tests that survive are the ones designed with this new reality in mind: original rather than generic, structured to reward real reasoning over recited answers, delivered in formats that are genuinely hard to fake, and always kept under human oversight. Done carelessly, assessment now measures little more than which candidate had the better chatbot. Done deliberately, it still does what it was always meant to do, which is to find the people who can actually do the work, and to know, with confidence, that it was really them.

Stay in the know

Subscribe and get Thrive's latest updates and articles straight to your inbox.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Website-Resoures-Signup