
For the last two years, most hiring teams have handled AI in assessments the same way: block it, detect it, penalize it. Locked browsers, "AI detectors", and a growing list of things candidates are not allowed to do.
We think this is the wrong fight. Not because cheating isn't real, but because the premise is already out of date.
Everyone will use AI at work. That includes the person you're about to hire.
Between May and July 2026, 90% of professional developers used AI coding agents at work at least weekly, and 68% used them daily. The same shift is running through sales, finance, design, support and ops. Not as a novelty, as the default way work gets done.
So a test that bans AI measures a job that no longer exists. You learn how someone performs in conditions they will never work in again.
And a test that "detects" AI treats the most common professional tool of the decade as contraband. It tells candidates one thing: this company is behind.
The question hiring teams need answered has changed:
Not "did they use AI?" but "how well did they use it?"
Give two people the same model and the same task and you get very different results. One writes a vague prompt, accepts the first output, and ships it. The other gives real context, checks the result against what they know, catches the bug the model introduced, and can explain every decision in the final version.
Same tool. Completely different skill. That difference is what you're hiring for, so it should be measured, not policed.
What good AI use looks like
When we say AI use is a skill, we mean something concrete. Strong AI users tend to:
- Know what to hand off and what to keep for themselves
- Give the tool enough context to do the job properly
- Check the output instead of trusting the first answer
- Catch the errors the model introduces
- Own the result: explain why it's built the way it is
Weak AI use is the mirror image: hand off the whole task, take what comes back, submit it, and hope nobody asks a hard question.
The hard question is the whole point.
How to measure it
Once you accept that AI use is a skill, the measurement gets simpler. You measure it the way you'd measure any skill: look at the work, then ask the person about it.
Three things carry the measurement, and none of them depend on which tool the candidate used.
Let them use AI where the work happens. If the job is done in a code editor, a document, a diagram or a spreadsheet, the assessment should happen there too, with an AI assistant available whenever the hiring team allows it. That gives a first layer of evidence: what they asked, what they took, what they changed. Useful, but not the foundation.
Score the work on its own. Whatever was submitted gets judged on its merits. Good AI use produces good work. Careless AI use produces work that looks fine and falls apart on inspection.
Then have them defend it. After the work is done, the candidate walks through their solution in a live conversation that probes their reasoning, their decisions, and how deep the AI use went. Someone who delegated blindly can't go one level deeper than the output. Someone who used AI well can explain what the tool did, what they changed, and why.
You can't defend work you didn't do. That holds whether the work was done by a friend, a contractor, or a model.
What comes out is a read on AI proficiency with the evidence behind it, sitting next to the role skills in the report. Not a cheating flag.
Why a walled AI sandbox measures the wrong thing
There's a tempting alternative, and some assessment tools are heading there: build a controlled AI environment. One model, one interface, every keystroke logged. Put every candidate through it and compare the prompts.
It produces clean data. It also measures the wrong thing.
Environments change every few months. The tools people actually use today: Cursor, Claude Code, Copilot, ChatGPT, custom agents, internal wrappers built by their own team. Next quarter it will be a different list. Nobody can capture all of them, and any sandbox is out of date the day it ships.
Everyone works differently. A senior engineer with a tuned agent setup, a marketer who lives in a chat window, and an analyst with a notebook plugin are all "using AI". They are not using it the same way.
A forced environment is an assumption about the candidate. The moment you push a specific harness onto candidates, you're deciding for them how they should work with AI. You end up testing how fast they can learn your tool, not how well they do their job. And you penalize the people with the strongest workflows: the ones who've built their own.
So the design principle is simple:
Measure what's stable, not what's volatile.
The tool will change. The quality of the work and the depth of understanding behind it won't. That's why an assistant inside the working environment should be a convenience and a source of evidence, not a cage. And it's why the defense shouldn't care which tool the candidate used. It only cares whether they can stand behind the result.
What this changes for hiring teams
- Decide per challenge whether AI is off or allowed, based on how the role actually works
- Read AI use as a signal next to the other skills, not as a red flag
- Compare candidates on outcomes and understanding, not on which tool they picked
- Ask the hard question, and hire the people who can answer it
AI didn't make skill harder to measure. It changed what the skill is.
Sources
- JetBrains, AI Coding Agent Adoption 2026: https://blog.jetbrains.com/research/2026/08/ai-coding-agent-adoption-2026/
- Gallup, AI at Work Indicator: https://www.gallup.com/699797/indicator-artificial-intelligence.aspx
- Stack Overflow, 2025 Developer Survey: https://survey.stackoverflow.co/


