
Google's CEO now says 75% of the company's new code is AI-generated, up from 25% in 2024 and 50% last autumn. Most companies are nowhere near that. The more careful industry-wide measurements put AI-authored code at roughly a quarter to a third of what gets written.
But the direction is the same everywhere: less of a developer's day goes into typing code, more goes into deciding what to build, directing the tool, and checking what came back.
That changes what a good developer looks like. It has not changed how most companies test them.
The old test measures the thing AI does best
The standard technical screen is still a timed coding exercise: an algorithm puzzle, a blank editor, a clock. It was built to answer one question: can this person write working code from scratch, quickly?
Two problems with that now.
- It is the easiest thing to fake. Any model solves the puzzle in seconds. If the candidate has a second screen open, your test is measuring copy-paste speed.
- It is the least relevant thing to measure. Writing code from a blank file is exactly the part of the job moving to AI. You are screening for the skill that matters least.
The timed puzzle tests the one thing you will never ask them to do at work.
What the job actually is now
Watch a strong developer work with AI for an hour. The loop looks like this:
- Frame the problem: turn a vague ticket into a precise task with the right constraints.
- Decide the shape: where the boundaries go, what the data looks like, which trade-off to take. The AI does not see your system, your users or your roadmap.
- Direct the tool: what to ask for, what context to give, when to stop asking and write it yourself.
- Read critically: the code compiles, the tests are green, and it is still wrong in some small way.
- Prove it works: write the test that matters, not the test the AI wrote for itself.
- Own it: explain every line that ships, because someone will ask.
The "read critically" step is where the pain is. In Stack Overflow's 2025 survey, two thirds of developers named AI solutions that are almost right but not quite as a problem they run into, and 45% said debugging AI-generated code takes more time. More developers now actively distrust the accuracy of AI output than trust it.
There is a twist, and it matters for hiring: developers are poor judges of their own AI use. In a 2025 METR study, experienced developers working on their own repositories took 19% longer with AI tools, while believing they had been about 20% faster.
Review is the new writing. The scarce skill is knowing whether the code in front of you is right, and being able to say why.
Why you cannot test this with a quiz
Judgment does not show up in a multiple-choice question. Neither does taste in trade-offs, or the instinct to distrust a plausible answer.
These things show up in two places: in real work, and in how someone explains their choices afterwards. So an assessment has to produce both: an artifact, and a conversation about the artifact.
A practical way to do it
1. Give real work in the real stack. Not a blank editor. A small codebase with a problem in it: a legacy function to refactor, a bug that only appears under load, an endpoint to design together with its data model, a pull request to review. The task should look like the ticket they would get in week two.
2. Decide about AI up front, then measure it. Switch it off if you want to check fundamentals. Switch it on if you want to see how they actually work. When it is on, do not try to detect it. Look at how it was used: what they asked for, what they rejected, what they verified before accepting. (We wrote about why AI use is a skill to measure, not a threat to detect, here: [link].)
3. Go beyond code. A system design diagram tells you more about senior judgment than another function does. A code review task, where the code is mostly right and subtly wrong, tests exactly the skill the survey numbers say is scarce. A debugging task shows you what happens when generated code fails inside a real system.
4. Then have them defend it. After the solution is submitted and frozen, hold a live conversation about their own work:
- Why this shape and not the other one?
- What breaks at ten times the traffic?
- What did the AI suggest that you threw away?
- What did you not test, and why?
Someone who did the work answers quickly and concretely. Someone who pasted it in stalls on the second question.
You cannot defend work you did not do.
5. Read the two signals separately. The artifact tells you what they produced. The defense tells you what they understood. Keep both, because the combinations mean different things:
- Strong artifact, strong defense: the hire signal.
- Strong artifact, weak defense: borrowed output, or a tool used without understanding. Dig deeper.
- Weak artifact, strong defense: the thinking is there. Maybe the time was short, maybe the stack was new. Often worth a second look.
- Weak artifact, weak defense: a clear no, learned in an hour instead of a month.
What this looks like in practice
A backend candidate is asked to refactor a legacy PHP function so it can be tested. The submitted code is cleaner. The score lands mid-range: output is still coupled to presentation, and it is not really testable yet.
In the defense, the candidate explains separation of concerns correctly, then admits they did not apply it because they ran out of time. That is a real signal: they understand the principle and did not execute it.
A different candidate with the same code, who cannot say what "testable" means, is a completely different hire.
Same artifact. Opposite conclusions. You only see the difference because you asked.
What changes in your rubric
The short version
The resume told you what they claimed. The coding puzzle told you they could type. Neither tells you the one thing that matters when AI writes the code: can this person tell right from almost right, and stand behind what ships?
Real work. Then defend it.
Sources
- Semafor on Google's AI-generated code: https://www.semafor.com/article/04/24/2026/google-ceo-says-75-of-companys-new-code-is-ai-generated
- Stack Overflow 2025 Developer Survey: https://survey.stackoverflow.co/2025
- METR study on AI and developer productivity: https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- METR paper on arXiv: https://arxiv.org/pdf/2507.09089


