Blog
/
AI & the Future of Hiring

From Claims to Proof

Generative AI has made hiring signals cheap to produce and hard to verify. This paper proposes an evidence-based alternative: candidates complete job-specific work, then defend it live, with the two scored separately.

Published on
September 25, 2026

AI made every hiring claim cheap. A resume, a test score, a confident interview answer: all of them can now be produced in seconds by someone who cannot do the job. Tendent replaces claims with proof. Candidates do real work built for the job, then defend it live. The work and the defense are scored apart, and a person makes the call.

What broke

Every signal hiring runs on used to be expensive to fake. A strong resume took experience. A passed test took knowledge. A good answer took understanding. Now all three are free, and checking them still costs a skilled person’s time. That imbalance is the whole problem.

39% of candidates used AI in their application. 29% of those used it to answer assessment questions (Gartner, 3,290 surveyed)
1 in 4 candidate profiles will be fake by 2028. 6% already admit to interview fraud (Gartner)
11,000 applications a minute on LinkedIn, up 45% in one year (The New York Times, via eWeek)
26% of candidates trust AI to evaluate them fairly (Gartner)
50%+ drop in new-graduate hiring at large tech companies since 2019 (SignalFire)

The result is a loop. Companies cannot assess at volume, so strong candidates cannot stand out, so candidates apply more with AI, so volume grows, so assessment gets worse. Each turn of the loop makes the next one harder.

The tools built for the old world do not stop it. Resume filters rank claims. Test libraries leak. Proctoring is beaten by a second device. AI interviewers score how well someone talks about a skill, which is exactly what a language model is best at. And AI detectors get worse as models get better.

The idea: you cannot defend work you did not do

Stop trying to catch AI. Start measuring what a person can actually do, and whether they understand it.

Tendent works in three steps.

Solve The candidate does real work built for this job, in the tools the job uses: code, spreadsheets, diagrams, slides, simulated conversations, voice, video. AI is allowed where the job allows it. How good is the work?
Defend After the work is submitted and scored, the candidate defends it in a live, adaptive interview grounded in what they produced. Why this decision? What breaks if the budget halves? Where is your solution weakest? Do they understand and own it?
Verify Identity and continuity evidence confirms the same verified person did all of it. Everything lands in one report with a written reason behind every score. Was it really them?

The work score and the defense score are never mixed into one number. That is where the signal lives. Work 91 and Defense 89 is a strong hire. Work 94 and Defense 52 is a conversation you need to have. One blended score would hide the difference.

AI use is a skill, not a crime. The Defense asks how the candidate used it: what they asked for, what they accepted, what they threw out, and why. A candidate who directed a model well and can explain every line has shown something the job needs. A candidate who pasted an answer cannot explain it. Same mechanism, no detector, no accusation.

And a person makes the call. Every score comes with a reason a hiring manager can check, change, and teach the AI from.

Why it works

None of this is a guess. Each part of the design rests on published research.

.42 Structured interviews are now the best predictor of job performance in selection research, with work samples and job knowledge tests close behind. General tests of cognitive ability and personality rank below them. Solve-then-defend combines the two strongest method families in the literature (Sackett et al., 2023)
Explain it People rate their understanding of everyday things highly, then cut that rating sharply once they try to write out how they work. Psychologists call it the illusion of explanatory depth. A defense is built to find that gap (Rozenblit and Keil, 2002)
61% of essays by non-native English writers were wrongly flagged as AI-generated by seven leading detectors. Meanwhile a one-line prompt dropped detection of real AI text to near zero. Detection is not a foundation (Liang et al., 2023)
13,342 people across twelve studies described themselves as more analytical and less intuitive when they believed AI was judging them. Score what candidates say and you score a performance. Score what they did and you do not (Goergen et al., PNAS, 2025)
80%+ agreement between strong AI judges and human evaluators, the same rate humans reach with each other. AI can grade well, with a rubric, a written reason, and a human who can overrule it (Zheng et al., NeurIPS 2023)

Put simply: the research says test the real work, then interview about that work, and never trust a detector. That is the whole design.

What we have seen so far

Early results from evaluations run with design-partner customers in real hiring processes:

93% agreement between Tendent recommendations and decisions by expert human vetters on the same candidates
79% of candidates who start an evaluation finish it
91% of candidates rate the experience positively afterwards

Want the full picture?

The detailed paper covers what this one leaves out:

  • the ten requirements any modern skill evaluation has to meet, and how common tools score against them
  • Solve, Defend and Verify in detail, including how evaluations are generated from the job and how the Defense chooses its questions
  • identity and continuity: how Tendent builds confidence that the same person did all of it
  • how the system learns from hiring managers and real hiring outcomes
  • how the design maps to the EU AI Act, and what it does not claim

Request the full whitepaper: victor@tendent.ai

About the author
Victor Cazacu

20+ years in HRTech and AI. Currently leading product at tendent.ai and serving as CEO of upper.co. Previously at WPP and N26.

LinkedIn
Table of Contents