Blog
/
Skills Assessment

What happens when the hiring manager disagrees with the AI?

Most AI tools score candidates and stop. Tendent treats your disagreement as the input: correct one score, and every candidate after is graded differently.

Published on
September 25, 2026

Every AI assessment tool will eventually score a candidate in a way the hiring manager thinks is wrong. The question is what happens next.

In most tools, nothing. You can note your disagreement, maybe change a number, and the next candidate gets scored exactly the same way. The tool never learns what "good" means for your team.

In Tendent, the disagreement is the input. This post explains the mechanism: how a correction on one candidate changes the grading for every candidate after them, how the same corrections shape the questions the AI asks in the Solution Defense, and why the AI starts to sound like your hiring manager after a handful of evaluations, not a hundred.

The first score is an opening position

Tendent evaluates candidates in two steps: they solve a real challenge, then defend it live to an AI interviewer. Each question in the challenge gets a score and a written rationale. So does the Defense.

That first score comes from a role-specific rubric built from the job description and the challenge itself. It is a good starting point. It is not your team's definition of good.

Take a refactoring exercise. The rubric rewards separation of logic from output, type hints, and testability. A candidate fixes the structure but skips the type hints and scores 40%. Your senior engineer reads the rationale and disagrees: for this team, structure is what matters. Type hints are a code review comment.

That gap between the rubric and the manager's judgment is exactly the signal Tendent is built to capture.

The AI's first score is an opening position. The hiring manager's correction is the calibration.

Three ways to disagree

Every score in the Evaluation Report has a "Coach the AI" action next to it. From there, the hiring manager can do three things, in any combination:

Action What you do What changes
Change the score Move 40% to 70% for this candidate. The override becomes the score of record for this candidate. The original AI score and rationale stay visible underneath.
Explain the why "Ignore missing type hints on this exercise. Score structure and testability only." The rubric for this question is updated. Every future candidate on this test is graded against the corrected version.
Tell it what to ask "If they use a static cache, ask how they would invalidate it." The AI interviewer adds the probe to its plan: as a follow-up during the challenge if it clarifies the answer, or in the Solution Defense if it is about the why.

The first action fixes one candidate. The other two fix the evaluation.

The correction carries forward

Here is what happens after a manager explains the why:

  • The reasoning is stored against the question, not the candidate. It becomes part of the rubric.
  • The next candidate who submits the same challenge is graded with the updated rubric. No retraining, no waiting for a batch job.
  • Scores already given to earlier candidates stay as they were. Tendent does not rewrite history. The manager can re-score earlier candidates, but that is a choice, not a side effect.
  • The correction is logged: who changed what, when, and why. The evidence behind the original score (code, transcript, telemetry) stays untouched.
  • Every correction is visible and reversible. The team can see what has been taught and undo it.

That last point matters more than it sounds. Hiring assessment falls under the high-risk category of the EU AI Act, and human oversight has to be real, not decorative. A logged override with a written reason is what "human in the loop" looks like in practice.

The interviewer aligns with the hiring manager

Grading corrections change how answers are scored. Follow-up guidance changes what the AI asks.

In the Solution Defense, the AI interviewer walks the candidate through their own submission: why they chose this approach, what they would change, what breaks under load. The default line of questioning comes from the challenge and the candidate's answers.

When a hiring manager adds guidance ("if they mention Doctrine, ask about N+1 queries"; "don't spend time on naming, dig into error handling"), the interviewer folds it into its plan for that challenge. After three or four candidates, the pattern is clear: it probes what this manager probes, and it moves on where this manager would move on.

Nothing about the candidate's experience changes. The questions are still generated from their own work. What changes is the taste behind them.

Coach it on the first few candidates. Everyone after gets interviewed by someone who knows what you care about.

Where the loop closes

Corrections and follow-up guidance make the AI match the hiring manager. Hiring outcomes tell you whether the hiring manager was right.

When Tendent is connected to the ATS, it sees who was hired and who was not. That closes the loop: corrections that predicted hires get more weight, corrections that did not get flagged. This is how Tendent Role Models are trained: on what hiring teams actually corrected and what actually led to an offer, not on a rubric written once. Next step: preferences that carry across roles, so a team that has coached one evaluation does not start the next one from generic defaults.

Assess, calibrate, compound. The disagreement is step two.

The test for any AI assessment tool

If you have tried AI screening and switched it off because the scores felt off, that is the right instinct. A score you cannot correct is not a signal. It is noise with a decimal point.

So disagree with the tool, then watch what it does with your disagreement. If the next candidate gets the same generic score, it is grading. It is not learning.

‍

Sources

  1. Zheng et al., "Judging LLM-as-a-Judge": https://arxiv.org/abs/2306.05685
    Strong LLM judges reach over 80% agreement with human preferences, about the same level as agreement between humans, which leaves a consistent gap for the hiring manager to close. arxiv
  2. Shankar et al., "Who Validates the Validators?" (UIST 2024): https://arxiv.org/pdf/2404.12272
    Documents "criteria drift": people need criteria to grade outputs, but grading outputs is what helps them define the criteria, so evaluation rubrics have to be refined against real examples rather than written once. arxiv
  3. EU AI Act, Annex III (point 4): https://ai-act-service-desk.ec.europa.eu/en/ai-act/annex-3
    Lists AI systems used for recruitment or selection, in particular to analyse and filter job applications and to evaluate candidates, as high-risk. Europa
  4. EU AI Act, Article 14 (Human oversight): https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-14
    Requires that the people overseeing a high-risk system can correctly interpret its output, stay aware of automation bias, and decide in any given case to disregard, override, or reverse it. presencis
  5. EU AI Act, Article 12 (Record-keeping): https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-12
    Requires high-risk systems to automatically record events over their lifetime so that their functioning can be traced. europa
  6. Fink, "Human Oversight under Article 14 of the EU AI Act" (Leiden University): https://papers.ssrn.com/abstract=5147196
    Argues that human oversight is primarily aimed at output correction, and that cognitive limits and automation bias mean it only works if it is implemented carefully rather than treated as a standalone safeguard. ssrn
About the author
Victor Cazacu

20+ years in HRTech and AI. Currently leading product at tendent.ai and serving as CEO of upper.co. Previously at WPP and N26.

LinkedIn
Table of Contents