
After every strong technical round, hiring teams now ask the same question: was that really them?
They have reason to ask. CodeSignal's own research shows that in 2025, cheating and fraud attempt rates for proctored assessments more than doubled, rising from 16 percent in 2024 to 35 percent. Junior hiring is hit hardest: entry-level assessment cheating and fraud attempt rates nearly tripled year-over-year, jumping from 15 percent to 40 percent. And hiring managers know their tools are not keeping up. A 2025 Checkr survey of 3,000 hiring managers found 59% had already suspected a candidate of misrepresenting themselves with AI, and only 19% felt confident their own process would catch it.
The assessment industry's answer has been more surveillance: more snapshots, more keystroke tracking, more suspicion scores. This post looks at what candidates actually use to cheat, how the main platforms respond, where each one falls short, and why the fix is a different kind of assessment, not a better detector.
Part 1: The cheating toolkit in 2026
Cheating is no longer a second tab. It is a funded product category with pricing tiers and marketing built around one promise: nobody will see it.
Invisible AI overlays
Cluely is the best-known example. Its founders were suspended from Columbia after creating Interview Coder, a tool that provided AI-generated coding solutions for LeetCode exercises during technical job interviews. They turned it into a company instead. Cluely officially launched on April 20, 2025, generating 70,000 user signups within the first week, and raised $15M from Andreessen Horowitz in June 2025 at a $120M valuation.
The mechanism is simple. The product reads your screen via OCR, captures system audio with speech-to-text, feeds that context to an LLM, and surfaces answers in a floating overlay that is invisible to screen-share in Zoom, Google Meet, and Microsoft Teams. Cluely later softened its message: by June 2025, all references to cheating on job interviews were removed from Cluely's website. The technology stayed the same.
The rest of the market sells the same thing, often with invisibility as a paid upgrade:
- Interview Coder does not appear in the dock or taskbar, the overlay remains hidden when sharing your screen, the process name is disguised, and clicks pass through to underlying windows. Listed at Pro $299/mo; Lifetime $799.
- Final Round AI markets a desktop Stealth Mode billed as "100% Invisible & Undetectable". Its own marketing names the assessment platforms it targets: LeetCode, HackerRank, CodeSignal, CoderPad, Mercor, and claims 10M+ users (self-reported).
- LockedIn AI charges $54.99/mo general; $119.99/mo with Stealth.
- Cluely lists Pro $19.99/mo; Pro + Undetectability $149.99/mo.
- Parakeet AI, Sensei AI, Verve AI and others fill out the category. One review notes that all of these run as native desktop apps, not browser extensions, which is the key technical reason they avoid detection.
The most important fact is how cheap the core trick has become. Open-source projects on GitHub document it openly: one uses the Windows API SetWindowDisplayAffinity with the WDA_EXCLUDEFROMCAPTURE flag to make the window invisible to screenshots, screen recording software, and video conferencing screen share. Invisibility is a single operating system setting. Anyone can build a free clone in a weekend.
The second device
The oldest trick still works. A phone on a stand, a tablet by the keyboard, a friend on a call. TestGorilla's own guide admits that candidates can place notes or hide other devices out of the webcam's view.
Leaked and public questions
Many assessments still reuse problems that exist online word for word. Among CodeSignal's flagged sessions in 2025, 15 percent demonstrated elevated similarity to known answers or leaked content.
Proxy candidates and fake identities
In a Gartner survey, 6% said they participated in interview fraud, either by posing as someone else or having someone else pretend as them. Gartner expects that by 2028, 1 in 4 candidate profiles worldwide could be fake. At the extreme end, a single DOJ indictment described defendants who compromised the identities of more than 80 US persons to obtain remote jobs at more than 100 US companies, including many Fortune 500 firms.
Everyday AI help
The most common version is also the simplest. Among Gartner respondents who used AI while applying, 29% used it to generate text for answers to questions on the assessment.
Part 2: How the incumbents respond, and where each one falls short
Every major assessment vendor now sells an integrity layer. Look closely and they share one flaw: they watch the candidate instead of fixing the test.
HackerRank: a strong classifier on a narrow window
What it does. HackerRank's proctoring includes tab monitoring, copy-paste tracking, image analysis, and keystroke pattern recognition, and it publicly reports 93% accuracy for its AI plagiarism classifier.
Where it falls short. None of these layers can see what is happening outside the browser tab running the assessment, so a desktop overlay that never pastes into the editor sits outside what HackerRank can observe. And even a 93% classifier leaves a 7% false positive and false negative rate, which is significant at scale. That means honest candidates wrongly flagged, and cheaters waved through, in every large hiring round. The system only flags cases and hands the decision to the hiring team, so your recruiters inherit the investigation work.
CodeSignal: trust that depends on surveillance
What it does. CodeSignal's Suspicion Score combines solution similarity, telemetry data, and copy-paste activity. For higher assurance, it records video, audio, and screen activity for the entire assessment session, with government ID checks.
Where it falls short. CodeSignal's own data tells the story: unproctored assessments showed score increases more than 4X larger than proctored ones. In plain terms, the standard test is not trustworthy unless the candidate is filmed from start to finish. That is expensive, invasive, and still limited: independent reviewers note it does not detect desktop AI overlay tools, browser extensions running outside the assessment tab, or local AI models.
Codility: the right instinct, bolted on
What it does. Codility tracks paste events, tab switches, unusually fast completion, and attempts to copy the task description, plus webcam snapshots. Its detection of hidden cheating apps is still marked as a preview. It has also added AI follow-up questions: after a take-home is submitted, the AI generates three targeted questions unique to that submission, and each question is time-boxed to 2 minutes.
Where it falls short. Codility is closest to the right idea, and its own content admits it: "The most effective defence against AI-assisted cheating is not better detection." But the follow-ups are an add-on to a detection product, not the core of the evaluation. Six minutes of short questions cannot carry a hiring decision, and Codility's own customers show it. One talent acquisition manager explains that the QA team currently spends the first part of follow-up interviews testing a candidate's understanding of what they completed in Codility. If your team still has to re-test understanding by hand, the assessment did not do its job.
TestGorilla: browser rules against desktop tools
What it does. TestGorilla monitors webcam snapshots (every 30 seconds), full-screen tracking, tab switching, webcam status, developer tools, IP-based location and paste detection.
Where it falls short. Every one of these signals lives in the browser or the webcam frame. A desktop overlay leaves no trace in any of them, and a phone below the camera is invisible between snapshots. TestGorilla's own guidance admits that VPNs can prevent IP leakage and that shared Wi-Fi in offices or universities makes cheating hard to detect. On top of that, its core product is a library of off-the-shelf tests, which is exactly the kind of content LLMs answer best.
HireVue: denial as a strategy
What it does. HireVue adds periodic candidate snapshots and transcript analysis to its video interviews.
Where it falls short. HireVue's position is that cheating in pre-employment assessments remains rare, at the same time CodeSignal reports fraud attempts doubling in a year. A vendor that does not see the problem will not redesign for it. Tellingly, HireVue's own best advice is a design fix, not a detection fix: asking candidates how they arrived at their answers is what reveals whether the answers are their own.
Karat: humans stretched too thin
What it does. Karat relies primarily on interviewer judgment, with interviewers watching for suspicious pauses, eye movement, and inconsistent responses.
Where it falls short. Those interviewers are also evaluating problem-solving, communication, and depth of understanding at the same time. Humans are not good at this job either: in interviewing.io's experiment, where candidates secretly used ChatGPT, no interviewee was caught cheating.
AI interviewers: conversation without work
What they do. A new wave of AI interviewers analyzes behavior during a conversation. Fabric, for example, tracks gaze tracking for reading patterns, response timing variance, LLM-typical language patterns.
Where they fall short. Two problems. First, detection is still probabilistic: Fabric claims to catch cheating in 85% of cases, which by its own number lets the rest through. Second, these products interview people about skills without watching them use those skills. Talking well about work is not the same as doing it, and a conversation-only format is exactly where an audio-listening overlay is strongest.
The pattern
Every approach in this table answers the same question: did this candidate get outside help? Not one of them is built to make outside help useless.
Part 3: Why detection is a losing strategy
It watches the wrong place. The cheating has moved to the desktop, a second device, or another person. Browser monitoring cannot follow it there, and installing monitoring software on candidates' personal machines drives good candidates away.
The economics favor the cheater. Stealth is one operating system flag. Detection is years of R&D. Worse, every published detection method becomes a training manual. Final Round coaches its users that unnatural pauses and reading behavior are what interviewers actually notice, so they practice hiding exactly that.
Probabilities hurt honest people. Suspicion scores are guesses. At scale, even small error rates wrongly flag real candidates, and webcam-based methods carry extra risk: one study found facial detection was significantly more likely to flag women with darker skin tones for review. Every flag then needs a human to investigate.
It punishes the skill you want. AI is part of the job now. At Canva, almost half of their frontend and backend engineers are daily active users of an AI assisted coding tool. A test that treats every AI interaction as fraud measures how people work without the tools they will use every day.
A clean report on a weak test is still a weak test. This is the real problem. Perfect detection would only prove the candidate solved the problem alone. It would not prove the problem was worth solving. The interviewing.io experiment shows what matters: verbatim questions had a 73% pass rate compared to just 25% for custom questions, while professional interviewers caught no one. Changing the test cut the cheaters' advantage by two thirds. Watching harder did nothing.
Part 4: Redesign the interview so cheating stops helping
The companies with the most at stake have already moved on from detection. Meta said its AI-enabled coding interview is more representative of the developer environment that our future employees will work in, and also makes LLM-based cheating less effective. Canva made AI part of its interviews on purpose, stating "We want to see the interactions with the AI as much as the output". Google chose the most expensive fix of all: at least one in-person round will be mandatory for certain roles. That works, but most companies cannot fly every candidate in.
The lesson is clear. The answer is not a better camera. It is a better assessment, one where outside help simply stops paying off. We call this approach solve, then defend.
1. Real work, built from the actual job
Every challenge is generated from the full job context: the stack, the responsibilities, the level. Candidates work in role-native formats: code, diagrams, spreadsheets, written work, voice, video. There is no public question bank to leak and no generic puzzle for an LLM to recognize. The test looks like the job, because it is built from the job.
2. AI is optional, and the hiring team decides
Some roles need to verify raw skill. Others depend on how well someone works with AI. So AI use is a setting, not a rule. The hiring team chooses per challenge whether AI is off or allowed.
When AI is allowed, it is not treated as a violation. It is measured as a skill: what the candidate asks, what they check, what they catch and fix. Hidden AI loses its value when open AI use is already part of the evaluation.
When AI is off, hidden help still has to survive the next step.
3. A live Solution Defense, always on video
Once the full challenge is submitted, the candidate defends their solution live on video with an AI interviewer that adapts to their answers. Why this approach? What did you reject? What breaks if this constraint changes? Where did AI help, and where did you overrule it?
This is not a few timed text questions. It is a real conversation about the candidate's own work, held after the solve phase ends. They have to own decisions they made earlier, in their own words, on camera, with follow-ups shaped by what they just said. An overlay can write code. It cannot give someone the memory of decisions they never made. You can't defend work you didn't do.
4. Two separate signals, not one guess
The work and the defense are scored separately. Strong work and a strong defense means real, owned capability. Strong work and a weak defense tells the hiring team something important, without anyone having to accuse anyone of anything. A suspicion score gives you a percentage to argue about. Two signals give you evidence to act on.
5. Authenticity read in context
Identity is still verified and integrity signals are still captured. But instead of a suspicion score to investigate, the result is a clear authenticity verdict with a one-line reason, with identity and work reported separately. Work signals are read through the defense. A paste event from someone who then explains every line on video means something very different from the same event from someone who can't.
Detection-first vs. design-first
The better question
For years, hiring teams have asked: how do we catch people who cheat? That question leads to an arms race you pay for and cannot win.
The better question is: how do we design an evaluation where cheating stops helping? That question leads to real work, a clear choice on AI, and a defense on video that only the person who did the work can give.
Stop watching harder. Start testing better.
Sources
- CodeSignal research on assessment fraud (PR Newswire): https://www.prnewswire.com/news-releases/codesignal-detection-systems-identify-and-stop-record-high-cheating-attempts-as-assessment-fraud-more-than-doubled-in-2025-302696534.html
- Checkr survey of hiring managers (via Truffle): https://hiretruffle.com/blog/ai-tools-video-interview-dishonesty
- Gartner on candidate fraud (HR Dive): https://www.hrdive.com/news/fake-job-candidates-ai/757126/
- interviewing.io ChatGPT cheating experiment: https://interviewing.io/blog/how-hard-is-it-to-cheat-with-chatgpt-in-technical-interviews
- Cluely (Wikipedia): https://en.wikipedia.org/wiki/Cluely
- Final Round AI, undetectable interview tools: https://www.finalroundai.com/blog/best-undetectable-ai-interview-tools
- GhostBar on GitHub: https://github.com/tolutally/ghostbar
- HackerRank on plagiarism detection: https://www.hackerrank.com/writing/plagiarism-detection-accuracy-2025-hackerrank-93-percent-vs-codesignal
- Codility assessment integrity features: https://support.codility.com/hc/en-us/articles/47800858996497-Assessment-Integrity-Features
- Codility on AI in technical assessment: https://www.codility.com/ai-technical-assessment/
- TestGorilla on cheating detection: https://www.testgorilla.com/blog/cheating-detection-skills-assessments/
- HireVue on mitigating cheating: https://www.hirevue.com/blog/hiring/mitigating-cheating-enhancing-candidate-experience-ai-hiring
- Canva engineering blog on AI in interviews: https://canva.dev/blog/engineering/yes-you-can-use-ai-in-our-interviews
- Meta AI-enabled interviews (TechRadar): https://www.techradar.com/pro/meta-is-going-to-let-candidates-use-ai-in-job-interviews


