Every agency owner has had the same week. Forty applications arrive, and thirty of them are excellent. The language is crisp, every requirement in the job description is addressed, the skills sections read like a mirror of the advert. Then the phone screens start, and the picture collapses. Candidate after candidate cannot describe the work they claimed. Two days of billable attention has gone into reading fiction. Learning how to detect ChatGPT-inflated CVs has become a core commercial skill, not a curiosity.
This piece explains what generative AI actually changed about CV screening, why AI-detection tools are the wrong response, and how evidence-based evaluation identifies inflation without penalising honest candidates who used a tool to write more clearly.
What Generative AI Actually Changed
It is worth being precise, because the popular framing is wrong. Candidates lying on CVs is not new. What changed is the cost of producing plausible, well-structured, requirement-matched prose, which fell to approximately zero.
For decades, screening leaned on a quiet correlation: candidates who wrote clearly about relevant work usually had done relevant work. Articulacy was a cheap proxy for competence because producing polished, tailored text took effort and understanding. That correlation is now broken. A candidate can paste a job advert into a chatbot and receive a CV that addresses every requirement in confident, idiomatic English, in about fifteen seconds, with no underlying experience required.
The implication is uncomfortable but clarifying: fluency is no longer evidence of anything. Any screening method that rewards well-written alignment with the job description is now measuring access to a free tool.
Why AI Detectors Are the Wrong Answer
The instinctive response is to buy a detector that flags machine-generated text. It is the wrong move, for three reasons that are worth understanding before spending money on one.
They are unreliable. Detection of machine-generated text rests on statistical properties of word choice that are not stable, and that shift with every model release and every round of human editing. A confident percentage from a detector is a far softer number than it appears.
They discriminate. Detectors disproportionately flag text written by people whose first language is not English, because non-native writing often shows the lower lexical variation that detectors read as machine-like. They also flag text produced with assistive technology, which disabled candidates may legitimately rely on. Screening on a detector score therefore imports a straightforward indirect discrimination risk into your process, on a protected characteristic basis, with no job-related justification.
They answer the wrong question. Here is the part most detector conversations miss. A candidate who did the work and used a language model to describe it more clearly has done nothing wrong. Plenty of excellent practitioners write poorly, and using a writing tool is no more objectionable than using a spellchecker or asking a friend to proofread.
The problem is not AI-authored text. The problem is unevidenced claims. Those are different things, and only one of them is your business. Once you frame it correctly, the solution follows: stop trying to determine how the document was produced, and start evaluating whether the claims inside it are supported.
The Anatomy of an Inflated CV
Inflated CVs share structural characteristics, and they are visible without any detector. They arise because a language model asked to make a candidate look suitable will produce a document optimised for apparent coverage rather than for verifiable substance.
Claims Without Specificity
This is the strongest single signal. Genuine experience carries incidental detail that nobody invents because nobody thinks to: the particular version that caused the problem, the awkward constraint imposed by a legacy system, the number of records that made the naive approach fail, the reason a sensible plan was abandoned in month three.
Inflated text is smooth and abstract. "Led the migration of critical systems to cloud infrastructure, improving performance and reliability" contains no information. Compare: "moved a 400GB reporting database off a failing on-premise server over two weekends, and rebuilt the overnight batch because the new latency broke the 6am deadline." One of these was written by someone who was there.
Uniform Register
Real CVs are uneven. People describe recent work in detail and older roles in summary; they write vividly about projects they enjoyed and flatly about ones they did not. Fully generated documents tend toward consistent tone and consistent density across a fifteen-year history, which is not how memory works.
Implausible Density
Count claimed competencies against elapsed time. Eleven technologies at expert level acquired in two years, in roles that would not plausibly have touched most of them, is an arithmetic problem rather than a judgement call.
Timeline Arithmetic That Does Not Survive Checking
Generated CVs frequently contain quiet inconsistencies: a qualification dated before the training began, overlapping full-time roles, a seniority jump with no intervening period, or total claimed experience in a specific tool that exceeds how long the tool has existed. That last one is a reliable catch and takes seconds.
Metrics That Are Suspiciously Tidy
Real measured outcomes are ugly: 37%, 18 months, £2.3m. Generated metrics cluster on round, pleasing numbers: 50% faster, 30% cost reduction, doubled efficiency. A CV where every achievement improved something by exactly a quarter or a half was not measuring anything.
Seniority and Outcome Mismatch
Claimed leadership scope should be consistent with the outcomes described. A candidate presenting as having led a department while every described achievement is individual task delivery is describing two different jobs, one of which is real.
Evidence-Based Skill Validation: Scoring Context, Not Frequency
The durable defence is a screening method that treats a skill mention as the weakest possible form of information and scores what surrounds it. In practice that means resolving every claimed capability to a position on an evidence hierarchy.
- Mention only. The term appears in a skills list. Evidential value: effectively nil. Free to produce, free to fabricate.
- Contextual mention. The term appears inside a role description, so at least it is anchored to a time and an employer.
- Applied context. The candidate describes doing something specific with it.
- Outcome-linked. The application is tied to a result, ideally one with a number attached.
- Scale and constraint. The description carries volume, complexity, team size or the specific difficulty that had to be worked around. This tier is very hard to fabricate convincingly because it requires knowing how the work actually behaves.
- Corroborated. The claim is consistent across the document, and consistent with the candidate's role chronology and seniority.
A candidate who lists a technology sits at tier one. A candidate who explains that they adopted it, hit a specific limitation at a specific scale, and what they did about it, sits at tier five. A keyword parser scores both as a match. An evidence-based approach scores them as what they are, which is not remotely the same candidate.
The commercial effect is direct: inflation becomes unrewarding. A candidate optimising a CV against an evidence-weighted screen cannot get far by adding keywords, because keywords carry almost no weight. They would have to invent detailed, internally consistent, technically coherent project narratives, which is substantially harder, much easier to unpick at interview, and beyond what a generic prompt produces.
Closing the Loop at Interview
Screening should hand the interviewer something more useful than a score. The most valuable output of evidence-based screening is a set of targeted probes derived from the weakest claims.
If a candidate claims a capability at tier one while the role requires genuine depth, the screen should say so and propose the question: ask them to describe the last time they used it, what went wrong, and how they resolved it. Unevidenced claims fail that question in under a minute, and genuine practitioners answer it enthusiastically because it is the interesting part of their job.
This is also the fair way to handle suspicion. Rather than rejecting on a hunch about how a document was written, you ask a job-related question and let the answer decide. That is defensible to the candidate, to your client, and to anyone reviewing your process later.
Frequently Asked Questions
Can You Reliably Detect Whether a CV Was Written by ChatGPT?
Not reliably, and it is the wrong objective. Detection tools produce unstable results and disproportionately flag non-native English writers and users of assistive technology, creating discrimination risk without job-related justification. Assess whether claims are evidenced instead of how the text was produced.
Should Candidates Be Banned from Using AI to Write Their CV?
A ban is unenforceable and largely misses the point. A candidate who did the work and used a tool to describe it clearly has not misrepresented anything. Policy effort is better spent on evidencing claims and probing them at interview than on policing composition.
What Is the Single Fastest Inflation Check?
Pick the two most important requirements for the role and look for concrete specificity behind each: scale, constraint, or something that went wrong. Generated text is smooth and abstract at exactly that point. It takes about thirty seconds per CV and catches most of it.
Does Evidence-Based Screening Disadvantage Candidates Who Write Badly?
Less than fluency-based screening does, which is the relevant comparison. Scoring concrete specifics rewards a plainly written account of real work over an elegant paragraph that says nothing. Candidates whose English is imperfect but whose experience is real do better under this model, not worse.
How Do You Handle a Candidate You Suspect Has Inflated Their CV?
Ask job-related questions derived from the specific claims in question and record the answers. Never reject on a detector score or a stylistic hunch. If the claims do not survive a competent conversation, you have an evidenced, defensible reason. If they do, you nearly discarded a good candidate for writing well.
The Shift Worth Making
Generative AI did not create dishonest candidates. It removed the disguise that used to protect fluency-based screening, and exposed how much of that screening was measuring writing ability rather than capability. The teams adapting fastest are not buying detectors. They are rebuilding their screens around evidence, which has the useful property of working regardless of what tool produced the document.
CVSense was built on evidence-based skill validation for exactly this reason: it evaluates the context and substance behind a claim rather than counting how often a keyword appears. If your team is losing hours to keyword-perfect profiles that evaporate on the first call, that is the specific problem it exists to solve.
Sources
Information Commissioner's Office. Guidance on AI and Data Protection.
https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/guidance-on-ai-and-data-protection/
Chartered Institute of Personnel and Development.
https://www.cipd.org/uk/
Equality Act 2010.
https://www.legislation.gov.uk/ukpga/2010/15/contents
Information Commissioner's Office. Rights Related to Automated Decision Making Including Profiling.
https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/individual-rights/individual-rights/rights-related-to-automated-decision-making-including-profiling/
InsightCircle



