Resource
The tells expired. The checklist didn’t.
Every list of ways to spot AI writing is a snapshot of the models that existed when somebody wrote it down. The models have moved. Most of the lists have not.
The Economist tested this properly. It asked four frontier models to write versions of its own articles without web access, then compared 55,940 sentences and 1.2m words against its own prose, against CNN, the New York Times and the Washington Post, and against sixty years of bestselling fiction.
Half the checklist still circulating is stale. The single most-cited item on it now points the wrong way. And the two strongest remaining tells appear on almost no list at all.
Eight tells, and where the evidence lands
Reversed direction No longer holds Still holds
-
Inverted
Em-dashes
Only Claude now exceeds the human rate. ChatGPT uses markedly fewer em-dashes than any other writer in the study, human or machine. Its use spiked in 2025, then fell below the human baseline.
-
Expired
delve, tapestry
The overused vocabulary moved on. These two are no longer characteristic of any model in the study.
-
Live
Long, Latinate words
Eight letters or more. All four models sit well above every human baseline, Gemini and Claude highest. Rare words, scientific register, and nouns made out of verbs.
-
Live
Scientific vocabulary
Parameter, methodology. Every model ranges above The Economist, general news and fiction.
-
Live · uneven
“Not X but Y”
ChatGPT and Claude are the outliers. Gemini sits close to human news writing, so the tell identifies a model rather than a machine.
-
Live · uneven
The rule of three
ChatGPT far above everything else, including the other models. The rest cluster just above the human range.
-
Live · missed
Sparse punctuation
Fewer commas and semicolons, hardly any parentheses, longer sentences. A stronger signal than any banned word, and almost nobody is looking for it.
-
Live · missed
No quoted experts
Models do not quote people. On a page of reported prose this is the most structural tell available, and it appears on no checklist.
The em-dash is the one to watch, because it is the item everybody repeats. It was a fair tell in 2024. ChatGPT’s use of it climbed through 2025, then fell below the human baseline entirely. Anyone still treating a dash as evidence is now reading the signal backwards, and doing it with more confidence than they had when it was true.
There is a deeper problem underneath the individual items. Every model release moves AI prose closer to human prose. The tells are not being defeated by anyone. They are being trained out, release by release, by the same human feedback that produced them — models learn what people find impressive and drop what they do not. A list of tells has a shelf life whether or not anyone writes an expiry date on it.
Three questions before detection becomes a control
None of this matters much while detection is a parlour game. It matters a great deal where a detector already sits inside a decision — hiring screens, admissions, procurement, vendor attestation. At that point it is a control, and it is the only control I know of whose most-cited signal reversed direction inside twelve months.
- 01
What does a positive result authorise?
Name the decision it triggers, not the tool that produced it.
- 02
Who re-validates it after the next model release?
A named person with a date, not a policy review cycle.
- 03
What happens to the people it flags wrongly?
Detectors are black boxes. They give no reason for the call.
Every item on that checklist was true when someone wrote it down. Nothing on it has been re-checked since.
Source: The Economist, “How to spot AI writing”, 30 July 2026 — 55,940 sentences and 1.2m words of output from ChatGPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash and Grok 4.5, benchmarked against The Economist, an average of the New York Times, Washington Post and CNN (2018–2022), and NYT bestselling fiction (1950–2022). Models were prompted without web access. Detection vendor Pangram claims 99.98% accuracy; The Economist notes detectors are black-box and can return false positives without giving reasons. The status labels above are this page’s reading of the published findings, not the newspaper’s framing.
Related
The agent halt matrix — the same failure in a different control. Capability checklists grade agents on one axis and leave off the one that decides who is accountable.
What each AI tier actually decides — regulatory exposure tracked against the consequence of the decision, with the reversibility column most charts leave off.
The Governance Memo carries this work monthly for boards and CISOs — one breach post-mortem and two or three governance items.