FL Fredrik Lindstrom


Every list of ways to spot AI writing is a snapshot of the models that existed when somebody wrote it down. The models have moved. Most of the lists have not.

The Economist tested this properly. It asked four frontier models to write versions of its own articles without web access, then compared 55,940 sentences and 1.2m words against its own prose, against CNN, the New York Times and the Washington Post, and against sixty years of bestselling fiction.

Half the checklist still circulating is stale. The single most-cited item on it now points the wrong way. And the two strongest remaining tells appear on almost no list at all.

A status table of eight claimed tells for AI-written text. Em-dashes are marked inverted; delve and tapestry are marked expired; long Latinate words and scientific vocabulary are live; the not-X-but-Y construction and the rule of three are live but uneven across models; sparse punctuation and the absence of quoted experts are live but missing from most checklists.
Download the status board PDF, text selectable, prints for board packs. No email required.

Eight tells, and where the evidence lands

Reversed direction No longer holds Still holds

  1. Inverted

    Em-dashes

    Only Claude now exceeds the human rate. ChatGPT uses markedly fewer em-dashes than any other writer in the study, human or machine. Its use spiked in 2025, then fell below the human baseline.

  2. Expired

    delve, tapestry

    The overused vocabulary moved on. These two are no longer characteristic of any model in the study.

  3. Live

    Long, Latinate words

    Eight letters or more. All four models sit well above every human baseline, Gemini and Claude highest. Rare words, scientific register, and nouns made out of verbs.

  4. Live

    Scientific vocabulary

    Parameter, methodology. Every model ranges above The Economist, general news and fiction.

  5. Live · uneven

    “Not X but Y”

    ChatGPT and Claude are the outliers. Gemini sits close to human news writing, so the tell identifies a model rather than a machine.

  6. Live · uneven

    The rule of three

    ChatGPT far above everything else, including the other models. The rest cluster just above the human range.

  7. Live · missed

    Sparse punctuation

    Fewer commas and semicolons, hardly any parentheses, longer sentences. A stronger signal than any banned word, and almost nobody is looking for it.

  8. Live · missed

    No quoted experts

    Models do not quote people. On a page of reported prose this is the most structural tell available, and it appears on no checklist.

The em-dash is the one to watch, because it is the item everybody repeats. It was a fair tell in 2024. ChatGPT’s use of it climbed through 2025, then fell below the human baseline entirely. Anyone still treating a dash as evidence is now reading the signal backwards, and doing it with more confidence than they had when it was true.

There is a deeper problem underneath the individual items. Every model release moves AI prose closer to human prose. The tells are not being defeated by anyone. They are being trained out, release by release, by the same human feedback that produced them — models learn what people find impressive and drop what they do not. A list of tells has a shelf life whether or not anyone writes an expiry date on it.


Three questions before detection becomes a control

None of this matters much while detection is a parlour game. It matters a great deal where a detector already sits inside a decision — hiring screens, admissions, procurement, vendor attestation. At that point it is a control, and it is the only control I know of whose most-cited signal reversed direction inside twelve months.

  1. 01

    What does a positive result authorise?

    Name the decision it triggers, not the tool that produced it.

  2. 02

    Who re-validates it after the next model release?

    A named person with a date, not a policy review cycle.

  3. 03

    What happens to the people it flags wrongly?

    Detectors are black boxes. They give no reason for the call.

Every item on that checklist was true when someone wrote it down. Nothing on it has been re-checked since.

Source: The Economist, “How to spot AI writing”, 30 July 2026 — 55,940 sentences and 1.2m words of output from ChatGPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash and Grok 4.5, benchmarked against The Economist, an average of the New York Times, Washington Post and CNN (2018–2022), and NYT bestselling fiction (1950–2022). Models were prompted without web access. Detection vendor Pangram claims 99.98% accuracy; The Economist notes detectors are black-box and can return false positives without giving reasons. The status labels above are this page’s reading of the published findings, not the newspaper’s framing.


Related

The agent halt matrix — the same failure in a different control. Capability checklists grade agents on one axis and leave off the one that decides who is accountable.

What each AI tier actually decides — regulatory exposure tracked against the consequence of the decision, with the reversibility column most charts leave off.

The Governance Memo carries this work monthly for boards and CISOs — one breach post-mortem and two or three governance items.