Resource
Not all bad data is the same kind of bad.
“Fix the data first” treats one word as one problem. It is at least four, they have different handlers, and one of them cannot be handled where the AI runs at all.
The readiness checklist prices all four the same and resolves none of them. A reasoning model genuinely absorbs the first — inconsistent formats, renamed fields, the exception no rule covered — and that is where it earns its cost against a rules engine that snapped whenever a column was renamed. The second, the implausibly wrong record, it catches by reading internal consistency and expected distributions against a working picture of the world nobody had to encode in advance.
The third is the honest limit, and it is worth being blunt about it. A stale price that is still a normal price is internally consistent, distributionally ordinary and wrong. No model surfaces it, no boundary surfaces it, and neither does a human reviewer — they see a normal price and a valid address. There is no detection answer to this one. There is a process answer, and it sits upstream of everything else on the sheet: make the customer confirm their details at login every six months, flag any price nobody has touched for twelve. Freshness as a control, rather than accuracy as a target.
The fourth is not a data problem at all. Five spreadsheets carrying five revenue figures is the record of a decision nobody made, and the answer is not in the data to be reasoned out. A bounded agent halts and escalates. An unbounded one fills in — confident, plausible, silent — which is a worse failure than the one it replaced, because RPA at least broke loudly.
The four kinds, and what handles each
-
Absorbed
Interpretive mess
The model reads around it — RPA snapped when a column got renamed. Lever: model selection — match capability to how readable one item is without a human, not to how many there are.
-
Caught
Implausible error
Caught by reading internal consistency and expected distributions against a working picture of the world. A rules engine catches only what somebody configured — a difference in coverage, not cleverness. Lever: the halt boundary and its sensitivity dial, which someone owns, because every turn of it adds queue volume.
-
Process redesign required
Plausible error
Nothing catches this — and a human reviewer misses it too, because they see a normal price and a valid address. Internally consistent, distributionally ordinary, and wrong. No detection lever exists. The only lever is upstream, in the process that creates the record: confirm-your-details at login every six months, a flag on any price untouched for twelve. Freshness as a control, not accuracy as a target.
-
A human decides
Unmade decisions
Not a data problem — the answer is not in the data at all. A bounded agent halts and escalates. An unbounded one fills in: confident, plausible, silent. Lever: a named owner with a scoped veto, a funded budget line and an unfiltered route upward.
None of these levers cleans a data estate, and the sheet does not claim they do. Model selection matches capability to how readable a single item is without a human. The halt boundary relocates a thousand silent per-transaction guesses into one explicit limit set at design time — enforced in code, because a boundary written into a prompt is one more input the model weighs. Escalation surfaces the decision nobody made and puts it in front of someone with the authority to make it. Three of the four route the mess to the thing that can act on it. The fourth has no detection answer at all, and its only lever runs upstream — periodic re-confirmation and staleness flags, which prevent the bad record rather than catch it. That last lever extends the argument in the source article rather than being drawn from it. All of which is worth knowing before you budget for eighteen months of remediation.
Three things to settle before it runs
- 01
Which of the four is actually in front of you? Three can be handled where the agent runs. One can only be handled upstream, in the process that creates the record.
- 02
Where is the boundary enforced — in code, or in a prompt? Specify it as an allowlist of what the agent may settle alone, with everything outside it halting by default.
- 03
Who answers the queue, and by when? Name them, fund the line, and put a clock on the route up: unanswered inside five working days, it goes to the risk committee unfiltered.
Reasoning resolves ambiguity. It does not resolve unmade decisions.
Source: Fredrik Lindstrom, AI Doesn’t Need Perfect Data, 21 August 2026 — sections “What an agent absorbs, and what it cannot” and “Build the agent to escalate”. The four-way split is this sheet’s framing of that argument and is not a taxonomy drawn from a cited study. Structural over prompt-layer: the preference for rule-based controls enforced outside the model is stated in Singapore IMDA’s Model AI Governance Framework for Agentic AI. Human oversight: EU AI Act Article 14 requires that oversight can be exercised, and Article 26(2) puts the competence, training, authority and support to exercise it on the deployer — a four-second review satisfies neither. The upstream lever on the plausible-error row — periodic re-confirmation at login, staleness flags on unmaintained fields — is this sheet’s addition and does not appear in the source article. None of the levers above cleans a data estate; three route each kind of wrong to the thing that can act on it, and the fourth prevents the record rather than catching it. Row groupings and verdict labels above are this sheet’s reading of the cited source, not the source’s own framing.
Related
The agent halt matrix — the boundary this sheet says to enforce, drawn against the capability axis buyer checklists grade on.
Governance is not the bottom brick — the corrective on the picture this argument is answering — data drawn as a foundation you finish before AI rests on it.
The Governance Memo carries this work monthly for boards and CISOs — one breach post-mortem and two or three governance items.