
By Fredrik Lindstrom · ~19 minute read · August 2026
Data Quality Is the Output, Not the Gate
Why the AI-agent readiness checklist has the sequence backwards — and what a bounded agent tells you that no assessment can.
The graphic runs in a dozen variations, and almost all of them borrow from civil engineering. A tower rising on cracked ground, upper floors already leaning. A house with a fissure running from the basement to the roofline, the basement labeled data quality. A bridge halted mid-span, the far pier never poured. A pyramid with AI at the apex and clean data as the bottom course, because a pyramid cannot be built from the top. And the version I have seen most this month: two queues at a trade show, one crowded in front of a booth selling agents, one nearly empty in front of the booth selling the things that make agents work.
Different picture, identical caption. Clean data, connected systems, clear workflows, governance, trained people. Build those five first. Then deploy.
Every item on that list is correct. The metaphor is doing the arguing, and each of the structural versions borrows the same property from masonry: a foundation must be complete before anything rests on it, because concrete cannot survey the ground it was poured on. An agent can.
Which turns the picture inside out. Perfect all five and you have written the specification for deterministic automation — RPA will do that work in a heartbeat, at volume it is cheaper per transaction, and there is no inference cost and nothing probabilistic to govern. That is a readiness checklist for RPA. Somebody put an AI label on it.
Almost no enterprise has clean data and clear workflows across the whole estate. Plenty have one well-governed domain sitting next to a supplier master nobody owns, which is the normal condition and not a failure of effort. That unevenness is where AI earns its cost.
The hole in the advice
Start with what the checklist is asking for. One authoritative source instead of five spreadsheets with five different answers. Working connections between systems. A redesigned process rather than an automated broken one. Someone who owns each agent and decides what it may touch. People who know when to step in.
Now read it as a precondition. One authoritative source and a defined decision boundary describe a system predictable enough to be automated by rule, and rules are cheaper than inference at volume.
Deterministic is not the same as governed, and it is worth saying so before I lean on the comparison. An unattended bot running on a shared service account produces a flawless record of what happened with no name attached to it. That is a log, not accountability, and it is why attributing an automated action to an identity is its own control rather than a property you get free from determinism.
That is the circularity, and it is the reason these programs stall. The readiness assessment becomes an eighteen-month workstream that produces a maturity score and no working capability, and by month twelve the business has bought its own agents anyway.

And the target moves while you chase it. The study this genre reaches for most often is the Friday Afternoon Measurement: seventy-five managers each scoring the last hundred records their own department completed, collected in Ireland between 2014 and 2016, published in Harvard Business Review in 2017. It is self-scored, it is pre-LLM, and I would not hang a business case on any decimal in it. What it does show, and what a loose sample cannot manufacture, is that 47% of newly created records carried at least one critical, work-impacting error. The error enters at the point of creation, inside the processes the checklist told you to redesign first. That makes it a flow rate rather than a backlog, and a backlog is the only thing a cleanup program can finish.
The shape is more useful than the headline. Those seventy-five measurements do not form a bell curve. The largest single group scored between 71% and 80% error-free records, and the second-largest scored between 0% and 10%. A quarter of the scores fell below 30%, half below 57%. Some departments are in decent shape and some are in a state nobody has looked at in years, and the average describes neither.
A gate implies a moment when you are through it. At that creation rate the moment never arrives. What you have instead is a program with a burn rate and no terminal date.
Gartner’s Data Quality Market Survey reports that nearly 60% of organizations do not measure the annual financial cost of poor data quality at all — which means the fifteen-million-dollar average it publishes alongside that finding is an average of estimates from the minority who believe they can produce one. Confidence in the asset inventory was always inversely proportional to whether anyone had run a scan.

The business does not wait
The people making the readiness-first argument are mostly data leaders, and the parallel runs further than they would like. Security spent most of the last decade as the department of no, and that era did not end with security winning the argument. It ended with the business routing around it. Shadow IT was never a failure of policy enforcement. It was the predictable result of a function that had made itself a gate. The security leaders who came out of it with their authority intact were the ones who moved from approval to instrumentation, putting controls in the path of the work rather than a committee in front of it. A data office positioning itself as the precondition for AI is running that experiment again, and we know the outcome.
One more thing, offered from the side that has already served the sentence. Cybersecurity spent most of a decade earning the title of the department of no, and the rest of it climbing out. Nobody in security is sorry to see the label move somewhere else, and if the data office wants it, it is available. The terms are well documented, because we wrote them. You get consulted late. You get told after the decision. You get handed the incident when it lands. And you get asked why you did not stop it.
And there is a cost the readiness argument never books, because it does not appear on any project plan. Holding up progress in the name of data first is the fastest way to lose relevance. Not lose money — lose relevance. The competitor who deployed narrowly eighteen months ago is not ahead because their data was cleaner. It is ahead because it now knows things about its own operation that you cannot know yet, has staff who have handled a thousand escalations, and has an owner who has learned where the boundary actually needs to sit. None of that was purchasable at the start and none of it can be caught up by procurement. Meanwhile the function that held the gate discovers what security discovered: the organization stopped asking.
What the counting rule does to the answer
Here is the part that cuts against my argument. Real distributions do exist where a regulator compelled the measurement. NHS England publishes a machine-computed data quality score for every healthcare provider in England, monthly, under a statutory duty. Across the 206 NHS trusts, every one an organization well above a thousand employees, the median is 91.9% and the lower quartile 88.9%. That is managed operational error, not chaos. So which is it, 53% or 92%? Both. The Friday Afternoon Measurement is record-level — one bad attribute out of ten to fifteen kills the record, judged by eye. The NHS index is field-level, roughly 180 fields each checked against a published rule by machine. An 8% field-level failure rate compounded across twelve fields lands in the low sixties, which is the Friday Afternoon answer. The numbers are not in conflict. The counting rules are.
Which produces the sentence I would put in front of a board. There is no such thing as the percentage of your data that is clean. That number is a property of your measurement rule, and anyone who quotes you one without naming the rule has told you nothing. Ask which rule, who applies it, and when it last ran. It holds at the top of the market too: of the 31 global systemically important banks, assessed by their own supervisors against the Basel Committee’s principles for risk data aggregation rather than by self-assessment, only two were fully compliant with all eleven, and no single principle had been fully implemented across all of them. The European Central Bank calls progress on the same ground slow and insufficient. Those are the most scrutinized data estates in the economy, and they are the ceiling rather than the average.
The final 20% comes from deployment
In my experience, readiness has never once arrived before the thing that required it.
Cybersecurity discovers before it enforces. CIS Control 1 is inventory, and every competent network access control rollout runs months in monitor mode before a single port is enforced. What discovery never does is finish. Passive scanning finds what answers. It does not find the service account nobody would admit to owning, or the shadow SaaS tenant paid for on somebody’s card. Those surface the moment something starts enforcing policy, and they announce themselves by breaking. Monitor mode is the point rather than the exception: you put the instrument into the live network first and let the network tell you what is on it. In my experience the final 20% of an asset inventory has never come out of an asset-management project. It came out of enforcement.
That is one practitioner’s pattern and not a study, offered as a prior rather than a proof. If agents turn out not to behave like enforcement points, delete this section — the argument still rests on the circularity, which needs no analogy at all.
I wrote the underlying version of this in June 2016, about information security being treated as a department: “Considering that technology is at best one third of the equation in people, process, and technology, that leaves two thirds of the solution without proper support or investment.” The shape has not changed. The technology term was never the constraint. What the deployment exposes is the other two thirds.
Agents are the same instrument, pointed at data and decision rights instead of endpoints. A narrowly scoped agent running against a real process finds the conflicting sources, the undefined authority, the approval nobody can name an owner for, and the step everyone thought was documented — faster and more honestly than any readiness questionnaire, because it is testing behavior rather than asking people to describe it.
Which is where the bridge picture inverts on itself. A half-built span tells you nothing about the far bank. An agent working a live process is the survey.
Build the agent to escalate
This is where the refutation has to become a design instruction, or it is just contrarianism.
An agent that meets two conflicting sources has two available behaviors. It can pick one, which is the failure the checklist team is right to fear. Or it can stop, state that the conflict sits outside its authority, and escalate. The second behavior is not emergent. You get it only if someone specified where the authority ends, which means the undefined-decision problem does not disappear. It moves from a thousand silent per-transaction guesses to one explicit boundary set at design time.
One condition decides whether that trade is real. A boundary written into a prompt is not a boundary — it is one more input the model weighs, and “is this inside my authority?” becomes another per-transaction probabilistic judgment carrying the failure signature I just described: confident, plausible, silent. The relocation only happens when the limit is enforced where the model cannot reach it. A tool it cannot call. A write it cannot make. A reconciliation check that hard-fails when two sources disagree, in code, before the agent is asked what it thinks. Singapore’s IMDA framework states the preference plainly, favoring structural rule-based controls over prompt-layer ones, and that is the difference between a control and a suggestion.
Specify it as an allowlist rather than a list of conflicts to watch for. You enumerate what the agent may settle on its own, and everything outside that halts by default. Enumerating the conflicts you can foresee only ever catches the conflicts you can foresee, and the entire value of the exercise is the decision nobody anticipated becoming visible the first time it occurs.
That is the trade worth taking, and it is the whole of what governance needs from you in order to start. Not five readiness bands in front of the program. A specified halt boundary and a named owner with the authority and budget to answer what comes over it. Then a record of how each escalated decision was resolved. Discovery, pilot, production, continuously — the same column at every stage. The rest of the obligation set is real: technical documentation, logging, post-market monitoring, incident reporting. None of it is a reason to wait, because none of it gets easier while you score yourself.
None of this is exotic, and it is worth naming where it already sits. NIST’s AI Risk Management Framework puts accountability structures in GOVERN 2, inside the only function that operates at the organization layer rather than the system layer. ISO/IEC 42001:2023 requires roles, responsibilities and authorities under Clause 5.3 and internal organization under Annex A.3, inside a management system defined by Clauses 4 to 10 — a system that operates, not a policy document describing one. The EU AI Act requires human oversight by design for high-risk systems under Article 14, and Article 26(2) obliges the deployer to assign that oversight to natural persons with the competence, training, authority and support to exercise it. All three describe a named person with real authority over a running system. None of them describe a readiness score.
The owner specification carries three elements and the third is the one that gets dropped. A scoped veto over launch. A ring-fenced budget line, so answering the queue is funded rather than donated. And an unfiltered route to the board risk committee, because nobody running an agent overrules the CFO and the Chief Commercial Officer about which revenue figure is real. If no name is attached to that escalation, it ages while the process stays halted. So the route upward carries a clock: an escalation the owner cannot answer inside five working days goes to the risk committee unfiltered, or you have not bought governance. You have bought an outage with a governance label on it.
That is a gate on scale, not a gate in front of the first deployment, and the distinction is the whole argument. Nothing here contradicts the governance-before-scale rule in my own toolkit. You still start narrow, build the capability the queue reveals you need, and widen deliberately.

Which is the argument I made in Columns Not Layers in August 2026, arrived at from the other direction. That piece starts from the frameworks and shows why governance drawn as a horizontal tier has no authority in it. This one starts from a trade-show graphic and arrives at the same design. Two routes by the same author prove nothing on their own; what they show is that the conclusion does not depend on where you enter. An agent that halts makes the shape legible: the halt is a governance event, the queue is a human-capability problem, and the conflict that triggered it belongs to whichever column produced the record — governance where your own systems generated it, supply chain where you bought the feed.
One honest limit, and it is narrower than it first appears. An agent can be given what good looks like — internal consistency and expected distributions, read against a working picture of the world nobody had to encode in advance. That is what catches an invoice for consulting services from a vendor whose registered activity is scaffolding, or a unit price that is perfectly ordinary in general and absurd for that commodity, in that country, in that season. A rules engine catches only what somebody thought to configure and kept maintained. That is the actual difference, and it is a difference in coverage rather than in cleverness.
What it cannot catch is the plausible error: the stale price that is still a normal price, the valid address for a customer who moved. Those are visible only against reality. And every turn of the sensitivity dial adds queue volume, so the tuning is itself a decision someone has to own. Treat the escalation log as the best inventory you have ever had, and not a complete one — it is a map drawn by behavior and bounded by the halt you specified, so widen the boundary and the map redraws.
What an agent absorbs, and what it cannot
Two kinds of mess get flattened into one word on that graphic, and agents only dissolve one of them.
The first kind is interpretive. Inconsistent formats. Renamed fields. Unstructured documents. Ambiguous requests. Exceptions that used to fall into a human queue because no rule covered them. This is where reasoning models earn their cost. RPA snapped when a column got renamed; an agent reads around it. That is real uplift, and it is the part the checklist genre quietly concedes when it admits current models are strong enough today.
The second kind is different. Five spreadsheets with five different revenue figures is the record of a decision nobody made. No amount of reasoning tells the model which number is real, because the answer is not in the data at all. Same with permissions: what an agent is allowed to touch is not a fact waiting to be inferred. A well designed agent escalates to a human here rather than choosing.
So the line runs here: reasoning resolves ambiguity; it does not resolve unmade decisions. Where the mess is interpretive, the agent bridges it. Where the organization never decided who owns the number, who can halt the process, or what happens after a bad call, an unbounded agent does not bridge anything. It fills in, confidently, and returns something plausible.
That failure mode is worse than the one it replaced. RPA broke loudly. An agent guessing at an unmade decision looks exactly like an agent doing its job.

Match the capability to the data
So far this has been an argument about sequence. There is a second decision underneath it, and the readiness framing hides that one too, by treating AI as a single thing you either deploy or do not.
It is at least five things. At one end, data that is schema-stable and validated on entry, where the answer is already in the field. Rules handle that, and rules are cheaper per transaction at volume with nothing probabilistic to govern. A step along, data that is structurally sound but semantically inconsistent, where fields drift, duplicates survive and the same customer appears four ways. Rules still work until they meet the fifth way, and classical machine learning earns its training cost at the edges.
Then the ground the readiness checklist tells you to eliminate. Records that are readable once something reads them: classify, extract, route, normalize, each item easy on its own and none of them structured. Small fast models do that work at a fraction of frontier cost. Above it, material that needs reading rather than parsing: contracts, call notes, tickets, where the answer depends on context across a whole document. That is mid-tier workhorse territory. And at the far end, the exception no rule covered and no pattern fits, together with the analysis nobody can specify in advance, where a frontier model chooses the method and reads the result rather than reading a billion rows itself. High consequence, and the reasoning has to be inspectable.
The failure runs in both directions and only one of them gets discussed. Sending an unstructured contract to a rules engine is the error everyone recognizes. Sending clean, schema-stable data to a frontier model is the same error inverted, and it is the more expensive one, because it looks like ambition rather than negligence. Nobody is criticized for over-specifying. That asymmetry is what makes the top tier the default, and it is the same asymmetry that makes a readiness program feel politically safe while a narrow deployment feels risky.
One caution about reading this as a ladder of volume. It is not one. Volume is a modifier on every tier rather than a position on any of them, and the axis is how readable a single item is without a human. A tier-three workload can run to millions of items a day and a tier-five workload to forty a month.

Where this breaks in week six
Set the boundary conservatively, as you should — which means a rule rather than a feeling: anything touching capital, safety or reputation halts, and anything irreversible halts until something other than optimism has moved it. Do that and the first month produces an escalation queue nobody staffed for. Humans start clearing it at a glance. Within six weeks you have approval theater with an audit trail, which is worse than no escalation at all, because now the rubber-stamping is documented.
This is where the frameworks stop being abstract. Article 14 of the EU AI Act is not satisfied by a person holding a title; it requires that oversight can be exercised, and Article 26(2) puts the authority and support to exercise it on the deployer. A four-second review satisfies neither. NIST’s AI RMF locates this in MEASURE and MANAGE, and the point of both is a metric someone acts on rather than a number someone reports. ISO/IEC 42001 Clause 9.1 asks what you monitor, how, and by whom — and an unread queue is precisely what fails to evidence that the management system operates. And where the queued decisions produce legal or similarly significant effects for a person, GDPR Article 22 makes that log the finding rather than the defense: it documents that the human involvement was decorative. A log showing that the human involvement was decorative is precisely what an authority looks for when it asks whether a decision was meaningfully reviewed.
This is what the AI Oversight Override Rate is for, and it is the number most often read backwards. A near-zero override rate on a busy agent is not a success signal, and it is not proof of failure either. On its own it is unreadable. It looks identical for a well-calibrated agent working inside a narrow scope and for a queue nobody is reading, and no threshold separates them.
What separates them is a second signal, and there are two worth having. First, evidence that the owner has exercised the authority they hold: a hold placed, a budget refused, an escalation raised past the sponsor, in the last quarter. Attendance at a governance forum does not count. Second, ten overrides you can tie to a consequence outside the log — an output that shipped differently, a release that moved. Zero overrides against a live owner is plausibly a quiet, well-scoped deployment; sample it and confirm. Zero overrides against an owner who has exercised nothing is the pairing that should stop a board meeting.
Report the rate against escalation volume and against the severity of what the overrides changed. Volume alone catches suppression, which shows up as a quiet queue with a disciplined-looking rate. It does not catch the other direction, because a four-second review that makes a one-word edit is an override, and a flooded queue full of them reads as vigilance on both numbers. Only the severity log separates a reviewer who changed an outcome from one who changed a word.
What it costs
One boundary this argument does not cross. Where the law puts a gate in front of you, the gate is not a maturity band and nothing here moves it. A high-risk system under the EU AI Act carries a risk management system under Article 9, data governance under Article 10, technical documentation under Article 11 and Annex IV, and conformity assessment under Article 43, all before it goes to market. Processing likely to result in high risk to individuals carries a data protection impact assessment under GDPR Article 35, before the processing starts. Those are conditions of acting, and an assessment produced afterward is evidence of a breach rather than a late receipt. Narrow scope is not an exit from those categories — they classify on what the system is used for, not on how ambitious you were, and the assessment that you sit outside them is itself a document you write before you deploy. What I am arguing against is the voluntary five-band readiness program that no regulator asked for.
That argument fails on its own terms. It asks you to reach a state that would obviate the tool, and while you work toward that state the shadow deployments accumulate without any of the controls the checklist was arguing for. The voluntary gate does not hold. It never has in any technology cycle I have worked through.
Deploy narrowly. Specify the halt. Name the owner. Then read what comes back, because the escalation log is the first honest map of your data and your decision rights that anyone in the organization has ever produced.
So the next time the cracked tower comes across your feed, or the bridge stopped mid-span, and the caption tells you to fix the foundation before you deploy — ask the question the picture is built to prevent. If my data were clean and my workflows were clear, why would I need AI here at all?
Organizations that ask why they need AI when this graphic comes across their desk will be running agents in production, bounded, instrumented, and governed by a named owner. The others will have a data maturity score, an expensive slide deck, and agents running anyway — implemented by the business, governed by no one.
Sources
Friday Afternoon Measurement — Tadhg Nagle, Thomas C. Redman and David Sammon, “Only 3% of Companies’ Data Meets Basic Quality Standards,” Harvard Business Review, 11 September 2017; peer-reviewed as “Assessing data quality: a managerial call to action,” Business Horizons 63(3), 2020. The 47% error rate and the distribution quoted here — the bimodal shape, the quarter below 30% and the half below 57% — belong to the original seventy-five measurements, collected in Ireland between 2014 and 2016 from managers scoring their own departments’ records. Self-scored, not a random sample of enterprises, and entirely pre-LLM. The 195-measurement extension published by Redman in January 2020 does not carry those percentiles.
NHS England Data Quality Maturity Index — published monthly under section 266 of the Health and Social Care Act 2012. Trust-level percentiles here are computed from the May 2025 raw file and replicated against the November 2024 file. The index measures completeness and validity against a published national standard, not whether a valid value is the correct one, so it is a floor on error rather than a ceiling. One month, one country, one sector.
Basel Committee on Banking Supervision — Progress in adopting the Principles for effective risk data aggregation and risk reporting, d559, November 2023, on data as at June 2022. Supervisor-assessed rather than self-assessed, across 31 global systemically important banks. European Central Bank supervisory assessment of the same ground: Aggregated results of SREP 2025, 105 significant institutions.
Gartner — “How to Stop Data Quality Undermining Your Business,” 18 January 2018, reporting its Data Quality Market Survey. Gartner has never published the sample size or the fieldwork dates for that survey, which is why the 60% who do not measure is the usable half of the finding and the dollar average is not.
The framework behind this piece
Columns Not Layers — why AI governance frameworks specify the human and never staff one, and why governance drawn as a horizontal tier carries no authority. Includes the Column Test.
The Decision-Rights Register names which decisions an agent may settle and which stay human, including the one this article leads with: declaring which of several conflicting records is the record of truth.