Data readiness: why AI projects die on data
In the autopsies of failed AI projects, technology rarely appears among the causes, data almost always does. RAND Corporation (2024) places inadequate data foundations among the primary reasons behind the more than 80% of projects missing their intended value, MIT Project NANDA (The GenAI Divide,...
In the autopsies of failed AI projects, technology rarely appears among the causes, data almost always does. RAND Corporation (2024) places inadequate data foundations among the primary reasons behind the more than 80% of projects missing their intended value, MIT Project NANDA (The GenAI Divide, July 2025) identifies poor data readiness and broken workflow integration as the root of the 95% of GenAI pilots with no P&L impact, and Gartner (2025) formulated the sharpest prediction: through 2026, organisations will abandon 60% of AI projects not supported by AI-ready data.
The paradox is that data is also the most predictable failure cause of all, because unlike the sponsor who changes roles or the market that moves, its condition can be verified in full before spending the first euro. It takes wanting to do it, and knowing what to look at.
The illusion of the demo dataset
The mechanism by which data kills projects follows a recurring dynamic we described when writing about stalled pilots, and it deserves to be seen from the data side. The pilot runs on a curated extract: someone who knows a system well exported a table, cleaned it by hand, resolved the duplicates and filled the holes. On that material the model works, the demo convinces, the project accelerates towards production.
In production, though, the model must drink from the real source, and the real source is another matter: customer master records duplicated between CRM and ERP with keys that do not talk to each other, free-text fields filled ten different ways over ten years, histories interrupted by migrations, fresh data arriving days later than the process needs it. The manual work that made the demo possible cannot be repeated every night, and the project enters the phase that consumes budget and patience, the one where everyone discovers that the real construction site was never the model, it was everything upstream of it.
The questions that measure readiness
Data readiness lets itself be interrogated with concrete questions, and they should be asked for the specific use case, because data that is excellent for one application can be useless for another. The first concerns access: is the necessary data reachable programmatically and repeatedly, or does it live behind manual extractions, personal spreadsheets, suppliers to be begged? The second concerns measured quality, which is a different thing from perceived quality: do completeness, accuracy and consistency metrics exist on those tables, or does the judgement rest on the fact that "the ERP works"?
The third question concerns ownership: for every source, is there a person accountable for its quality who can authorise its use, or is the only expert a colleague who "knows that system"? The fourth concerns time, in its two forms, enough historical depth to train and enough freshness to serve the process. The fifth, often forgotten, concerns legitimacy: can that data be used for that purpose, on that legal basis, or is the project building on a processing operation nobody ever assessed?
A use case that answers all five well is rare, and that is fine: readiness does not demand perfection, it demands knowing the gaps before starting and putting them into the budget and the plan, instead of discovering them at the worst possible moment.
Fixing data without the alibi of the endless programme
Faced with a mediocre data landscape, organisations swing between two mirror-image mistakes. The first is starting anyway, trusting that "we will fix the data along the way", which is the exact script of the failures catalogued by the research above. The second is indefinite postponement, the great three-year data quality programme that must complete before any AI, and that usually starves to death because no business case keeps it alive.
The path we see working is surgical: you fix the data the chosen use case needs, in depth and with reuse criteria, and let the sequence of use cases progressively extend the perimeter of data in good order. Every AI project thus becomes an incremental investment in foundations too, with a visible return financing the next one, and data readiness stops being an abstract prerequisite to become a measurable by-product of the roadmap.
Measuring before betting
In our AI Rating the condition of data runs through the whole evaluation: it weighs on Readiness (quality, accessibility and ownership of sources), conditions Delivery (pipelines and integration) and touches Risk, where data oversight is one of the non-compensable critical gates, consistent with what the EU AI Act asks on data for high-risk systems and with quality standards such as ISO 8000. The score says, before the investment, whether the next pilot rests on rock or sand, and what consolidating costs.
If you have an AI project starting in 2026, the most profitable exercise of the week costs one hour: take the use case and ask it the five questions above, in writing, with source names next to each answer. If you prefer doing it with method and a comparable score, you will find us at calendly.com/fabiolalli/zerofive or hello@zerofive.ai. The uncomfortable version of the question, the one worth the entire budget, remains a single one: your next model, where will it drink from every night?