Back to blogApprofondimenti

    Data governance for AI: roles, policies and measurable quality

    Data governance lived for years the thankless life of insurance: everyone agrees it is needed, few buy it until the damage arrives. AI changed its status in a couple of years, for a structural reason: a data error inside a traditional process produces an error, while inside a learning system it p...

    ZeroFive.AI March 10, 2026Updated on September 18, 2026 4 min

    Data governance lived for years the thankless life of insurance: everyone agrees it is needed, few buy it until the damage arrives. AI changed its status in a couple of years, for a structural reason: a data error inside a traditional process produces an error, while inside a learning system it produces a behaviour, which replicates across every subsequent decision and is discovered late. The European legislator drew the consequence, and Article 10 of the AI Act dedicates explicit requirements to the data of high-risk systems, covering governance, relevance, representativeness and bias management, turning into obligation what used to be good practice.

    The operational question, for anyone who must build or upgrade it, is what a data governance worthy of AI actually contains. The answer sits in three components, and in one design mistake to avoid.

    Names before committees

    The first component is responsibility, and the practical rule is that governance exists when every relevant data asset has two names next to it. The data owner answers for the data on the business side: authorises its uses, defines its rules, is accountable for its correctness towards the organisation. The data steward oversees it operationally: monitors quality, handles anomalies, maintains documentation. Without these two roles assigned in writing, every AI project restarts the negotiation from scratch, and the scene we see most often in assessments, a team spending weeks searching for "someone who can tell us whether this data can be used", is the exact symptom of their absence.

    The data governance committee, so beloved by frameworks, comes afterwards and exists to arbitrate conflicts between owners, never to replace them. A committee without underlying owners is a recurring meeting about nobody's data.

    Policies written for builders

    The second component is the rules, and the quality criterion is that they must be consultable by whoever is building at the moment of building, not only by audit after the fact. The policies that matter for AI answer project-shaped questions: which categories of data may feed which categories of systems, under which approvals; how the provenance (the lineage) of a training dataset is documented, so that a year later it is possible to reconstruct what a model learned from; which quality thresholds a dataset must pass before entering a pipeline; how retention and deletion are handled when data lives inside models and not only inside tables.

    On this last point, the GDPR and AI Act pair we covered elsewhere enters data governance as a design requirement, because a deletion requested by a data subject is manageable if lineage exists, nearly impossible if nobody knows which trainings that data passed through.

    Quality that gets measured, against a standard

    The third component is quality made measurable, and it is the one separating real governance from a statement of intent. The ISO 8000 standard offers the vocabulary: completeness, accuracy, consistency, timeliness, uniqueness, validity, each measurable with concrete metrics on the tables that matter. The maturity leap lies in moving from judgement ("the CRM data is good") to a number ("customer master completeness at 94%, duplicates at 3%, average refresh 36 hours"), because only the number allows thresholds, trends and owner accountability.

    For AI, measurement must be tuned to the use case, with a declared threshold below which the dataset does not enter the pipeline, and monitoring that continues in production, since data degrades and models degrade with it. It is the same principle as the critical gates we apply in the AI Rating: below certain conditions you do not compensate, you stop.

    The mistake of the total build

    The most widespread design mistake is conceiving governance as a total installation to complete before any AI, with years of catalogues, committees and universal policies. The total project almost always fails for the same reason as the grand data quality programmes, no visible return keeps it alive, and in the meantime it blocks everything else.

    The approach we see working proceeds by perimeters: start from the data domains the first AI use cases touch, assign owners and stewards there, write the policies that are genuinely needed, measure the six quality dimensions, and let the AI roadmap extend the governed perimeter one domain at a time. Governance thus grows at the speed of the value it enables, which is also the best way to defend its budget.

    In our AI Rating, data governance runs through Readiness and Risk, with precise evidence: roles assigned or vacant, policies existing or presumed, quality measured or perceived, and every gap enters the roadmap with a priority and an owner. If you wanted to apply the minimal test today, take the three data sources most critical to your next AI project and write the owner's name next to each: if one of the three lines stays blank, you know where to start. For the full version: calendly.com/fabiolalli/zerofive, or hello@zerofive.ai. Data governance, tested against AI, boils down to an almost civil-registry question: whose data, first and last name, are you about to build on?

    Want to discuss this for your company?

    30 minutes with us to figure out where to start, or an AI Rating to measure your starting point.

    #data governance AI#data governance#data owner data steward#ISO 8000 data quality#article 10 AI Act data
    Share

    Keep reading