Clean Data Isn’t a Project. It Is a Precondition.
Every conversation about AI in Business Central eventually arrives at the same place. The data. Here is why that conversation needs to happen first, not last.
When organisations begin evaluating AI for their Business Central environment, they typically start with the use case. What do we want the AI to do? Which workflows do we want to automate? Where is the biggest opportunity for the system to act on our behalf and return time to the team?
These are the right questions to be asking. They are also, in most cases, the second questions. The first question is one most organisations have not yet answered with enough specificity to move forward confidently.
What is the quality of the data the AI will be reading?
We have been thinking about this for long enough to have watched the pattern repeat across multiple client conversations. The enthusiasm for what AI can do is genuine. The readiness of the data environment it will operate in is consistently overestimated. And the gap between those two things is where AI deployments underperform.
Not because the AI is wrong. Because the AI is right about the wrong data.
What clean data actually means
Clean data is not the absence of errors. Every live business system accumulates some degree of data imperfection over time. Duplicate records that were never merged. Pricing structures that were updated in one entity but not another. Item descriptions that reflect a product catalogue from three years ago.
Clean data, in the context of AI in Business Central, means something more specific. It means that the data the AI will read to make a decision is accurate, current, and consistently structured across the entities and environments the AI will operate in.
For a pricing AI, clean data means that the customer’s pricing agreement is recorded accurately in BC, that it reflects the current commercial terms, and that it is consistent with what the sales team has on record in the CRM. A pricing AI reading a price list that is three months out of date is not making a pricing decision. It is making a historical one.
For a subscription management AI, clean data means that the contract terms recorded in BC reflect what the customer actually agreed to, that the renewal dates are accurate, and that the billing entity is correctly assigned. A subscription AI processing a renewal against incorrect contract terms does not produce a wrong output. It produces a precisely correct output of the wrong input.
For an access governance AI, clean data means that the user records in BC accurately reflect the current team, that former employees have been deactivated, and that role assignments reflect current responsibilities rather than historical ones. An access governance AI analysing an environment where twenty percent of the user records are outdated is not producing an access review. It is producing an access review with a twenty percent blind spot.
Why this is a precondition and not a project
The reason we describe clean data as a precondition rather than a project is that the distinction changes how organisations approach it.
A project has a timeline, a budget, a delivery date, and a defined scope. When data quality is treated as a project, it is planned, resourced, executed, and then considered done. The AI deployment follows. Two years later, the data quality has degraded back to its pre-project state because the underlying processes that generate the data have not changed, and nobody has maintained the work that was done.
A precondition is different. A precondition is a standard the environment must meet before the AI can be trusted to operate in it. And maintaining that standard is an ongoing operational responsibility, not a one-time project.
The practical implication is that before an AI deployment begins, the organisation needs to establish not just that the data is clean enough to start, but that the processes which generate the data will keep it clean enough to sustain. Who is responsible for data quality in the domains the AI will operate in? What is the process for identifying and correcting data errors when they occur? How frequently is the data reviewed? Who is accountable when the AI produces an output that traces back to a data quality failure?
These are operational governance questions. Answering them is the precondition work. It is less exciting than configuring the AI. It is more important.
What this means in practice for BC environments
For most mid-market BC environments, the precondition work involves three things.
The first is an audit of the data the AI will act on, focused on the specific domains relevant to the intended use cases. Not a full data quality assessment of the entire BC environment, but a targeted review of customer records, pricing structures, contract terms, user assignments, or whatever data the AI will be reading.
The second is a remediation pass. Duplicates resolved, outdated records corrected, inconsistencies between entities aligned. This is the one-time project component. It needs to happen before go-live.
The third is a governance design for ongoing data quality. Who owns the data in each domain? What is the process when a discrepancy is identified? What monitoring is in place to surface data quality issues before they reach the AI? This is the operational component. It needs to be in place before the AI goes live, and it needs to be maintained after.
The organisations that get this right treat data quality as part of the AI governance framework, not as a separate workstream that happened before the AI project started. The AI and the data are one system. Governing one without governing the other is not a complete governance programme.
The sequence that works
At Bluefort, the sequence we have found that works is: clean data first, connected data second, governed AI third.
Clean data means the records the AI reads are accurate and current. Connected data means the systems holding relevant data are synchronised in real time across the full Dynamics ecosystem. Governed AI means the AI operates within a policy framework the business defines, with a complete audit trail and the approval mechanisms to make that governance real.
Each step in that sequence depends on the one before it. Governed AI on unclean data is governance of the wrong decisions. Connected data feeding unclean records is a wider distribution of the same inaccuracies. Clean data without connection is an accurate but incomplete picture.
The sequence matters as much as each individual step.
Let’s chat further.
"*" indicates required fields
