
The screen looked impressive. Predictive maintenance alerts, demand curves refreshing in near real time, a confidence score printed beside every recommendation, all of it in front of a leadership team that had spent eleven months and a seven figure budget getting there.
Two floors down, the production manager opened the third alert of the week on an asset he had personally rebuilt in March. He ignored it. He had ignored the two before it as well.
That is how most AI and ERP programs in manufacturing actually end. Not in a cancellation meeting. In quiet non-use.
The model was competent. The data underneath it was not, and everyone close to the work knew that long before the board did.
What sits beneath the model is a decade of inconsistent part numbering, bill of material revisions nobody controlled, duplicate vendor records, and inventory balances the plant reconciles by hand every Friday afternoon. ERP systems expose operational maturity. They don’t create it, and neither does an AI layer bolted on top.
Failure here is rarely a single event. It arrives as four separate costs.
The direct write-off is the easy number. Licenses, integration work, contractors, plus the internal hours of capable people pulled off operational improvement to sit in workshops. For a mid-sized manufacturer, a pilot that produces nothing usable typically consumes six to eighteen months of scarce technical capacity, which is what Ultra observes in the field rather than a figure lifted from a survey.
Credibility costs more. A CFO who stood up in front of the board and committed to AI driven forecasting accuracy now has to explain why the plant still schedules from a spreadsheet on a shared drive. That explanation is rarely forgiven quickly.
It changes how the next technology proposal gets received, however sound that proposal happens to be.
Organizational fatigue never makes the budget. Every failed initiative teaches the people doing the work that corporate technology arrives, disrupts, and disappoints. When a necessary program comes along later, a system replacement or a business process improvement effort, resistance starts higher.
The model learns your errors. A model trained on unreliable data doesn’t simply fail to help. It produces confident, specific, wrong recommendations, and a few of them get acted on before anyone catches the error.
Bad decisions made faster are still bad decisions.

Executives evaluating AI rarely see the data as it actually exists. They see a curated extract prepared for a demonstration. Production looks nothing like that.
BOM accuracy is the most common failure point, by a wide margin. Revisions get applied in one plant and not the other.
Engineering changes are approved in PLM and reach ERP weeks later, sometimes by manual keying. Phantom assemblies exist to make a routing work rather than to describe how anything is actually built.
We watched a data scientist at a fabricator spend three weeks tuning a material cost model before he walked out to the shop and compared a BOM against the physical build. Two of the five components on the parent record hadn’t been used since 2019. Nobody could tell him when that record had last been right.
A model consuming that structure inherits every distortion in it. The output looks precise, because software output always looks precise.
Most mid-sized manufacturers carry two or three generations of part numbering convention stacked on top of each other, the residue of acquisitions, migrations, and a rationalization somebody abandoned in year two.
The same physical item shows up under multiple identifiers. Consumption history splits, and a forecasting model sees three intermittent low volume parts where the business has one steady mover.
Spend analytics and supplier risk scoring rest on one simple question. How much do we buy from this supplier?
When that supplier exists as four records with different addresses and payment terms, no model can answer it, and no amount of algorithmic sophistication compensates.
The customer side is the same problem wearing a revenue label. It distorts segmentation and churn signals, and it quietly ruins any attempt at predictive pricing.
If the plant runs a shadow reconciliation before each planning cycle, that is a declaration that ERP inventory isn’t the system of record in practice, whatever the org chart says.
Undocumented workarounds fill the gap. A scheduler keeps her own lead time table. A controller has run a parallel close in Excel for six years.
Those spreadsheets are the interesting part. They often feed the very reports an AI initiative gets trained on, which means the model is learning the workaround instead of the process.
The knowledge lives in individuals, and it never reaches the training set in any structured form.
This is an uncomfortable thing to say to a leadership team that already announced an AI strategy at the sales kickoff. It is usually correct anyway.
Where part numbering, BOM control, and vendor master governance are unresolved, the highest return technology investment available isn’t a model. It’s the unglamorous work of getting master data under control, and that work tends to pay for itself in planning accuracy before anyone applies AI to anything.
A pilot that never happens costs time. A pilot that produces visibly wrong output in front of operations costs trust, which is the scarcest resource in any transformation.
Once a plant manager decides the system lies, that judgment survives leadership changes and system upgrades. Delaying a pilot two quarters is cheap next to spending three years rebuilding confidence.
The models available through mainstream ERP platforms and cloud services are more than adequate for the problems mid-sized manufacturers actually have. Nobody in this market is constrained by forecasting algorithms.
They’re constrained by whether shipment history is clean, whether the product hierarchy holds together, and whether anyone owns the definition of an active customer. AI amplifies whatever it is given. Amplifiers work in both directions.

The requirement is more specific than “clean data,” a phrase vague enough to be useless in a budget conversation. ERP data integrity is shorthand for holding five properties at once.
Completeness comes first. Models are unforgiving about missing fields in a way that reports never are, because a person reading a report compensates for the blank cell without noticing.
Consistency of definition matters more than most executives expect. If three business units define on time delivery differently, an aggregated model learns the average of three incompatible things.
Standard definitions have to exist before the data is pooled. Afterward is too late.
History and depth are a hard constraint. Demand forecasting generally needs two to three years of transaction history at the level you intend to forecast, which is what Ultra typically sees work in practice. If you replatformed eighteen months ago and left the transactional history behind, the capability isn’t available yet, whatever the demo showed.
Granularity has to match the decision. A model forecasting at product family level can’t drive item level replenishment, and stretching it to do so produces exactly the confident error that erodes trust on the floor.
Traceability closes the set. When a recommendation gets questioned, and it will be, someone has to trace it back to source records. Systems that can’t answer “why did it say that” don’t survive contact with an experienced operations team.
Three claims deserve real skepticism in a vendor conversation. I’ll admit to being harder on this than most people in my field.
The first is accuracy stated without context. A ninety-five percent forecast accuracy claim means nothing without the aggregation level, the time horizon, the demand pattern, and the error measure.
Aggregate high enough and long enough and accuracy becomes trivially achievable. It also becomes operationally worthless.
The second is the suggestion that AI will compensate for poor data. Some vendors now market data cleansing capability as an AI feature.
Pattern matching genuinely helps identify duplicate records and suggest merges, and it saves real hours. It cannot decide which of two conflicting BOM revisions reflects how the product is built today. That takes an engineer, a site visit, and governance.
The third is the implicit promise that AI capability arrives with a license. In most deployments the feature ships switched off and needs configuration against your master data.
The software is capable. The gap between capable and useful is organizational, and that gap is where independent judgment during ERP technology selection earns its keep.

Realism isn’t pessimism. Several applications deliver measurable value right now, without exotic infrastructure.
What they share is a tolerable data requirement and a named owner for the output.
The most mature application, and usually the first worth attempting. It works when shipment history is clean and somebody in planning is accountable for acting on the signal.
High volume repeat order items with real seasonality are where it shines. Engineer to order demand is where it struggles, because statistical methods have almost nothing to learn from it.
Detecting drift in process parameters or inspection results suits machine learning well, since the data is machine generated and therefore consistent. That sidesteps most master data problems in one step.
The value depends on integration back into the quality process. An anomaly flagged into a dashboard nobody owns achieves nothing. An anomaly routed into a nonconformance workflow with a named responder changes outcomes.
Genuinely valuable on high criticality assets with sensor instrumentation and reliable maintenance history. It fails when applied broadly across an asset base with poor work order discipline, which is the scenario in the opening scene.
Where maintenance history is recorded inconsistently, the model has no reliable definition of failure. It generates false positives until the floor stops reading the alerts.
The most underrated application on this list. Pulling line items out of supplier invoices or inbound quality certificates removes measurable clerical effort at low risk, because a person still confirms the result before posting.
Payback is short. That makes it a sensible first project for organizations that need an early win while the master data work grinds on.
The sequence matters more than the components. The common failure is running these tracks alongside AI rather than ahead of it.
What follows is the path Ultra typically walks a manufacturer through, in order.
Most programs that go wrong compress the middle steps on the theory that the model will sort it out. It will not.
Ask what data the capability requires, at what granularity, and how much history. Ask whether the feature is generally available today or sitting on a roadmap, then request the release note. Below is the short version I hand clients before a demo.
| What the vendor claims | What to ask | What a good answer sounds like |
|---|---|---|
| “Our forecasting is ninety-five percent accurate” | At what level, horizon, and error measure? | A specific level, horizon, and named metric |
| “AI cleans your data automatically” | Which conflicts does it resolve without a human? | Duplicates yes, BOM revision judgment no |
| “It works for manufacturers like you” | Can we see it live at a comparable reference? | A named customer of similar size and mix |
| “AI is included in the license” | Is it generally available in our version today? | A release note and a configuration estimate |
| “The recommendations are explainable” | Trace one back to source records for us | A drill path to the underlying transactions |
Ask to see it running on data like yours, at a reference customer of comparable size. Vendor experience across discrete and process manufacturing industries varies far more than the marketing suggests. A capability proven in high volume consumer goods may mean nothing in a low volume, high mix plant.
Ask how the model explains its recommendations, and who is accountable when it’s wrong. Ask what the feature costs beyond the base license, services included.
Last, ask what the vendor’s own assessment is of your data readiness. A credible partner will tell you where you aren’t ready. A vendor who says your data is fine without ever examining it has told you something useful about the rest of the conversation.
AI in manufacturing is real, it’s improving, and it will matter. None of that changes the sequence.
The organizations that benefit are the ones already running disciplined master data, standard processes, and a stable core system, because those conditions are what make any analytical capability trustworthy in the first place.
That work has independent value. Accurate BOMs improve margin visibility whether or not a model ever reads them.
Clean supplier masters improve your negotiating position. Reliable inventory cuts expedite freight. None of it depends on AI delivering anything.
Here is the calm position. Treat AI as a capability you are deliberately preparing for, not a race you’re losing.
Preparation is measurable, fundable, and defensible to a board. Waiting two quarters to start a pilot won’t disadvantage a mid-sized manufacturer. Starting one on data the organization doesn’t trust very likely will.
Mostly yes, with one exception. Training models on data from a partially migrated environment produces results that describe the transition rather than the business, and the model usually has to be rebuilt afterward on the new structures anyway.
Run master data governance and cleansing as part of the ERP implementation itself, so the new system starts clean, then apply AI once several full operating cycles have run. Document extraction is the exception. It sits at the edge of the transaction flow and can proceed on its own timing.
In the mid-sized manufacturing engagements Ultra runs, we typically see sixty to eighty percent of total effort go to data preparation and definition alignment. That share doesn’t drop much with better tooling, since the underlying problems are organizational rather than technical.
A project plan allocating two weeks to data preparation and four months to model development has the ratio inverted. We say so early, and it isn’t always welcome.
Partially. Pattern matching is effective at identifying probable duplicates and outlier values, and it takes real manual effort out of a deduplication exercise.
What it can’t do is resolve conflicts requiring domain judgment, such as which BOM revision reflects current production. It also can’t prevent recurrence.
Without governance and ownership, a cleansed master file degrades again inside eighteen months, which is what we typically observe. We’ve watched that happen at the same client twice.
Embedded features are usually a reasonable starting point, since they’re pre-integrated and carry no separate infrastructure cost. The caveats are worth knowing.
They run against your master data and inherit its quality. Capability varies widely between vendors, and “built in” frequently means a higher tier or a later release.
Verify what is generally available now, in your version, and confirm it with a reference customer rather than a demo.
Choose something with a bounded data requirement and a clear owner. Document extraction for accounts payable is often the best candidate, because the data is self-contained, a person validates the output, and the saving is measurable within a quarter.
Demand forecasting on a subset of high volume items works too, where shipment history holds up. Avoid anything needing accurate BOMs or complete maintenance history unless you’ve verified both yourself.
Run a structured assessment before committing to a use case. Measure duplicate rates in the item and supplier masters, sample BOMs against physical builds, and compare cycle count variance to book inventory.
Then ask each function to list the spreadsheets they maintain outside the system, and treat that list as the most honest data quality report you will receive.
The technology isn’t overhyped. The timeline and the effort are.
Manufacturers get told value follows a license purchase, when in practice it follows data discipline most organizations have deferred for a decade. That gap is where budgets disappear, and an independent, vendor-neutral perspective is often the difference between an AI investment that changes operations and one that produces a dashboard.
Data integrity isn’t preparation for AI. It’s an operational capability that AI happens to expose, and it returns value on its own terms regardless of what you build on top of it.
For more on the foundation this depends on, manufacturing AI and ERP data integrity covers the data side in detail, and why data accuracy matters in ERP implementations addresses how these problems originate during deployment.
If your leadership team is weighing an AI initiative against the state of the current system, that assessment is worth having with someone who has no software to sell. Ultra’s work across manufacturing and distribution sectors shows how others have approached the same question. Manufacturers who understand all of this find the conversation about AI and ERP gets easier, because the hard part is already behind them.