Your AI problem is a data problem, and it started ten years ago
Every AI pilot works beautifully until it touches the real data estate. The constraint isn't the model, it's the twenty years of permissions sprawl, uncatalogued documents and archived applications nobody planned to read again.
Every Microsoft 365 Copilot deployment I have been close to has the same week two.
Week one goes fine. The demo lands, a few people summarise meetings they did not want to sit through, the pilot group is pleased with itself. Then in week two somebody asks a perfectly reasonable question and gets an answer built out of a document they should never have been able to open. Usually a spreadsheet. Occasionally a spreadsheet with names and salaries on it.
I have written about this before and it has not stopped being true. That is not a Copilot bug. Copilot surfaces what the user's permissions already allow. The document was always readable by that person; what changed is that finding it no longer required knowing it existed. Ten years of "share with everyone in the organisation" sat harmlessly in SharePoint because no search tool was motivated enough to go looking. Now one is.
The incident is not the interesting bit. What the incident tells you about everything else is.
The model was never the constraint
Almost every organisation I speak to is under pressure to do something with AI, and almost all of them are arriving at the same conclusion by different routes: the model is not the problem. The pilot works flawlessly on a curated dataset and then falls over the moment it meets the actual estate.
That is not a failure of the technology. It is a decade of deferred data work presenting its invoice. The permissions problem people hit with Copilot is the smallest, most visible version of it, because Copilot is where most organisations pointed AI first. The same problem exists at enterprise data scale, and it is a governance problem wearing an AI costume.
There are three versions of it worth separating out.
Permissions
Covered above, and the one my readers have lived through. Worth adding: the remediation is genuinely unglamorous. Sharing link audits, SharePoint site permission reviews, sensitivity labelling, killing off the "Everyone except external users" grants that somebody added in 2016 to unblock a project that finished in 2017. None of it demos well. All of it has to happen before, not after.
Dark and unstructured data
The large majority of enterprise data is unstructured. The figure people quote is around 80%, which I would treat as directional rather than precise, but nobody seriously argues the shape of it. Contracts, policies, board packs, scanned documents, PDFs of PDFs. Scattered across file shares, mailboxes and departmental systems, unindexed and uncategorised.
Nobody knows what is in there. Nobody knows what retention or compliance obligations attach to it. And an AI cannot reason over material that has never been catalogued. This is the part where "AI-ready data" as a phrase starts to earn its keep, right before the marketing departments wear it out.
Legacy and archived data, which almost nobody thinks about
This is the one I find genuinely interesting, and it gets the least attention.
Enterprises have spent twenty years retiring applications. Mainframes, old ERPs, the finance system that came with an acquisition in 2011, the claims platform that was decommissioned when the group standardised. That data did not get deleted, because you cannot delete it. It got moved into an archive whose design goals were cheap retention and the ability to answer a legal request once a quarter.
Which means it is sitting in a proprietary format, at rest, functionally invisible to anything modern.
And it is frequently the most valuable material the organisation owns for AI purposes. Decades of transactions. Claims history. Trial data. Customer behaviour across two recessions and a pandemic. The kind of longitudinal record you cannot buy and cannot recreate.
The decisions organisations made about archiving five and ten years ago now determine whether that history is an asset or a write-off. Almost none of those decisions were made with this in mind, because there was no reason to. The requirement was "retrievable if the regulator asks", and that requirement was met.
If you retire an application into a format only a lawyer can query, you have preserved the data and destroyed its usefulness.
Why the "AI on your ERP" demos disappoint
There is a technical reason those demos underwhelm, and it is worth understanding because it applies well beyond ERP.
SAP and Oracle E-Business Suite have thousands of tables with cryptic names and relationships that exist mainly in the heads of a small number of specialists who are all quite busy. Point a general-purpose model at that schema and it will guess. It will produce an answer that is fluent, confident and wrong, which is considerably worse than no answer, because somebody will act on it.
Getting past that requires a semantic layer specific to the application: something that encodes what the tables actually mean, how they relate, and which business terms map to which fields. Solix calls theirs the Application Knowledge Graph, and describes it as encoding "the business objects, relationships, terms, and tested query patterns", shipped pre-built and modelled across four value streams: Procure-to-Pay, Order-to-Cash, Record-to-Report and Make-to-Order.
The branding matters less than the concept. Any credible approach to natural-language querying over enterprise systems needs some version of this. If a vendor cannot tell you what theirs is, ask harder.
The worked example
Solix has been doing this unfashionable work since 2002, out of Santa Clara. For most of that time the product was enterprise archiving: application retirement, database archiving, mainframe and email and file archiving. Sold on storage cost, database performance and compliance. Nothing about it was exciting.
The public case studies read exactly as you would expect from that. American Tire Distributors reduced their Oracle EBS database by 25%. Overstock.com took 1TB out of their Oracle database with a performance improvement alongside it. LG Electronics improved query time and cut backup costs. Forbes Marshall reduced database size and improved Oracle E-Business Suite performance. Among the unnamed ones, a Fortune 50 food and beverage company retired 95 legacy applications and 75TB of data; a clean energy utility saved over $1M using SOLIXCloud; the US Department of Health and Human Services consolidated data masking and subsetting. There are twenty-five customer logos on the homepage, including PepsiCo, Amazon, Kaiser Permanente, Wells Fargo, Swiss Re, Unilever and BAE Systems.
The platform is now described across five capabilities: Govern, Classify, Build, Access and Preserve. The language is all about grounding and citation, "cited, grounded, governed", a trust perimeter, answers tied back to source.
The line that made me stop, though, is aimed at their existing customers: "Already running Solix? You're already most of the way there." The argument being that twenty years of preservation work means those customers start their AI journey at step three rather than step zero.
That is the whole point of this post, and it applies whether or not Solix is your vendor. The organisations best placed to do something useful with AI over their business data are the ones who, for entirely unrelated reasons, did the boring work properly. They catalogued things. They kept archived data queryable instead of merely stored. They knew where it was.
Where I would push back
I have not used the product. This is analysis of public positioning and published customer stories, not a review, and you should weight it accordingly.
The category is also crowded. Informatica, IBM, OpenText, Databricks and Microsoft Purview are all making some version of the "AI-ready data" claim, and the ambition is not the differentiator. The differentiator, if there is one, is the twenty-year archive estate and the application-specific semantic layer sitting on top of it.
I would also note that "AI-ready data" is well on its way to becoming a phrase that means nothing. Be sceptical of the language while taking the underlying problem entirely seriously. Those are compatible positions.
And a fair amount of the public proof is older than the current positioning. Archiving engagements from the last decade are evidence of durability, not evidence that the AI layer works. Ask for something recent.
What to actually do
Audit permissions before you deploy Copilot, not after the spreadsheet incident. It is the same work either way; only the reputational cost differs.
Find out what is in your archives and what format it is in. Not how much of it there is. What format.
If you have an application retirement programme running now, ask whether the plan preserves queryability or just bytes. That question costs nothing to ask today and is very expensive to ask in 2031.
Treat governance as the enabling layer rather than compliance overhead. It is the thing that decides what AI is allowed to see, which makes it product work, not paperwork.
Start with one system you understand well. Resist the enterprise-wide data programme. They do not finish.
The reframe
For twenty years, cataloguing, archiving, permissions hygiene and retention policy were the work you did when there was nothing more interesting to do, staffed accordingly and funded worse.
That work is now the thing standing between an organisation and anything useful coming out of AI. It did not get more interesting. It just got load-bearing.
Public sources: solix.com and the public case study library, verified August 2026.
