In short
- An AI assistant inherits the quality of the knowledge it reads, including the parts that are out of date.
- The expensive problem is several versions of the same knowledge with nothing on any of them to say which one holds.
- Provenance and a single current version do more for answer quality than a larger model.
- Gartner's own analysis puts data readiness among the reasons AI projects stall, and that matches what we found.
We began by asking which model to use. The question that mattered was what the model would be reading.
In any business over a certain age, what the organisation knows is spread across document libraries, mail, chat threads, presentations built for one meeting, and the memory of the people who did the work. Most of it was correct when it was written. A good part of it is no longer correct and carries nothing on the surface to say so.
Give a capable assistant access to that and it will do exactly what you would expect. It will read the wrong version, find a statement that was true eighteen months ago, and present it in fluent, confident prose. The answer is not hallucinated. It is worse than that. It is accurately retrieved and out of date.
Three things had to be true before anything else
One current version. Where a document has been revised, the revision has to supersede the original everywhere it can be read from. An assistant that can see version two and version four has no basis for choosing between them, and neither does a person who finds both.
Provenance on every answer. A statement without a source is an opinion with good grammar. If an answer cannot name the document it came from, the reader cannot check it, and if the reader cannot check it they will either trust it too much or ignore it entirely. Both are failures.
A home for everything. The quiet killer is work that is complete, correct and filed nowhere in particular. It exists, it cost money to produce, and it is invisible to everyone who did not attend the meeting it was made for.
None of that is glamorous. All of it had to come before the interesting work.
Why this sequence is the right way round
Gartner has been saying for some time that data readiness is among the reasons AI initiatives stall, alongside unclear business value and cost. Our experience matches the finding, with one addition: readiness is usually discussed as a data-engineering problem, and in a professional-services business most of the value sits in documents and decisions rather than in tables.
The remedy is different. Cleaning a data warehouse is a project with a scope. Getting an organisation to file its thinking where the organisation can find it again is a change in habit, and habits do not respond to architecture diagrams.
What worked for us was making the correct path the easy one. If filing a document where it belongs is harder than emailing it to three people, it will be emailed to three people, forever. If the system takes it, classifies it, and puts it where the next person will look, the habit follows the convenience.
The part that surprised us
We expected the benefit to be better answers. We got that, and we also got something we had not anticipated.
Once the knowledge was structured and current, the gaps became visible. Not gaps in the data, gaps in the thinking: questions that had been raised and never resolved, and positions held for reasons nobody could locate. That is uncomfortable reading and it is genuinely valuable, because a gap you can see is a gap you can close.
It also changes the first week of a client engagement. We are building the same governed record into our discovery conversations. The intent is that when a client states an objective, it is checked against what we have delivered before and what went wrong the last time, with the source document shown next to the answer. A proposal used to depend on who happened to be in the room; it now starts from a structured record of the client's inputs and our standards, with every open question written down, and the record will improve with every engagement that closes.
If you are starting this now
If you are choosing a model before you have looked at what it will read, you are optimising the wrong end of the problem. The order we would recommend, having done it in the wrong order first:
- Find out where your institutional knowledge actually lives, including the parts that live in people.
- Decide what a current version means and make it enforceable.
- Require a source on every answer and refuse the ones that cannot produce one.
- Then choose your model, and expect to change it, because that decision ages faster than any of the above.
For clients, this is usually the first phase of an AI-ready data foundation, and the platform underneath it is covered under governed data platforms and analytics. Where most of the knowledge sits in documents, the work starts with classification and version control, and a model comes later.
Once you have knowledge worth protecting, one question follows immediately, and in this region buyers raise it earliest in a procurement conversation. Where does all of this physically run?