The AI worked. Our assumptions did not.
First of three. We believed more output would create more value. Measuring outcomes rather than activity changed what we believed, and the biggest lesson turned out not to be about AI.
Second of three. We compared what we had produced against what anyone had actually asked for. A large share of it could not be matched to a question, and that changed how we choose what to work on.
The first piece in this series described the change of question. This one is about what the first answer showed, and it was not what we expected.
We had two things: a body of research and analysis the business had produced, and an archive of the questions people had actually brought to us.
Comparing them is an obvious exercise. It had never been done, which is its own finding.
Most of what we had produced could not be matched to any question in the archive.
That needs saying carefully, because it is easy to overstate. It does not mean nobody read it. It does not mean it was wrong, or wasted in every case. What it means precisely is that we could not connect most of our supply to visible demand. Somebody decided each piece was worth producing. In most cases we could not point at the person who had asked.
We are documenting the method and the period properly before publishing the number, because a figure of that kind without its basis is exactly the sort of claim this series is arguing against.
The mechanism is not incompetence. It is closer to the opposite.
A team with good tools and good judgement can see useful work everywhere. The tools make production cheap, the judgement makes the output genuinely worth reading, and nothing in the process asks who is waiting for it. So it gets produced and filed.
Meanwhile the questions people actually have arrive informally, in a corridor or a message thread, get half-answered by whoever is nearest, and never reach the team that could have answered them properly.
Both halves of that are invisible from inside. The producing team sees an organisation slow to take up what it offers. The rest of the business sees a team producing things nobody needed. They are describing one failure from two ends.
People do not want more information. They want an answer good enough to stop turning the problem over, so that the decision can move.
Read against that, a comprehensive piece of analysis is not automatically better than a two-line answer. Often it is worse, because it transfers the work of extracting the answer back to the person who asked. The instinct to be thorough, which is a virtue in most professional contexts, becomes a way of not committing.
In the accounts we can see properly, the teams getting most out of AI are the ones whose work starts from a problem somebody named.
We start from the question, in the words of whoever asked it. Written down, with a name attached. This connects to the point made earlier in this series about what a well-formed request looks like, and the measurement is what forced us to take it seriously rather than agree with it in principle.
We keep the archive of questions, and we use it. It is the demand signal. Without it you are guessing.
We accept that some good work will not get made. This was the hardest change. Capacity spent on something nobody asked for is capacity not spent on something somebody is waiting for, and the second one is almost always worth more even when it is less interesting.
The same discipline now governs what we build for clients. A dashboard, a report or an agent is scoped from a named decision and the person who takes it, and an executive decision intelligence engagement starts with the list of questions the leadership team is currently answering late or not at all.
Take a month of what your team produced and try to trace each piece back to somebody who asked for it, by name.
If that is easy, you have a demand-led operation and this finding will not apply to you. If it is difficult, the difficulty is the result.
That leaves the speed finding, which is where this series ends.
Part of Insights series. Get new articles by email or follow the RSS feed.
First of three. We believed more output would create more value. Measuring outcomes rather than activity changed what we believed, and the biggest lesson turned out not to be about AI.
Most AI initiatives stop at the pilot. The reason is usually not the model. It is that nothing connects an AI output to a decision anyone is willing to be accountable for.
The first correction in our AI work had nothing to do with AI. What an organisation knows is usually scattered, undated and contradictory, and a model working from that will answer confidently and wrongly.