Our documents are a mess. Clean them up before AI, or after?
Data quality is the top obstacle for 43% of organizations, yet 95% of AI pilot programs fail to deliver measurable impact. Deploying AI over inconsistent documentation risks quiet degradation and untraceable hallucinations.
The question
Our internal documentation is inconsistent, outdated in places, and was never structured for retrieval. Do we invest in cleaning and structuring it before deploying AI over it, deploy first and let real usage reveal what needs fixing, or scope the cleanup narrowly to the one workflow we are trying to prove?
Counsel's position
Scope cleanup narrowly to a critical workflow, deploy AI, and then iteratively expand based on demonstrated value and learned needs.
Verdict
The verdict: Scope cleanup narrowly to a critical workflow, deploy AI, and then iteratively expand based on demonstrated value and learned needs.
How the criteria decide
3 of 5 criteria resolved on cited evidence. 2 had none either way.
| Criterion | Favours | Evidence |
|---|---|---|
| data quality | Clean & structure all docs | Data quality is the top obstacle to AI success for 43% of organizations 43% of organizations named data quality and readiness as their top obstacle to AI success. Not model performance. Not tooling. Data. Unstructured document workflows require a unified, governed data foundation The issue, however, is not AI automation itself, but the fragmented, incomplete data foundations these early tools sit on. |
| user experience | Scope cleanup narrowly | AI systems degrade quietly when the retrieval layer struggles The uncomfortable truth is that AI systems don’t break loudly. They degrade quietly through slightly worse answers, slightly slower responses, slightly higher costs. |
| time to deploy | No evidence either way | |
| resource cost | No evidence either way | |
| workflow impact | Scope cleanup narrowly | Most internal document AI projects stall before reaching production usage Most in-house "ask our docs" projects stall out somewhere between the proof of concept and the version anyone actually uses AI pilot programs fail to deliver measurable impact Ninety-five percent of artificial intelligence (AI) pilot programs fail to deliver measurable impact, according to research from MIT’s NANDA initiative. |
Data quality is the top obstacle to AI success for 43% of organizations
Given your inconsistent documentation, deploying a model before fixing the data layer will result in confident but untraceable hallucinations.
Unstructured document workflows require a unified, governed data foundation
While pursuing a scoped cleanup, establishing a structured extraction pipeline ensures your downstream agents have reliable context.
Most internal document AI projects stall before reaching production usage
Given your inconsistent documentation, scoping the cleanup to a single workflow allows you to validate the impact before scaling.
AI systems degrade quietly when the retrieval layer struggles
Deploying AI over uncurated documentation will cause your retrieval pipeline to feed noisy, unstructured context to the reasoning engine.
AI pilot programs fail to deliver measurable impact
To avoid becoming part of this failure rate, scope your documentation cleanup to a small, manageable project that demonstrates value before scaling.
Read another verdict
- Which process should we point AI at first?
- Put one person in charge of AI — or is a Head of AI premature for us?
- Buy a tool for this process, or build around our own knowledge?
- Centralize AI strategy under CEO or distribute ownership?
- Adopt new AI ROI tools or refine existing methods?
- Invest in pre-build costing or post-deployment ROI tracking?
- How do we measure the return on an AI workflow — and what baseline is honest?
- Our best people's know-how isn't written down — can AI even use it?