I write documentation, so what caught my attention in a recent Pragmatic Programmers newsletter about AI and code quality was not the proposed score for “AI-ready” code. It was the pattern underneath it.
Code for Machines, Not Just Humans, a 2026 study accepted for FORGE, found that AI systems were more likely to preserve behavior when they refactored healthier, more maintainable code. CodeScene’s AI-Ready Code: How Code Health Determines AI Performance made the argument more bluntly: AI accelerates what is already present. Give it a strong foundation and it can help improve that foundation. Give it a tangled system and it can produce more breakage, faster.
That is a code example of a documentation problem I have been writing about for months. AI does not remove the consequences of weak source material, unclear authority, missing context, neglected maintenance, or overloaded reviewers. It operates on top of those conditions and can amplify them.
What the Code Research Makes Visible
The FORGE study tested AI-assisted refactoring across 5,000 Python files from competitive programming. The researchers found a meaningful association between CodeHealth and whether the refactored programs preserved their behavior. Refactorings on healthier files broke tests less often.
The finding builds on the earlier Code Red: The Business Impact of Code Quality, which studied 39 proprietary production codebases. That research linked lower CodeHealth categories with more reported defects and slower, less predictable issue resolution. AI did not create the cost of an unhealthy foundation. It made that cost newly relevant to automated change.
Documentation teams should notice the systems lesson. Code and documentation play different roles, but each supplies part of the context an AI system uses. Structure, boundaries, consistency, and visible relationships affect what the system can interpret safely. Hidden exceptions and accumulated debt increase the chance that a plausible output will be wrong.
Documentation Has a Foundation Too
Organizations often talk about using AI to generate, summarize, migrate, personalize, or retrieve documentation. The conversation usually starts with the model and the use case, while the condition of the documentation system receives less attention.
What sources will the AI use? Which one is authoritative? Which version applies? Who owns the content? What context must remain attached when a passage is reused? Which instructions are obsolete? What happens when two sources disagree? How will anyone test whether an answer remains correct after the product changes?
A healthy documentation system helps people and machines determine:
- Which source is authoritative.
- Who owns the information.
- Which product, version, environment, audience, role, or permission set the guidance applies to.
- When someone last verified the content.
- What has been deprecated, superseded, or retired.
- Which dependencies, warnings, and exceptions must travel with reused content.
- What evidence supports an important claim or procedure.
- How a product or policy change triggers review.
- How reviewers detect contradictions and decide which source wins.
- Whether critical tasks and answers still work after a change.
Structured content, metadata, version control, reusable components, taxonomy, and automated checks can support those capabilities. None of them can manufacture a decision the organization never made. If two teams disagree about the supported workflow, better retrieval only finds the disagreement faster. If nobody owns the answer, generation produces another version that nobody owns.
AI Amplifies Documentation Debt
I have already written about both sides of this pattern. In The Best Documentation Is Easy to Miss—Until You Look at What It Prevented, I argued that documentation’s value often appears in problems other teams never have to solve. In AI Writing Is Not Cheap: It Moves the Cost to Reviewers, I argued that AI generation often looks inexpensive because organizations leave verification, correction, and maintenance out of the cost.
The code-health research connects those arguments. AI makes the condition of the underlying system more consequential. Healthy documentation can support useful answers whose success gets credited to the AI interface. Contradictory, obsolete, or weakly owned documentation shifts work to reviewers, specialists, support teams, and maintainers.
The point here is not to restate those cases. It is to recognize that an AI documentation initiative inherits both: the value created by a strong foundation and the hidden cost of compensating for a weak one.
Test the Foundation Before Scaling
Organizations do not need perfect documentation before experimenting with AI. They do need to know where the foundation is weak before they scale generation. That means identifying the highest-consequence domains, canonical sources and owners, known contradictions, applicability rules, review and retirement triggers, provenance, and the real capacity to verify and maintain what the system produces.
Documentation does not have to become code, though it can benefit from the same operating disciplines: versioned sources, structured content, automated checks, traceable review, explicit ownership, and repeatable publishing. These practices do not guarantee that an AI initiative will work. They make its risks, dependencies, and decision points visible.
That is why I created AI Fit & Foundation. AI Fit turns broad interest or executive pressure into one bounded, defensible use case—or a clear decision not to use AI. AI Foundation tests whether the knowledge, ownership, architecture, controls, oversight, delivery practices, and capacity can support that use case. For documentation, the question is not whether a model can produce acceptable prose in a demonstration. It is whether the system behind that prose can sustain trustworthy output.
The answer may be to proceed, reshape the use case, repair the foundation first, keep critical work human-led, or stop. AI can restructure, adapt, retrieve, and accelerate documentation work. It cannot create authority, resolve decisions the organization has not made, restore context the source never recorded, or assign ownership where none exists. The code-health research makes the same point in another domain: AI reveals and amplifies the quality of the system it enters.
Before scaling AI-generated documentation, ask whether the use case fits the real problem and whether the knowledge, people, controls, and review capacity can support it. If your organization needs both decisions, start with AI Fit & Foundation.
Sources
- Code for Machines, Not Just Humans, accepted for FORGE 2026, studied 5,000 Python files and found a meaningful association between CodeHealth and semantic preservation after AI refactoring.
- AI-Ready Code: How Code Health Determines AI Performance is a CodeScene whitepaper. It reports the commercial benchmark and makes clear that outcomes below a CodeHealth score of 7 are projected.
- Code Red: The Business Impact of Code Quality studied 39 proprietary production codebases and linked lower CodeHealth categories with more defects and slower, less predictable work.
- The Pragmatic Programmers newsletter received October 7, 2026, supplied the timely code-health comparison and promotional framing.
