AI Won't Fix Your Technical Debt. It May Make It More Expensive

Author
Ravi Prajapati

AI coding tools speed up output, not code quality. See what GitClear, DORA, and METR research reveals about AI and technical debt, plus a practical framework to manage it.
Quick answer: AI coding tools speed up how fast code gets written, not how well it fits into an existing system. Evidence from Google's DORA research, GitClear's analysis of 211 million lines of code, and a 2025 METR randomized trial all point the same direction: AI adoption correlates with more duplicated code, lower refactoring rates, higher churn, and in some settings, slower delivery. Technical debt does not disappear under AI. It compounds faster, and it compounds in places teams are not yet measuring.
Every engineering leader has heard some version of the pitch: AI coding assistants will finally let the team outrun its backlog of technical debt. Write code faster, ship more, and use the reclaimed time to pay down the mess accumulated over the last decade.
The data collected since generative coding tools went mainstream tells a more complicated story. AI does increase the volume of code a team produces. It does not, on its own, improve the structural health of the codebase that code lands in. In several of the largest studies available, the opposite trend shows up: more duplication, less refactoring, and higher rates of code that gets rewritten within weeks of being merged.
This matters because technical debt is not primarily a coding problem. It is a compounding-cost problem, similar to financial debt. Left unmanaged, small shortcuts accrue interest in the form of slower delivery, harder onboarding, and more fragile systems. A tool that makes it cheaper to write code without making it cheaper to understand, integrate, or maintain that code does not reduce the principal. It can increase the rate at which interest accrues.
What Is Technical Debt, and Why Does AI Change the Calculus?
Technical debt is the implied cost of additional rework created when a team chooses an expedient solution now instead of a better approach that would take longer. The term originated with Ward Cunningham in the early 1990s as a metaphor: like financial debt, some technical debt is a reasonable, even strategic, choice, as long as someone tracks it and pays it down before interest overwhelms the budget.
Two things distinguish debt taken on by AI-assisted teams from debt taken on by human engineers alone.
First, the volume problem. AI tools lower the cost of producing code, which means more code gets written per unit of engineering time. If even a stable percentage of that code is low quality, the absolute amount of debt added per sprint goes up simply because output has gone up.
Second, the comprehension problem. A human engineer who writes a shortcut usually retains some contextual memory of why the shortcut exists. An AI-generated shortcut, accepted quickly during a review, often has no such memory attached to it anywhere in the organization. Nobody fully understands why the code is shaped the way it is, which makes it more expensive to touch safely later. Engineering teams have started calling this comprehension debt, separate from the traditional categories of design debt, code debt, and test debt.

What Does the Evidence Actually Show About AI and Code Quality?
The most detailed dataset on this question comes from GitClear, a code analytics company that studied 211 million changed lines of code authored between January 2020 and December 2024 across repositories including ones owned by Google, Microsoft, and Meta. GitClear's research classifies changed code into categories such as added, moved (refactored), copy-pasted, and churned, which lets it track structural code health over time rather than relying on self-reported productivity surveys.
The findings, published in the 2025 AI Copilot Code Quality report, show a consistent pattern across the five-year window:
Refactoring collapsed. The share of changed lines associated with refactoring, which GitClear calls "moved" code, fell from roughly 25% of changed lines in 2021 to under 10% in 2024.
Duplication surged. Copy-pasted code rose from 8.3% to 12.3% of changed lines over the same period, and 2024 was the first year on record where copy-paste volume exceeded refactoring volume.
Duplicated code blocks rose sharply. Blocks of five or more duplicated lines increased roughly eightfold during 2024 alone, according to GitClear's analysis.
Churn nearly doubled. GitClear defines churn as code revised or reverted within two weeks of being written, a proxy for low-quality first drafts. That figure rose from a pre-AI baseline near 3% to close to 6% by 2024, and later GitClear updates put 2025 churn above 7%.
This pattern is not proof that AI causes worse code in a strict causal sense. GitClear's own researchers and independent reviewers note the design is correlational, tracking a period when AI adoption and these quality signals both moved together, not a controlled experiment. But the size of the shift, its consistency across metrics, and its concentration in AI-heavy repositories make alternative explanations, such as changing team composition or unrelated process shifts, harder to sustain as the sole cause.
Google's DORA research group reached a related, and independently derived, conclusion using a completely different methodology: an annual survey of more than 39,000 engineering professionals. The 2024 State of DevOps report found that a 25% increase in AI adoption was associated with modest gains in code quality (3.4%) and code review speed (3.1%), but also with a 1.5% decrease in software delivery throughput and a 7.2% decrease in delivery stability. In other words, teams got a little better at writing individual pieces of code and noticeably worse at shipping reliable systems.
DORA's own researchers point to a specific mechanism rather than blaming AI-generated code quality outright: AI makes it easier to produce larger changesets, and DORA's decade of prior research has consistently linked larger batch sizes to more delivery risk, independent of AI. The tool did not have to write bad code to hurt stability. It only had to make it easier to ship more code per change.
What Does This Mean for a Reader Deciding Whether to Trust AI Output?
A concise answer: Across the largest available datasets, AI adoption correlates with more code volume, less refactoring, more duplication, and higher short-term rework, while individual developers report feeling faster. The gap between felt productivity and measured system-level outcomes is the central risk engineering leaders need to manage, not a reason to avoid AI tools altogether.
Does AI Coding Actually Make Developers Faster?
Even the productivity premise deserves scrutiny before an organization builds a debt-reduction strategy on top of it. The most rigorous controlled study available comes from METR, an AI evaluation research organization founded by former OpenAI alignment researcher Elizabeth Barnes. In a 2025 randomized controlled trial, METR recruited 16 experienced open-source developers working on their own repositories and had them complete 246 real backlog tasks, randomly assigned to allow or disallow AI tool use.
Before the study, participants forecast that AI tools would cut their task completion time by 24%. The measured result was the opposite: developers using AI tools took 19% longer to complete their tasks than developers working without them. Strikingly, even after finishing the study and experiencing the slowdown directly, participants still estimated they had been about 20% faster with AI.
That gap between perceived and measured productivity is directly relevant to technical debt management. A team that believes it is moving 20% faster than it actually is will systematically underinvest in review, testing, and refactoring, because the felt cost of those activities looks proportionally higher against an inflated sense of output. METR is explicit that this is a snapshot of one setting, early-2025 tools, experienced developers, complex existing codebases, rather than a universal verdict on AI coding, and that results may shift as tools improve. But it is the strongest causal evidence available, and it points the same direction as the correlational GitClear and DORA data: individual-level speed does not reliably translate into system-level throughput.

Why Is AI-Generated Code Riskier to Maintain?
Three mechanisms explain why AI-assisted code tends to accumulate debt faster than the productivity gains would suggest.
AI models do not have a persistent, verified model of your specific codebase. A large language model generates code based on patterns learned from its training data and whatever context fits in its current context window. It does not inherently know your team's established conventions, the reason a particular abstraction exists, or which parts of the system are load-bearing versus disposable. Each generation is, in effect, a fresh guess informed by general patterns rather than institutional memory, which is a structural reason duplication and inconsistent style tend to increase rather than decrease as AI usage grows.
Security is not improving at the same pace as functional correctness. Application security firm Veracode evaluated more than 100 large language models across 80 curated coding tasks and found that AI-generated code introduced a security vulnerability from the OWASP Top 10 in 45% of the test cases, according to its 2025 GenAI Code Security Report. Risk varied sharply by language: Java code failed security tests more than 70% of the time, while Python, JavaScript, and C# failed between 38% and 45% of the time.
Notably, Veracode found that newer, more capable models were not meaningfully better at avoiding these flaws than older ones. Functional accuracy improved across model generations; secure coding practice did not. That gap is a form of technical debt that is invisible until an audit, penetration test, or breach surfaces it, which is exactly why it tends to be the most expensive kind.
Trust in AI output is declining even as usage rises, which is itself a signal worth taking seriously. Stack Overflow's 2025 Developer Survey, which gathered responses from more than 49,000 developers, found that 84% of developers now use or plan to use AI tools, up from 76% the prior year, while 46% said they don't trust the accuracy of AI tool output, up sharply from 31% the year before.
Nearly half of respondents, 45%, said debugging AI-generated code takes longer than they expected, and the top-cited frustration, reported by two-thirds of developers, was that AI output is often "almost right, but not quite." Code that is almost right is arguably more dangerous than code that is obviously wrong, because it passes a casual review and fails somewhere less visible later.
How Much Could AI-Amplified Technical Debt Actually Cost You?
Technical debt was already a large line item before generative AI. McKinsey's research on the topic, based on a survey of CIOs at financial-services and technology firms, found that organizations typically divert 10 to 20 percent of the technology budget earmarked for new products toward resolving issues related to tech debt, and that CIOs estimate tech debt amounts to 20 to 40 percent of the value of their entire technology estate before depreciation. In a related McKinsey analysis, 60% of surveyed CIOs said their organization's tech debt had risen over the prior three years, and that survey predates the widespread deployment of AI coding assistants.
There is not yet a definitive, audited dollar figure for how much AI-specific debt adds on top of that baseline; the field is too new and no standards body has published an accepted methodology. What can be stated with reasonable confidence, based on the mechanisms documented above, is the shape of the risk:
Volume effect. If AI increases code output per engineer, and even a stable share of that code carries defects, duplication, or security gaps, the absolute volume of debt added per quarter rises even if the defect rate per line stays flat.
Detection lag effect. Comprehension debt and security debt are often invisible at merge time. They surface later, during an incident, an audit, or an onboarding cycle, at which point the cost of fixing them is materially higher than it would have been at the point of creation, consistent with the long-established finding in software engineering that defects caught late cost far more to resolve than defects caught early.
Trust-erosion effect. As Stack Overflow's data shows, developers are increasingly aware that AI output needs more scrutiny, not less. Teams that skip that scrutiny because velocity metrics look good are borrowing against detection lag, which is exactly how debt becomes expensive.
None of this means AI-generated code is inherently worse than human-written code in every case. It means the debt AI creates behaves less like a fixed cost and more like variable-rate debt: the interest rate depends entirely on how rigorously a team reviews, tests, and refactors what the model produces, and most teams have not yet adjusted those practices to match the new volume of code entering the system.
AI Coding Tools vs. Traditional Development: What Changes for Technical Debt
Dimension | Traditional human-written code | AI-assisted code (unmanaged) |
Code volume per engineer | Lower, bounded by typing and thinking speed | Higher, bounded by review and verification speed |
Refactoring share of changes | Roughly 25% of changed lines historically | Under 10% by 2024, per GitClear |
Duplicate code blocks | Occasional, usually caught in review | Rose roughly 8x in 2024 alone, per GitClear |
Short-term churn (rework within 2 weeks) | Lower, near 3% baseline | Nearly doubled, approaching 6-7% |
Security vulnerability rate | Varies by team maturity | 45% of AI-generated task completions introduced an OWASP Top 10 flaw, per Veracode |
Developer trust in output | High, based on direct authorship | 46% of developers report low trust, per Stack Overflow |
System-level delivery stability | Baseline | Down roughly 7% per 25% increase in AI adoption, per DORA |
This is not an argument that AI coding tools should be avoided. It is evidence that the debt profile changes shape, and that teams measuring only velocity or lines-of-code output are missing the metrics that actually predict cost.
A Practical Framework: The Debt Velocity Ratio
Because no single, universally validated metric yet captures AI's net effect on technical debt, engineering teams need a working model they can apply now, using data most already collect. The following is an original analytical framework, not a scientifically validated methodology, offered as a starting point for internal measurement rather than a benchmark to compare across companies.
Debt Velocity Ratio (DVR) compares the rate at which new code enters a system to the rate at which that code is stabilized, tested, and integrated with existing patterns.
DVR = (Lines of new or AI-generated code merged per sprint) ÷ (Lines refactored, deduplicated, or brought under test coverage per sprint)
A rising DVR over consecutive sprints signals that code is entering the system faster than the team is folding it into a maintainable structure, which is the leading indicator of compounding debt, well before it shows up in defect rates or incident counts.
Three practical thresholds teams can use as a starting discussion point, not a rigid rule:
DVR trending flat or declining: New code is being integrated at a sustainable pace. Refactoring and consolidation are keeping up with generation.
DVR rising moderately over 2 to 3 sprints: Early warning sign. Worth a targeted review of the highest-churn files and modules before the pattern becomes structural.
DVR rising sharply or sustained over multiple quarters: Strong signal that debt is compounding faster than the team can service it, similar in spirit to a company whose debt-service ratio is deteriorating even while revenue grows.
Pair DVR with the churn and duplication signals GitClear tracks, and with change failure rate and time to restore from the standard DORA metrics, to avoid over-indexing on a single number.

When Does AI Actually Help Reduce Technical Debt?
The evidence above should not be read as a blanket case against AI in engineering workflows. There are specific, well-supported use cases where AI assistance measurably helps rather than hurts debt management.
Test coverage and documentation
DORA's 2024 research found AI adoption correlated with a 7.5% increase in documentation quality, the largest positive effect it measured, likely because AI lowers the effort barrier to writing and updating documentation that would otherwise be skipped under deadline pressure.
Legacy code comprehension
Understanding an unfamiliar, undocumented codebase is one of the more time-consuming tasks in software engineering. AI models are generally strong at summarizing and explaining existing code, even when they are less reliable at extending it correctly, which makes them useful for the diagnostic phase of a debt-reduction project even when a human should still own the actual refactor.
Mechanical, well-scoped refactors
Tasks with a narrow, verifiable scope, such as updating a deprecated API call across a codebase or converting a test suite to a new framework, are a better fit for AI assistance than open-ended feature work, because the correctness criteria are clear and easy to verify automatically.
Code review augmentation, not replacement
DORA's data showing a 3.1% increase in code review speed under AI adoption suggests AI tools can help reviewers move faster on style and pattern-matching issues, freeing human attention for the architectural and security judgment calls that current models still handle unreliably.
The common thread is that AI helps most when the task has a narrow, verifiable scope and a human retains responsibility for the parts of the decision that require judgment about the specific system, not general pattern-matching.
What Can AI Not Solve, No Matter How Good the Model Gets?
A concise answer: AI cannot decide what a codebase should be, only generate plausible continuations of what it already is. Architectural decisions, prioritization tradeoffs, and judgment about which debt is worth paying down first all require context and accountability that current models do not carry.
Three categories of debt sit outside what any current AI coding tool can meaningfully address:
Architectural debt. Decisions about system boundaries, data ownership, and service structure require understanding business tradeoffs and organizational constraints that are not encoded in a code repository. AI can implement a decision once it is made; it cannot reliably make the decision.
Prioritization debt. Knowing which of a hundred possible refactors will actually reduce incident rates or unblock a roadmap requires institutional knowledge of where the real pain points are, not just where the code looks messiest.
Accountability debt. When a shortcut causes a production incident, someone needs to own the postmortem and the fix. An AI model has no stake in the outcome and cannot be held accountable for a bad tradeoff, which means humans still carry full responsibility for reviewing what gets shipped, even as more of it is machine-generated.
Read Also: Chris Ross on Building AI-Native Startups That Beat the Odds
A Realistic Scenario: What Unmanaged AI Debt Looks Like in Practice
The following is a hypothetical, illustrative scenario, not a documented case study, built to show how the mechanisms above combine in a plausible team setting.
A mid-sized SaaS engineering team of 40 developers adopts an AI coding assistant broadly, with leadership tracking pull requests merged per week as the primary success metric. Output rises noticeably within two months. Leadership treats this as validation and reduces planned headcount for a deferred platform modernization project, reallocating that budget to new feature work.
Over the following two quarters, code review times lengthen slightly, though not enough to trigger concern on its own. Bug reports tied to edge cases in less-frequently-touched modules increase. A security review ahead of a customer audit surfaces several input-validation gaps in recently generated code, consistent with the kind of OWASP-category issue Veracode's research found in AI output at meaningful rates. Fixing them requires stopping feature work for two sprints. Onboarding time for new engineers increases because duplicated, inconsistent implementations of similar logic make it harder to identify the canonical pattern to follow.
None of these individual events looks catastrophic in isolation. Together, they resemble exactly the pattern DORA's research describes: higher individual output, declining system-level stability, and a debt load that was invisible in the metrics leadership was actually watching.
Risks, Limitations, and Open Questions in the Evidence
Honest analysis requires acknowledging what this evidence does not establish. GitClear's dataset, while large, draws on a specific mix of enterprise and open-source repositories and uses proprietary classification logic for what counts as refactoring versus churn; other analysts could draw the boundaries differently.
DORA's findings are based on self-reported survey data at scale, which is a well-established methodology but still subject to the biases inherent in any survey. METR's RCT is rigorous in design but small, 16 developers and 246 tasks in one specific setting, early-2025 tools on complex, unfamiliar-to-the-model open-source codebases, and the organization itself cautions against treating it as a permanent verdict, since model capability is improving quickly.
It is also worth stating plainly where the research disagrees with intuition: several vendor-sponsored studies and internal case studies report substantial productivity gains from AI coding tools, particularly for less experienced developers and for greenfield projects without deep legacy context.
The consistent pattern across the independent, non-vendor research cited here is narrower and more specific: gains concentrate at the individual task level and in specific use cases, while system-level stability and long-term maintainability show more mixed to negative signals, especially in complex, established codebases maintained by experienced teams, which is precisely the population METR studied.
Practical Takeaways for Engineering and Technology Leaders
Track code health metrics, not just output metrics. Deploy frequency and PR volume tell you almost nothing about whether debt is compounding; churn rate, duplication rate, and refactoring share tell you much more.
Treat AI-generated code as requiring the same or greater scrutiny as code from a new hire, not less, given the trust and debugging-time data from Stack Overflow's survey.
Budget explicit time for refactoring and consolidation in every sprint that uses AI-assisted development heavily, since AI tools do not perform this work unprompted and the data shows teams are doing less of it, not more, as adoption rises.
Run security scanning specifically calibrated for AI-output patterns, given the consistent 45% vulnerability introduction rate Veracode documented across models and languages.
Separate "faster" from "verified faster." Before reallocating headcount or deferring modernization work based on perceived AI productivity gains, validate those gains against delivery stability metrics, not developer sentiment alone.
Frequently Asked Questions
Does AI coding actually increase technical debt, or does it just make existing debt visible faster?
Both, based on current evidence. GitClear's data shows measurable increases in duplication and churn that represent genuinely new debt, not just faster detection of debt that was already there. At the same time, the higher volume of code changes does surface some pre-existing fragility in codebases sooner than it might have otherwise appeared.
Should engineering teams stop using AI coding tools because of these findings?
No. The evidence points to a governance gap, not a case against the technology. Teams that pair AI adoption with dedicated refactoring time, stronger security scanning, and metrics beyond raw output tend to avoid the worst outcomes described in this research. The risk comes from adopting AI while leaving review, testing, and debt-tracking practices unchanged.
How is AI-generated technical debt different from debt written by humans?
AI-generated debt tends to accumulate at higher volume because AI lowers the cost of writing code, and it often lacks the contextual memory a human author would carry about why a shortcut was taken. This makes it harder to locate and more expensive to safely modify later, a pattern sometimes called comprehension debt.
What metrics should a CTO track to know if AI is adding unmanaged technical debt?
Refactoring share of changed code, duplicate code block frequency, short-term churn rate, change failure rate, and time to restore service are the strongest available leading indicators, based on the GitClear and DORA research cited above. A rising ratio of code generated to code stabilized, the Debt Velocity Ratio described earlier, is a useful composite signal.
Does using a more advanced or newer AI model reduce these risks?
Not necessarily for security. Veracode's testing across more than 100 models found that newer models improved at producing functionally correct code but showed no meaningful improvement in avoiding common security vulnerabilities. Model capability and secure-by-default behavior appear to be separate problems.
Is the productivity slowdown found by METR likely to apply to every team?
Not universally. METR studied experienced developers working on complex, unfamiliar-to-the-model open-source codebases, a scenario where AI tools have less context to draw on. Less experienced developers, greenfield projects, and well-documented codebases may see different results, which is why the researchers frame this as one data point rather than a general law.
What is the single most useful first step for a team worried about AI-driven technical debt?
Start measuring refactoring share and code churn alongside existing velocity metrics before making any policy change. Most teams currently have no visibility into whether their debt is compounding, which makes it impossible to know whether a governance intervention is even needed yet.
The Bottom Line
AI coding tools change how fast code gets written. They do not, by themselves, change whether that code fits cleanly into the system it joins. The research available so far, from GitClear's 211-million-line analysis to DORA's delivery-stability findings to METR's controlled productivity trial, points toward the same conclusion from different angles: unmanaged AI adoption tends to widen the gap between how much code a team produces and how well that code is understood, tested, and integrated.
Technical debt was never really a coding problem. It has always been a governance problem, a question of whether an organization tracks what it owes and pays it down deliberately, rather than letting it accrue. AI does not resolve that governance question. It raises the stakes of getting the answer wrong, because it can add debt at a pace few review processes were designed to catch.
Read Also:
Why AI Strategies Fail When the Technology Works
If AI Does the Junior Work, Where Do Experts Come From
Comments (0)
No comments yet. Be the first to share your thoughts!