software complexity: how it is measured and what real code bases show (authored by agents unless marked đ§)
takeaway
- đ§ question: âhow to avoid the monotonic growth of software complexityâ
- source: research notes
- this file: how complexity and its growth are measured, and what the data says
- causes and cures: sibling file
- fact: size grows almost everywhere it was measured, but the shape differs
- Linux: faster than linear in the 1990s, about linear since 2.6
- whole collections (Debian, all public code): exponential, mostly because more projects exist
- one third of 1,519 Android apps shrank over their life
- fact: per-function complexity in Linux and Unix did not grow; it fell
- mostly because many small functions were added
- so âcomplexity grows monotonicallyâ is true for size, options and dependencies, not for the classic per-function scores
- fact: most code-level scores add little once you know the size
- 121 scores failed to predict how well people understood code
- I think size plus a few structure counts is the honest baseline for any new score
- fact: most installed code is not used, by every definition tried
- reported unused share runs from 20% to over 99%
- the definitions differ, so these numbers cannot be compared
- fact: LLM agents add code faster; whether the code is worse per line is disputed
- best study so far: Cursor adoption gave +28.6% lines added, and about +9% complexity after controlling for code size
- agents rarely add new dependencies in the one study that checked: 1.3% of pull requests
- opinion: the best openings are measurements nobody has redone or joined up
- Linux size, options and deletions for 2008 to 2026
- Rust dependency growth and its build cost since 2022
- which code properties make coding agents fail on later changes
- details in âresearch we can doâ
how to read the source entries
- each source has a
read:line- full: whole text read by me or by a reading agent I started
- part: abstract, method, results and limits read, not every page
- abstract: abstract or landing page only
- quotes are verbatim and were checked by program against the downloaded text
- exception: the 8 sources kept from the first version of this file, and Wirth (a scan)
- âIâ is the agent that wrote this file
quantities to keep separate
- size: amount of maintained code and supported behavior
- source lines, tokens, functions, dependencies and configuration options
- distinguish handwritten, generated, vendored and test code
- dependency complexity: what must be considered together when making a change
- direct imports and calls
- indirect dependencies through other components
- shared state, schemas and compatibility rules
- files repeatedly changed together
- a historical association, not proof of a runtime dependency
- cognitive difficulty: how hard one person finds one understanding task
- time to a correct explanation or prediction
- wrong answers and missed cases
- maintenance cost: resources needed to complete a correct change
- implementation, validation and review time
- regressions and rework
- elapsed issue time includes waiting, so it is not engineering effort
- technical debt: extra future cost blamed on a present decision
- warning counts and tool-estimated repair time are stand-ins for this cost
- I would call tool output âstatic-analysis findingsâ unless cost was checked separately
part 1: do the scores measure anything beyond size?
- Thomas J. McCabe, A Complexity Measure, 1976
- read: full
- cyclomatic complexity: number of branches in a function plus one
- âcomplexity depends only on the decision structure of a programâ
- says nothing about names, unfamiliar APIs or coordination between machines
- Landman, Serebrenik, Bouwers, Vinju, CC and SLOC in Java methods and C functions, JSEP 2016
- read: part
- summary, results on trimming, conclusion
- 17.6 million Java methods, 6.3 million C functions
- âlinear correlation between SLOC and CC is only moderateâ
- âCC summed over larger code units measures an aspect of system size rather than internal complexity of subroutinesâ
- best fit when trimming large functions: RÂČ 0.60 for Java, 0.67 for C
- so a per-file or per-repository branch total is mostly a size count
- a 2017 corrigendum exists
- still unread; the publisher blocked the download
- read: part
- El Emam, Benlarbi, Goel, The Confounding Effect of Class Size on the Validity of Object-Oriented Metrics, 1999 report; TSE 2001
- read: part
- one large C++ telecom system; class scores against field faults
- âAfter controlling for size none of the metrics we studied were associated with fault-proneness anymore.â
- âfuture validation studies should always control for sizeâ
- limit: one system; fault severity ignored
- Chowdhury, Holmes, Zaidman, Kazman, Revisiting the debate: are code metrics useful for measuring maintenance effort?, EMSE 2022
- read: part
- about 730,000 Java methods, 47 projects
- outcome: how often a method later changes
- âthe widely adopted size normalization approach fails to neutralize the size influenceâ
- âcode metrics can in fact help estimate maintenance effort, such as change proneness, even when the confounding influence of size is eliminatedâ
- limit: âchanges oftenâ is not cost
- this is the main counterweight to El Emam
- Scalabrino and colleagues, Automatically Assessing Code Understandability, TSE 2019
- read: part
- 444 human evaluations from 63 developers, 121 scores
- ânone of the 121 experimented metrics is able to capture code understandability, not even the ones assumed to assess quality attributes apparently related, such as code readability and complexityâ
- limit: short snippets
- Peitek, Apel, Parnin, Brechmann, Siegmund, Program Comprehension and Code Complexity Metrics: An fMRI Study, ICSE 2021
- read: part
- 19 people reading short snippets in a brain scanner, more than 41 scores
- âa codeâs textual size drives programmersâ attention, and vocabulary size burdens programmersâ working memoryâ
- âthere is no single metric that predicts the overall cognitive effortâ
- Muñoz Barón, Wyrich, Wagner, An Empirical Validation of Cognitive Complexity, ESEM 2020
- read: full
- Cognitive Complexity: SonarSourceâs score that adds a penalty for nesting
- about 24,000 evaluations of 427 snippets from ten studies
- âCognitive Complexity positively correlates with comprehension time and subjective ratings of understandabilityâ
- correctness gave mixed results
- data: Zenodo
- Lavazza, Abualkishik, Liu, Morasca, An empirical evaluation of Cognitive Complexity, JSS 2023
- read: abstract and publisher excerpts
- âthe performance of models that use âCognitive Complexityâ is extremely closeâ
- compared with models using only older scores
- data: Zenodo
- Gopstein and colleagues, Understanding Misunderstandings in Source Code, FSE 2017
- read: full
- tiny C patterns that people misread, such as assignment inside a condition
- âa significantly increased rate of misunderstanding versus equivalent code without the patternsâ
- 73 participants on snippets, 43 on larger programs
- spelling out steps can help even when code gets longer
- SjÞberg, Yamashita, Anda, Mockus, DybÄ, Quantifying the Effect of Code Smells on Maintenance Effort, TSE 2013
- read: abstract
- six paid developers, three tasks, four equivalent Java systems, measured hours
- âNone of the 12 investigated smells was significantly associated with increased effortâ
- after adjusting for file size and number of changes
- rare: real effort, not a stand-in
- Nagappan and Ball, Use of Relative Code Churn Measures to Predict System Defect Density, ICSE 2005
- read: abstract
- churn: lines added, deleted or changed
- âabsolute measures of code churn are poor predictors of defect densityâ
- churn relative to component size works better; Windows Server 2003
- Banker, Datar, Kemerer, Zweig, Software Complexity and Maintenance Costs, 1990 working paper; CACM 1993
- read: part
- maintenance projects at one large COBOL site, real project cost
- high-complexity projects âcost approximately 35% more than similar projects dealing with less complex codeâ
- limit: 1980s, one site; size and branching mixed together
- still one of the few studies with money as the outcome
- Besker, Martini, Bosch, Software developer productivity loss due to technical debt, JSS 2019
- read: abstract
- âdevelopers waste, on average, 23% of their time due to TDâ
- limit: 43 developers reporting their own time
- Tornhill and Borg, Code Red: The Business Impact of Code Quality, TechDebt 2022
- read: part
- 39 company code bases, 30,737 files, issue-tracker time per file
- âlow quality code contains 15 times more defects than high quality codeâ
- limit: the authors work for CodeScene, which sells the score
- âCode Health is a proprietary metric that is automatically calculated in the CodeScene tool.â
- I did not find a size control
- Baldwin, MacCormack, Rusnak, Hidden Structure: Using Network Methods to Map System Architecture, Research Policy 2014
- read: part
- 1,286 releases of 17 systems as file dependency graphs
- core: the largest group of files that all depend on each other, directly or indirectly
- âwe find that the majority of releases possess a âcore-peripheryâ structureâ
- âopen, distributed organizations develop systems with smaller Cores, while closed, co-located organizations develop systems with larger Coresâ
- limit: describes structure, does not test cost
- Mo, Cai, Kazman, Xiao, Feng, Decoupling Level: A New Metric for Architectural Maintenance Complexity, ICSE 2016
- read: part
- 108 open source and 21 industrial projects
- âwe still cannot reliably measure if one design is more maintainable than anotherâ
- own limit: âthe maintenance measures we proposed in Section 4 may not reflect the true maintenance effortâ
- Arif, Kuutila, Ralph, Assessing the Construct Validity of Object-Oriented, Class-Level Code Quality Metrics, arXiv, September 2026
- read: abstract
- not peer reviewed
- âTen metrics did not correspond to any known dimension of software quality and were removed in the exploratory analysis.â
- the remaining 24 scores group into 6 things: size, cohesion, coupling in, coupling out, and two about inheritance
- Kudrjavets, Rastogi, Thomas, Nagappan, On Quantifying the Benefits of Dead Code Removal, ICSME 2022
- read: abstract
- one page; âHowever, not all LOC are equalâ
- asks for a way to rank removal work; gives none
- what I take from part 1
- any claim âX raises complexityâ must show the effect per line of code or with size controlled
- scores for single functions say little about a whole system
- structure scores (core size, decoupling) are promising but were never tested against real effort in what I read
- real effort or money was measured only by SjĂžberg 2013 and Banker 1993
part 2: how real code bases grow
- my own check: Linux release archive sizes on kernel.org, listed 7 October 2026
- read: full (directory listings; sizes rounded to MB by the server)
.tar.xzsize: 25 MB (2.6.0, Dec 2003), 61 MB (3.0, Jul 2011), 78 MB (4.0, Apr 2015), 100 MB (5.0, Mar 2019), 128 MB (6.0, Oct 2022), 153 MB (7.2, Aug 2026)- about 5 to 8 MB added per year in every span since 2003
- 2011 to 2026: 2.5 times larger, about 6% per year
- so growth is about linear, never negative across these releases
- limit: compressed archive size is a rough stand-in for lines
- Godfrey and Tu, Evolution in Open Source Software: A Case Study, ICSM 2000
- read: full
- 96 Linux versions, 1994 to 2000
- Linux has been âgrowing at a super-linear rate for several yearsâ
- âmore than half of the code consists of device drivers, which are relatively independent of each otherâ
- any compiled kernel is âlikely to include only fifteen to fifty percent of the source files in the full source treeâ
- so most of the growth is optional code a given user never builds
- Israeli and Feitelson, The Linux kernel as a case study in software evolution, JSS 2010
- read: part
- 810 versions, 1994 to 2008
- growth is faster than linear up to 2.5, then âcloser to linearâ in 2.6
- median branches per function: âThis was 4 in 1994, 3 from 1995 to the beginning of 2003, and 2 since then.â
- âthe average complexity of functions is decreasing with time, but this is mainly due to the addition of many small functions.â
- configuration options âseem to be growing at an ever increasing rate.â
- Robles, Amor, Gonzalez-Barahona, Herraiz, Evolution and Growth in Large Libre Software Projects, IWPSE 2005
- read: part
- 18 large open source projects
- âsuper-linearity occurs only exceptionally, that many of the systems follow a linear growth pattern and that smooth growth is not that common.â
- dips come from removed or restructured code
- the Evolution mail client shrank at least twice after heavy refactoring
- Herraiz, Rodriguez, Robles, Gonzalez-Barahona, The Evolution of the Laws of Software Evolution, ACM CSUR 2013
- read: part
- a review of how Lehmanâs laws were tested
- on a study of 8,621 SourceForge projects: âaround 40% of the projects showed a superlinear pattern, incompatible with the laws.â
- results depend on the level you measure: whole system, subsystem or file
- Hatton, Spinellis, van Genuchten, The long-term growth rate of evolving software, JSEP 2017
- read: part
- 2,118 projects, 404 million lines, 9 closed source systems
- âsoftware source code in systems doubles about every 42 months on average, corresponding to a median compound annual growth rate (CAGR) of 1.21 ± 0.01.â
- âThere was no evidence to suggest any obvious relationship between either project duration in years and CAGR, or project size in LOC and CAGR.â
- limit: one growth rate from first to last release hides the shape
- compare: Linux since 2011 grew about 6% per year by archive size, well under 21%
- Gonzalez-Barahona, Robles, Michlmayr, Amor, German, Macro-level software evolution: a case study of a large software compilation, EMSE 2009
- read: part
- Debian 2.0 to 4.0, 1998 to 2007
- âstable releases double in size (measured by number of packages or by lines of code) approximately every two years.â
- âthe mean size of packages has remained almost constantâ
- so Debian grew by adding packages, not by packages growing
- limit: six releases
- Rousseau, Di Cosmo, Zacchiroli, Software provenance tracking at the scale of public source code, EMSE 2020
- read: part
- all of Software Heritage: 4 billion distinct files, 1 billion commits
- âWe find the growth rates to be exponential over a period of more than 40 years.â
- new distinct files double about every 22 months, new commits about every 30 months
- this counts how much code exists in public, not how big one system is
- Dorner, Capraro, Barcomb, Wnuk, Quo Vadis, Open Source? The Limits of Open Source Growth, arXiv 2022
- read: part
- 172,833 projects tracked by Open Hub
- âAfter an initial exponential growth, all measurements show a monotonic downwards trend since its peak in 2013.â
- limit: Open Hub is a hand-picked sample; half the projects had no activity data
- Kuiter and colleagues, How Configurable is the Linux Kernel? Analyzing Two Decades of Feature-Model History, TOSEM manuscript, 2025
- read: part
- build options of Linux, 2002 to 2024, all processor families
- âThe total number of features grows linearly over time for both extractors (r = 0.99, p < 0.001).â
- a typical release âadds 221 featuresâ and âremoves 69 featuresâ
- about 20,000 options in September 2024
- a drop in 2018 came from removing several processor families
- the number of possible kernel builds grows exponentially
- Lotufo and colleagues, Evolution of the Linux Kernel Variability Model, SPLC 2010
- read: part
- x86 build options, 2.6.12 to 2.6.32
- 3,284 options grew to 6,319
- âthe number of features had doubled, and still the structural complexity of the model remained roughly the same.â
- âremoving features is a rare motive for edits.â
- Bagherzadeh, Kahani, Bezemer, Hassan, Dingel, Cordy, Analyzing a Decade of Linux System Calls, EMSE 2018
- read: part
- 2005 to 2014
- â76 system calls were added to and 6 system calls were removed from the kernelâ
- â40 out of 76 (53%) new system calls were sibling callsâ
- sibling: a near copy of an old call that fixes its interface
- public interfaces pile up because old ones cannot be removed
- Zhou, Chen, Mockus, Wu, On the Scalability of Linux Kernel Maintainersâ Work, FSE 2017
- read: part
- 2009 to 2016
- âthe number of files does not appear to be increasing for a median maintainerâ
- âadding more maintainers to a file yields only a power of 1/2 increase in productivityâ
- Linux held load per person flat by adding maintainers
- Spinellis and Avgeriou, Evolution of the Unix System Architecture: An Exploratory Case Study, TSE 2021
- read: part
- Unix from 1970 to todayâs FreeBSD
- âthe systemâs source code grew by three orders of magnitude, from 13 thousand to more than ten million lines of code.â
- âcyclomatic complexity has been religiously safeguarded.â
- per-function complexity rose steeply, then slowly fell
- Nayebi, Kuznetsov, Chen, Zeller, Ruhe, Anatomy of Functionality Deletion, MSR 2018
- read: part
- 1,519 open source Android apps, 14,238 releases
- â98.8% of apps had decreased their size at least once over their lifetime.â
- â33.3% of apps even had a decreasing size trend over timeâ
- the clearest evidence that growth is a choice, at least for small apps
- Azad, Laperdrix, Nikiforakis, Less is More: Quantifying the Security Benefits of Debloating Web Applications, USENIX Security 2019
- read: part
- across the versions studied: â82% LLOC increase for phpMyAdmin, 99% for MediaWiki, and 171% for Magentoâ; WordPress went down 2%
- Bijlani, Ramachandran, Campbell, Where did my 256 GB go?, POMACS 2021
- read: part
- Android apps, 2014 to 2019
- âaverage app size in each category grew at least by 50% in five yearsâ
- Prokhorenko and colleagues, Analyzing the Evolution of Inter-package Dependencies in Operating Systems: A Case Study of Ubuntu, ECSA 2023
- read: part
- 84 Ubuntu images, 2005 to 2023
- âthe live Ubuntu image size grew from 600MB (version 5.04) to 3.7GB (version 23.04)â
- âwhereas the average total number of dependencies largely stays the same, developer-facing complexity tends to decrease over timeâ
- HTTP Archive, Web Almanac 2025, Page Weight
- read: part
- âYear over year, the median home page size grew 7.8% to 2.7 MB.â
- chart: mobile median 505 KB in October 2014, 2,559 KB in July 2025
- median mobile home page carries 632 KB of JavaScript
- the chapter text gives 2,362 KB in one place; I think that is a typo, the chart is consistent
- Gerard Holzmann, Code Inflation, IEEE Software 2015
- read: full
- opinion column with numbers
- the shell grew âFrom about 11K bytes in 5th Edition Unix in 1974, to 2.1M bytes for bash forty years later: an increase of 191 times.â
- âSo, why does software grow with time? The answer seems to be: because it can.â
- Niklaus Wirth, A Plea for Lean Software, IEEE Computer 1995
- read: part (scanned pages 1 to 4; quote typed from the scan)
- opinion
- âSoftware is getting slower more rapidly than hardware becomes faster.â
- Ben Kero, Trends in Mozillaâs central codebase, blog, 2015
- read: full
- not peer reviewed
- âFirefox 5 is about 3.4 million lines of code while Firefox 35 is almost exactly 6.6 million linesâ
- 2011 to 2015
- I found no peer-reviewed size series for any browser
- what I take from part 2
- growth of one mature system looks linear, set by how many people work on it
- exponential numbers come from whole collections, where the count of projects grows
- what grows without limit is the set of things kept alive: options, drivers, system calls, packages
- removal happens (69 Linux options per release, a third of apps shrink) but is always smaller than addition in big systems
- per-function scores miss all of this
part 3: dependencies and unused code
growth of dependencies
- Kikas, Gousios, Dumas, Pfahl, Structure and Evolution of Package Dependency Networks, MSR 2017
- read: part
- npm, RubyGems, and Rust projects on GitHub, to 2016
- transitive dependency: a package you get because something you use needs it
- âthe number of transitive dependencies for JavaScript has grown 60% over the last yearâ
- mean transitive dependencies per project: JavaScript 54.6, Ruby 34.1, Rust 9.3
- Decan, Mens, Grosjean, An Empirical Comparison of Dependency Network Evolution in Seven Software Packaging Ecosystems, EMSE 2019
- read: part
- data to early 2017
- âWe observe that Cargo and CPAN reveal a linear growth for both size metricsâ
- npm and CRAN grew exponentially
- half the packages that have dependencies in Cargo, npm and NuGet âhave at least 41, 21 and 27 transitive dependencies, respectively, where their median number of direct dependencies is only 2.â
- Zimmermann, Staicu, Tenny, Pradel, Small World with High Risks: A Study of Security Threats in the npm Ecosystem, USENIX Security 2019
- read: part
- âthe number of transitive dependencies of an average package has increased to a staggering 80 in 2018â
- maintainers able to affect more than 10,000 packages: 59 in 2015, 391 in 2018
- Schueller, Wachs, Servedio, Thurner, Loreto, Evolving collaboration, dependencies, and use in the Rust Open Source Software ecosystem, Scientific Data 2022
- read: full
- a dataset, not an analysis: 91,437 crates, 5.6 million commits, to September 2022
- growth numbers are only in its figures and data
- Li and colleagues, Demystifying Compiler Unstable Feature Usage and Impacts in the Rust Ecosystem, ICSE 2024
- read: abstract
- âWe have analyzed the whole Rust ecosystem with 590K package versions and 140M transitive dependencies.â
- useful as a method for resolving every dependency tree on crates.io
- Arafat, How Deep Does Your Dependency Tree Go?, arXiv, December 2025
- read: abstract and method
- one author, not peer reviewed, 50 popular packages per ecosystem
- Maven projects pull in 24.7 times their direct dependencies on average, npm 4.3 times
how much is unused
- Soto-Valero, Harrand, Monperrus, Baudry, A Comprehensive Study of Bloated Dependencies in the Maven Ecosystem, EMSE 2021
- read: part
- 9,639 Java packages, 723,444 dependency relations
- â(75.1%) of all dependencies are bloated, they are not needed to compile and run the code.â
- 18 of 21 answered removal requests were merged
- Drosos, Sotiropoulos, Spinellis, Mitropoulos, Bloat beneath Pythonâs Scales, FSE 2024
- read: part
- 1,302 Python projects
- âmore than 50% of dependencies are bloatedâ
- âon average, 87% of the dependency source files are bloated.â
- 28 of 36 removal requests were merged
- Jokƫbauskas, Investigating dependency code reuse using callgraphs, TU Delft MSc thesis
- read: part
- thesis, year not checked (around 2020), not peer reviewed
- all of crates.io
- âon average 91% of lines of code come from external dependenciesâ
- âin 95% of packages, 71% of their callgraphs are never usedâ
- the only whole-ecosystem Rust number I found; old and never repeated
- Latendresse, Mujahid, Costa, Shihab, Not All Dependencies are Equal, ASE 2022
- read: part
- 100 JavaScript projects
- âless than 1% of the installed dependencies are released to productionâ
- installed includes build and test tools, hence the tiny share
- Liu, Tiwari, Bogdan, Baudry, Detecting and removing bloated dependencies in CommonJS packages, arXiv 2024
- read: abstract
- 91 packages; 50.6% of 50,488 dependencies never loaded when tests run
- Quach, Prakash, Yan, Debloating Software through Piece-Wise Compilation and Loading, USENIX Security 2018
- read: part
- âonly 5% of libc is used on average across the Ubuntu Desktop environment (2016 programs); the heaviest user, vlc media player, only needed 18%.â
- Kurmus and colleagues, Attack Surface Metrics and Automated Compile-Time OS Kernel Tailoring, NDSS 2013
- read: part
- building Linux for one workload removes much of what an attacker can reach: the reduction âranges from about 50% to 85%â
- Kupoluyi and colleagues, Muzeel: A Dynamic JavaScript Analyzer for Dead Code Elimination in Todayâs Web, arXiv 2021
- read: abstract
- about 40,000 web pages
- â70% of JavaScript functions on the median page are unusedâ
- unused means not triggered by a botâs clicks
- HTTP Archive, Web Almanac 2024, JavaScript
- read: part
- 206 KB, â44% of bytes deliveredâ, unused during page load on the median mobile page
- Zhang and colleagues, Machine Learning Systems are Bloated and Vulnerable, arXiv 2024
- read: abstract
- 15 container images
- âbloat accounts for up to 80% of machine learning container sizesâ
- Brown and colleagues, A Broad Comparative Evaluation of Software Debloating Tools, USENIX Security 2024
- read: part
- 10 removal tools on 20 programs
- âonly 13% of our debloating attempts produced a sound and robust debloated programâ
- automatic removal after the fact mostly fails on real programs
what dependencies cost
- Weeraddana and colleagues, Dependency-Induced Waste in Continuous Integration, FSE 2024
- read: part
- 20,743 commits in 1,487 npm projects
- â55.88% of the CI build time that is associated with dependency updates is only triggered by unused dependenciesâ
- update bots cause most of it, âcontributing 92.93% of the CI build timeâ
- the only paper I found that prices a dependency in time; it is machine time
- Pashchenko, Plate, Ponta, Sabetta, Massacci, Vulnerable Open Source Dependencies: Counting Those That Matter, ESEM 2018
- read: part
- 200 Java libraries used at SAP
- âabout 20% of the dependencies affected by a known vulnerability are not deployedâ
- Bogart, KĂ€stner, Herbsleb, Thung, How to Break an API, FSE 2016
- read: part
- 28 interviews in Eclipse, npm and CRAN
- ecosystems differ in who pays for a breaking change
- âlong-term stability is a key value of the Eclipse community: this shifts costs to the developers making the changeâ
- no hours measured
- what I take from part 3
- dependency counts grew fast in npm; the data for Rust stops at 2017 for analysis and 2022 for raw data
- âunusedâ has at least five meanings
- not needed to build (Maven 75%)
- not reachable in the call graph (Python over 50%, Rust 71%)
- not loaded when tests run (JavaScript 51%)
- not shipped to production (JavaScript over 99%)
- not run in a page visit (web 44% of bytes, 70% of functions)
- nobody applied two definitions to the same projects
- cost in human hours is unmeasured
part 4: LLM-written code
studies of whole projects
- He, Miller, Agarwal, KĂ€stner, Vasilescu, Speed at the Cost of Quality: How Cursor AI Increases Short-Term Velocity and Long-Term Complexity in Open-Source Projects, MSR 2026
- read: full (I read the main text, not the appendix)
- 806 GitHub projects that committed a Cursor rules file, 1,380 similar projects that did not
- method: compare each groupâs change before and after adoption, month by month, January 2024 to August 2025
- âLines added increase by about 28.6% (Table 2). There is no statistically significant effect for the volume of commits.â
- âThe only significant development velocity gain is in the first two months post Cursor adoption.â
- âstatic analysis warnings increase significantly by 30.3%, and code complexity increases by 41.6%. The effect on duplicate line density is insignificant.â
- complexity here is SonarQubeâs Cognitive Complexity summed over the code base
- size matters most: âincreases in codebase size are a major determinant of increases in static analysis warnings and code complexity, and absorb most varianceâ
- after controlling for size, Cursor still gives âa 9% baseline increaseâ in complexity; the warning effect is no longer significant
- their second model: âA 100% increase in code complexity and static analysis warnings causes a 64.5% and 50.3% decrease in development velocity as measured by lines addedâ
- limits the authors state
- adoption is seen only through a committed file: âour sample represents repositories with observable Cursor adoption rather than all possible Cursor-adopting repositoriesâ
- they do not know how much Cursor was used
- mostly TypeScript, Python and JavaScript
- limits I see
- the discussion says complexity rose 25.1% and cites the same table that says 41.6%
- velocity is lines added, so âcomplexity slows velocityâ partly says big projects add proportionally less
- data: Zenodo
- GitClear, Coding on Copilot, 2024
- read: part
- vendor report, not peer reviewed, 153 million changed lines
- churn here: code âpushed to the repo, then subsequently reverted, removed or updated within 2 weeksâ
- 2020 to 2023: moved lines 25.0% to 16.9%, copied lines 8.3% to 10.5%, churn 3.3% to 5.5%
- limit: no label says which code an LLM wrote; it is a trend over years
- its 2024 numbers are a projection, not data
- the 2025 report says copied lines reached 12.3% in 2024
- read: landing page only
- Daniotti, Wachs, Feng, Neffke, Who is using AI to code?, Science 2026
- read: part
- a classifier guesses which Python functions on GitHub an LLM wrote
- âAI writes an estimated 29% of Python functions in the USâ
- measures share and output, not quality
- Mao and colleagues, A Large-Scale Comprehensive Measurement of AI-Generated Code in Real-World Repositories, arXiv 2026
- read: part
- finds LLM code through comments that admit it
- âHuman-written code shows substantially higher duplication rates than AI-generated code, mainly due to cross-file duplication rather than within-file clones.â
- LLM-involved code gets more follow-up changes later
studies of agent pull requests
- Li, Zhang, Hassan, The Rise of AI Teammates in Software Engineering 3.0, arXiv 2025
- read: part
- the AIDev dataset; the 2026 version has 932,791 agent pull requests in 116,211 repositories
- most studies below use it
- Huang and colleagues, More Code, Less Reuse, MSR 2026
- read: part
- âWhile traditional metrics show minimal differences between agentic-PRs and human-PRs, redundancy metric analysis shows code in agentic-PRs contain significantly more redundancyâ
- limit: the redundancy result uses 617 pull requests from one repository
- Popescu and colleagues, Investigating Autonomous Agent Contributions in the Wild, MSR 2026
- read: part
- about 110,000 pull requests; do lines survive 3 weeks?
- âthe fraction of commits where all lines survived is consistently higher for humans than for any agentâ
- the gap is small
- Xia, Miller and colleagues, Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code, arXiv, July 2026
- read: part
- 182 repositories, not peer reviewed
- âagentic code receives a 46% higher corrective maintenance rate and a 45% higher bug-fixing rate on averageâ
- worse where pull requests merge without review
- Sawada and colleagues, To What Extent Does Agent-generated Code Require Maintenance?, EASE 2026
- read: abstract and limits
- âAI-generated files receive less frequent maintenance than human-authored codeâ
- only 508 agent-written files
- Horikawa and colleagues, Do AI Agents Really Improve Code Readability?, MSR 2026
- read: abstract
- 403 agent commits that claim to improve readability
- âthe Maintainability Index decreased in 56.1% of commits, while Cyclomatic Complexity increased in 42.7%â
- Hasan, Rabbi, Zibran, The Quiet Contributions, MSR 2026
- read: abstract and first result
- 4,762 agent pull requests merged without discussion
- 59.89% leave branch count unchanged, 36.88% raise it, âonly 3.23% of the SPRs reduce complexityâ
- Twist and Zhang, A Study of Library Usage in Agent-Authored Pull Requests, arXiv 2025
- read: abstract
- 26,760 pull requests
- âAgents often import libraries (29.5% of PRs) but rarely add entirely new dependencies (1.3% of PRs).â
- Cotroneo, Improta, Liguori, Human-Written vs. AI-Generated Code, ISSRE 2025
- read: part
- over 500,000 single functions written from a description
- âAI-generated code is generally simpler and more repetitive, yet more prone to unused constructs and hardcoded debuggingâ
benchmarks with many changes in a row
- Orlanski and colleagues, SlopCodeBench, arXiv v2, May 2026
- read: part
- agents extend their own code over 196 checkpoints in 36 problems
- erosion: share of all branch count sitting in functions with more than 10 branches
- âstructural erosion rising in 77% of trajectories and verbosity in 75.5%â
- âagent code is 2.3x more verbose and 2.0x more erodedâ
- than human repositories
- âExplicit quality guidance reduces initial verbosity and erosion by up to a third, without affecting degradation rates.â
- limit: small made-up command-line tasks; human baseline is whole repositories of other sizes
- Chen and colleagues, SWE-CI, arXiv 2026
- read: part
- 100 tasks replaying months of real project history
- âmost models achieve a zero-regression rate below 0.25â
- âall 20 LLMs underperform on MI scoreâ
- MI: Maintainability Index, a formula from lines, branches and operator counts
- the paper never checks that MI means anything here
- Deng and colleagues, EvoClaw, ICML 2026
- read: part
- replays sequences of real milestones
- scores drop âfrom >80% on isolated tasks to at most 38% in continuous settingsâ
- measures tests passed and broken, not code properties
speed and delivery
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, arXiv 2025
- read: part
- 16 developers, 246 tasks in their own large projects, AI allowed or not at random
- âallowing AI actually increases completion time by 19%â
- METR, uplift update, blog, 24 February 2026
- read: full
- later cohorts were 4% to 18% faster with AI, with wide error bars
- âour data is only very weak evidence for the size of this increaseâ
- many developers no longer submit tasks they would have to do without AI
- DORA, 2024 report and 2025 report
- read: part (AI chapter of 2024; summary of 2025)
- industry surveys of self-reported numbers
- 2024: delivery stability falls by âan estimated 7.2% reduction for every 25% increase in AI adoptionâ
- 2025: throughput now improves, but âit still increases delivery instability.â
- two more are covered in the sibling file
- Debt Behind the AI Boom: 22.7% of tool findings from AI commits still alive later
- Echoes of AI: 151 developers, no later slowdown from AI-written code
- what I take from part 4
- agents add more code; that part is solid
- âworse per lineâ rests on one +9% number, tool warnings, and custom scores
- duplication results conflict: GitClear and Huang say more, Mao says less, He says no change
- maintenance results conflict: Xia says more fixes, Sawada says fewer changes
- no study follows a real code base under agents for more than about a year
- no study uses structure measures (core size, dependency graph) on agent code
- every âqualityâ outcome here is a score that part 1 says is weak
where the sources disagree
- shape of growth
- one system: linear (Robles, Israeli, Kuiter, my archive sizes)
- many systems pooled: exponential (Hatton, Debian, Software Heritage)
- I think both hold; the pooled numbers count new projects
- is open source still growing?
- Software Heritage says exponential, Open Hub says decline since 2013
- different populations: everything public versus tracked team projects
- is code ever removed?
- Lotufo: rarely; Kuiter: 69 options per release; Nayebi: a third of apps shrink
- does a score add anything beyond size?
- no: El Emam, Scalabrino, Peitek
- yes, somewhat: Chowdhury, Landman, Muñoz Barón
- how many options does Linux have?
- Kuiter counts about 20,000; an LWN article reportedly counts over 32,000 for one build target
- I only saw a summary of the LWN article and did not trace the difference
- do dependency counts keep growing?
- npm: yes to 2018; Ubuntu binaries: flat over 18 years; Rust: unknown after 2017
what nobody has measured
- Linux lines, functions and per-function complexity after 2008, in a paper
- code per feature over time
- Lotufo found constant lines per option over 5 years; nobody extended it
- lines deleted as a share of lines added, per year, in mature systems
- Rust dependency counts after 2022, and build time or binary size caused by dependencies
- a dependencyâs cost in human hours
- whether unused share per project rises over time
- every study is one snapshot
- peer-reviewed build-time or browser size series
- whole-project growth and structure under coding agents over years
- whether any score predicts that a coding agent fails on the next change
research we can do
- đ§ wants: âsignificant & popular, easy sellâ and âeasy to implementâ (research notes)
- what makes code hard for agents to change?
- my pick for the best fit
- idea: the reader of code is now often an agent, so measure complexity by agent failure
- data: SWE-CI and EvoClaw replays, both public
- at each step compute size, branch counts, Cognitive Complexity, file coupling, core size
- test which ones predict that the agent breaks something in the next step
- baseline: size alone, as part 1 demands
- new because: benchmarks report scores next to failures but never link them
- nearest work: SWE-CI, SlopCodeBench, Scalabrino for humans, Complexity Backpressure (in the sibling file; I have not read its full paper)
- cost: compute only, no human subjects
- risk: size alone may explain everything; that is still a publishable negative result
- Rust dependency growth and what it costs
- data: the daily crates.io database dump, Schuellerâs dataset to 2022
- resolve the full tree for each release, 2015 to 2026; plot median and worst cases per year
- add build time and binary size caused by dependencies for the top 1,000 crates
- add wasted CI time from unused crates, repeating Weeraddana for Rust
- new because: Rust analysis stops in 2017, and nobody priced Rust dependencies
- nearest work: Decan 2019, Kikas 2017, the Delft thesis, Weeraddana 2024
- risk: Rust builds are cached and the linker drops dead code, so cost may be small
- Linux 2008 to 2026: lines, options, system calls, deletions
- data: the kernel git history
- redo Israeli and Feitelson, then add two new things
- lines per option and per system call over time
- lines deleted over lines added per release, and what got deleted
- new because: last full measurement ended in 2008
- nearest work: Israeli and Feitelson 2010, Kuiter 2025 for options
- risk: reviewers may call it a replication; the deletion and per-feature parts must carry it
- one set of projects, five meanings of âunusedâ
- take 200 projects in one language and apply all five definitions from part 3
- shows how much of the 20% to 99% spread is definition and how much is real
- nearest work: each definition has its own paper, none compares
- risk: tooling for five analyses in one language; JavaScript or Java is the practical choice
- agent adoption and whole-project structure
- data: AIDev repositories and Heâs Cursor dataset
- monthly snapshots: lines, deletion share, dependency graph, core size
- compare before and after adoption against matched projects, always per line
- new because: nobody used structure measures or deletion share on agent code
- nearest work: He 2026, Popescu 2026, Xia 2026, Baldwin 2014
- risk: crowded area; several MSR 2026 papers and a planned study by Coppola are close
- scripts and bundled packages per web page, 2016 to 2026
- data: HTTP Archive monthly crawls
- count script sources, bundled packages and unused bytes per page over time
- nearest work: Web Almanac (bytes only), Swierzy 2025 (update speed), Lauinger 2017 (one snapshot)
- risk: the crawlâs site sample changed over the years
- I would not build a new complexity score first
- part 1 shows new scores rarely beat size
- the sibling fileâs plan, measuring whether removal makes later changes cheaper, fits with idea 1
advice for any of these studies
- outcome first: correct completion of a later change, by a person or an agent
- define correct before running
- count failures and timeouts, not only time
- always report size next to any score
- totals, distributions and the worst components, not only averages
- averages fall when many small functions are added
- separate prediction from cause
- prediction: train on some projects, test on others, compare against size alone
- cause: assign the treatment at random when possible; otherwise matched projects, before and after
- count units honestly
- functions in one project are not independent projects
- traps
- lower function scores by splitting functions, with no gain in later work
- a warning that vanishes because the file was deleted or the rule changed
- tests that pass because removed behavior was never tested
- keeping only attempts that passed
- a âhumanâ baseline that used unmarked AI help
- newer code had less time to be fixed or removed; compare equal time windows
coverage
- researched 7 October 2026 UTC
- 80 sources with a
read:line- 10 read in full: McCabe, Muñoz BarĂłn, Gopstein, He, Godfrey, Holzmann, Schueller, Keroâs blog, METRâs 2026 blog, the kernel.org listings
- 54 read in part, 16 abstract only
- three reading agents did most of the reading; I read He and the Landman results myself and rechecked 80 quotes by program
- wanted but not obtained
- Potvin and Levenberg on Googleâs repository; Koch on SourceForge growth
- Gil and Lalouche 2017 and Fenton and Neil 1999 on size and score validity
- Hassan 2009 on change entropy; MacCormack and Sturtevant 2016 on coupling and defects
- Abdalkareem 2017 on trivial npm packages; Kuo 2020 on kernel debloating
- the GitClear 2025 report itself; the Landman corrigendum
- not a systematic review
- ChatGPT was not consulted for this file; the tool needed a sign-in
Last edited: