Review status
This summary is based on the full working paper text and verified article records.
The DOI has been verified through SSRN. The source URL has been verified through the official Harvard Business School working paper page.
The article is a working paper rather than a final journal publication. It is an empirical economic valuation study using large-scale OSS usage datasets, repository-level code data, wage data, and developer contribution data. It does not test formal causal hypotheses. Its contribution is measurement: estimating the supply-side and demand-side value of widely used open source software.
Research question
What is the economic value of open source software?
More specifically, the paper asks how valuable widely used OSS is when measured not only by the cost of recreating the code once, but also by the value firms receive from using OSS instead of building equivalent software internally.
Hypotheses
Not specified.
The paper is not structured as a hypothesis-testing study. Instead, it estimates the economic value of OSS using supply-side and demand-side replacement-cost logic.
Method
The paper uses an economic valuation design based on large-scale archival data.
The authors estimate two different forms of OSS value.
The first is supply-side value. This asks how much it would cost society to recreate widely used OSS once if the existing code disappeared but the concept of OSS still existed.
The second is demand-side value. This asks how much it would cost if OSS did not exist and every firm using an OSS package had to recreate that package internally.
The authors combine three main datasets.
The first dataset is Census II of Free and Open Source Software. This dataset was created by the Linux Foundation and the Laboratory for Innovation Science at Harvard using data from three software composition analysis firms. These firms scan client codebases to detect OSS use, often for license-compliance or due-diligence purposes. The Census data capture inward-facing OSS use, meaning OSS embedded in products that firms create. The Census contains more than 2.7 million observations of OSS package use from 2020. The final Census analysis covers 1,840 matched OSS packages after repository matching and cleaning.
The second dataset is BuiltWith. BuiltWith scans public websites and identifies technologies used by those websites. The authors use this dataset to capture outward-facing OSS use, meaning OSS used in company websites that customers or users directly interact with. The BuiltWith data include scans of 8.8 million unique websites and 72.8 million observations of OSS usage from January 1 to November 16, 2020. After matching websites to registered firms in Orbis, Compustat, and PitchBook, the authors obtain around 3.4 million firm websites. The final BuiltWith package analysis covers 741 matched OSS packages.
The third dataset is GHTorrent. This dataset records activity on GitHub and is used to study how OSS value creation is distributed across developers. After matching repositories and removing likely bot accounts, the final developer contribution sample contains around 60,000 developers and 2.3 million commits.
The authors identify code repositories for OSS packages, count lines of code using pygount, and identify programming languages using linguist. They classify languages into three buckets. The main estimates use bucket 1 only, which contains programming and markup languages most likely to be human-written. This makes the main valuation conservative.
The supply-side valuation uses the Constructive Cost Model II, also known as COCOMO II. This model estimates the person-month effort required to recreate software from scratch based on lines of code. The authors then multiply estimated effort by programmer wages.
The wage assumptions include three scenarios:
- a low-wage scenario based on Indian programmer wages;
- a global weighted wage based on the top 30 countries by GitHub developer activity;
- a high-wage scenario based on United States programmer wages.
The demand-side valuation multiplies the package-level replacement value by firm usage. The authors avoid counting multiple uses of the same package within the same firm more than once, because a firm could reuse internally recreated software as a club good.
The paper also analyzes heterogeneity by programming language, data source, industry, and developer contribution concentration. Developer contribution inequality is measured using Lorenz curves based on value contributions and repository contributions.
Results / key findings
The central result is that the demand-side value of OSS is vastly larger than the supply-side cost of recreating the code once.
Using the main labor-market valuation and bucket 1 languages only, the estimated supply-side value of widely used OSS is $4.15 billion under the global wage assumption. The low-wage estimate is $1.22 billion, and the high-wage estimate is $6.22 billion.
The estimated demand-side value is much larger. Under the global wage assumption, firms would need to spend about $8.8 trillion to recreate the widely used OSS they rely on. The low-wage estimate is $2.59 trillion, and the high-wage estimate is $13.18 trillion.
The paper compares this demand-side estimate to global software spending. The authors estimate that firms spent about $3.4 trillion on software used in 2020. Adding the estimated $8.8 trillion OSS replacement value implies that firms would need to spend about $12.2 trillion, or roughly 3.5 times their current software spending, if OSS did not exist.
The paper also shows that the supply-side value of widely used OSS is much smaller than estimates of the full OSS universe. Prior work estimated the global supply-side value of all OSS at around $78 billion. The authors’ $4.15 billion estimate for widely used, firm-relevant OSS is about 5.5% of that total. This difference reflects the paper’s focus on OSS that is actually used by firms rather than the full long tail of OSS projects.
Table 1 reports the underlying code and usage data. The Census sample contains 261,653,728 lines of code across 1,840 packages, with a mean of 142,203 lines per package. The top five programming languages plus Go account for 189,673,184 lines of code across 1,668 packages. Census package usage totals 2,709,155 observations, with 2,497,785 observations associated with the top five languages plus Go.
For BuiltWith, Table 1 reports 82,504,613 lines of code across 741 packages, with a mean of 111,342 lines per package. The top five languages plus Go account for 58,664,935 lines across 734 packages. BuiltWith package usage totals 142,794.4 observations, with almost all of this usage associated with the top five languages plus Go.
Figure 1 shows large differences across programming languages. On the supply side, Go has the highest estimated value at about $803 million, followed by JavaScript at $758 million and Java at $658 million. C accounts for about $406 million, TypeScript for $317 million, and Python for about $55 million.
On the demand side, Go is especially dominant. The paper reports that Go has more than four times the demand-side value of the next highest language, JavaScript. TypeScript ranks third, followed by C, Java, and Python.
Figure 2 separates the Census and BuiltWith data. The Census results are driven heavily by inward-facing code used in firm products. In the Census supply-side estimates, Go and Java are especially important. In the Census demand-side estimates, Go dominates the value distribution. In BuiltWith, which focuses on outward-facing website use, JavaScript dominates both supply-side and demand-side value, with TypeScript second.
The paper also estimates industry-level demand-side value using BuiltWith data. Figure 3 shows that Professional, Scientific, and Technical Services receives the highest outward-facing OSS value, at around $43 billion. Retail Trade follows at around $36 billion, and Administrative and Support and Waste Management and Remediation Services follows at around $35 billion. Industries such as Mining, Utilities, and Agriculture receive much smaller estimated outward-facing OSS value.
A major finding is that OSS value creation is highly concentrated among developers. Figure 4 shows Lorenz curves for developer contribution value. The top 5% of developers, about 3,000 developers in the sample, generate more than 93% of the supply-side value and more than 96% of the demand-side value.
This concentration is not only because a few developers contribute to a few extremely valuable repositories. The paper shows that high-value contributors also contribute to substantially more repositories than lower-value contributors. This suggests that a small group of developers plays a broad and central role in maintaining widely used OSS infrastructure.
The appendix reports robustness checks using broader language buckets. Including bucket 1 and bucket 2 languages raises the global supply-side estimate slightly from $4.15 billion to $4.18 billion and the global demand-side estimate from $8.80 trillion to $8.84 trillion. Including all three buckets raises the global supply-side estimate to $6.41 billion and the global demand-side estimate to $11.96 trillion.
The paper also includes an alternative goods-market valuation approach, but the authors place more weight on the labor-cost approach. The goods-market method requires stronger assumptions because it uses proprietary software substitutes and market prices. Depending on assumptions, this approach produces much lower estimates and is treated as less central to the paper’s contribution.
Overall, the paper shows that OSS is economically massive despite being priced at zero. Its value is hidden because standard economic measures struggle to capture free digital goods whose quantity of use is not centrally tracked.
Practical implications
For managers, the paper shows that OSS is not a minor technical convenience. It is a core input into modern software production.
The most important implication is that firms receive enormous value from OSS without paying market prices for it. If firms had to recreate the OSS they use internally, the authors estimate the global cost at about $8.8 trillion under the global wage assumption. This makes OSS a major source of cost avoidance and productivity.
Managers should therefore treat OSS as strategic infrastructure. Many firms rely on OSS in products, websites, analytics systems, AI tools, cloud infrastructure, and internal operations. If that infrastructure becomes insecure, undermaintained, or unavailable, the operational and financial exposure can be large.
The findings also challenge pure free-riding behavior. Because OSS is usually available at zero price, firms may underinvest in the projects they depend on. The paper’s value estimates suggest that contributing back to OSS can be far cheaper than recreating software internally or dealing with failures in critical dependencies.
For software, technology, and digital transformation managers, the paper supports more systematic OSS governance. This includes tracking OSS dependencies, understanding which packages are business-critical, contributing to important projects, funding maintainers, and managing license and security risks.
The concentration of value among developers is especially important. If 5% of developers create 96% of demand-side value, then firms and policymakers should pay close attention to maintainer sustainability. A small number of contributors may be supporting infrastructure used by millions of firms.
The paper is also relevant for boards and executives. OSS dependency is not only an IT issue. It is part of intangible capital, operational resilience, cybersecurity, innovation capability, and cost structure. Firms that depend heavily on OSS but lack visibility into that dependency may underestimate their true digital supply-chain risk.
For practitioners, useful diagnostic questions include:
- Which OSS packages does the organization depend on most?
- Which OSS dependencies are embedded in products, websites, and internal systems?
- Are critical OSS packages actively maintained?
- Does the firm contribute financially, technically, or organizationally to critical OSS infrastructure?
- Are OSS risks treated as part of digital supply-chain risk?
- Does the firm know which maintainers or communities support its most important dependencies?
- Would the firm be able to replace critical OSS if a project became unavailable, insecure, or abandoned?
- Are OSS contributions treated as strategic investment rather than charity?
Theoretical implications
The paper contributes to research on intangible capital by showing how large amounts of economic value can be missed when a good has a zero market price.
Traditional valuation logic often multiplies price by quantity. OSS creates a measurement problem because price is usually zero and quantity is hard to observe. The paper addresses this by combining large-scale usage data with replacement-cost methods.
The paper also contributes to research on public goods and digital commons. OSS is a global public good that many firms use without direct payment. This creates a classic commons problem: the resource is widely valuable, but maintenance and contribution may be underfunded because individual users can free ride.
The study also contributes to innovation management by showing that innovation inputs are often shared, cumulative, and distributed across organizational boundaries. Firms innovate using code created by external communities, volunteers, firms, foundations, and maintainers. This makes innovation less firm-contained than traditional R&D models suggest.
The paper contributes to strategy by making hidden software infrastructure more visible. Competitive advantage in digital markets often depends on software capabilities, but many of those capabilities are built on shared OSS foundations. This means firms compete partly on top of a common technological base.
The paper also contributes to the literature on IT productivity. If OSS reduces software production costs at massive scale, then some productivity gains may be hidden in cost avoidance rather than directly observed in revenue or expenditure data.
Finally, the developer concentration finding contributes to research on distributed production. OSS may look decentralized, but value creation is highly concentrated. This suggests that open production systems can combine broad participation with extreme dependence on a small group of central contributors.
Limitations
The paper is a working paper, so the findings should be interpreted as draft research rather than final journal-published evidence.
The estimates focus on widely used OSS, not the full universe of OSS. This is a strength for measuring economically relevant OSS, but it excludes long-tail projects and may miss value created by less visible packages.
The authors argue that the demand-side estimates are likely conservative because no dataset can capture 100% of OSS use globally. Important omitted categories include operating systems and other types of OSS not captured by the Census and BuiltWith data.
The valuation depends on the COCOMO II replacement-cost model. This is a standard software cost-estimation model, but it still requires assumptions about effort, productivity, and wages.
The demand-side valuation assumes that if OSS did not exist, each firm using a package would need to recreate it internally. This is useful as a replacement-cost thought experiment, but real markets might respond differently. Firms could buy proprietary alternatives, reduce software use, collaborate on private shared infrastructure, or redesign products.
The BuiltWith data focus on websites and are especially relevant for outward-facing web software. This may overrepresent JavaScript-related OSS in outward-facing use while underrepresenting non-web OSS.
The Census data are based on software composition analysis customers and the top packages observed by SCA vendors. This gives strong insight into firm product code, but the sample may not perfectly represent all global firms.
The paper measures economic replacement value, not necessarily consumer surplus, social welfare, security value, innovation spillovers, or strategic option value.
The developer contribution analysis uses commits as a proxy for contribution share. Commit counts are useful but imperfect because commits can vary greatly in complexity, importance, and quality.
Future research
Future research could estimate OSS value using additional sources of usage data, especially for operating systems, cloud infrastructure, AI libraries, embedded systems, and enterprise back-end systems.
Researchers could study how OSS dependency varies by firm size, country, industry, and digital maturity.
Future studies could examine whether firms that contribute more to OSS receive measurable advantages in productivity, innovation speed, security, recruitment, or reputation.
Another useful direction would be to study maintainer sustainability. If a small number of developers create most OSS value, researchers should examine how those developers are funded, governed, retained, and protected from burnout.
Future research could connect OSS dependency to cybersecurity risk. Firms may rely on packages without understanding who maintains them or how quickly vulnerabilities are patched.
Researchers could also compare OSS value across different governance models, such as foundation-led projects, firm-led projects, community-led projects, and hybrid models.
Another promising direction is policy research. Governments may need better tools to support critical digital infrastructure, especially where OSS creates large public value but receives insufficient private investment.
Future work could also refine valuation methods by comparing labor replacement value, goods replacement value, consumer surplus, productivity effects, and macroeconomic spillovers.
Finally, researchers could examine how generative AI changes the economics of OSS. AI may reduce the cost of creating software, but it may also increase dependency on OSS training data, package ecosystems, developer tools, and shared infrastructure.