Essay · Innovation & the economics of discovery
The Price of Being Wrong
Humanity spends three trillion dollars a year trying to discover things. Almost none of it buys ideas. This is an investigation into what it buys instead — and why the answer is stranger, and more precisely wrong, than it first appears.
Humanity spends three trillion dollars a year trying to discover things. Almost none of it buys ideas. This is an investigation into what it buys instead — and why the answer is stranger, and more precisely wrong, than it first appears.
In 2024 the world spent an estimated $2.87 trillion on research and development, nearly three times what it spent, in real terms, twenty-five years earlier.1 Measured in current rather than constant dollars, the OECD’s accounting puts the figure closer to $3.8 trillion — not a competing estimate so much as the same mountain photographed in a different light.2 Either number exceeds the GDP of every country on Earth except the United States and China. Somewhere between a fortieth and a thirtieth of everything humanity produces in a year, it produces in order to find out something it does not yet know.
And no institution on the planet can tell you, in advance, what that money will buy. Boards approve R&D budgets the way governments approve defense budgets: as an article of faith, priced by benchmark rather than expected return. Nicholas Bloom and his co-authors spent a decade measuring what the spending produces, and their answer, published in the American Economic Review, was unwelcome. Research effort is rising sharply across nearly every domain they examined, while research productivity — ideas per researcher — is falling just as sharply. Sustaining the rate of improvement in chip density that Moore’s Law describes now takes more than eighteen times the researchers it took in the early 1970s. Across a panel of individual firms, research productivity was declining in more than 85 percent of the sample, at roughly 10 percent a year.3
This is the mystery worth sitting with, because the easy explanations do not survive contact with it. It is not that we have stopped trying; effort is up. It is not obviously that we have run out of money; global R&D intensity has climbed from under 1.5 percent of world output in 2000 to nearly 2 percent today.4 And it is not simply that the low-hanging fruit is gone, though that is part of it — “the ideas are harder to find” is a description, not an explanation. It tells you that productivity is falling, not why an extra dollar buys less discovery, or why a wartime aircraft division, a patent-poor telephone monopoly, and a fourteen-person skunkworks produced outsized shares of the twentieth century’s most consequential ideas on what were, by today’s standards, modest sums.
So: what is a trillion dollars actually for? If it is not reliably purchasing ideas, what is it purchasing?
I want to warn the reader early that the tidy answer this essay first arrived at turned out to be wrong — or rather, right in a way so unqualified that it was false. Getting to the version that survives means letting the tidy one break.
Follow the money
Start with where the money concentrates, because concentration is a clue. Five American technology companies — Amazon, Alphabet, Meta, Apple, and Microsoft — spent $229.1 billion on R&D in the twelve months to March 2024, more than the next five largest corporate spenders combined.5 That is now the smaller of the two numbers that matter. Layer on what the same companies are committing to capital expenditure — the data centers, chips, and power contracts to train and run frontier AI models — and the picture changes registers. The largest hyperscalers are expected to spend on the order of $725 billion on capex in 2026, up roughly 77 percent from an already record 2025, en route to a combined $5.3 trillion between 2025 and 2030 by one investment-bank estimate.6
Here the first tempting error appears, and it is worth naming because the earlier draft of this essay committed it. It is easy to call that capex the purchase of “falsification infrastructure” — the machinery for testing which ideas are any good. Mostly it is not. The large majority of that spend is inference-serving capacity: chips to deliver a shipped product to paying users, production capital no different in kind from a factory or a fleet of trucks. A data center is not a peer reviewer because a model happens to run on it.
But strip out the serving capacity and something real remains: the training runs, the evaluation harnesses, the red-team budgets, the synthetic-data pipelines. That residual — the part of the spend that tests rather than serves — is the interesting quantity, and it points at the same structural fact the corporate R&D numbers do. Across the OECD, the business sector now accounts for 73 percent of total R&D, up from 67 percent in 2010,7 and within corporate budgets the overwhelming majority goes not to basic or even applied research but to development: turning an already-known idea into a shippable, verified product.8 Discovery is the smallest line item in the discovery budget.
That reframes the mystery. If most of the three trillion is not chasing new ideas so much as refining, verifying, and shipping ideas that already exist in rough form, then the productivity puzzle is not “why are we bad at having ideas.” It is “why has it become so expensive to find out whether an idea is any good” — to develop it far enough, and test it hard enough, to know.
The wrong unit of analysis
Every organization in the standard innovation literature — Bell Labs, DARPA, Xerox PARC, Skunk Works, Google X, DeepMind, OpenAI — gets discussed as a case study in creativity: culture, autonomy, “20 percent time,” brilliant hires. That framing is not wrong, but it hides the more useful comparison. Put the labs side by side and the variable that best tracks their output is not funding, headcount, or freedom. It is how cheaply, and how quickly, each one could find out that an idea was wrong.
Bell Labs ran for decades as a regulated monopoly’s research arm, funded by a fixed slice of every American’s phone bill, insulated from quarterly earnings. It generated more than 28,000 patents and work recognized by ten Nobel Prizes: the transistor, the theoretical basis of the laser, information theory, the charge-coupled device, Unix.9 Its horizon was measured in decades, so it could afford slow testing of any single hypothesis — but once a hypothesis was ready, Bell Labs had internal peer challenge dense and literate enough that a wrong idea rarely survived contact with a colleague down the hall. Slow but cheap, because the reviewers were free and world-class.
DARPA inverts this. Program managers serve a few years, arrive with a mandate to fund specific bets, and hold explicit authority to kill programs — indeed, killing them fast is treated as a feature. DARPA does not need decades because it has built the habit of admitting error quickly and reallocating.
Xerox PARC is the case that should trouble the optimists most, and it is where the tidy version of this essay first cracked. PARC’s researchers built the graphical user interface, the mouse, Ethernet, and laser printing, years or decades ahead of the market, and Xerox commercialized almost none of it as well as the firms that copied it. The convenient reading is that PARC had a high market-test cost: it could not cheaply find out whether a customer would pay. But that reading quietly redefines “innovation” as “money Xerox captured.” By any honest accounting PARC did innovate, spectacularly; the world runs on what it made. What PARC lacked was not a falsification loop but appropriability — the ability, in David Teece’s sense, to capture the value it created, which flows to whoever controls the complementary assets of manufacturing, distribution, and brand.10 Hold that distinction; the whole argument will turn on it.
Skunk Works, Lockheed’s small-team aircraft division, ran at the opposite extreme: a handful of engineers, one demanding customer, brutal deadlines, authority to bypass the parent’s approval chain. The test — does the plane fly, does it evade radar, does it meet the date — was short, physical, and impossible to fudge. Expensive per attempt, but fast and unambiguous per finding.
Google X, DeepMind, and OpenAI are contemporary versions of the same trade-off, run at software’s far lower marginal cost of experiment. Google X talks openly about “rapid evaluation” — killing a moonshot within months if a small team can find the one reason it won’t work — precisely because cheap tests let it attempt hundreds of ideas rather than a handful. The labs that produced the most, across a century of cases, were not the ones with the most money or the most freedom. They were the ones that made being wrong cheap.
But notice what that sentence smuggles in. These are the labs we remember. For each, there were dozens with dense internal review and cheap tests that produced nothing anyone can name. Cheap falsification did not make them famous; it is common, and fame is rare. So the honest claim is weaker than “cheap testing causes great labs.” It is that cheap, legible testing is a necessary condition — a thing every one of the winners had, and a thing that, on its own, guarantees nothing. PARC proves the point: it had the cheap internal test and failed anyway, on a variable the test could not see.
Table 1 — Innovation labs compared
| Lab | Funding | Falsification loop | Speed | Cost per test | What it captured | What it created |
|---|---|---|---|---|---|---|
| Bell Labs | Monopoly rents, decades-long horizon | Internal peer review by elite staff | Slow | Low (staff time) | Licensing revenue, later dispersed at divestiture | Transistor, information theory, CCD, Unix11 |
| DARPA | Government appropriation, PM authority | External deliverables; hard kill/continue | Fast | Medium–high | Nothing directly; handed to industry | Seeded ARPANET, GPS, stealth |
| Xerox PARC | Corporate subsidy, no market discipline | Internal only; no route to a market test | Slow, then never | Low internally | Almost none — appropriability ≈ 0 | GUI, mouse, Ethernet, laser printing |
| Skunk Works | Single-customer contract, bypassed bureaucracy | Physical test flights vs. a hard spec | Very fast | High per attempt | Direct to one customer | U-2, SR-71, F-117 |
| Google X / DeepMind / OpenAI | Corporate balance sheet + capital markets | Rapid-evaluation kill gates; model self-critique | Fastest | Falling toward marginal compute | Folded into parent product lines | Waymo, AlphaFold, frontier models |
What the psychology actually says
The individual-level research gets summarized, in the popular literature, as a call to “think outside the box.” The findings are less inspirational and more mechanical. Functional fixedness — the documented tendency to see a tool only in terms of its familiar use — is not a character flaw; it is a byproduct of expertise, the same efficient pattern-matching that makes an expert fast at everything except the case that doesn’t fit. Divergent thinking, generating many candidate solutions, and convergent thinking, selecting correctly among them, are not two ends of one dial that creative people turn up. They are separate skills, and most institutional processes reward the second while claiming to want the first.
There is also a structural reason ideas are getting harder to find. The economist Ben Jones argues that as a field’s knowledge deepens, the training needed to reach its frontier lengthens, teams grow because no one can hold the whole field in their head, and the easy cross-disciplinary insight — the physicist who wanders into biology and sees what the biologists missed — gets rarer because the entry cost to either field keeps rising.12 The Renaissance man is not extinct because people got less curious. He is extinct because the distance between “informed layperson” and “frontier researcher” now takes a career to close.
This matters for the argument in a specific way: as the burden of knowledge rises, the cost of checking whether a new idea is even coherent rises with it, because fewer people are qualified to check and checking takes them longer. The bottleneck and the burden are the same phenomenon seen from two sides.
Table 2 — Psychological mechanisms and their organizational analogue
| Mechanism | Effect on an individual | Organizational parallel |
|---|---|---|
| Functional fixedness | Blinds expert to non-standard uses of a known tool | Sunk-cost attachment to existing product lines |
| Divergent thinking | Generates many candidate solutions | Hackathons, 20%-time, moonshot factories |
| Convergent thinking | Selects and refines among candidates | Stage-gates, red teams, peer review |
| Burden of knowledge | Lengthens time-to-frontier expertise | Larger teams, longer PhDs, narrower specialists |
| Recombination | Connects distant domains analogically | Cross-functional labs, “T-shaped” hiring |
Why the test is never clean
Before going further it is worth confronting the philosophical hole under the whole enterprise, because a reader who knows the terrain will have spotted it. This essay leans on falsification — the idea, associated with Popper, that you make progress by ruling wrong things out. But Popper’s clean picture was dismantled decades ago. The Duhem–Quine thesis holds that no hypothesis can be tested in isolation: every experiment rides on a bundle of auxiliary assumptions, so when a prediction fails, all you learn is that something in the bundle is wrong, not what.13 And Thomas Kuhn — awkwardly, since the innovation literature likes to cite him and Popper in the same breath — argued that theories rarely die by falsification at all. They persist until the people who hold them retire. Science advances, in the paraphrase of Planck, one funeral at a time.
If that is how discovery actually works, a model that treats falsification as a clean, purchasable step is describing a world that does not exist. So the model has to survive the objection or be abandoned. It survives — by absorbing it. Ask why falsification is expensive, and Duhem and Quine hand you the answer: because to isolate which assumption failed, you need many people, much apparatus, and repeated independent tests. The holism of testing is not an argument against the cost of testing; it is the source of that cost. And Kuhn supplies the sharpest case rather than the refutation. Sometimes the auxiliary assumption that has to be rejected is the one held by the person who owns the field — and then the cost of disproving it is measured in career lengths. The most expensive falsification in the history of science is a funeral. Understood properly, the philosophy does not weaken the claim that disconfirmation is costly. It explains why the cost is irreducible, and why it is social rather than technical.
Security as evidence, not metaphor
It would be easy to treat threat modeling, red teaming, zero trust, and chaos engineering as a colorful analogy for what labs do. They are not an analogy. They are the same activity, done explicitly under a different name: the deliberate manufacture of falsification.
A red team’s whole job is to find, cheaply and on your schedule, the ways your system is already wrong, before an adversary finds them expensively on theirs. Zero-trust architecture assumes every credential and request may be compromised and forces continuous re-verification — the same move as refusing to accept a hypothesis because it passed one test. Chaos engineering, Netflix’s discipline of deliberately breaking production to see what fails, is functional fixedness’s direct opponent: it forces engineers to meet the ways a system behaves outside its familiar range, instead of waiting for a customer to find out at 2 a.m.
What security engineering has that most corporate innovation processes lack is an honest price on being wrong. A vulnerability found by your own red team costs an afternoon; the same one found by an attacker costs a breach disclosure, a regulatory inquiry, and a share-price drop. Security teams have done the arithmetic most R&D organizations avoid: it is nearly always cheaper to disprove yourself than to be disproven by the world. Nearly — the qualifier will matter.
AI: buying disconfirmation directly
This is the lens through which the AI buildout stops looking merely enormous and starts looking legible. Frontier labs spend a conspicuous share of effort on activities that produce no new capability at all: red-teaming their own models, building evaluation suites, running reinforcement learning from human feedback, generating synthetic data to probe edge cases no real dataset covers. None of that is optional. It is the disconfirmation infrastructure a model needs before its outputs can be trusted, and its cost has grown comparable to the cost of the model.
Reasoning models make the point almost literal. A chain-of-thought or test-time-compute system generates a candidate answer, checks it against constraints, revises, and checks again — an internal falsification loop bought with compute at inference time. Frontier labs are, in effect, selling falsification cycles by the token. That is a genuinely new thing for an R&D budget to buy: not researchers, not equipment, but the disconfirmation step itself, unbundled and metered.
Which brings the essay to the question it least wants to get wrong: can a model innovate, or only recombine patterns already in its training data? The comfortable answer — machines do convergent search in a space humans define — was true until recently and is now too tidy. AlphaTensor discovered matrix-multiplication algorithms faster than the one Strassen found in 1969.14 AlphaDev found sorting routines now shipping inside the C++ standard library.15 FunSearch produced genuinely new mathematics on the cap-set problem, the first time a language model generated a result no human had.16 These are not retrievals. They are discoveries.
And yet the line did not vanish; it moved. In every one of those cases a human still specified the objective and the search space — minimize multiplications, maximize the set, define “stable material.” The machine then searched that space with a divergence no human could match. So the frontier between human and machine is migrating from “generate versus test” toward “choose the question versus search the answer.” AlphaFold did not decide that protein folding mattered; a generation of structural biologists did, and AlphaFold then searched the space they had defined, superbly. That may be a temporary limit or a durable division of labor. Either way it fits the pattern: what is scarce is not the search but the judgment about which space is worth searching — and the cost of the search itself is falling toward the price of electricity.
Table 3 — What AI labs are actually buying
| Spend category | What it produces | Falsification function |
|---|---|---|
| Pretraining compute | Base capability | Generates the candidate hypothesis space |
| RLHF / fine-tuning | Alignment with human preference | Convergent selection among outputs |
| Red-teaming | Adversarial failure cases | Deliberate, cheap self-disconfirmation |
| Evaluation suites | Benchmarked capability claims | Standardized, repeatable falsification |
| Synthetic data | Coverage of rare edge cases | Falsification of scenarios too rare to observe |
| Test-time / reasoning compute | Higher-quality single answers | Falsification purchased per query, not per model |
| Inference-serving capex | Delivery of a shipped product | None — this is production, not testing |
The last row is the honest one, and the one the tidy version of this essay left out.
Where the thesis breaks
If the argument were simply “cheaper falsification is better,” the historical record would embarrass it, and it is worth letting the record try.
Consider the highest-return biomedical innovation of the century. Katalin Karikó spent decades on messenger RNA in precisely the conditions this framework is supposed to disdain: chronically underfunded, unfashionable, slow, largely unsupervised basic research. The University of Pennsylvania demoted her in 1995 when the grants would not come; the 2005 paper that cracked the immune-response problem was rejected by Nature and by Science before it appeared in Immunity, and took over a decade to be believed.17 By any “cheap, fast falsification” scorecard this is a disaster — and it produced the mRNA vaccines. Or take the Manhattan Project and Apollo: enormous, bureaucratic, slow, expensive, rule-bound, roughly $30 billion and $260 billion in today’s money respectively,18 and among the most concentrated bursts of innovation ever recorded. If cheap disconfirmation were the whole story, none of these should exist.
They force two admissions. First, sometimes the binding constraint is not the price of the test but the patience to survive a long generative phase before any test is worth running — the very generation-side scarcity a testing-centric model tends to ignore. Second, and more damaging, the relationship between falsification cost and innovation is not monotone. The replication crisis is the proof. When the p < 0.05 significance test made “confirmation” cheap and mechanical, the result was not more knowledge but less: the Open Science Collaboration reproduced roughly a third of 100 psychology studies, and Amgen could reproduce six of 53 landmark cancer papers.19 Falsification made too cheap does not sharpen a field; it fills it with findings that were never really tested. Cheapness, past a point, is indistinguishable from not testing at all.
So the naive thesis is false, and we know it is false because reality ran the experiment. What replaces it is not a retreat but a correction: there is an optimum. Innovation is maximized not where the cost of being wrong is lowest but where it is correctly priced — high enough that a “confirmed” result means something, low enough that testing actually happens. Call the target Minimum Viable Doubt: the least expensive test that still produces a result you would stake a decision on. Below it lies the replication crisis, hype cycles, move-fast-and-break-things applied to things that should not be broken. Above it lies bureaucratic freeze. The pathology at both ends is the same — a falsification cost that has decoupled from the value of what is being decided.
This is also why the security engineer’s arithmetic — always disprove yourself first — has a limit. Pharmaceutical regulation sits exactly on the knife-edge. The randomized controlled trial is one of the most rigorous falsification machines ever built, and the same apparatus that makes a drug’s efficacy claim trustworthy is why a new drug now costs, on the most-cited industry estimate, around $2.6 billion and a decade — a figure whose critics rightly note that roughly half is capitalized opportunity cost rather than cash.20 The rule and the tax are the same rule. You cannot lower that CPF without lowering the trustworthiness that was the entire point. The goal was never cheap falsification. It was calibrated falsification.
Table 4 — When rules lower the cost of being wrong, and when they raise it
| Rules that fix the variables under test (lower CPF) | Rules that fix who may test, and how often (raise CPF) |
|---|---|
| Pre-flight checklist — frees attention for the one unsolved problem | Multi-layer approval committees — a project that can’t run can’t fail |
| Toyota standardized work — makes a defect attributable to one change21 | Patent thickets — a pure tax on testing whether an idea infringes |
| Chess rules, jazz chord changes — make a wrong move legible instantly | Program-manager churn — a promising bet dies when its champion leaves |
| Open-source CI and code review — thousands of free, parallel tests | Liability-driven compliance — optimized to avoid an inconvenient finding |
| The RCT protocol — makes an efficacy claim trustworthy | The RCT’s cost — makes that same claim take a decade and a billion dollars |
The RCT appears in both columns on purpose. It is the cleanest evidence that calibration, not cheapness, is the variable that matters.
The original contribution: the Calibrated Falsification Engine
Put all of this together and a single mental model falls out — the one this essay wants a reader to keep after every statistic has been forgotten. The first version of it was too simple, and its correction is the point.
An organization’s capacity to innovate is not a function of how many ideas it generates, nor of how cheaply it can disprove them. It is a function of how well the cost of being wrong is matched to the value of the decision at stake — and, separately, of whether it can capture the value of what survives the test.
Two quantities, not one. The first is Cost Per Falsification (CPF): the time and money to move a hypothesis from “untested” to “a result you would act on.” It is not “lower is better.” It has a target — Minimum Viable Doubt — and the failure modes are symmetric. And it is measurable, at least in proxy: cycle time from proposal to kill-or-scale, dollars per decisive experiment, and the false-discovery rate — how often a “confirmed” result later reverses — standing in for legibility. Any team can ask those three questions on a Monday.
The second quantity is appropriability: whether the organization controls the complementary assets to capture what it discovers. This is the axis PARC fell off. CPF governs discovery; appropriability governs capture; and the two are independent, which is why an organization can be brilliant at one and bankrupt at the other.
Plot CPF against appropriability and the famous cases sort themselves without special pleading:
Table 5 — Discovery vs. capture
| High appropriability (can capture value) | Low appropriability (cannot capture value) | |
|---|---|---|
| CPF near Minimum Viable Doubt | Disciplined Frontier — Toyota, Skunk Works, modern AI eval pipelines, Bell Labs at its peak. Tests are calibrated and the org keeps the winnings. | Gift to the Commons — Xerox PARC, much publicly funded science. Discovery is real and cheap to verify; the value flows to whoever owns distribution. |
| CPF far from Minimum Viable Doubt | Expensive Certainty — pharma, Apollo, Manhattan. Breakthroughs are real but ruinously priced; only an unlimited or mission budget sustains them. | Bureaucratic Freeze — patent thickets, compliance-first committees. Tests are costly and the value leaks. Worst quadrant on both axes. |
Notice what this fixes. PARC is no longer a “failure”; it is a Gift to the Commons — a triumph of discovery and a catastrophe of capture, which is exactly what the history shows. Apollo is no longer a “freeze”; it is Expensive Certainty, a real breakthrough bought at a price only a superpower could pay. And the quadrant no longer pretends to predict success. It predicts two different things — whether you can discover, and whether you can keep — and refuses to collapse them into one.
The AI labs are racing toward the top-left corner at a falsification speed no prior institution could attempt, which is the real argument for why so much of the buildout aims not at abstract intelligence but at shortening the loop between a model’s guess and a verified answer. But the same map warns them: a lab can drive CPF below Minimum Viable Doubt — ship benchmarks that certify nothing, mistake a passed eval for a tested claim — and land in the replication crisis’s quadrant without leaving the office.
A better question than the one this essay started with
The mystery this essay opened with — why can no one reliably manufacture innovation, despite spending more on it than ever — has a less romantic answer than “ideas are hard.” Ideas were rarely the scarce resource, at least for the well-resourced organizations that spend most of the money. What is scarce, and getting scarcer even as the dollars pour in, is correctly priced proof of which ideas are worth keeping — cheap enough to actually run, dear enough to actually mean something, and legible enough that the answer survives the person who found it. The trillions the world spends on R&D are mostly a bet on the infrastructure of finding out: laboratories, red teams, checklists, evaluation suites, chip fabs that let a model test ten thousand hypotheses before breakfast. The return on that bet was never about how many ideas the infrastructure generated. It was, quietly, the whole time, about how honestly it priced the cost of being wrong.
That reframing leaves a sharper question than the one on most boards’ slides. Not “how do we generate more breakthrough ideas” — that machinery is closer to abundant than scarce. And not even “how do we test them faster” — because faster, past a point, is how a field lies to itself. The harder question is the one the replication crisis and the billion-dollar drug pose from opposite directions at once: for a given decision, what is the right price of doubt — and who in the organization is allowed to be wrong about that price, cheaply enough to find out before the world charges full freight?
Notes
Word count: approx. 4,750
WIPO, “End of Year Edition — Despite the Odds, Global R&D Spending Grew Again in 2024,” Global Innovation Index blog, 2025. Figure in constant 2015 PPP terms; R&D intensity from the same source.↩︎
OECD, “OECD Overall R&D Growth Stable; Government R&D Budgets Decline and Reorient Towards Defence,” Statistical release, March 2026. The ~$3.8T figure is in current PPP dollars; the WIPO and OECD numbers differ by price base, not by disagreement.↩︎
Bloom, Jones, Van Reenen, and Webb, “Are Ideas Getting Harder to Find?” American Economic Review 110, no. 4 (2020): 1104–1144; NBER Working Paper No. 23782. A scholarly rebuttal argues part of the measured decline reflects dilution of the researcher pool by less-productive marginal entrants rather than a fundamental scarcity of ideas; the direction of the finding is not seriously disputed.↩︎
WIPO, “End of Year Edition — Despite the Odds, Global R&D Spending Grew Again in 2024,” Global Innovation Index blog, 2025. Figure in constant 2015 PPP terms; R&D intensity from the same source.↩︎
Trendline, “Big Tech’s Big R&D Bill,” 2024, compiling trailing-twelve-month 10-Q filings to March 2024. Note that Amazon reports R&D as “technology and content,” a broader line than a strict R&D definition.↩︎
Goldman Sachs Research, “Why AI Companies May Invest More Than $500 Billion in 2026” and “Tracking Trillions,” 2026. Figures are investment-bank estimates as of early 2026 and are volatile.↩︎
OECD, March 2026 statistical release.↩︎
NSF/NCSES Business Enterprise R&D (BERD) tables; development is consistently the largest component of U.S. business R&D, with basic and applied research a minority. Summarized in UIDP, “R&D in Review: Global Investment Trends.”↩︎
Nokia Bell Labs, “Awards” (official count: 10 Nobel Prizes, 5 Turing Awards) and “History”; patent count “more than 28,000 since 1925.” Some sources count 9–11 Nobel-recognized researchers depending on method.↩︎
David J. Teece, “Profiting from Technological Innovation,” Research Policy 15, no. 6 (1986): 285–305.↩︎
Nokia Bell Labs, “Awards” (official count: 10 Nobel Prizes, 5 Turing Awards) and “History”; patent count “more than 28,000 since 1925.” Some sources count 9–11 Nobel-recognized researchers depending on method.↩︎
Benjamin F. Jones, “The Burden of Knowledge and the ‘Death of the Renaissance Man,’” summarized in Nicholas Decker, “Are Ideas Getting Harder to Find?” 2025.↩︎
“Underdetermination of Scientific Theory,” Stanford Encyclopedia of Philosophy (holist / Duhem–Quine underdetermination). Kuhn: The Structure of Scientific Revolutions (1962).↩︎
Fawzi et al., “Discovering Faster Matrix Multiplication Algorithms with Reinforcement Learning,” Nature 610 (5 Oct 2022): 47–53.↩︎
Mankowitz et al., “Faster Sorting Algorithms Discovered Using Deep Reinforcement Learning,” Nature 618 (7 Jun 2023): 257–263.↩︎
Romera-Paredes et al., “Mathematical Discoveries from Program Search with Large Language Models,” Nature 625 (2023/2024): 468–475.↩︎
Penn Today and Forbes coverage of the 2023 Nobel Prize in Physiology or Medicine; Karikó and Weissman, “Suppression of RNA Recognition by Toll-like Receptors,” Immunity (2005).↩︎
Manhattan Project ~$1.9–2.0B (1940s) ≈ ~$30B today (Brookings, NPS). Apollo $25.8B nominal (1960–73) ≈ ~$260–310B today (Dreier, Space Policy, 2022; The Planetary Society).↩︎
Open Science Collaboration, “Estimating the Reproducibility of Psychological Science,” Science 349 (2015); Begley and Ellis, “Raise Standards for Preclinical Cancer Research,” Nature 483 (2012): 531–533. Begley–Ellis has been critiqued for a small non-random sample; the later Reproducibility Project: Cancer Biology (eLife, 2021) corroborates the direction with substantially reduced effect sizes.↩︎
DiMasi, Grabowski, and Hansen, “Innovation in the Pharmaceutical Industry: New Estimates of R&D Costs,” Journal of Health Economics 47 (2016): 20–33. Critics (Public Citizen, KEI) note ~half the $2.6B is capitalized opportunity cost and that the underlying data are confidential.↩︎
Toyota Motor Corporation, “Toyota Production System: Vision & Philosophy”; The Systems Thinker, “Pulling the Andon Cord.”↩︎