Shwetha Devanga  ·  ← The Price of Being Wrong

Sources & reading list · every number, mapped to its origin

Where it all came from

The research behind "The Price of Being Wrong," the Minimum Viable Doubt framework, and the Doubt product concept.

Each entry gives the citation, the specific number or finding it supports (in rust), and links — an open PDF or author copy where one exists, otherwise the journal/DOI page. A few classic books are cited without a link. Most links point to the paper's canonical page or an open PDF; publisher pages may be paywalled and a few institutional/landing links may have moved — but the full citation always lets you find the source (Google Scholar, DOI). The core empirical papers (bold numbers) were surfaced directly during the research.
A · The core puzzle: spend up, discovery down B · Innovation labs & value capture C · Philosophy of science: falsification → severity D · The formal core: making doubt measurable E · Metascience: the replication evidence F · Failure-tolerance & experimentation economics G · Where the thesis breaks H · AI discovery I · Case histories J · Decision science (the product layer) K · Market & competitive sources

A The core puzzle: spend up, discovery down

The founding tension — record R&D spend, falling productivity per researcher.

Bloom, Jones, Van Reenen & Webb, "Are Ideas Getting Harder to Find?" American Economic Review 110(4), 2020 (NBER WP 23782). 18× more researchers to sustain Moore's Law vs. the early 1970s; research productivity falling ~10%/yr, declining in >85% of firms. The spine of the whole essay. Open PDF (Stanford)AERNBER
Jones, "The Burden of Knowledge and the 'Death of the Renaissance Man,'" Review of Economic Studies 76(1), 2009. → Why ideas get harder: time-to-frontier lengthens, teams grow, cross-disciplinary insight rarer. The structural companion to Bloom et al. REStud
WIPO, Global Innovation Index — Global R&D spending (2024/2025 editions). → Global R&D $2.87 trillion (constant 2015 PPP); R&D intensity climbing toward ~2% of world output. Global Innovation Index
OECD, Main Science & Technology Indicators / R&D statistical releases (2024–2026). ~$3.8T global R&D in current PPP; business sector now 73% of total R&D (up from 67% in 2010). OECD R&D data
Big-tech R&D (10-Q compilation, TTM to March 2024) & Goldman Sachs Research, "Why AI Companies May Invest More Than $500 Billion" / "Tracking Trillions" (2026). → Five US tech firms spent $229.1B on R&D (TTM Mar 2024); hyperscaler capex heading to ~$725B in 2026 (+77%), ~$5.3T across 2025–30 (bank estimates, volatile). Goldman Sachs Insights
NSF / NCSES, Business Enterprise R&D (BERD) survey tables. Development, not research, is the largest component of US business R&D — "discovery is the smallest line item in the discovery budget." NCSES BERD

B Innovation labs & value capture

Teece, "Profiting from Technological Innovation," Research Policy 15(6), 1986. Appropriability & complementary assets — the frame that reclassifies Xerox PARC as a capture failure, not an innovation failure. The essay's second axis. ScienceDirect
Nokia Bell Labs — Awards & History. 10 Nobel Prizes, 5 Turing Awards, "more than 28,000 patents since 1925": transistor, information theory, CCD, Unix. Bell Labs awards

C Philosophy of science: falsification → severity

From Popper's clean picture to the rigorous modern version of "calibration."

Deborah Mayo & Aris Spanos, "Severe Testing as a Basic Concept in a Neyman–Pearson Philosophy of Induction," BJPS 57(2), 2006; Mayo, Statistical Inference as Severe Testing (Cambridge, 2018). Severity: a result is evidence only if the test would probably have failed had the claim been wrong. The rigorous form of "correctly priced doubt," and the two symmetric fallacies = the two ends of the curve. Severity notes (PDF)errorstatistics.com
Peirce, "Note on the Theory of the Economy of Research" (1879); Rescher, "Peirce and the Economy of Research," Philosophy of Science 43(1), 1976; Wible, Journal of Economic Methodology 1(1), 1994. → The 145-year-old ancestor of Cost Per Falsification: spend inquiry where knowledge-per-dollar is highest. Cost-of-testing as part of the logic of science. Wible (JEM)Overview
Lakatos, The Methodology of Scientific Research Programmes (1978); Popper, The Logic of Scientific Discovery (1959); Kuhn, The Structure of Scientific Revolutions (1962); Duhem–Quine underdetermination. → Progressive vs. degenerating programmes (the "is calibration just hedging?" test); paradigm persistence ("one funeral at a time"); no hypothesis tests in isolation — why the cost of falsification is irreducible and social. SEP: LakatosSEP: UnderdeterminationSEP: Kuhn

D The formal core: making doubt measurable

The five literatures that turn "Minimum Viable Doubt" into a real quantity with an interior optimum.

Howard, "Information Value Theory," IEEE Trans. SSC 2(1), 1966; Raiffa & Schlaifer, Applied Statistical Decision Theory (1961). Value of information: run a test iff expected value of what you learn > its cost; never pay more than perfect information. The ceiling on any test. IEEE
Wald, Sequential Analysis (Wiley, 1947). → The SPRT: optimal stopping — minimum expected observations for a set error rate. "Cycle time to a decision-grade result," solved. Book (1947)
Weitzman, "Optimal Search for the Best Alternative," Econometrica 47(3), 1979. Pandora's Box / reservation-price rule: which hypotheses to test, in what order, when to stop. The portfolio ("what to test next") logic. Open PDF (Harvard)
Gittins, "Bandit Processes and Dynamic Allocation Indices," JRSS-B 41(2), 1979. → The multi-armed bandit / explore–exploit optimum. Minimum Viable Doubt = the optimal switch from testing to committing. JRSS-B (1979)
Ioannidis, "Why Most Published Research Findings Are False," PLoS Medicine 2(8), 2005. → The positive predictive value identity — calibration as a computable number; six corollaries = the anatomy of a too-cheap test. PLoS Medicine
Benjamini & Hochberg, "Controlling the False Discovery Rate," JRSS-B 57(1), 1995. FDR control — the operational tool for keeping "legibility" honest across many parallel tests. Wiley (JRSS-B)
Manso, "Motivating Innovation," Journal of Finance 66(5), 2011. → The optimal contract for innovation "exhibits substantial tolerance — or even reward — for early failure." Failure-tolerance as a theorem. Open PDF (MIT)
Brier (1950), "Verification of Forecasts…," Monthly Weather Review; Murphy (1973), "A New Vector Partition of the Probability Score," J. Applied Meteorology. → The Brier score and its decomposition (reliability − resolution + uncertainty) — the math behind the deep MVP's "calibration vs. discrimination" engine. Brier 1950

E Metascience: the replication evidence

The natural experiments proving cheap confirmation ≠ knowledge.

Open Science Collaboration, "Estimating the Reproducibility of Psychological Science," Science 349(6251), 2015. ~36% of 100 psychology studies replicated. ScienceOpen PDF
Begley & Ellis, "Raise Standards for Preclinical Cancer Research," Nature 483, 2012. → Amgen reproduced 6 of 53 landmark cancer papers. (Prinz/Bayer 2011: ~25% reproduced.) NaturePubMed
Scheel, Schijen & Lakens, "An Excess of Positive Results: Comparing the Standard Psychology Literature with Registered Reports," AMPPS 4(2), 2021. → Positive results: 96% in standard papers vs. 44% in Registered Reports — the calibration natural experiment. AMPPS
Franco, Malhotra & Simonovits, "Publication Bias in the Social Sciences: Unlocking the File Drawer," Science 345(6203), 2014. → 221 TESS studies: only ~20% of nulls published (vs ~60% of strong results); ~65% of nulls never even written up. Science

F Failure-tolerance & experimentation economics

Causal evidence that pricing "being wrong" correctly changes what gets discovered.

Azoulay, Graff Zivin & Manso, "Incentives and Creativity: Evidence from the Academic Life Sciences," RAND J. Economics 42(3), 2011 (HHMI vs. NIH). → Failure-tolerant funding: +39% papers, +55% top-5%, +97% top-1% — and +35% bottom-quartile flops. The essay's causal spine. Open PDF (Berkeley)NBER
Azoulay, Fuchs, Goldstein & Kearney, "Funding Breakthrough Research: Promises & Challenges of the 'ARPA Model,'" Innovation Policy and the Economy 19, 2019. ~half of ARPA-E projects funded below the peer-review cutoff; DARPA runs ~$3B with ~100 program managers; "ARPA-able" scope conditions. Working paper (PDF)
Kohavi et al., "Online Controlled Experiments and A/B Testing" (2017); Kohavi, Tang & Xu, Trustworthy Online Controlled Experiments (Cambridge, 2020). → Only ~⅓ of ideas succeed at Microsoft, ~10% at Google; Netflix considers ~90% wrong. "We are poor at assessing the value of ideas." exp-platform.com
Ewens, Nanda & Rhodes-Kropf, "Cost of Experimentation and the Evolution of Venture Capital," J. Financial Economics 128(3), 2018. → After AWS (2006): initial rounds −20%, "spray & pray," and the streetlight effect — a shift toward innovations "revealed quickly and cheaply," away from complex tech. The deepest counter-attack. Summary (Harvard)
Kremer, Levin & Snyder, "Designing Advance Market Commitments for New Vaccines" (NBER WP 28168); GAVI Pneumococcal AMC. → $1.5B AMC → 150M+ doses/yr, ~700k lives saved, PCV ~5 yrs faster to low-income countries. Calibration applied to the capture axis. NBERGAVI

G Where the thesis breaks

March, "Exploration and Exploitation in Organizational Learning," Organization Science 2(1), 1991. Faster convergence on the shared view lowers long-run knowledge — the robust version of "too much legibility hurts." Attack on the third leg. Organization ScienceOpen PDF
Zollman, "The Communication Structure of Epistemic Communities," Philosophy of Science 74(5), 2007; Rosenstock, Bruner & O'Connor, "In Epistemic Networks, Is Less Really More?" Philosophy of Science 84(2), 2017. → Sparse networks can converge on truth more reliably (Zollman) — but the effect is fragile (vanishes at pB≥0.525; 100-agent nets: 99.12% in 20 rounds vs 100% in 1,977). The honest, bounded caveat. Zollman (PDF)Rosenstock (PDF)
Strathern (1997) on Goodhart's Law; Campbell, "Assessing the Impact of Planned Social Change" (1979). → "When a measure becomes a target, it ceases to be a good measure" — why any CPF metric is gameable (benchmark overfitting). The reflexivity attack. "When a Measure Becomes a Target" (PMC)

H AI discovery

Fawzi et al., "Discovering Faster Matrix Multiplication Algorithms with Reinforcement Learning," Nature 610, 2022 (AlphaTensor). Nature
Mankowitz et al., "Faster Sorting Algorithms Discovered Using Deep Reinforcement Learning," Nature 618, 2023 (AlphaDev). Nature
Romera-Paredes et al., "Mathematical Discoveries from Program Search with Large Language Models," Nature 625, 2023/24 (FunSearch). → AI doing genuine, superhuman search inside a human-specified space — the migration of the human/machine line from "generate vs. test" to "choose the question vs. search the answer." Nature

I Case histories

Karikó & Weissman, Immunity (2005); 2023 Nobel Prize in Physiology or Medicine. → Demoted in 1995, paper rejected by Nature and Science — decades of unfashionable work → the mRNA vaccines. The "patience before any test" case. Nobel 2023
The Planetary Society (Dreier), Apollo cost analysis; Brookings, "The Costs of the Manhattan Project." → Apollo ~$260B, Manhattan ~$30B in today's money — "Expensive Certainty," breakthroughs at ruinous cost. Apollo costManhattan cost
DiMasi, Grabowski & Hansen, "Innovation in the Pharmaceutical Industry," J. Health Economics 47, 2016 (with Public Citizen / KEI critiques). → The ~$2.6B-and-a-decade drug estimate — the RCT as calibration you can't cheapen without losing the trust that was the point. J. Health Economics

J Decision science — the product layer

The forecasting/decision literature behind the Doubt MVP (calibration, resulting, base rates).

Tetlock & Gardner, Superforecasting (2015); Mellers, Tetlock et al., "Psychological Strategies for Winning a Geopolitical Forecasting Tournament," Psychological Science, 2014; "Forecasting Tournaments," Current Directions in Psych. Science, 2014. → Calibration is a trainable skill; base-rate anchoring and track-record scoring are the biggest accuracy levers. The basis for the calibration engine. Mellers et al. (PDF)Forecasting Tournaments
Duke, Thinking in Bets (2018). "Resulting" — judging a decision by its outcome rather than its process. The basis for the MVP's decision-quality-vs-outcome separation. Book (2018)
Arrow, "Economic Welfare and the Allocation of Resources for Invention" (1962); Nelson, parallel-path R&D. → The information paradox & parallel-path search — background for the portfolio/appropriability logic. Arrow 1962 (NBER)

K Market & competitive sources (Doubt)

Grand View Research, Decision Intelligence Market (2025). ~$17.8B (2025) → ~$53.2B by 2033, ~14.4% CAGR (adjacent-market context). Grand View
Research and Markets, A/B Testing Software Market (2026). ~$1.3B (2025) → ~$2.7B by 2032, ~11% CAGR. Research and Markets
Datadog acquires Eppo (May 2025). $220M — experimentation consolidating into the analytics/observability giants. TechCrunchDatadog
Competitors profiled in the competitive analysis. → Direct/adjacent products: forecasting trackers, experimentation platforms, and enterprise decision tools. FatebookCloverpopMetaculusManifoldGood JudgmentStatsigGrowthBook

How this maps to the deliverables. Sections A–I underpin the essay "The Price of Being Wrong" and the research dossier; C–D and J underpin the Minimum Viable Doubt framework and the Doubt MVP; J–K underpin the product brief, strategy brief, and competitive analysis. Books (Wald, Gittins-era texts, Superforecasting, Thinking in Bets) are cited without links. Every figure quoted in the deliverables traces to an entry above; the one claim flagged as unverified throughout — that Mayo's severity function misbehaves under sequential stopping (via Yanofsky, encountered in C. Shalizi's review of Mayo) — is not in this list because it was never confirmed to a primary source, and should be checked before use.