Part IWhat a Framework Actually Is
At bottom, a framework is a compression algorithm for a domain. Reality presents too many variables to hold at once; a framework is a claim that a small subset of those variables carries most of the explanatory or decision-relevant weight, together with a structure for how they relate. That is the whole thing. Porter's Five Forces is a claim that industry profitability is mostly explained by five pressures. The scientific method is a claim that reliable knowledge mostly emerges from a particular loop of conjecture and test. Maslow's hierarchy is a claim that human motivation is mostly ordered by a stack of needs.
This definition immediately separates a framework from its neighbors — terms people constantly conflate. A theory makes causal, falsifiable claims about how the world works. A model is a simplified, often formal representation you can run. A methodology sequences actions; a checklist enumerates without structure; a taxonomy sorts without implying relationships between the categories. The table below is the fastest way to see where the boundaries fall — read the "test of quality" column especially, because that is where the confusion does real damage: people routinely judge frameworks by a theory's standard (is it true?) when the right question is a tool's standard (does it change decisions for the better?).
| Artifact | Core claim | Output when applied | Test of quality | Example |
|---|---|---|---|---|
| Framework | "These few variables, related this way, carry most of the weight." | Structured attention; a map of where to look | Does it change decisions, and can it say no? | Five Forces, Cynefin, SWOT |
| Theory | "X causes Y, by this mechanism." | Predictions and explanations | Falsifiability; survives attempts at refutation | Natural selection, disruption theory |
| Model | "This simplified representation behaves like the real thing." | Computed outcomes ("if X rises 10%, Y falls 3%") | Predictive accuracy within stated bounds | DCF valuation, epidemiological SIR |
| Methodology | "Do these steps in this order." | A sequenced course of action | Reliability of outcome when followed | Scrum, double-diamond design |
| Checklist | "Don't forget any of these." | Verified completeness | Error rate of users vs. non-users | Surgical safety checklist |
| Taxonomy | "Everything here belongs to exactly one of these kinds." | Sorted categories | Exhaustiveness; low ambiguity at the edges | Species classification, bug severity levels |
A framework sits between all of these. It tells you where to look and how the places you look relate to each other — but not what you will find when you get there. It is a structure for attention. That is both its power and its danger. The power is transferability: because a framework carries no case-specific conclusions, it travels across situations. The danger is that using one feels like understanding while quietly smuggling in unexamined assumptions about what matters and, more importantly, what doesn't.
A framework's omissions are its real content. Every framework asserts two things at once: these variables matter — and everything left out matters less.
When a framework is the wrong tool entirely
An earlier edition of this essay skipped a question its own method demands: it told you how to build frameworks without ever saying when not to. That is a says-no failure at the top level — the same flaw it accuses bad frameworks of. So, before the anatomy: a framework is one tool among five, and it wins only under specific conditions. If the decision is recurring and the failure modes are known, a checklist beats a framework — aviation and surgery figured this out decades ago, and "richer thinking" is exactly what you don't want at 3 a.m. in a cockpit. If the key variables are quantifiable and their relationships are stable, build a model and compute the answer instead of framing it. If what you need is prediction or explanation, borrow a theory. And if the domain gives fast, unambiguous feedback to a practiced individual — firefighting, chess, trauma triage — trained intuition outperforms explicit structure, and imposing a framework slows the expert down. A framework earns its place in the remaining territory: many variables, mostly qualitative, judgment that must transfer across people and cases.
Part IIThe Six Elements Every Framework Contains
Dissect any framework that has survived contact with reality and you find the same six elements, whether or not its author stated them. The diagram below shows how they fit together; the unstated ones — usually the boundary conditions and the mechanism — are where frameworks quietly fail.
1 · A scoped question and unit of analysis. What decision or understanding is this for, and about what kind of object — a firm, a product, a team, a person's week? Frameworks fail most often right here, applied with confidence to units they were never built for. Five Forces analyzes an industry; pointed at a single product line, it produces fluent nonsense.
2 · The dimensions. The handful of variables the framework claims matter. The implicit assertion is always double-sided: these matter, and everything omitted matters less. A framework that names customer, competitor, and capability is also asserting that regulation, timing, and luck are second-order. Whether that's true is an empirical question most framework users never ask.
3 · The structure. How the dimensions relate — and this should mirror the domain's actual topology, the subject of Part III. A large share of bad framework design is topology mismatch: forcing a feedback-loop reality into a linear pipeline, or carving a smooth continuum into four tidy quadrants because quadrants present well.
4 · Boundary conditions. Where it applies and where it breaks. These are almost always left implicit, which is precisely how frameworks get misused. Lean-startup logic breaks where experiments are expensive and irreversible; agile breaks where requirements genuinely are fixed. The framework isn't wrong — it's out of bounds.
5 · Decision logic. What you do differently after using it. If running a case through the framework changes no downstream action and no belief, the framework is decoration — an elaborate way of restating what you already thought.
6 · A mechanism. The deepest element: why these dimensions and not others. Five Forces works because it is microeconomics — bargaining power, entry barriers, substitution — wearing a business suit. A consultant's ad-hoc two-by-two usually has no mechanism underneath, which is why it cannot tell you when it is wrong. A framework without a mechanism can only ever be pattern-matched, never reasoned with.
The dissection habit is worth practicing on frameworks you already use. Here are three familiar ones taken apart along the six elements — notice how much of each framework's real content sits in the two rows their popular versions never mention:
| Element | Porter's Five Forces | Eisenhower Matrix | Cynefin |
|---|---|---|---|
| Question / unit | Is this industry structurally profitable, and where is the pressure coming from? | What should I do with this task right now? | What kind of problem situation am I in, so which response strategy fits? |
| Dimensions | Rivalry, new entrants, substitutes, buyer power, supplier power | Urgency; importance | Knowability of cause and effect: clear, complicated, complex, chaotic (+ disorder) |
| Structure | Hub — five pressures converging on one center | 2×2 matrix — two independent axes | Taxonomy with movement — situations migrate between domains |
| Boundary conditions | Stable industries with definable borders; breaks in fast-moving platform and ecosystem markets | Separable, individually-ownable tasks; breaks when everything is urgent-and-important or tasks interdepend | Diagnosis of response strategy only; breaks when treated as a source of answers rather than of approach |
| Decision logic | Enter, avoid, or reposition; where to build bargaining power | Do now / schedule / delegate / delete | Clear → best practice; complicated → analyze; complex → probe and sense; chaotic → act first |
| Mechanism | Microeconomics: bargaining power and entry barriers determine who captures value | Urgency bias: salience crowds out value unless separated explicitly | Systems theory: how knowable causality is determines which strategy can work at all |
Part IIIA Field Catalog of Shapes
The structure element deserves its own section, because structure is where a framework makes its strongest hidden claim: that the domain has a shape, and that this is it. A previous edition of this essay claimed there were exactly five shapes — and then dissected Cynefin, whose structure ("a taxonomy with movement between domains") fit none of them. That is an out-of-sample failure inside the essay's own pages, and it earns the honest correction: this is a field catalog, not a closed taxonomy. Seven shapes cover most frameworks in circulation; hybrids are common and legitimate; and a shape you haven't seen yet is always possible. Each entry asserts something different about how the world works, and each lies in a characteristic way when misapplied.
Sequence
Fits — true processes with ordered, gated stages: hiring funnels, sales pipelines, double-diamond design.
Breaks — when reality iterates. A "pipeline" with constant backflow is a loop wearing a pipeline costume.
Matrix (2×2)
Fits — genuine tradeoffs between two independent axes: urgency × importance, growth × share.
Breaks — when the axes correlate, or when a smooth continuum is carved into quadrants because quadrants present well.
Layers
Fits — things that genuinely build on each other, where lower levels enable upper ones: tech stacks, capability maturity.
Breaks — when the "prerequisite" claim is false. Maslow's stack is famous partly because people routinely violate its ordering.
Loop
Fits — systems with feedback, where output re-enters as input: build-measure-learn, flywheels, OODA.
Breaks — when the loop never actually closes in practice, or when "everything affects everything" becomes an excuse not to specify which link is the constraint.
Hub
Fits — many independent pressures converging on one object of interest: Five Forces on an industry, stakeholder maps.
Breaks — when the spokes interact with each other. If suppliers and substitutes are colluding, the hub picture hides the alliance.
Tree
Fits — decomposition and branching choices: issue trees, MECE breakdowns, decision trees, org charts, taxonomies like Cynefin's domains.
Breaks — when branches share causes. A cost tree that splits "people" from "process" hides that one drives the other; trees assert independence between siblings.
Network
Fits — domains where the relationships are the content: stakeholder influence maps, systems maps, dependency graphs.
Breaks — as a working framework it often carries too little compression: "everything connects to everything" is a map, not a tool, until you name which links dominate.
Hybrids are not a special case; they are the norm among frameworks that survive. Cynefin is a tree (a taxonomy of domains) crossed with a loop (situations migrate between them, and the response strategy feeds back into which domain you're in). The Business Model Canvas is a hub with an implied flow through it. When you meet a hybrid, dissect each shape's claim separately — the tree part of Cynefin asserts the domains are distinct; the loop part asserts they're transient. Both claims can be checked, and they fail independently.
The practical use of this catalog is diagnostic, in both directions. When designing, ask which shape the domain really has before drawing anything — the honest answer is often "a loop," and loops are the shape people most resist because they present badly on slides. When consuming someone else's framework, ask what the chosen shape is asserting and whether the domain agrees. The growth world's shift from the AARRR "funnel" to flywheel models was exactly this correction: retention feeds acquisition through referral, so the domain was a loop all along, and the funnel shape was actively hiding the most important link.
Every shape is a claim. A pipeline claims stages don't feed back. A 2×2 claims the axes are independent. A stack claims lower levels are prerequisites. Check the claim, not the diagram.
Part IVTwo Routes to Designing One
An earlier edition of this essay taught exactly one design method — induce dimensions from cases — while holding up Five Forces as the exemplar of good design. But Five Forces was not induced from cases at all: Porter deduced it from industrial-organization economics, a body of theory about bargaining power and entry barriers, and then checked it against industries. The essay's procedure contradicted its own favorite example. The correction matters, because there are two legitimate routes into a framework, they fail differently, and the strongest frameworks run both.
The inductive route starts from cases: gather instances including failures, and extract dimensions from the variance between good and bad outcomes. Its strength is contact with reality; its characteristic failure is overfitting — small samples, hindsight-labeled outcomes, and dimensions that describe your last twenty cases rather than the domain. The deductive route starts from a theory you already trust: derive what should matter from the mechanism, then check that the derived dimensions actually discriminate among real cases. Its strength is a built-in mechanism (element six comes free); its characteristic failure is elegance without traction — dimensions that follow beautifully from the theory and don't move any real decision. Triangulation is the discipline: dimensions that both routes produce independently are the ones to trust.
Whichever route supplies the dimensions, the shared pipeline runs in six steps, and the order matters: each step constrains the next, and each has a characteristic failure when skipped.
Anchor on the decision not the domain
Name the choice someone will make differently with this in hand. "Understand innovation" sprawls; "which of our projects should we kill this quarter" cuts. Skipped: you get a mood board, not a tool.
Gather cases including the failures
Real instances of the phenomenon — and crucially the ones that went badly. A framework induced only from successes inherits survivorship bias at birth. Skipped: you describe what winners share without noticing losers shared it too.
Induce dimensions from variance
What actually differed between good and bad outcomes? This is informal factor analysis, and it demands a researcher's discipline: resist dimensions that are fashionable, symmetrical, or alliterative. If your three factors all start with the same letter, be suspicious of how they got there. Skipped: imported dimensions, borrowed authority.
Match structure to topology
Sequence, matrix, layers, loop, or hub — whichever shape the domain really has (Part III). Skipped: topology mismatch, the most common design flaw in circulation.
Operationalize every construct
Define each dimension precisely enough that two independent users put the same case in the same box. If your dimensions are vibes — "high vs. low innovation culture" — the categorization step is where users' priors sneak in, and the framework will confirm whatever they already believed. Skipped: the framework outputs its user.
Compress, then bound
Remove each element; if the output never changes, it was dead weight — and dead weight in a thinking tool dilutes attention. Three dimensions that discriminate beat seven that overlap. Treat MECE as a smell test, not a fetish. Then write the boundary conditions down yourself, before someone else discovers them for you in production. Skipped: bloat now, misuse later.
1 · Decision. A product lead must decide, each quarter, which of ~15 running projects to stop. Not "assess the portfolio" — choose what dies.
2 · Cases. She lists the last two years of projects: eight shipped and mattered, five shipped and didn't matter, six were killed — three of those too late.
3 · Variance. What separated "shipped and mattered" from the rest? Not team quality or effort — those were uniform. The variance sat in two places: whether a named customer was waiting for the output, and whether the project's core assumption had been tested in the last 90 days. Two dimensions, both earned from the cases, neither fashionable.
4 · Topology. The two dimensions are independent (a project can have an eager customer and an untested assumption), so a 2×2 is honest here: demand evidence × assumption freshness.
5 · Operationalize. "Demand evidence" = a named external user who asked for it unprompted, in writing, within the quarter — not "the roadmap says." "Fresh" = the riskiest assumption was tested against reality in the last 90 days. Two reviewers, same project, same box.
6 · Compress and bound. A third candidate dimension — strategic alignment — was cut by the collapse test: it never changed a verdict, because misaligned projects always failed one of the other two gates first. Boundary conditions, written down: applies to funded product bets; does not apply to infrastructure, compliance, or research, where demand evidence is structurally absent. Decision logic: weak on both axes → kill this quarter; weak on one → 90-day test with a named owner; strong on both → fund fully.
Total apparatus: two operational questions and a rule. That is what a framework that has been compressed properly looks like — small enough to remember, sharp enough to say no.
Now the honest part — what this tidy story hides. Nineteen cases is a small sample; with two candidate dimensions surviving from a handful tried, there is real risk the "variance" is noise. The outcome labels ("shipped and mattered") were assigned in hindsight by the same person building the framework — hindsight bias with a pen. And the claim that the two dimensions are independent was assumed, not checked: eager customers may be why assumptions get tested. A designer taking the inductive route owes three mitigations: label outcomes before analyzing drivers (or have someone else label them blind); check the candidate dimensions for correlation across the cases before declaring a 2×2; and treat the first year of use as the real test — the framework is a hypothesis until its quarterly verdicts have been scored against what actually happened. Real induction is this messy. A worked example that doesn't say so is modeling exactly the success theater this essay warns against.
Part VHow to Pressure Test One
The master question is falsifiability-in-use: does the framework ever say no? Run live cases through it and check whether it can produce the verdicts "don't do this" or "this case doesn't fit here." A framework that blesses every option it is shown is not an analytical tool; it is a rationalization engine with a diagram. Beyond the master question, seven tests do most of the work — each aimed at a specific way frameworks fake competence.
| Test | The question you ask | Failure mode it catches | Pass looks like |
|---|---|---|---|
| Says-no | Can it output "don't" or "doesn't fit"? | Rationalization engine — blesses everything it is shown | Documented cases where the framework vetoed the popular option |
| Out-of-sample | Does it discriminate on cases it wasn't built from — especially adversarial ones chosen to break it? | Overfitting to its own founding anecdotes | New cases sort cleanly; the framework generates surprise, not just confirmation |
| Retrodiction | Does it explain known failures better than base rates would? | Survivorship laundering — "explaining" only successes | Past failures land in the boxes the framework says they should |
| Inter-rater | Do two independent users put the same case in the same box? | Vibes constructs — the framework outputs its user's priors | High agreement without discussion; disagreements traceable to the case, not the constructs |
| Collapse | Delete each element and rerun old cases — does anything change? | Ornamental dimensions kept for symmetry or politics | Every surviving element has changed at least one real verdict |
| Inversion | Could a reasonable person justify the opposite conclusion from the same facts and framework? | The conclusion was decided before the framework was opened | The framework's verdict is stable across users with opposite priors |
| Goodhart | If people are measured against it, what gets gamed and what gets neglected? | Metric capture — the framework warps the behavior it was built to describe | Known gaming vectors listed; off-framework blind spots monitored separately |
| Boundary probe | Where, specifically, does it break? | Unexamined universality — "applies everywhere" means "tested nowhere" | A written failure map; the designer can name the breaking point faster than a critic can |
Run the battery in roughly that order — it is sequenced from cheapest to most expensive. The says-no test costs an afternoon with old cases. The inter-rater test costs two colleagues and an hour. The Goodhart test only pays off once the framework has users who didn't build it, which is also exactly when it becomes dangerous.
An earlier edition of this essay demanded operationalized constructs and then left its own tests as vibes — "high agreement," "discriminates cleanly." Practicing what it preaches: for the inter-rater test, compute Cohen's kappa across ≥ 10 cases and treat κ ≥ 0.6 as passing, 0.4–0.6 as "rewrite the definitions," below 0.4 as failure. For the says-no test, audit the last 20 verdicts; if fewer than ~20% were some form of "no" in a domain where the base rate of bad options is plainly higher, the framework is blessing, not filtering. For the collapse test, each surviving element must have flipped at least one real verdict in the audit window. For the out-of-sample test, hold back a third of your cases during design and score the framework on them like a forecaster — before showing it to anyone. These numbers are defaults, not laws; the point is that a threshold you can argue with beats an adjective you can't.
A framework you cannot locate the breaking point of is not robust. It is unexamined.
Part VIHow Frameworks Die
The subtlest failure mode is that frameworks shape perception, not just analysis. Once adopted, you stop seeing the variables they omit — the framework becomes the territory. This is why pressure testing cannot be a one-time gate. Frameworks are instruments, and instruments drift: the world changes under them, their users learn to game them, and their boundary conditions silently move. A framework's life is a loop, and the loop has an exit no one plans for.
Death, when it comes, arrives through a small number of recognizable routes. Each has a symptom you can watch for, a root cause in one of the six elements, and a repair — which is the practical payoff of the anatomy from Part II: knowing which element failed tells you what to fix instead of throwing the whole tool away.
| Failure mode | Symptom you'll notice | Element at fault | Repair |
|---|---|---|---|
| Unit transplant | Fluent analysis that never quite lands — the framework "works" but decisions don't improve | Question / unit of analysis | Re-derive the unit; either re-scope the framework or stop using it here |
| Topology mismatch | Constant exceptions: backflow in the "pipeline," cases straddling quadrant lines | Structure | Re-draw with the domain's real shape — usually a loop that was flattened for slides |
| Vibes constructs | Meetings argue about which box a case belongs in, not what to do about it | Dimensions (not operationalized) | Write pass/fail definitions until two independent raters agree |
| Survivorship laundering | The framework explains every success story and no failures | Mechanism (induced from winners only) | Rebuild the case base with failures; rerun the variance analysis |
| Rationalization engine | Every proposal that enters comes out approved, with framework citations attached | Decision logic (cannot say no) | Add explicit kill criteria; audit past verdicts for a base rate of "no" |
| Goodhart capture | Scores improve while the underlying reality doesn't; savvy teams "framework-wash" their work | Decision logic (became a metric) | Separate the diagnostic use from the evaluative use; rotate the operationalizations |
| Framework-as-territory | Nobody can remember the last time the team discussed a variable the framework omits | Dimensions (omissions forgotten) | Schedule an omissions review: list what the framework excludes and check each against reality |
| Zombie framework | Used from habit and slide inertia; its confident verdicts and reality have visibly diverged | Boundary conditions (world moved) | Retire it publicly, or re-bound it to the niche where it still holds |
So the final discipline is periodic and uncomfortable: actively hunt for cases where the framework pointed confidently one way and reality went the other. Treat each one not as an exception to wave off but as data about the boundary conditions. The organizations that get lasting value from frameworks are not the ones with the best frameworks; they are the ones with a working recalibration loop.
Part VIIThis Essay, Tested Against Itself
An essay that hands you a pressure-test battery and doesn't run it on itself is asking you to do something it won't. So: the essay is itself a framework — six elements, seven-shape catalog, two-route pipeline, eight tests — and here is how it scores against its own battery. This is also a changelog: every revision in this edition came out of a failed test, which is the recalibration loop of Figure 4 running on its own author.
| Test | Verdict on this essay | What was done about it |
|---|---|---|
| Says-no | Failed in the previous edition — it never said when a framework is the wrong tool entirely | Figure 1 added: four questions that route you to a checklist, model, theory, or intuition first |
| Out-of-sample | Failed — the "five shapes" claim broke on Cynefin, an example inside the essay's own Table 2 | Part III rebuilt as an open catalog with seven shapes plus explicit hybrids |
| Retrodiction | Failed — the design procedure couldn't retrodict Five Forces, the essay's own flagship example, which was deduced, not induced | Part IV rebuilt around two routes with triangulation |
| Inter-rater | Failed — the battery's own pass criteria were adjectives ("high agreement"), the exact vibes-construct flaw it names | Operational thresholds added after Table 3 (κ ≥ 0.6, ~20% no-rate, one flipped verdict per element) |
| Collapse | Passes, with a cut — the worked example's "clean" version added nothing the caveats didn't teach better | The tidy story now carries its own contradiction: small n, hindsight labels, unchecked independence |
| Inversion | Partial pass — the six elements are stable across readers, but "which shape does the domain really have" still admits motivated answers | Flagged honestly; the shape question is the essay's least operationalized construct |
| Goodhart | Untestable yet — the essay has no users being scored against it | Predicted gaming vector, stated in advance: teams will "run the battery" as theater, one soft case per test |
| Boundary probe | Now stated — the essay's examples are Western business frameworks; its transfer to scientific, engineering, or policy frameworks is asserted, not demonstrated | Treated as the essay's open out-of-sample frontier — reader beware, and reader welcome to break it |
A first-principles essay owes you a label on its load-bearing claims. Well-documented elsewhere: the surgical-checklist and aviation results; expert intuition beating explicit analysis in fast-feedback domains; the weak empirical support for Maslow's strict ordering; Five Forces' roots in industrial-organization economics. The author's synthesis, offered for testing: the six-element anatomy, the shape catalog, the two-route design claim, the test battery and its thresholds, and every entry in the failure-mode gallery. The kill-list example is a constructed illustration, not a case study. Nothing in the second list should be believed because it is written confidently; it should be believed if it survives your cases.
CodaOne Caveat on This Essay's Own Framing
In the spirit of the exercise, this essay should disclose its own optimization target: everything above is tuned for decision-usefulness — frameworks as tools that change what a practitioner does next. The strongest alternative lens is the academic one, which treats frameworks as precursors to theory and judges them by whether they generate testable hypotheses rather than better decisions. If your context is research rather than practice, weight the mechanism element and the falsifiability tests more heavily, and the compression-for-usability criteria less. The anatomy is the same; the grading rubric is not.
And that disclosure is itself the final lesson. Every framework — including a framework about frameworks — is a bet about what matters. The honest ones state the bet, publish their test results, and revise in public, the way Part VII tries to. The dangerous ones let you forget a bet was ever placed.