Chapter 12
Selection Bias: Decomposing the Distortions That Enter Observation
12.1 The problem
Applying the framework of Part I to real firms introduces systematic distortion at the very start of any comparison. The distortion is not a single “bias” but a mixture of at least eight problems of different kinds. They are decomposed below, and those that can be handled are separated from those that cannot be, in principle.
12.2 Conditioning on survival: a circularity in the framework
An observable firm is the realization of a path that did not violate the constraint (4.3). Writing the insolvency time as , what is observed is
| (12.1) |
and not . And what fixes is precisely , the object of estimation. The structure is shown in Figure 12.1.
is a common consequence of and the exogenous shock , that is, a collider. Conditioning on a collider induces correlation between two causes that are independent (Berkson’s paradox).
12.2.1 The direction of the distortion reverses with what is conditioned on
The distinction matters in practice.
Conditioning on survival. Among surviving firms, those that persisted despite an unfavourable are those whose was favourable. In a surviving sample the disadvantage of therefore appears offset by an advantage in , and the effect of structure is understated. The disadvantage of looks smaller than it is when only survivors are examined.
Conditioning on success (fame). Here the situation differs. Collecting only successful firms removes from the sample every firm that failed with the same , so there is no control group. The problem is not understatement but that estimation is impossible at all: the sample contains no information that separates from .
Treating both as one and the same “survivorship bias” hides this difference.
12.3 Selection into disclosure: correlated with the axis of interest
Listing and disclosure are not random selections. By (5.6), a firm with has no and can finance growth internally, so in principle it has little need to go to outside capital markets. Conversely a firm with a large positive needs outside capital and is pushed towards listing.
That is, the probability of selection is itself a function of the object of estimation. The selection probability tends to be increasing in . The distribution of observed in a database of listed companies is therefore displaced from the population distribution systematically and in a predictable direction.
12.4 Age: a problem of domain, not of bias
That a firm on its first day cannot be compared with one several years old is a problem of a different kind. Equation (5.1) was derived assuming a steady state and is undefined at , where it divides by zero.
This is therefore not a bias but a question of the framework’s range of application. The remedy differs accordingly: a bias is corrected, whereas a question of domain can only be declared. Concretely, one of the following is taken.
- (1)
- Restrict the objects to a region where is stable, and state the threshold.
- (2)
- Model the transient explicitly with a separate system of variables.
- (3)
- Lower the unit of analysis from the firm to and judge maturity per (Section 12.10.3).
Introducing a correction factor and forcing the steady-state formula onto the data is the remedy most to be avoided.
12.5 Age, period and cohort are not identified
Three clocks act on an observed firm at once. Writing for the founding cohort, for the year of observation and for age (this section follows the notation of the Age–Period–Cohort literature; have these meanings in this section only and are unrelated to their use elsewhere),
| (12.2) |
holds identically. By (12.2) the three variables are linearly dependent, and identifying the three effects simultaneously is impossible without adding an exogenous restriction. This has been discussed at length in demography and epidemiology as the Age–Period–Cohort problem, and no general solution exists.
Hence the question “is a given industry’s favourable in a given period due to , to youth, or to the financial environment of the time?” cannot be answered from data alone. All one can do is state an identifying assumption — setting one of the effects to zero, say — and show how much the conclusion depends on it.
12.6 Retrospective narrative
The description of a successful firm’s is itself contaminated. The structure is retold so as to explain the success, so bias enters the measurement of the explanatory variable . No correction on the dependent-variable side removes it.
As Part I stated, most are not products of design but arise as properties of a line of business, their financial function recognized after the fact. Distinguishing primary sources — contemporaneous disclosures, contracts, reporting — from retrospective accounts is the minimum remedy.
12.7 Mistaking the unit: what a row denotes versus what its values denote
The preceding sections concerned distortions in the selection of the sample. Failures of a different kind occur at the stage of acquisition.
What a row denotes and what the values in that row denote need not coincide. When a published list has rows at a lower unit while the figures are reported at a higher unit, the value is replicated across every row belonging to the same higher unit. Treating the list as observations lets an object with many lower units count as many times over.
Missing values can be detected in the output; this mistake cannot. The values are complete, so it passes any post-collection check. Section 15.3.5.0 records a case in which it reversed the conclusion.
There is only one remedy: confirm the unit at which the values are reported and consolidate at that unit. Consolidation is not preprocessing; it is a step that decides the conclusion.
12.8 Selection of variables: the gaps are not random
The foregoing concerned distortions in the sample and in identification. The same kind of problem arises for which variables are observable.
There can be a correlation whereby the areas of greatest theoretical interest are the hardest to measure. A field in which institutions have intervened heavily is evidence that the authorities judged there to be a problem there, and theoretical interest is drawn to the same place. Yet in a field whose transaction forms are unstable enough to require intervention, that very instability can make the required quantity impossible to compute from published aggregates (Remark 15.3).
Treating gaps in the data as random systematically removes from the sample the area one most wants to know about. The other sections of this chapter treat selection of the sample; this is selection of variables, and no correction on the sample side addresses it. All one can do is describe where the gaps are concentrated.
12.9 The observation frame in practice
Some orders of magnitude for Japan.
Population |
Size | Source and notes |
Firms (SME Agency basis) |
about 3.6 million | recompiled from the Economic Census; excludes primary industries |
Firms (National Tax Agency basis) |
about 6.4 million | includes about 3.79 million sole proprietors |
Listed companies |
3,975 | as of 27 August 2026 |
Firms mentioned routinely |
a few hundred | of order of the whole |
Listed companies are about 0.1% of the whole, well-known firms less than 0.01%. When one says “a variety of firms”, which population is meant decides everything.
But exhaustiveness does not guarantee size. Even with a complete frame, if the population itself is small there is no power. Section 15.3.5.0 treats a population that can be obtained exhaustively yet numbers only 60 to 70 entities. Exhaustiveness and size are independent requirements, and both must be checked before a survey’s prospects can be assessed.
12.10 How to proceed
What can be handled is listed separately from what cannot.
12.10.1 Declare the frame in advance and track forward
Rather than “collect well-known firms”, define “all operators in a given industry existing at a reference date ” and track forward from there. Survival then becomes an observed dependent variable rather than a hidden filter, and the collider conditioning of Section 12.2 is avoided.
Available frames in Japan include the Economic Census and the databases of Teikoku Databank and Tokyo Shoko Research. In Europe, Orbis covers unlisted firms broadly. Disclosure obligations in the United States are looser, so its unlisted data are relatively weak; in this respect Japan and Europe are better placed.
12.10.2 Collect the failure side explicitly
Insolvencies are observable through the official gazette and the tabulations of credit research firms. If a denominator can be constructed, survivorship bias becomes an estimable selection model.
12.10.3 Lower the unit of analysis from the firm to
This is the most effective remedy. By the definitions of Chapter 2, is an object at the level of a contract and a firm is a bundle of them. A large firm normally runs several in parallel whose differ even in sign. Computing at the level of the firm is looking at an average that mixes them.
Lowering the unit of analysis to
- partially resolves the age problem (a new inside a mature firm can be handled)
- increases the number of observations
- reduces the influence of the institutional artefact called “the firm”
12.10.4 Group by the shape of the constraint, not by industry classification
Industry classifications fix a particular era’s view of industry and differ by country. Grouping instead by the structural features the framework of Part I names —
- presence of a capacity constraint (Section 2.7)
- transaction frequency ((6.4))
- observability of effort (Chapter 4)
- capital intensity (the weight of in (3.8))
makes comparison meaningful across industries and eras. This is the return on having constructed the theory.
12.10.5 Where a distortion cannot be removed, state its sign
The non-identification of Section 12.5 and the retrospective narrative of Section 12.6 cannot be removed. Their direction can still be stated. Saying “this estimate understates the effect of structure” or “this description of likely contains after-the-fact rationalization” makes the robustness of a conclusion assessable.
12.10.6 A rule for using well-known firms
Hence the following rule.
Well-known firms are used for illustration, not for inference.
Citing an individual firm to explain a structure is permitted. Concluding that a structure is superior from the existence of a successful example is not.