English日本語|PDF (EN)PDF (JA)
v0.28.9 — This text is under construction. The structure of the theory, the propositions, and the empirical conclusions may all change. Overview

Chapter 12
Selection Bias: Decomposing the Distortions That Enter Observation

12.1 The problem

Applying the framework of Part I to real firms introduces systematic distortion at the very start of any comparison. The distortion is not a single “bias” but a mixture of at least eight problems of different kinds. They are decomposed below, and those that can be handled are separated from those that cannot be, in principle.

12.2 Conditioning on survival: a circularity in the framework

An observable firm is the realization of a path that did not violate the constraint (4.3). Writing the insolvency time as τ = inf ⁡ {t : M(t) < 0}, what is observed is

ℙ(X∣τ > t) (12.1)

and not ℙ(X). And what fixes τ is precisely Φ, the object of estimation. The structure is shown in Figure 12.1.

Figure 12.1: Survival S is a common consequence of Φ and 𝜀 — a collider. Conditioning on S induces a spurious correlation between them.

S is a common consequence of Φ and the exogenous shock 𝜀, that is, a collider. Conditioning on a collider induces correlation between two causes that are independent (Berkson’s paradox).

12.2.1 The direction of the distortion reverses with what is conditioned on

The distinction matters in practice.

Conditioning on survival. Among surviving firms, those that persisted despite an unfavourable Φ are those whose 𝜀 was favourable. In a surviving sample the disadvantage of Φ therefore appears offset by an advantage in 𝜀, and the effect of structure is understated. The disadvantage of CCC > 0 looks smaller than it is when only survivors are examined.

Conditioning on success (fame). Here the situation differs. Collecting only successful firms removes from the sample every firm that failed with the same Φ, so there is no control group. The problem is not understatement but that estimation is impossible at all: the sample contains no information that separates Φ from 𝜀.

Treating both as one and the same “survivorship bias” hides this difference.

12.3 Selection into disclosure: correlated with the axis of interest

Listing and disclosure are not random selections. By (5.6), a firm with CCC < 0 has no g⋆ and can finance growth internally, so in principle it has little need to go to outside capital markets. Conversely a firm with a large positive CCC needs outside capital and is pushed towards listing.

That is, the probability of selection is itself a function of the object of estimation. The selection probability ℙ(Sdisc = 1∣CCC) tends to be increasing in CCC. The distribution of CCC observed in a database of listed companies is therefore displaced from the population distribution systematically and in a predictable direction.

12.4 Age: a problem of domain, not of bias

That a firm on its first day cannot be compared with one several years old is a problem of a different kind. Equation (5.1) was derived assuming a steady state and is undefined at r ≈ 0, where it divides by zero.

This is therefore not a bias but a question of the framework’s range of application. The remedy differs accordingly: a bias is corrected, whereas a question of domain can only be declared. Concretely, one of the following is taken.

(1)
Restrict the objects to a region where r is stable, and state the threshold.
(2)
Model the transient explicitly with a separate system of variables.
(3)
Lower the unit of analysis from the firm to Φ and judge maturity per Φ (Section 12.10.3).

Introducing a correction factor and forcing the steady-state formula onto the data is the remedy most to be avoided.

12.5 Age, period and cohort are not identified

Three clocks act on an observed firm at once. Writing c for the founding cohort, p for the year of observation and a for age (this section follows the notation of the Age–Period–Cohort literature; a,p,c have these meanings in this section only and are unrelated to their use elsewhere),

a = p − c (12.2)

holds identically. By (12.2) the three variables are linearly dependent, and identifying the three effects simultaneously is impossible without adding an exogenous restriction. This has been discussed at length in demography and epidemiology as the Age–Period–Cohort problem, and no general solution exists.

Hence the question “is a given industry’s favourable CCC in a given period due to Φ, to youth, or to the financial environment of the time?” cannot be answered from data alone. All one can do is state an identifying assumption — setting one of the effects to zero, say — and show how much the conclusion depends on it.

12.6 Retrospective narrative

The description of a successful firm’s Φ is itself contaminated. The structure is retold so as to explain the success, so bias enters the measurement of the explanatory variable Φ. No correction on the dependent-variable side removes it.

As Part I stated, most Φ are not products of design but arise as properties of a line of business, their financial function recognized after the fact. Distinguishing primary sources — contemporaneous disclosures, contracts, reporting — from retrospective accounts is the minimum remedy.

12.7 Mistaking the unit: what a row denotes versus what its values denote

The preceding sections concerned distortions in the selection of the sample. Failures of a different kind occur at the stage of acquisition.

What a row denotes and what the values in that row denote need not coincide. When a published list has rows at a lower unit while the figures are reported at a higher unit, the value is replicated across every row belonging to the same higher unit. Treating the list as observations lets an object with many lower units count as many times over.

Missing values can be detected in the output; this mistake cannot. The values are complete, so it passes any post-collection check. Section 15.3.5.0 records a case in which it reversed the conclusion.

There is only one remedy: confirm the unit at which the values are reported and consolidate at that unit. Consolidation is not preprocessing; it is a step that decides the conclusion.

12.8 Selection of variables: the gaps are not random

The foregoing concerned distortions in the sample and in identification. The same kind of problem arises for which variables are observable.

There can be a correlation whereby the areas of greatest theoretical interest are the hardest to measure. A field in which institutions have intervened heavily is evidence that the authorities judged there to be a problem there, and theoretical interest is drawn to the same place. Yet in a field whose transaction forms are unstable enough to require intervention, that very instability can make the required quantity impossible to compute from published aggregates (Remark 15.3).

Treating gaps in the data as random systematically removes from the sample the area one most wants to know about. The other sections of this chapter treat selection of the sample; this is selection of variables, and no correction on the sample side addresses it. All one can do is describe where the gaps are concentrated.

12.9 The observation frame in practice

Some orders of magnitude for Japan.

Population

Size

Source and notes

Firms (SME Agency basis)

about 3.6 million

recompiled from the Economic Census; excludes primary industries

Firms (National Tax Agency basis)

about 6.4 million

includes about 3.79 million sole proprietors

Listed companies

3,975

as of 27 August 2026

Firms mentioned routinely

a few hundred

of order 10−4 of the whole

Table 12.1: The size of the corporate population in Japan. The order of magnitude changes with each stage of definition.

Listed companies are about 0.1% of the whole, well-known firms less than 0.01%. When one says “a variety of firms”, which population is meant decides everything.

But exhaustiveness does not guarantee size. Even with a complete frame, if the population itself is small there is no power. Section 15.3.5.0 treats a population that can be obtained exhaustively yet numbers only 60 to 70 entities. Exhaustiveness and size are independent requirements, and both must be checked before a survey’s prospects can be assessed.

12.10 How to proceed

What can be handled is listed separately from what cannot.

12.10.1 Declare the frame in advance and track forward

Rather than “collect well-known firms”, define “all operators in a given industry existing at a reference date t0” and track forward from there. Survival then becomes an observed dependent variable rather than a hidden filter, and the collider conditioning of Section 12.2 is avoided.

Available frames in Japan include the Economic Census and the databases of Teikoku Databank and Tokyo Shoko Research. In Europe, Orbis covers unlisted firms broadly. Disclosure obligations in the United States are looser, so its unlisted data are relatively weak; in this respect Japan and Europe are better placed.

12.10.2 Collect the failure side explicitly

Insolvencies are observable through the official gazette and the tabulations of credit research firms. If a denominator can be constructed, survivorship bias becomes an estimable selection model.

12.10.3 Lower the unit of analysis from the firm to Φ

This is the most effective remedy. By the definitions of Chapter 2, Φ is an object at the level of a contract and a firm is a bundle of them. A large firm normally runs several Φ in parallel whose κ differ even in sign. Computing CCC at the level of the firm is looking at an average that mixes them.

Lowering the unit of analysis to Φ

12.10.4 Group by the shape of the constraint, not by industry classification

Industry classifications fix a particular era’s view of industry and differ by country. Grouping instead by the structural features the framework of Part I names —

makes comparison meaningful across industries and eras. This is the return on having constructed the theory.

12.10.5 Where a distortion cannot be removed, state its sign

The non-identification of Section 12.5 and the retrospective narrative of Section 12.6 cannot be removed. Their direction can still be stated. Saying “this estimate understates the effect of structure” or “this description of Φ likely contains after-the-fact rationalization” makes the robustness of a conclusion assessable.

12.10.6 A rule for using well-known firms

Hence the following rule.

Well-known firms are used for illustration, not for inference.

Citing an individual firm to explain a structure is permitted. Concluding that a structure is superior from the existence of a successful example is not.