English日本語|PDF (EN)PDF (JA)
v0.28.9 — This text is under construction. The structure of the theory, the propositions, and the empirical conclusions may all change. Overview

Appendix A
Supplement: proofs concerning method

Those claims about method made in the body that call for proof are gathered here. Each section also states why the formalization takes the form it does, which implications were not treated in the body, and the limits within which the proof applies.

The implications developed here are: the relation between the strength of conditioning and the distortion (Remark A.3); that collecting only failures does not remove the distortion (Remark A.4); that curvature is identified in the APC problem (Remark A.8); that the APC problem is latent in the CCC series of this work (Remark A.10); that the objective becomes discontinuous under a deadline constraint (Remark A.13); and that substitutability is confined to a static constraint (Remarks A.17 and A.18).

Limits of application are recorded in Remarks A.5 and A.14.

A.1 Correlation induced by conditioning on a collider

A.1.1 The choice of formalization

Section 12.2 treats the problem that observed firms are conditioned on survival. To formalize this one must decide what survival is a function of.

This work adopts a linear additive threshold S = 1{aΦ + 𝑏𝜀 > 𝜃}, where Φ is the advantage deriving from the contractual form and 𝜀 an exogenous shock. There are three reasons for this form.

First, additivity is essential. The structure in which a disadvantage in Φ can be compensated by an advantage in 𝜀 is the source of the correlation induced by conditioning. A multiplicative form S = 1{Φ𝜀 > 𝜃} yields results of the same kind, but loses the reading in terms of compensation.

Second, it must be a threshold. Survival is a binary event rather than a continuous quantity, so what is conditioned on is a set. Were S continuous, this would be a problem of regression, not of conditioning.

Third, the sign restriction a,b > 0 states that both contribute positively to survival. If one were negative the result would reverse. The statement in Section 12.2 that “the direction of the distortion reverses with what is conditioned on” depends on these signs.

Proposition A.1 (Collider bias). Let Φ and 𝜀 be independent random variables, and for a,b > 0 and a threshold 𝜃 define the survival indicator S = 1{aΦ + 𝑏𝜀 > 𝜃}. Then conditioning on S = 1 makes Φ and 𝜀 negatively correlated:

Cov ⁡ [Φ,𝜀∣S = 1] < 0.

Proof. Fix a realization x of Φ. The condition S = 1 is equivalent to 𝜀 > (𝜃 − 𝑎𝑥)∕b. The right-hand side is decreasing in x, so the conditional distribution 𝜀∣(Φ = x,S = 1) is stochastically decreasing in x. Hence

𝔼[𝜀∣Φ = x,S = 1]

is decreasing in x. If the joint distribution of Φ and 𝜀 is non-degenerate the decrease is strict, and in the definition of the covariance

Cov ⁡ [Φ,𝜀∣S = 1] = 𝔼[(Φ − 𝔼Φ)(𝔼[𝜀∣Φ,S = 1] − 𝔼𝜀)|S = 1]

the two factors inside the bracket move in opposite directions, so the covariance is negative. □

Corollary A.2 (Understatement of the structural effect). In a surviving sample, a disadvantage in Φ is partly offset by an advantage in 𝜀. Estimators of the effect of Φ on outcomes are therefore biased towards zero.

A.1.2 Implications not treated in the body

Remark A.3 (Strength of conditioning and size of the distortion). The smaller ℙ(S = 1), the larger the absolute conditional correlation. Raising the threshold 𝜃 requires ever more extreme combinations of Φ and 𝜀 for survival, which strengthens the compensating relation between them.

This has a practical implication: the more severely the sample is narrowed, the larger the distortion. The listed firms treated in Section 12.9 (0.11% of the whole) are subject to a double conditioning — listing as well as survival — so the distortion exceeds that of simple survival conditioning.

Remark A.4 (Conditioning in the other direction). The same mechanism operates when only exited firms are observed, that is S = 0. Since ℙ(𝜀 < (𝜃 − 𝑎𝑥)∕b) is increasing in x, Φ and 𝜀 are negatively correlated under S = 0 as well.

Hence collecting only failures does not remove the distortion. Only a sample containing both successes and failures — that is, an unconditioned sample — is unbiased. This explains the importance of the fact that in Part IV the media panel for family 2 contained growing and shrinking segments at the same time.

Remark A.5 (Limits of the proof). Proposition A.1 assumes unconditional independence of Φ and 𝜀. If the two are correlated to begin with, the negative contribution of conditioning mixes with the original correlation and the sign is indeterminate.

In reality contractual form and exogenous shocks may not be independent. A firm that chooses a Φ robust to the business cycle has, in effect, chosen an environment with small variance of 𝜀. In that case the proposition does not apply.

Example A.6 (Numerical confirmation). Let Φ,𝜀 ∼ N(0,1) be independent and S = 1{Φ + 𝜀 > 0}. Then Φ + 𝜀 ∼ N(0,2) and ℙ(S = 1) = 1∕2. It is a known result that under S = 1, Corr ⁡ [Φ,𝜀∣S = 1] ≈−0.57. A correlation that is 0 unconditionally swings sharply negative merely by conditioning on survival.

A.2 Non-identification of age, period and cohort

A.2.1 Why the argument is made in a linear model

The claim of Section 12.5 is stated here in the language of linear algebra. The linear setting is adopted not because non-identification arises from linearity, but because it appears most clearly in the linear case.

In a non-linear model (α,π,γ) can sometimes be identified: if, for instance, the age effect is quadratic and the period effect linear, the two can be distinguished. But such identification depends entirely on the assumed functional form. When changing the assumption changes the conclusion, one can hardly say the parameters are identified.

The linear case is presented to make clear that the core of the non-identification lies in the identity a = p − c and cannot be evaded by contrivances of functional form.

Proposition A.7 (Non-identification in APC). This proposition follows the notation of the age–period–cohort literature. Within this section a,p,c,π,y denote age, period, cohort, the period effect and the observation, and bear no relation to their use elsewhere in this work (allocation received, unit price, marginal cost, settlement, outcome).

Consider the model explaining an observation y as a linear sum of age a, period p and cohort c,

y = μ + αa + πp + γc + 𝜖.

When the identity

a = p − c (12.2)

holds, (α,π,γ) is not identified.

Proof. For any η ∈ ℝ set (α′,π′,γ′) = (α + η,π − η,γ + η). By (12.2) ,

α′a + π′p + γ′c = 𝛼𝑎 + 𝜋𝑝 + 𝛾𝑐 + η(a − p + c) = 𝛼𝑎 + 𝜋𝑝 + 𝛾𝑐.

That is, the fitted values agree for every η. The columns (a,p,c) of the design matrix are linearly dependent, the rank falling short by one, so the least-squares solution is not unique. □

A.2.2 Implications not treated in the body

Remark A.8 (What is identified and what is not). What is unidentified is the level of the three coefficients; not everything is unknown.

Any quantity left invariant by the transformation (α + η,π − η,γ + η) of Proposition A.7 is identified. For instance α + γ, and the second differences of the coefficients — that is, curvature — do not depend on η.

Hence “does CCC rise monotonically with age?” is not identified, but “is there curvature in the age effect?” is. Abandoning the separation of the three variables in Section 12.5 does not mean abandoning curvature as well. This work does not pursue that direction.

Remark A.9 (A constraint for identification). Identification requires adding one exogenous constraint. Typically one imposes α = 0, π = 0 or γ = 0, or assumes two adjacent coefficients equal. Every such constraint is untestable, and the conclusion depends on the assumption. This work imposes no constraint and adopts the policy of stating the non-identification explicitly (Section 12.5).

Remark A.10 (How it appears in this work). The APC problem is latently at work in several places here.

Chapter 16 treats the time series of CCC; if the age distribution of firms changes over time, the observed rise in CCC is a mixture of period and age effects. The shift-share decomposition of Section 16.8 separates industrial composition but does not separate the age composition.

If younger firms have shorter CCC, a rising share of older firms could account for the rise in CCC. This work has not tested that.

A.3 Preference for variance under a deadline

The proposition of Section 8.6 is proved here.

Proposition A.11 (Under a deadline, minimizing variance is not optimal). Let Θ > 0 be the time required to arrive, Tmax ⁡ the deadline, and define success as {Θ ≤ Tmax ⁡ }. When 𝔼[Θ] > Tmax ⁡ , an increase in variance holding 𝔼[Θ] fixed can increase the probability of success. In particular the probability of success converges to 0 in the limit Var ⁡ [Θ] → 0.

Proof. If Var ⁡ [Θ] = 0 then Θ = 𝔼[Θ] > Tmax ⁡ with probability one, so ℙ(Θ ≤ Tmax ⁡ ) = 0. On the other hand, giving the variable variance while preserving the mean — say a two-point distribution with Θ = T1 < Tmax ⁡ with probability q and Θ = T2 with probability 1 − q, chosen so that qT1 + (1 − q)T2 = 𝔼[Θ] — gives a success probability of q > 0. The increase in variance thus improves the probability of success. □

Remark A.12 (Contrast with the floor constraint). This proposition does not contradict the conclusion of Section 8.5, “lower the variance with many small amounts”. The latter holds under the floor constraint M ≥ 0, the former under the deadline constraint Tins ≤ E(0)∕c𝑙𝑖𝑣. Change the type of constraint and the optimal attitude to variance changes.

Remark A.13 (Implications of taking the probability of success as the objective). Proposition A.11 takes the objective to be ℙ(Θ ≤ Tmax ⁡ ), which differs from 𝔼[∑ ⁡ βtϕt] in Section 4.2.

The two diverge because the loss from exceeding the deadline does not depend on by how much it is exceeded. Missing Tmax ⁡ by a day and by a year give the same result, so the objective is binary. This discontinuity reverses the attitude to variance.

The insolvency time τ of Section 4.2 is likewise an absorbing state with the same property. That is, the framework of this work contains regions in which the objective is not smooth. The pathwise constraint of Proposition 4.7 is a manifestation of this.

Remark A.14 (Limits of the proof). The proposition shows that an increase in variance can improve the probability of success, not that it always does. If 𝔼[Θ] < Tmax ⁡ , an increase in variance lowers it instead.

The boundary is the comparison of 𝔼[Θ] with Tmax ⁡ . In practice this becomes the judgement play safe if the prospect suffices, gamble if it does not, but 𝔼[Θ] is not usually knowable in advance. This work does not treat that estimation problem.

A.4 Substitutability of capital and credit

Proposition A.15 (The substitution relation). When A = Inv = 0, the constraint M ≥ 0 is equivalent to

E + ∑ i max ⁡ {−κi,0}≥∑ i max ⁡ {κi,0}. (A.1)

That is, E and negative κ are perfect substitutes in the constraint.

Proof. Substitute A = Inv = 0 into equation (3.8) of Proposition 3.15 and rearrange M ≥ 0; the claim follows immediately. The two terms on the left enter with the same coefficient 1, so a decrease in one is exactly compensated by an equal increase in the other. □

Corollary A.16 (Implications in both directions). If E ≈ 0 then κ < 0 is required (Section 8.5). Conversely, if E > 0 is large, κ > 0 can be held (Part IV, the result on implication F).

A.4.1 The meaning and the limits of “perfect substitution”

Remark A.17 (The substitution holds only in the constraint). Proposition A.15 states that E and − κ are equivalent only within the single inequality (3.8). It does not mean the two are economically equivalent.

They differ in three respects.

First, there is a deadline. κ < 0 carries an obligation to deliver, and if the τi of Section 3.1.4 is finite, delivery equivalent to repayment occurs on the due date. E has no deadline.

Second, availability differs. E is raised from the capital market, κ < 0 from a counterparty. As Section 3.1.1 sets out, the two are explained by different theories and become available under different conditions.

Third, the claim on ϕ differs. The provider of E holds a residual claim (Definition 3.21); the provider of κ < 0 does not.

Remark A.18 (Where the limits of substitution appear). The one-person business of Chapter 8 presupposes E ≈ 0 and therefore treats only one endpoint of the substitution.

When E > 0 was introduced in Section 8.6, the type of constraint changed from a floor to a deadline. This is a manifestation of the asymmetry between E and κ < 0: E is consumed as living expenses and declines, whereas κ < 0 is not consumed but extinguished by delivery.

Proposition A.15 is thus a claim about a static constraint; behaviour through time differs.