Excursus 134 – On Correlation, Regression, and Causality: The Missing Counterfactual

Hossein Jorjani

First publication: 2026 – 08 – 08

The proposition that “correlation does not imply causation” is among the most frequently repeated warnings in empirical research. Properly understood, it states that an observed association is not by itself sufficient to establish a causal relation. The same association may arise from direct causation, reverse causation, a common cause, selection, measurement procedures, or coincidence.

The warning is valuable, but it is sometimes expressed too strongly. Correlation may provide evidence relevant to a causal hypothesis even though it does not establish that hypothesis. If correlation supplied no information relevant to causality, many empirical investigations could hardly begin. The methodological problem is therefore not that correlation and causation are wholly unrelated, but that an observed correlation ordinarily underdetermines the causal relation that may have produced it.

The warning has also been extended too freely from correlation to regression. Regression is sometimes treated as though it were merely an elaborate form of correlation and therefore inherently incapable of contributing to causal inference. This confuses a statistical procedure with the interpretation assigned to its results. A regression coefficient may receive a descriptive, predictive, or, under appropriate conditions, causal interpretation. The calculation itself does not determine which interpretation is justified. A causal interpretation depends upon the research design, temporal ordering, selection and measurement of variables, treatment of confounding, and assumptions connecting the statistical model to the process under investigation. Regression does not manufacture causality, but it may estimate a causal quantity once that quantity has been defined and the assumptions required for its identification have been defended.

Sewall Wright’s path analysis made an important part of this distinction unusually explicit. The direction and organization of proposed causal paths did not emerge from the correlations themselves; they depended upon substantive biological knowledge and a prior causal hypothesis. Statistical analysis could then determine whether the observed correlations were compatible with the proposed causal structure and estimate relations within it. Compatibility with the correlations, however, could show that a proposed causal scheme was possible without establishing that it was the correct causal explanation.

When path analysis and structural-equation modelling entered sociology, they encouraged researchers to state directional relations more explicitly than conventional regression ordinarily required. Nevertheless, the presence of arrows in a diagram could create an appearance of causal explanation exceeding what the research design supported. An arrow is not made causal by being drawn as an arrow. Its causal interpretation depends upon the substantive meaning, design, and assumptions attached to it.

The potential-outcomes tradition approached causality from another direction. Its origins include Jerzy Neyman’s work on randomized experiments, while Donald Rubin subsequently developed potential outcomes into a more general framework for causal inference in randomized and nonrandomized studies. Instead of beginning with a regression equation or a system of causal paths, this approach begins by defining the causal comparison itself.

For each person or other unit, one may imagine an outcome under one condition and another outcome under an alternative condition. Suppose, for example, that a student participates in a particular educational programme. One potential outcome represents that student’s educational attainment under participation; another represents the educational attainment the same student would have had without participation. The causal effect for that individual may be conceived as the difference between these two potential outcomes.

Only one of them, however, can ordinarily be observed. The student either participates or does not participate; a patient either receives a treatment or does not receive it; a policy is either implemented in a particular setting or it is not. Once one possibility has been realized, the outcome under the alternative possibility remains unobserved. Paul Holland later described this impossibility of simultaneously observing both potential outcomes for the same unit as the fundamental problem of causal inference.

Causal inference therefore contains a distinctive missing-data problem. Yet the language of missing data should not obscure an important ontological difference. An ordinary missing observation may correspond to a value that was actually realized but was not recorded: a questionnaire was lost, a measurement instrument failed, or a respondent omitted an answer. The missing potential outcome is different. It corresponds to an alternative condition that the unit did not actually experience. The first may in principle be recoverable through reconstruction or additional records; the second requires counterfactual inference.

No increase in the number of observations can remove this problem for the individual unit. Observing another thousand students does not reveal what would have happened to this particular student under the condition that the student did not experience. Additional observations may provide information with which the missing counterfactual can be estimated, but the counterfactual itself remains unobserved.

This structure creates a close connection between causal inference and missing-data theory, another area substantially developed by Rubin and later systematically treated by Roderick Little and Donald Rubin. In ordinary missing-data analysis, valid inference depends upon the process determining which values become observable and whether that process may legitimately be ignored. In potential-outcomes causal inference, an analogous question concerns the assignment process determining which potential outcome becomes observable for each unit. The analogy is powerful, although the two problems are not identical: an ordinary missing value may correspond to a value that was realized but not recorded, whereas a missing potential outcome corresponds to an alternative condition that was not realized.

Randomization provides one way of addressing the problem at the level of groups. The missing counterfactual outcome of each individual remains unobservable, but random assignment makes treatment groups comparable in expectation. The observed outcomes of one group can therefore provide information about the missing potential outcomes of the other. Randomization does not reveal the individual counterfactual; it creates a defensible basis for estimating causal effects across groups.

Observational research lacks this guarantee. Individuals, institutions, or societies exposed to different conditions may already differ in ways that influence their outcomes. Matching, regression adjustment, weighting, and related procedures may attempt to restore comparability, but their success depends upon assumptions concerning the assignment process and the relevant pretreatment variables. Statistical adjustment cannot automatically remove the effects of causes that were neither measured nor adequately represented in the design.

The potential-outcomes framework therefore proposes a particular methodological order. One does not first perform a regression and then decide whether its coefficients look causal. One first defines the alternative conditions, the potential outcomes, the causal quantity of interest, and the assumptions under which the missing counterfactual outcomes may be inferred. Statistical estimation comes afterward. In this sense, causal inference is a problem of design and definition before it becomes a problem of statistical calculation.

Potential outcomes do not exhaust the concept of causality. Structural causal models, mechanistic explanations, interventionist approaches, and other traditions formulate causal questions differently. Yet the potential-outcomes framework makes one difficulty exceptionally visible: causal claims frequently depend not only upon what was observed, but upon comparison with an alternative that was not observed.

This becomes especially important in historical and social situations described as singular and unrepeatable. The completed character of an event does not make its causes certain. Knowing what happened reveals the realized historical sequence; it does not reveal what would have happened if one presumed cause had been absent, delayed, weakened, or replaced.

A tightly connected historical narrative may reconstruct the observed sequence with considerable detail while leaving the missing counterfactual almost entirely unspecified. Suppose a sequence is reconstructed as:

A → B → C → D

Establishing that A preceded B, that B preceded C, and that C preceded D may produce an exceptionally detailed historical account. But the causal proposition that C would not have occurred without B requires something more. The observed sequence contains B and C; it does not contain the alternative history in which B is absent.

Indeed, the more insistently a historical phenomenon is treated as unique and irreplaceable, the more difficult it may become to identify comparable cases or otherwise justify the alternative outcome required for causal inference. Singularity may enrich historical understanding while weakening some of the grounds available for estimating a causal effect. Narrative density and causal certainty are therefore not the same thing.

This does not mean that historical causality is impossible. Counterfactual reasoning may draw upon comparisons among cases, differences in timing, natural experiments, institutional discontinuities, documentary evidence, and knowledge of causal mechanisms. Nor must every historical causal argument be translated into the formal language of potential outcomes. The methodological point is narrower: the certainty of a causal claim cannot be inferred merely from the continuity, detail, or narrative coherence of the observed sequence.

The broader lesson is that causality should be neither casually asserted nor prohibited by methodological slogan. Correlation does not establish causality, but it may provide evidence relevant to a causal question. Regression does not create causality, but it may estimate causal effects within an adequately specified design. A causal diagram does not transform an association into a cause merely by adding an arrow. And a detailed historical narrative does not eliminate the unobserved alternative against which many causal claims implicitly make their comparison.

The contribution of potential-outcomes and missing-data perspectives is therefore conceptual before it is statistical. They direct attention toward what has not been observed, why it has not been observed, and what assumptions permit information about the missing alternative to be inferred. Causal inference begins not with the tightness of the observed relations, but with recognition of the absent comparison upon which the causal claim depends.

Leave a comment