Position: Probabilistic Modelling is Sufficient for Causal Inference

Bruno Mlodozeniec, David Krueger, Richard E Turner

International Conference on Machine Learning 2025 · Oral

Overview

This talk, presented by Bruno Mlodozeniec and co-authored with David Krueger and Richard E Turner, challenges a foundational debate in machine learning and statistics: whether specialized "causal tools" are strictly necessary for answering causal questions. The central thesis is that standard probabilistic modeling, when applied rigorously and comprehensively, is entirely sufficient for performing causal inference. This contentious claim directly confronts the long-held views of prominent researchers like Judea Pearl, who has famously advocated for a distinct causal-statistical divide and the necessity of operators like the do-operator.

Watch on SlidesLive

Visual summary for Position: Probabilistic Modelling is Sufficient for Causal Inference by Bruno Mlodozeniec, David Krueger, Richard E Turner
Visual summary for Position: Probabilistic Modelling is Sufficient for Causal Inference by Bruno Mlodozeniec, David Krueger, Richard E Turner

Key moments

  1. 0:00 Introduction and the Pearl vs. Gelman debate
  2. 1:30 Thesis: Probabilistic modeling is sufficient for causal inference
  3. 2:00 Illustrative example: Aspirin efficacy and confounding
  4. 4:00 Key assumptions for probabilistic causal inference
  5. 5:00 Demonstrating causal inference with the probabilistic model
  6. 6:00 Do-operator as 'syntactic sugar' for probabilistic models
  7. 6:30 Analyzing Pearl's causal-statistical distinction

Probabilistic Modelling: A Sufficient Foundation for Causal Inference

Speakers: Bruno Mlodozeniec, David Krueger, Richard E Turner

Conference: ICML 2025

YouTube: https://slideslive.com/39043887

Overview

This talk, presented by Bruno Mlodozeniec and co-authored with David Krueger and Richard E Turner, challenges a foundational debate in machine learning and statistics: whether specialized "causal tools" are strictly necessary for answering causal questions. The central thesis is that standard probabilistic modeling, when applied rigorously and comprehensively, is entirely sufficient for performing causal inference. This contentious claim directly confronts the long-held views of prominent researchers like Judea Pearl, who has famously advocated for a distinct causal-statistical divide and the necessity of operators like the do-operator.

The work aims to unify disparate perspectives by demonstrating that the perceived unique capabilities of causal frameworks can be elegantly replicated and understood within a purely probabilistic paradigm. By illustrating this through concrete examples, the authors argue that the existing causal toolkit, while useful, can be re-interpreted as a set of convenient "shorthands" or higher-level abstractions that operate within a more fundamental probabilistic framework. This perspective promises to simplify the conceptual landscape of causal inference, making it more accessible and generalizable for practitioners and researchers alike, by advocating for a unified approach centered on explicit probabilistic model specification.

The talk asserts that this probabilistic unification offers a clearer and more general approach to causal inference, capable of addressing not only interventional queries but also counterfactual questions, all within the same modeling machinery. This has significant implications for how researchers approach problems involving interventions, policy evaluation, and understanding "what-if" scenarios, suggesting that a deep understanding of probabilistic models and their assumptions is the ultimate prerequisite, rather than mastery of a separate causal calculus.

Background

▶ Watch: Introduction and the Pearl vs. Gelman debate (0:00)

The field of machine learning and statistics has long grappled with the distinction between correlation and causation. While statistical models excel at identifying relationships within observed data, inferring causal effects – understanding how interventions change outcomes – has traditionally been seen as a more complex problem requiring specialized tools. This perceived divide has fueled a vigorous debate among eminent researchers, perhaps best exemplified by the disagreement between Judea Pearl and Andrew Gelman.

Judea Pearl, a pioneer in causal inference, notably articulated his framework in "The Book of Why," where he introduced causal graphical models and the do-operator. Pearl's perspective is that traditional statistics is "model-blind" and merely a "data reduction enterprise," inherently incapable of answering causal questions without enriching its language with explicit causal operators. He famously asserted, "No way... there is no way to answer causal questions without what I call statistical statistical apparatus vocabulary is helpless in solving causal problems." For Pearl, the essence of causality lies in dealing with "changing conditions such as those induced by an intervention," distinguishing it fundamentally from statistical analysis of static conditions.

In stark contrast, Andrew Gelman, a distinguished professor of statistics, argued in his review of Pearl's work that it is "baffling that Pearl and his colleagues keep taking statistical problems and to my mind complicating them by wrapping them in a causal structure." Gelman maintained that many causal questions could indeed be answered within a standard statistical vocabulary. This fundamental disagreement reverberates throughout the machine learning and deep learning literature, creating a schism between those who believe in the necessity of a distinct causal framework and those who see it as an extension or re-framing of existing statistical principles.

The problem, as highlighted by Mlodozeniec, is the existence of these "completely contradictory views on the technical foundations of their fields." This creates confusion, potentially leads to the development of redundant tools, and arguably imposes artificial boundaries on what constitutes "causal" versus "statistical" methodology. The authors contend that this debate can be put to rest by demonstrating that standard probabilistic modeling tools are not only capable but also sufficient for causal inference, thereby offering a unified and general approach that encompasses both observed data analysis and interventional reasoning.

Key Findings

▶ Watch: Illustrative example: Aspirin efficacy and confounding (2:00)

The core contribution of this work is the assertion that standard tools of probabilistic modeling are sufficient for causal inference. This directly challenges the notion that a separate "causal toolkit" or a distinct "causal-statistical distinction" is fundamentally necessary. The authors demonstrate this sufficiency by showing that by carefully and explicitly writing down all assumptions about a problem in the form of a probabilistic model, one can mechanically derive answers to causal questions.

Key findings and contributions include:

  1. Sufficiency of Probabilistic Modeling: The central finding is that any causal inference problem can be framed and solved purely within a probabilistic framework. This is achieved by defining a joint distribution over all variables of interest across all relevant settings – both the observed world and any hypothetical, intervened-upon worlds. Once this joint distribution is specified based on explicit assumptions, any causal query becomes a mechanical inference problem (e.g., marginalization or conditioning).
  1. Generality and Clarity: The resulting probabilistic approach is presented as both clear and general. By forcing explicit articulation of assumptions about observed data and the effects of interventions, the method provides a transparent pathway to causal conclusions. This clarity contrasts with the potential for obfuscation or specialized jargon that can arise from distinct causal frameworks.
  1. Causal Toolkit as Shorthands: The authors propose that the established "causal toolkit," including concepts like the do-operator and causal graphical models, can be reinterpreted not as fundamentally different mathematical constructs, but as useful shorthands or "syntactic sugar" within the broader probabilistic framework. For example, the do-operator implicitly defines a joint distribution over multiple intervened-upon settings, starting from a model of the observed setting. This re-framing positions existing causal methods as convenient abstractions rather than indispensable foundational elements.
  1. Unified Treatment of Counterfactuals: A significant finding is that the probabilistic framework can handle counterfactual questions (e.g., "What would have happened if this specific individual had received a different treatment?") under the same umbrella. While traditional causal frameworks often require transitioning to a new model class like Structural Causal Models (SCMs) to address counterfactuals, the probabilistic approach can incorporate them by sharing additional latent variables (e.g., individual characteristics) between the observed and counterfactual worlds within the same joint model. This highlights the unifying power of the proposed method.
  1. Challenging the Causal-Statistical Distinction: The talk directly disputes Pearl's strict causal-statistical distinction. Mlodozeniec argues that Pearl's definition of "statistical" – limited to static conditions – unfairly constrains the term and is historically inaccurate, ignoring decades of statistical and machine learning research on generalizing under changing conditions (e.g., covariate shift, distribution shift, off-policy reinforcement learning). The work implies that the "causal trademark" has been applied too broadly, encompassing concepts that are inherently part of a richer probabilistic understanding.

In essence, the key finding is a methodological unification, advocating for a "write down the probability of everything" approach as a universal solvent for causal problems, making specialized causal formalisms optional conveniences rather than necessities.

Technical Deep Dive

▶ Watch: Key assumptions for probabilistic causal inference (4:00)

The core technical argument revolves around the principle of explicitly modeling all relevant variables and their relationships across all settings of interest within a single probabilistic model. This approach contrasts with causal frameworks that often start with a model of the observed world and then apply specific operators (like the do-operator) to infer interventional effects. The proposed method makes the "interventional" world an explicit part of the initial model specification.

Let's dissect this using the concrete example provided: predicting the efficacy of aspirin on headache duration.

  1. The Problem with Naive Approaches:

The talk begins by illustrating a common pitfall: using observational data naively. If we plot aspirin dose against headache duration from survey data, we might observe a counterintuitive positive correlation (headache duration increases with aspirin dose). This is a classic case of confounding. In this scenario, the initial headache severity is the confounder:

  • More severe headaches lead to longer durations.
  • More severe headaches also lead people to take more aspirin.

This creates a spurious positive correlation between aspirin dose and duration in the observational data. Grouping by severity reveals the true negative correlation within subgroups, but a direct causal inference remains elusive with naive methods.

  1. Modeling the Hypothetical, Intervened-Upon World:

The central challenge of causal inference is to answer questions about a hypothetical world where an intervention has occurred (e.g., "What if everyone was assigned a specific aspirin dose, T-star, regardless of their initial severity?"). The probabilistic framework tackles this by modeling both the observed world (where people choose their own dose) and the intervened-upon world (where dose is assigned) simultaneously.

The authors use a probabilistic graphical model, specifically a Bayesian network, to specify their assumptions. This network describes the conditional independencies among variables. Crucially, the model needs to encompass variables from both the observed (obs) and intervened (int) settings.

  1. Explicit Assumptions in the Probabilistic Model:

The power of the approach lies in making all assumptions explicit:

  • Assumption 1: Shared Model Parameters (Latent Variables): The observed world and the intervened-upon world are assumed to be independent once we know an unknown set of model parameters, theta. These parameters represent the underlying mechanisms governing the relationships between variables. This implies that the fundamental laws of nature (or the model's parameters) are the same across both settings. If we were interested in counterfactuals, we might also share other latent variables, such as an individual's specific headache severity, across the observed and counterfactual worlds.
  • Assumption 2: Intervention as a Delta Function: In the intervened-upon world, the aspirin dose is assigned to a specific value, T-star, for everyone. This is modeled by stating that the conditional distribution of aspirin dose in the intervened world, P(A_int | ...) (where A is aspirin dose), is a delta function at T-star. This mathematically represents the intervention: the variable A_int no longer depends on other factors like initial headache severity; it is deterministically set.
  • Assumption 3: Invariance of Causal Mechanisms: The mechanism determining the headache duration is assumed to be the same in both worlds. This means the conditional distribution of headache duration given the aspirin dose and initial headache severity, P(H | A, S) (where H is headache duration, S is severity), is identical for both the observed and intervened settings. This is a crucial assumption, often referred to as invariance or transportability of mechanisms – the underlying biological process of how aspirin affects headaches doesn't change just because the dose was assigned rather than chosen.
  • Assumption 4: Same Population: The distribution over the initial headache severities, P(S), is assumed to be unchanged between the observed and intervened populations. This means we are considering the effect of intervention on the same population that we observed.
  1. Deriving the Joint Distribution and Inference:

Once these assumptions are explicitly written down, they allow for the construction of a joint probability distribution over all variables of interest across both worlds: P(H_obs, A_obs, S_obs, H_int, A_int, S_int | theta).

From this comprehensive joint distribution, inference becomes "purely mechanical." The goal is to find the hypothetical headache duration once aspirin dose T-star has been assigned, conditioned on the observed survey data. This translates to computing P(H_int | A_int=T_star, D_obs), where D_obs represents all the observed data. This is achieved through standard probabilistic operations: marginalization and conditioning.

The result of this inference, when plotted, correctly shows that headache duration decreases with the assigned aspirin dose, effectively removing the confounding bias observed in the naive approach.

  1. Comparison with Pearl's do-operator:

The authors argue that Pearl's do-operator (do(A=a)) can be seen as "syntactic sugar." Instead of explicitly writing out separate models for observed and intervened worlds, the do-operator provides a shorthand to implicitly define the joint distribution over many intervened-upon settings, starting from a model of just the observed setting. It encodes a specific set of assumptions (like those outlined above regarding intervention, mechanism invariance, and population stability) into a single, compact notation. This makes it a useful abstraction but not a fundamentally distinct mathematical concept.

  1. Unified Counterfactuals:

A key advantage highlighted is the ability to answer counterfactual questions within the same probabilistic machinery. If, instead of just sharing model parameters (theta), we also share other latent variables (e.g., an individual's specific headache severity s_i) between the observed and counterfactual scenarios, the framework naturally extends. This avoids the need to switch to a completely new model class, such as Structural Causal Models (SCMs), which is often required in traditional causal frameworks for counterfactual analysis. The probabilistic approach maintains a unified view, where counterfactuals are just another form of inference over an expanded joint distribution that includes hypothetical situations for specific individuals.

In summary, the technical deep dive reveals that the proposed method is a rigorous application of probabilistic principles, emphasizing explicit assumption specification and the construction of comprehensive joint distributions to mechanistically derive causal inferences, thereby subsuming specialized causal tools as convenient notational devices.

Experimental Setup & Results

▶ Watch: Do-operator as 'syntactic sugar' for probabilistic models (6:00)

The talk primarily illustrates the theoretical sufficiency of probabilistic modeling through a concrete, pedagogical example rather than presenting an extensive experimental setup with novel datasets or benchmarks. The "experiment" is a thought experiment designed to highlight the conceptual power of the probabilistic approach in overcoming a classic causal inference challenge.

Dataset:

The example uses hypothetical "observational survey data" on aspirin dosage, initial headache severity, and subsequent headache duration. No specific real-world or synthetic dataset is named, nor are its characteristics (size, features, distributions) detailed beyond what's necessary to illustrate the confounding problem.

Baselines:

The primary baseline is a "naive approach" of directly plotting observed aspirin dose against headache duration. This approach fails due to confounding, showing an incorrect positive correlation. Another implicit baseline is the traditional causal inference framework, which would typically employ specific causal graphs and the do-operator to address such problems. The talk aims to show that the probabilistic method achieves the same (or superior, due to unification) results without these specialized tools.

Hardware:

As this is a theoretical and conceptual demonstration, no specific hardware (GPUs, CPUs, TPUs) or computational resources are mentioned. The focus is on the mathematical framework.

Metrics:

The key "result" is the correct inference of the causal effect of aspirin on headache duration.

  • Naive Approach Result: The observed data shows headache duration increasing with aspirin dose, indicating a failure to capture the true causal effect due to confounding.
  • Probabilistic Model Result: After explicitly modeling the observed and intervened-upon worlds and performing the mechanical inference, the results show that the headache duration decreases with the assigned aspirin dose in expectation. This accurately reflects the hypothesized (and biologically plausible) causal effect, demonstrating that the confounding has been successfully addressed.

Ablations/Variations:

While not explicitly termed "ablations," the talk discusses how the framework can be extended to answer different types of causal questions:

  • Interventional Queries: Demonstrated by predicting the effect of assigning a fixed aspirin dose.
  • Counterfactual Queries: The talk mentions that by sharing other latent variables (e.g., a specific person's headache severity) between the observed and counterfactual worlds, the same probabilistic machinery can answer counterfactual questions. This is a conceptual extension rather than an empirical ablation, illustrating the generality of the framework.

In essence, the "experimental setup" is a carefully constructed illustrative scenario, and the "results" are the successful conceptual demonstration that the proposed probabilistic framework can correctly identify and quantify a causal effect that is obscured by confounding in naive observational analysis. The talk's power lies in its theoretical elegance and unifying potential, rather than in empirical benchmarks or performance numbers.

Practical Implications

▶ Watch: Analyzing Pearl's causal-statistical distinction (6:30)

The implications of adopting a purely probabilistic view for causal inference are far-reaching for practitioners, infrastructure teams, model builders, and deployers.

For Practitioners (Data Scientists, ML Engineers):

  • Unified Mental Model: The most significant implication is the potential for a unified mental model for statistical, causal, and counterfactual questions. Instead of learning a separate "causal grammar" with its own tools and concepts, practitioners can rely on their existing understanding of probabilistic modeling. This reduces cognitive overhead and the learning curve for engaging with causal problems.
  • Explicit Assumption Specification: The approach forces practitioners to explicitly write down all assumptions about the observed data generation process and the nature of any interventions. This rigor can lead to more transparent, auditable, and robust analyses, as hidden assumptions are brought to light.
  • Flexibility for Non-Standard Problems: When standard causal frameworks (e.g., do-calculus) don't perfectly fit a non-standard problem, the probabilistic framework provides a natural "lower level of abstraction" to fall back on. By making all assumptions explicit, practitioners can tailor models to unique scenarios without being constrained by the conventions of a specific causal toolkit.

For Infrastructure Teams and Model Builders:

  • Simplified Toolchain: If causal inference can be handled within standard probabilistic programming frameworks (e.g., PyMC, Stan, Edward, Pyro) or general-purpose deep learning libraries (e.g., TensorFlow Probability, PyTorch), it could simplify the required software stack. There might be less need for specialized causal inference libraries that operate on distinct conceptual foundations.
  • Generalizable Architectures: Model builders might design more generalizable architectures that can answer both predictive and causal questions, potentially using the same underlying probabilistic neural networks or generative models. This could streamline development and deployment pipelines.
  • Data Requirements: The need to model all settings of interest (observed and intervened) explicitly emphasizes the importance of data collection strategies that can inform these different components. This might encourage more thoughtful experimental design or the use of simulation to augment observational data.

Tradeoffs and Limitations:

  1. Increased Modeling Burden: While conceptually unifying, the explicit specification of joint distributions over all relevant worlds can be a substantial modeling burden, especially for complex scenarios with many variables and potential interventions. Pearl's do-operator, precisely because it's "syntactic sugar," can abstract away much of this complexity for common problems.
  2. Identifiability Remains Crucial: As discussed in the Q&A, the probabilistic framework does not inherently solve identifiability challenges. It's possible to write down a model where, even with infinite data, certain causal quantities cannot be uniquely determined. While identifiability is not unique to causal problems (statisticians have long dealt with it), its complexity can increase with the number of latent variables and settings being modeled. Practitioners still need to ensure their chosen model and available data allow for unique identification of the parameters or queries of interest.
  3. Loss of Abstraction for Specific Tasks: Dedicated causal frameworks often provide powerful abstractions and algorithms for specific tasks like causal discovery (learning the causal graph from data) or root cause analysis. While the probabilistic framework is sufficient, it might require more manual effort or custom derivations for these problems, potentially losing the efficiency and specialized guidance offered by established causal toolkits. The "syntactic sugar" and "do-calculus" can indeed speed up tackling causal inference problems for many people and provide a good level of abstraction for certain research problems.
  4. LLMs are not a Panacea: The talk explicitly states that merely training a Large Language Model (LLM) on observed data would not be sufficient for causal inference. While LLMs can be viewed as probabilistic tools, they would require "some kind of rules for extracting from it the answers to causal questions," implying that the explicit model specification and reasoning framework are still necessary, even if an LLM is used as a component.
  5. Challenging Established Paradigms: The position challenges a well-established and influential paradigm. Adopting this view may require a significant shift in thinking for many researchers and practitioners who have been trained in the Pearlian framework.

In conclusion, the probabilistic approach offers a powerful, unifying vision for causal inference, promising greater clarity and generality. However, its practical adoption will require a commitment to rigorous model specification and an understanding of its inherent tradeoffs, particularly regarding modeling complexity and the continued importance of identifiability. It reframes the "causal toolkit" as a set of valuable, but ultimately optional, abstractions within a more fundamental probabilistic reality.

Key Takeaways

  • Probabilistic modeling is fundamentally sufficient for causal inference. There's no need for a separate "causal toolkit" as a foundational requirement.
  • Explicitly define joint distributions over all settings of interest. This includes both observed data generation processes and hypothetical, intervened-upon worlds.
  • Causal tools like the do-operator are useful "syntactic sugar" or shorthands. They encapsulate specific assumptions and operations within the broader probabilistic framework, making them convenient abstractions rather than indispensable foundational elements.
  • Counterfactual questions can be answered within the same probabilistic machinery. By sharing relevant latent variables (e.g., individual characteristics) across observed and counterfactual scenarios, the framework avoids the need for new model classes like Structural Causal Models (SCMs).
  • The strict "causal-statistical distinction" is challenged. The talk argues that Pearl's definition of "statistical" is overly restrictive and historically inaccurate, ignoring decades of work in machine learning and statistics on generalizing under changing conditions.
  • The core principle is "always write down the probability of everything." This rigorous approach ensures all assumptions are explicit, leading to clearer, more general, and unified solutions for various causal problems.

About the Speaker(s)

Bruno Mlodozeniec was the presenter of this talk at ICML 2025. He introduced himself as the primary speaker for this presentation. The work discussed in the talk is a collaborative effort, co-authored with David Krueger and Richard E Turner. While specific biographical details for each speaker were not extensively provided within the transcript, their collective contribution to this paper indicates their involvement in foundational research at the intersection of machine learning, statistics, and causal inference, aiming to re-evaluate established paradigms in the field.

Reviews

Maya Iyer (Theoretical ML Researcher) — SOLID

A competent and clearly argued philosophical position paper that challenges Pearl's causal-statistical distinction by reframing causal inference as a special case of probabilistic modeling over expanded variable sets. The central thesis — that the do-operator and SCMs are 'syntactic sugar' over a sufficiently general joint distribution — is not technically wrong, but it is also not new. The work occupies a well-trodden space between Dawid's decision-theoretic approach, Rubin's potential outcomes framework, and the long-standing observation that interventional distributions can be represented as observational distributions over augmented spaces. The talk is well-constructed and…

Chen Zhao (Applied ML Researcher & Empiricist) — SOLID

Mlodozeniec, Krueger, and Turner argue that probabilistic modeling is sufficient for causal inference, reframing Pearl's do-operator as syntactic sugar over explicit joint distributions. The thesis is philosophically coherent and the pedagogical framing is clear, but the contribution is primarily conceptual reinterpretation rather than a new technical capability. There are no empirical experiments, no benchmarks, no ablations, and no falsifiable predictions — which is fine for a position paper, but limits the rating considerably. The core argument is not wrong, and it's a useful perspective, but it doesn't resolve the hard problems it gestures at (identifiability, causal discovery…

→ Top-rated talks at International Conference on Machine Learning 2025

All talks from International Conference on Machine Learning 2025