CausalPFN trains on simulated causal worlds to estimate what may happen when we intervene, not just predict what comes next.
Prediction has become one of the defining strengths of modern AI.
Across industries, models are used to forecast what may happen next: whether a patient is at risk, whether a transaction looks suspicious, or whether demand may rise. In many settings, prediction is already useful because it helps people anticipate the future.
But real decisions rarely stop at prediction.
A company does not only want to know whether a customer might leave. It wants to know whether sending an offer would change that outcome. A doctor does not only want to know whether a patient is at risk. They want to know whether a treatment would help. A policymaker does not only want to forecast what may happen under current conditions. They want to know what might happen if a rule, program, or intervention is implemented.
Those are questions of cause and effect.
Whereas prediction asks what may happen if the world continues as it is, causal inference asks what may happen if we intervene.

That distinction is at the heart of our work on CausalPFN, a foundation-model approach to causal effect estimation. Our research explores whether causal inference, long one of the more specialized and expert-driven areas of machine learning, can benefit from the same broad shift that has changed predictive AI: train one model across many tasks in advance, then reuse what it has learned on new problems.
CausalPFN is an important first step on the longer road toward more capable causal AI systems, and answers an important question: can we move some of the hardest work in causal inference from bespoke, task-by-task modelling into a reusable model trained across many simulated causal worlds?
Why prediction is not enough
Predictive models are powerful because they learn patterns from observed data. If the same conditions continue, they can often estimate what is likely to happen next.
Intervention changes the problem. Once an organization acts, the original prediction no longer applies. The question changes, asking whether the action influenced the outcome.
A/B testing is one of the cleanest ways to answer a causal question. In a controlled experiment, people are randomly assigned to different options and the outcome is measured. That random assignment helps isolate whether the intervention caused a difference in the result.

But running an A/B test is not always possible. Sometimes it is too costly, too slow, impractical, or even inappropriate. Often, all that is available is observational data from the field, where confounding factors (outside influences that affect both who receives an intervention and what outcome follows) can make cause and effect harder to separate.
That is where causal inference becomes an important tool. It tries to estimate the effect of an intervention when the data was not necessarily created through a clean experiment. That is a key distinction from the more widely used predictive models.
If a model predicts that a customer has a high probability of leaving, the organization may decide to intervene with a discount, a message, or a different service experience. Once that action happens, the original prediction is no longer enough. The relevant question becomes whether the action caused a different outcome than the one that would have happened otherwise.
The same issue appears in many domains. In medicine, the question is not only whether a patient may recover, but whether a specific treatment will improve their outcome. In marketing, it is not only whether someone is likely to buy, but whether receiving a message will make them more likely to buy. In policy, it is not only whether a trend is moving in the right direction, but whether a change in policy can influence that trend.
These are causal questions because they involve a comparison between what happened and what might have happened under a different action. Comparisons like this make causal inference essential to decision-making, but they also make it difficult.
Observational data can show that two things are related, but it does not automatically show that one caused the other. A person who receives a treatment may have different health conditions from someone who does not. A customer who gets a marketing message may already be more engaged with the brand. A group affected by a policy may also be affected by other changes happening at the same time.
Those hidden or overlapping factors are the reason causal inference requires more than pattern recognition. It depends on assumptions about the data, the decision, and the variables that have or have not been observed.
Why causal inference has remained difficult to scale
Causal inference has a long and rigorous history, but it is harder to use at scale than predictive modelling.
One reason is that causal inference workflows are typically bespoke. Practitioners need to understand nuances of a particular dataset, decide which assumptions are appropriate, choose among many specialized estimators, tune methods on each dataset, and validate whether the result can be trusted. Even in relatively common settings, there are many possible approaches, and choosing the right one requires expertise.
That is very different from the direction much of AI has taken in recent years.
Foundation models, including large language models and other generative AI systems, changed expectations in language, vision, coding, and tabular prediction by showing that a model trained broadly can be reused across many downstream tasks. Instead of training a new model from scratch for every problem, practitioners can increasingly begin with a model that has already absorbed patterns from a large and diverse training process.
Causal inference has not moved as quickly in that direction because the task is harder, and interventional data is scarce. The task is not only about recognizing statistical associations, but about estimating the effect of an intervention under assumptions that may or may not hold. If important confounding factors are missing from the data, no model can simply infer the correct causal effect with certainty.
Compared to large language models where textual data for training can be scraped off the internet at enormous scale, diverse interventional datasets are not readily available for training a causal foundation model. The only appropriate datasets come from randomized controlled trials in scientific experiments, but these remain small-scale, scarce, and are not always permitted for use in model training.
These were the challenges we wanted to explore: could a model learn to account for confounding factors and estimate causal effects on new observational datasets without being rebuilt from scratch each time? Could we generate enough simulated causal settings to solve the lack of interventional data?
Researchers from Layer 6 worked jointly with Professor Rahul Krishnan’s lab at the University of Toronto and the Vector Institute to answer these questions. As part of a multi-year collaboration, the team drew on Prof. Krishnan’s depth of knowledge in causal inference, and Layer 6’s expertise in training predictive foundation models on tabular data. The resulting work, CausalPFN, was published at NeurIPS 2025, the world’s premiere academic conference on machine learning, and was recognized as a Spotlight paper, putting it in the top 15% of accepted papers.
Training on Simulated Causal Worlds
CausalPFN brings foundation-model thinking into causal effect estimation.
The core idea is to train a single transformer model across a large collection of simulated data-generating processes. These simulated environments act like causal training worlds. In each one, the model can observe data and learn how treatments, covariates, outcomes, and causal effects relate to one another.
In that sense, the model is not being handed one map for one world. It is being exposed to many different maps for alternate universes in advance.
Because these worlds are simulated, the true causal effects are known during training. That gives the model a learning signal it would rarely have in ordinary observational data, where we typically observe what happened but not what would have happened under a different intervention.
Rather than spending modelling effort from scratch on every new causal inference task, we invest that effort upfront during training. After this pre-training phase, the model can be applied to new observational datasets out of the box. It does not need to be retrained, fine-tuned, or adjusted with task-specific hyperparameters for each new problem.
To return to the map analogy: traditional causal modelling can resemble building a map for one place at a time. CausalPFN is closer to training a navigator across many simulated terrains. It does not mean the model has a perfect map of every future setting. But it may learn enough structure to orient itself when it sees a new one.
That is the conceptual shift.
The goal is not to remove human judgment from causal inference. Rather, the goal is to explore whether part of the modelling burden can become more reusable.
Read our introduction to causal foundation models here
How CausalPFN changes the workflow
In a traditional causal inference workflow, a domain expert or data scientist typically begins with a specific dataset and a specific causal question. They then need to reason about the data-generating process, select or design an estimator, tune it, and evaluate whether the estimates are credible.
CausalPFN changes where some of that work happens.
During training, we generate many causal settings and train the model to estimate causal effects from observational data. Because the training data is synthetic, there is no limit to how much can be generated, which helps achieve the scale needed for a foundation model.
At inference time, the model receives a new dataset and estimates the effect of an intervention directly. CausalPFN doesn’t need human experts to study and understand the causal relationships inherent in each new dataset. Instead, it uses its knowledge built up during pre-training to analyze causal relationships automatically in a fully data-driven process. CausalPFN leverages a labeled observational dataset as samples for in-context learning, and predicts causal effects on unlabeled data without any updates to the pre-trained model weights.

The causal settings generated for pre-training are identifiable by construction, which is a technical requirement that ensures there are no confounding effects in the data. In practical terms, that means the relevant confounding variables must be observed for inference datasets as well. If important hidden causes are missing, the causal question may not be identifiable from the data alone.
But within that setting, the workflow is meaningfully different. Instead of asking a practitioner to build or choose a new estimator for every problem, CausalPFN asks whether a broadly trained model can serve as a reusable causal estimator.
That is the sense in which CausalPFN is a foundation model: not because it is a language model, but because it applies the same broader principle of training once across many settings and reusing that training on new tasks.
Evidence That the Approach Can Transfer
Our experiments suggest that this approach is highly promising.
CausalPFN achieved strong performance across standard academic causal effect estimation benchmarks, including the best average rank for conditional treatment effect estimation across the IHDP, ACIC, and Lalonde tasks, along with competitive results for average treatment effect estimation. It also showed competitive out-of-the-box performance on uplift modelling tasks, where causal estimates are used to support policy or targeting decisions.

The important point is not only that the model performed well, but that it did so after being trained on simulated causal worlds, and purely relying on in-context learning rather than being built separately for each benchmark task.
That result supports the larger idea behind the research: causal inference may not always need to begin with a bespoke estimator designed from scratch for every dataset. A model trained across many simulated causal environments can learn patterns that transfer to new causal inference tasks.
At the same time, academic benchmarks are evidence of real-world readiness, but are not a guarantee.
Working with advice from Prof. Krishnan, the Layer 6 team also benchmarked CausalPFN against real-world causal inference problems that had been tackled by data scientists at TD in the past. In head-to-head comparisons between bespoke causal models built and used at TD, and CausalPFN applied out-of-the-box, CausalPFN once again outperformed. This is a significant milestone, as the bespoke models had been built by teams of experts over many months of trial and error. CausalPFN accomplished better results with no adjustment to the specific tasks, eliminating what would traditionally be months of effort.
Making Causal Inference Practical at Scale
If every causal model takes months to produce, causal inference must be reserved for a small number of high-value problems. A reusable foundation model changes that constraint, opening the doors to an expanded world of possibilities. It lowers the cost of asking causal questions repeatedly, allowing causal inference to be considered across a wider range of decisions.
A reusable foundation model changes the economics of causal modeling. Much of the work that once had to be repeated for each problem can move into the reusable model. While that does not make each new application cost-free, it can reduce the months of task-specific modelling effort that traditional workflows often require. This allows for a positive shift in the deployment of experts’ time. Expertise that once went
into selecting and training estimators can be redirected toward higher-value parts of the causal process: defining the intervention, testing whether the assumptions are reasonable, validating the outputs, and governing how estimates are used.
CausalPFN can also reduce operational costs compared to maintaining several purpose-built models. Since many different tasks can be completed with a single foundation model, only one model needs to be hosted and maintained over time. New causal tasks can reuse the same CI/CD pipeline, adding minimal marginal cost.
There is another potential benefit. By moving more of the modelling work into pre-training, CausalPFN reduces some of the task-by-task human choices that enter traditional causal workflows. That does not eliminate assumptions or biases — the simulated training process has its own design choices that influence the synthetic data — but it can make those choices easier to test and improve upon. Since inference becomes fully data-driven, CausalPFN may be better at identifying complex or unintuitive causal relationships that human designers would fail to consider.
Where the limits remain
CausalPFN does not remove all the conditions that make causal inference difficult to implement at scale.
First, the model depends on positivity and strong ignorability, both for pre-training data and at inference. This means CausalPFN assumes the relevant confounding factors have been observed in the data. If important hidden causes are missing, CausalPFN has no guarantee of producing a valid causal estimate. This is why human domain expertise remains essential: teams still need to decide whether the method is appropriate for the question, the data, and the decision at hand.
Second, our implementation focuses on binary treatments. That means the treatment is represented as one of two possibilities: an intervention happens or it does not. Many real-world settings involve more complex choices, including multiple treatment levels, continuous treatments, or sequences of decisions over time. Extending foundation-model-style causal inference to those settings remains an important direction we are exploring.
Additionally, CausalPFN inherits some of the context size constraints of PFN-style transformer models.1 In our experiments, performance degradation is observable on the largest data tables.
But those limits do not diminish the shift away from a bespoke modelling exercise and toward a reusable foundation-model workflow.
A Backbone Model for Causal Inference
It should be impressed that CausalPFN does not take people out of the loop. Rather, it gives domain experts a different role in the process. They spend less time building a new model for each problem and put more attention on whether causal estimates are credible enough to inform action.
The promise of a causal foundation model is not just better performance on individual tasks. It is a more consistent way to apply causal inference across many decisions.
At TD, that shift is already shaping how causal inference is approached. The longer-term ambition is to bring causal modelling to the kind of scale already emerging in predictive AI. TD’s predictive foundation model, TD AI Prism2, is operating across roughly 60 predictive tasks, demonstrating how a shared backbone can support many different modelling needs. The goal for CausalPFN is to bring that same logic to causal inference: one backbone model capable of supporting causal tasks across the institution.
References
- Müller et al. “Transformers Can Do Bayesian Inference” ICLR 2022. https://openreview.net/forum?id=KSugKcbNf9 ↩︎
- “3 things to know about breakthrough AI model, TD AI Prism” TD Stories. https://stories.td.com/ca/en/article/foundation-ai-model ↩︎