Benchmarking Causal Reasoning in Foundation Models
CausalArena introduces a four-regime testing protocol to evaluate whether models perform causal discovery based on structural logic or mere pattern memorization.
CausalArena evaluates whether models possess generalizable causal reasoning by testing them across four distinct regimes: synthetic, semantic, formula-grounded, and real-world structural causal models. Existing benchmarks often rely on static datasets that invite memorization. It becomes impossible to separate a model's grasp of a causal principle from its ability to recognize a familiar graph topology. The framework uses diverse structural causal models to ensure the evaluation covers structures unlikely to be represented in pretraining data.
Isolating Reasoning from Retrieval
The evaluation process treats a model as a discovery engine. For each instance in the benchmark, the model is provided with a dataset sampled from a specific structure. The model must perform structure learning by outputting an adjacency matrix or a structured description of the causal graph. This requires the model to infer the causal dependency directly from the statistical associations present in the provided samples. If the model simply relies on a lookup of pre-learned edge connectivity, it will fail to recover the graph topology when presented with new, synthetic causal families.
| Evaluation Regime | Primary Focus | Evaluated Capability |
|---|---|---|
| Synthetic SCMs | Structural diversity | General graph discovery |
| Semantic SCMs | Grounded environments | Human-auditable reasoning |
| Formula-grounded | Scientific mechanisms | Law-based discovery |
| Real-world Data | External validity | Handling observational noise |
Performance Gaps in Formal Logic
Current foundation models underperform in formula-grounded regimes compared to their synthetic graph results. In a formula-grounded test, the relationship between variables is defined by a specific mathematical function such as Y = f(X) + error. To succeed, the model must identify that X influences Y while accounting for the specific mathematical constraint governing that influence. The observed failure here is that models often capture the existence of a correlation but fail to respect the logic of the underlying formula when predicting outcomes under new conditions.
Consider a simple case of a structural model where a variable X directly influences Y through a linear function. A model might correctly predict the correlation between the two during initial observations. When the task shifts to an interventional prediction—where the value of X is forced to a specific level outside the original observation range—the model must apply the underlying functional constraint to predict Y. The benchmark shows that when models are forced to operate within these formal constraints rather than relying on familiar statistical patterns, their performance degrades. This suggests that the current gap is rooted in how models process formal constraints versus simple observational data.
It is not yet known how much of this performance gap is due to the inherent constraints of pretraining on observational data versus the limitations of current prompting strategies for structure learning. The benchmark demonstrates that high performance on traditional causal reasoning tasks does not guarantee success when the causal graph or the underlying functional mechanism is novel. Future investigations will likely focus on whether specialized fine-tuning or extended reasoning chains can bridge the gap between pattern matching and formal causal discovery.