Abstract
Reasoning requires adaptation to novel problem settings under limited data and distribution shift. This work introduces CausalARC: an experimental testbed for AI reasoning in low-data and out-of-distribution regimes, modeled after the Abstraction and Reasoning Corpus (ARC). Each CausalARC reasoning task is sampled from a fully specified causal world model, formally expressed as a structural causal model. Principled data augmentations provide observational, interventional, and counterfactual feedback about the world model in the form of few-shot, in-context learning demonstrations. As a proof-of-concept, we illustrate the use of CausalARC for four language model evaluation settings: (1) abstract reasoning with test-time training, (2) counterfactual reasoning with in-context learning, (3) program synthesis, and (4) causal discovery with logical reasoning.
Abstract (translated)
推理需要在数据有限和分布变化的情况下适应新的问题设置。这项工作介绍了CausalARC:一个实验平台,用于在低数据量和分布偏移环境下测试AI的推理能力,它基于抽象与推理语料库(Abstraction and Reasoning Corpus, ARC)构建。每个CausalARC推理任务都是从一个完全指定的因果世界模型中采样出来的,并且这个模型以结构化因果模型的形式正式表达。原则性的数据增强提供了关于这个世界模型的观察、干预和反事实反馈,形式为少量上下文学习示范。作为概念验证,我们展示了如何在四种语言模型评估场景下使用CausalARC:(1)带有测试时训练的抽象推理;(2)基于上下文学习的反事实推理;(3)程序合成;以及(4)结合逻辑推理进行因果发现。
URL
https://arxiv.org/abs/2509.03636