← Back to feed
2026-07-24agentsreasoningalignmentcode

CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

Jiyuan Tan, Vasilis Syrgkanis

PDF preview for CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference
Read on arXiv →

Key claim

CausalForge enhances reliability in automated theoretical research.

In plain English

Imagine you're a researcher trying to automate the process of developing and verifying new theories in causal inference. Currently, many rely on large language models to review research, but these models often struggle with reliability, sometimes accepting fabricated results or failing to detect errors effectively. This unreliability is a significant hurdle, as it can lead to the acceptance of incorrect scientific claims, a problem known as 'Bad Scientist.' To address this, a new framework called CausalForge has been developed, which integrates a foundational library for causal inference with a self-improving pipeline that not only proposes research topics and formalizes statements but also rigorously checks the accuracy of these statements against their intended scientific claims. This dual approach enhances the verification process by ensuring that formal proofs align with the original research intent, thereby improving the overall reliability of the automated research output. Compared to previous methods that relied heavily on LLM reviewers, CausalForge offers a more structured and reliable way to conduct theoretical research, making it a valuable tool for builders in the field of causal inference.

Novelty
8.5/10

CausalForge introduces a novel framework for automating theoretical research in causal inference.

Reliability
7.5/10

The evaluation of the system is based on artifacts from autonomous research runs, providing a solid basis for reliability.

Deep reliability assessment

The methodology supports the generation and formal verification of causal inference research artifacts, but it may overclaim by implying that machine-checked proofs fully capture the intended scientific claims without human oversight.

Reproducibility

yes, the source code, formal library, and run records are available at the provided GitHub URL.

Key figure

The key architectural diagram likely illustrates the interaction between Causalean and CausalSmith within the CausalForge framework.

GitHub1 repo
Jiyuan-Tan/CausalForgeOfficial