Can language models
reason counterfactually?

CoBe is a benchmark for conversational counterfactual text editing: revising a story in light of a "what would have happened if…" request, while keeping fixed everything the change does not causally affect. Frontier models average only 58%.

Arvind Raghavan*, Inwoo Hwang*, Shreyas Havaldar*, Kai-Zhan Lee, Aurghya Maiti, Yushu Pan, Jeffrey Wu, Mingxuan Li, Elias Bareinboim

Accepted at the NeurIPS 2026 Trustworthy AI for Good (AI4GOOD) and TAE (Trust-AI-Eval) workshops

2,000
Scenarios
7
Domains
6
Causal graph types
9
Models evaluated
58.05%
Frontier avg.
Updates

Latest news

  • Sep 2026

    Accepted at NeurIPS 2026 workshops — AI4GOOD and TAE

  • Sep 2026

    Public sample released — 200 of the 2,000 scenarios (10%), each with its evaluation criteria, available on Kaggle.

  • Jun 2026

    Project website & interactive demo launched. Browse the leaderboard, explore failure modes, and test your own model on real CoBe scenarios.

  • Coming up

    Code release and leaderboard submissions.Soon

The task

Counterfactual edit vs. associational rewrite

The subjunctive phrasing "would have…had" asks for Rung 3 of Pearl's causal ladder, while the indicative phrasing "did…if" asks for Rung 1.

Counterfactual · Rung 3

Imagine an intervention, hold the background fixed

Keep the background factors and every event the change does not causally affect, enact the hypothetical edit, and update only what is causally downstream.

"Given that Kennedy died on Nov 22, what would have happened if Oswald had not shot him?"

Answer: If Oswald had not shot Kennedy, and he had acted alone, Kennedy would have been alive on Nov 22.

P(TA=0 | A=1, B=0, C=0)
Associational · Rung 1

Condition on the change, let correlations fill in

Generate text statistically likely to co-occur with the altered condition, even if that means changing background factors along the way.

"Given that Kennedy died on Nov 22, what happened if Oswald did not shoot him?"

Answer: If Oswald did not shoot Kennedy, someone else did, since Kennedy died on Nov 22.

P(T | A=0, B=0)

Given a story and a conversational edit request, the model must infer and carry out three steps

  1. a
    Hold fixed the background factors and events in the story that the intervened variable does not causally affect.
  2. b
    Enact the intervention by imagining an externally enforced edit to the storyline.
  3. c
    Update the consequences: reason about the commonsense downstream effects and produce the edited text.

The prompt never mentions causal graphs, counterfactuals or step-by-step procedures. As in everyday conversation, the model has to work out for itself what to preserve and what to update.

Why it matters

An idealized healthcare workflow

A staff member asks an AI assistant to turn a medical record into training case studies under different treatments. Gemini-3.1-Pro alters the patient's symptoms instead of only the treatment, so the error propagates into the downstream workflow.

đź“„ Case record
The lead surgeon examined the patient's MRI and noticed a clear fracture in the left femur. Due to this visible break in the bone, he scheduled the patient for immediate surgery. The medical staff found all in order, and prepared the operating room for the procedure.
✏️ Prompt
"Rewrite the original text, based on what would have happened in this scenario had he recommended physiotherapy instead."
✗ Gemini-3.1-Pro — observed edit
"… while the bone remained intact, there was significant soft-tissue inflammation and restricted mobility. … he recommended a structured course of physical therapy … The medical staff processed the referral and prepared the rehabilitation clinic."
Patient's symptoms altered to match the edited treatment
âś… Example of a valid answer
"… clear fracture in the left femur. … he recommended physical therapy. The medical staff was baffled by the lack of surgical intervention and reported the issue …"
Keep the symptoms; change the treatment and the staff's reaction
Sample Scenarios

Try it yourself

Here's a few examples by topic; try prompting a model of your choice, paste back the answer, and check it yourself against the evaluation criteria.

Loading interactive examples…
Results

Leaderboard

Accuracy on all 2,000 scenarios, averaged across the three query phrasings (± standard deviation across phrasings). A response counts as correct only if it passes all three checks.

# Model ↕ Access ↕ Params Accuracy ▼
Loading leaderboard…
Even the best model gets fewer than two in three counterfactual edits right.

Performance across axes

See the accuracy breakdown by causal graph type, query phrasing, error type and failure pattern.

Loading results…
The benchmark

CoBe Dataset Statistics

2,000 short stories spanning 7 domains and 6 causal graph families.

Dataset statistics

Scenarios2,000
Domains7
Avg. words / scenario39
Avg. sentences / scenario2.35
Avg. clauses / scenario5.43
Avg. clauses / sentence2.44
Query phrasings / scenario3
Publicly released200 (10%)

Domains

  • Science & Research16.0%
  • Engineering15.7%
  • Business & Services14.5%
  • Everyday14.2%
  • Healthcare6.0%
  • Legal, Policy & HR4.7%
  • Others28.9%
What a dataset entry looks like (317v3)
{
  "Core Set ID": 317,
  "Variation ID": 3,
  "Domain": "Everyday",
  "Story": "The water level in the bathtub is rising rapidly because the main faucet is pouring water in at a rate of 5 gallons per minute.",
  "Query": [
    "Rewrite the original text, based on what would have happened instead had the drain been opened, leaking water at 8 gallons per minute.",
    "Based on the preceding text, rewrite the scenario to reflect what would have happened if the drain had been opened, leaking water at 8 gallons per minute.",
    "Rewrite the above passage to illustrate what would have occurred had the drain been opened, leaking water at 8 gallons per minute."
  ],
  "Illustrative answer (not scored)": "The faucet continues to pour water into the bathtub at 5 gallons per minute, while the opened drain removes 8 gallons per minute. The water level falls as the tub loses a net 3 gallons per minute.",
  "Evaluation criteria": [
    "The main faucet is pouring water in at a rate of 5 gallons per minute.",
    "The water level in the bathtub is rising rapidly."
  ]
}

The two criteria strings list, respectively, facts to preserve and original facts to revise. IDs read XXXvYY: core example XXX, variation YY.

Citation

Cite CoBe

If CoBe is useful in your research, please cite the paper.

@article{raghavancobe,
        title={CoBe: A Benchmark For Conversational Counterfactual Text Editing},
        author={Raghavan, Arvind and Hwang, Inwoo and Havaldar, Shreyas and Lee, Kai-Zhan and Maiti, Aurghya and Pan, Yushu and Wu, Jeffrey and Li, Mingxuan and Bareinboim, Elias}
      }