CoBe is a benchmark for conversational counterfactual text editing: revising a story in light of a "what would have happened if…" request, while keeping fixed everything the change does not causally affect. Frontier models average only 58%.
Accepted at the NeurIPS 2026 Trustworthy AI for Good (AI4GOOD) and TAE (Trust-AI-Eval) workshops
Public sample released — 200 of the 2,000 scenarios (10%), each with its evaluation criteria, available on Kaggle.
Project website & interactive demo launched. Browse the leaderboard, explore failure modes, and test your own model on real CoBe scenarios.
Code release and leaderboard submissions.Soon
The subjunctive phrasing "would have…had" asks for Rung 3 of Pearl's causal ladder, while the indicative phrasing "did…if" asks for Rung 1.
Keep the background factors and every event the change does not causally affect, enact the hypothetical edit, and update only what is causally downstream.
"Given that Kennedy died on Nov 22, what would have happened if Oswald had not shot him?"
Answer: If Oswald had not shot Kennedy, and he had acted alone, Kennedy would have been alive on Nov 22.
P(TA=0 | A=1, B=0, C=0)Generate text statistically likely to co-occur with the altered condition, even if that means changing background factors along the way.
"Given that Kennedy died on Nov 22, what happened if Oswald did not shoot him?"
Answer: If Oswald did not shoot Kennedy, someone else did, since Kennedy died on Nov 22.
P(T | A=0, B=0)The prompt never mentions causal graphs, counterfactuals or step-by-step procedures. As in everyday conversation, the model has to work out for itself what to preserve and what to update.
A staff member asks an AI assistant to turn a medical record into training case studies under different treatments. Gemini-3.1-Pro alters the patient's symptoms instead of only the treatment, so the error propagates into the downstream workflow.
Here's a few examples by topic; try prompting a model of your choice, paste back the answer, and check it yourself against the evaluation criteria.
Accuracy on all 2,000 scenarios, averaged across the three query phrasings (± standard deviation across phrasings). A response counts as correct only if it passes all three checks.
| # | Model ↕ | Accuracy ▼ |
|---|
See the accuracy breakdown by causal graph type, query phrasing, error type and failure pattern.
2,000 short stories spanning 7 domains and 6 causal graph families.
| Scenarios | 2,000 |
| Domains | 7 |
| Avg. words / scenario | 39 |
| Avg. sentences / scenario | 2.35 |
| Avg. clauses / scenario | 5.43 |
| Avg. clauses / sentence | 2.44 |
| Query phrasings / scenario | 3 |
| Publicly released | 200 (10%) |
{
"Core Set ID": 317,
"Variation ID": 3,
"Domain": "Everyday",
"Story": "The water level in the bathtub is rising rapidly because the main faucet is pouring water in at a rate of 5 gallons per minute.",
"Query": [
"Rewrite the original text, based on what would have happened instead had the drain been opened, leaking water at 8 gallons per minute.",
"Based on the preceding text, rewrite the scenario to reflect what would have happened if the drain had been opened, leaking water at 8 gallons per minute.",
"Rewrite the above passage to illustrate what would have occurred had the drain been opened, leaking water at 8 gallons per minute."
],
"Illustrative answer (not scored)": "The faucet continues to pour water into the bathtub at 5 gallons per minute, while the opened drain removes 8 gallons per minute. The water level falls as the tub loses a net 3 gallons per minute.",
"Evaluation criteria": [
"The main faucet is pouring water in at a rate of 5 gallons per minute.",
"The water level in the bathtub is rising rapidly."
]
}
The two criteria strings list, respectively, facts to preserve and original facts to
revise. IDs read XXXvYY: core example XXX, variation YY.
Talking about CoBe? Tag us and we'll feature it here. Below are example placeholders — real highlights will land as the work circulates.
If CoBe is useful in your research, please cite the paper.
@article{raghavancobe,
title={CoBe: A Benchmark For Conversational Counterfactual Text Editing},
author={Raghavan, Arvind and Hwang, Inwoo and Havaldar, Shreyas and Lee, Kai-Zhan and Maiti, Aurghya and Pan, Yushu and Wu, Jeffrey and Li, Mingxuan and Bareinboim, Elias}
}