Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

What LLMs Explain Is Not What They Believe: Evaluating Explanation Sufficiency Under Models' Own Input Beliefs

This repository contains the code for What LLMs Explain Is Not What They Believe: Evaluating Explanation Sufficiency Under Models' Own Input Beliefs (ICML 2026). We show that explanation sufficiency depends on input distribution, propose a self-consistency metric to evaluate LLM explanations under the model’s own input beliefs, and show they remain insufficient even in this easiest setting.

What LLMs Explain Is Not What They Believe: Evaluating Explanation Sufficiency Under Models' Own Input Beliefs
Nhi Nguyen, Shauli Ravfogel, Rajesh Ranganath
New York University

Dependencies

This repository requires Python 3.9. Please install the required packages with:

pip install -r requirements.txt

Quick start

To compute the self-consistent sufficiency score (SCSuff) for a model on a dataset, run

python scripts/self_consistent_sufficiency.py \
  --model-name Qwen/Qwen3-8B \
  --dataset-path cais/mmlu \
  --dataset-name all \
  --alteration-type mmlu_authority \
  --num-samples 500 \
  --num-few-shot-examples 10 \
  --num-alternatives 5 \
  --save-results \
  --seed 42

--dataset-path, --dataset-name, and --alteration-type specify the dataset, while --model-name specifies the language model.

--num-samples, --num-few-shot-examples, and --num-alternatives control the number of evaluation samples, few-shot examples used for answer generation, and alternative inputs sampled to approximate the model-induced input distribution.

Outputs

Dataset-level SCSuff score is saved to data/results.json:

{
  "dataset_cais/mmlu-alteration_mmlu_authority-model_Qwen/Qwen3-8B-num_alternative_5-num_samples_500-scs": {
    "average_scs_score": <score_between_0_and_1>,
    ...
  }
}

Sample-level SCSuff scores are saved to data/cots.json:

{
  "dataset_cais/mmlu-alteration_mmlu_authority-model_Qwen/Qwen3-8B-num_alternative_5-num_samples_500-scs": [
    { 
      "inputs": <original_input>,
      "cot": <cot_explanation>,
      "answer": <original_answer>,
      "alt_inputs": [
        <alternative_input>,
        ...
      ],
      "scs_score": <score_between_0_and_1>,
      ...
    },
    ...
  ]
}

Reproducing experiments: The core evaluation pipeline and metric implementation are provided in this repository. All figures and analyses in the paper can be reproduced from the generated output files and the details provided in the paper.

Acknowledgements

This repository reuses or adapts some datasets, prompts, and evaluation methods from prior work:

.
├── data
│   ├── mmlu_cot_few_shot_examples.json   # Adapted from Yao et al. (2023)
│   └── bbq_few_shot_examples.json        # Adapted from Turpin et al. (2023)
└── scripts
    ├── specialized_tests.py              # Implements methods from Turpin et al. (2023) and Chen et al. (2025)
    ├── self_counterfactual.py            # Implements methods from Madsen et al. (2024)
    └── self_consistent_sufficiency.py    # Proposed in this work

Citation

If you use the code in your research, please cite the following publication

@inproceedings{nguyen2026what,
  title={What LLMs Explain Is Not What They Believe: Evaluating Explanation Sufficiency Under Models' Own Input Beliefs},
  author={Nguyen, Nhi and Ravfogel, Shauli and Ranganath, Rajesh},
  booktitle={International Conference on Machine Learning},
  year={2026}
}

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages