<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://stanichor.net/feed.xml" rel="self" type="application/atom+xml" /><link href="https://stanichor.net/" rel="alternate" type="text/html" /><updated>2026-10-07T03:03:22+00:00</updated><id>https://stanichor.net/feed.xml</id><title type="html">Stanichor</title><subtitle>This is the website of Stanichor. I write.</subtitle><entry><title type="html">Revisiting the 2012 LessWrong Personality Results</title><link href="https://stanichor.net/rat-personality/" rel="alternate" type="text/html" title="Revisiting the 2012 LessWrong Personality Results" /><published>2026-10-04T00:00:00+00:00</published><updated>2026-10-04T00:00:00+00:00</updated><id>https://stanichor.net/rat-personality</id><content type="html" xml:base="https://stanichor.net/rat-personality/"><![CDATA[<p>If you ask an LLM what the personality profile of rationalists is, they’ll probably direct you to the <a href="https://www.lesswrong.com/posts/x9FNKTEt68Rz6wQ6P/2012-survey-results">2012 LW Survey results</a>, which state:</p>

<blockquote>
  <p>mean+standard_deviation (25% level, 50% level/median, 75% level) [n = number of data points]<br />
[…]<br />
Big 5 (O): 60.6 + 25.7 (41, 65, 84) [n = 453]<br />
Big 5 (C): 35.2 + 27.5 (10, 30, 58) [n = 453]<br />
Big 5 (E): 30.3 + 26.7 (7, 22, 48) [n = 454]<br />
Big 5 (A): 41 + 28.3 (17, 38, 63) [n = 453]<br />
Big 5 (N): 36.6 + 29 (11, 27, 60) [n = 449]</p>
</blockquote>

<p>So, it looks like rats are more Open than average, and less Conscientious, Extraverted, Agreeable, and Neurotic than average, but even then, the differences aren’t that large. If we convert the median percentiles to z-scores, we get:</p>

<table>
  <thead>
    <tr>
      <th>Trait</th>
      <th style="text-align: right">Median percentile</th>
      <th style="text-align: right">Corresponding z-score</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Openness</td>
      <td style="text-align: right">65th</td>
      <td style="text-align: right">+0.39</td>
    </tr>
    <tr>
      <td>Conscientiousness</td>
      <td style="text-align: right">30th</td>
      <td style="text-align: right">−0.52</td>
    </tr>
    <tr>
      <td>Extraversion</td>
      <td style="text-align: right">22nd</td>
      <td style="text-align: right">−0.77</td>
    </tr>
    <tr>
      <td>Agreeableness</td>
      <td style="text-align: right">38th</td>
      <td style="text-align: right">−0.31</td>
    </tr>
    <tr>
      <td>Neuroticism</td>
      <td style="text-align: right">27th</td>
      <td style="text-align: right">−0.61</td>
    </tr>
  </tbody>
</table>

<p>So, it doesn’t seem like rats are <em>that</em> different from the general population in personality. Except, I don’t buy that. Rats are <em>weird</em>. They talk about weird things, do weird things, and take weird ideas seriously. So, for a trait like Openness, I’d expect rats to be much higher than just 0.39 SD above the population mean. There must be something going on.</p>

<p>The problem is the online personality test’s norms. The <a href="https://docs.google.com/forms/d/e/1FAIpQLSdGPNzS7f25N2xh0HA9e8L41qW7EnR8TK67KKje3S95U6SJdQ/viewform">2012 survey</a> instructed respondents to take the <a href="https://www.outofservice.com/bigfive/">OutOfService BFI-44 test</a><sup id="fnref:out-of-service" role="doc-noteref"><a href="#fn:out-of-service" class="footnote" rel="footnote">1</a></sup> and enter the percentiles it gave them. If the test’s norms come from people who take personality tests online rather than the general population, those percentiles could make rats look more ordinary than they are. To check, we’d want to compare LessWrong’s raw scores with those of a more representative sample. Unfortunately, the survey only collected percentiles. Foretunately, <a href="https://www.greaterwrong.com/posts/bJiyYJeCyh4HcKHub/2012-less-wrong-census-survey/comment/yZzME5X5QgZPLzz7W">VincentYu found</a> the means and SDs the test used which allows me to recover most of the trait scores in the public survey data.<sup id="fnref:score-reconstruction" role="doc-noteref"><a href="#fn:score-reconstruction" class="footnote" rel="footnote">2</a></sup></p>

<p>For my representative sample, I’ll use <a href="https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2014.00370/full">Bogg &amp; Vo (2014)</a>. They administered the BFI-44 to a weighted, probability-based U.S. sample (N = 1,015). So, how do LessWrong’s raw scores compare with the test’s assumed means and the U.S. sample’s means? Here they are:</p>

<div style="text-align: center;">
    <figure>
        <img src="/assets/images/rat-personality/lesswrong-2012-big-five-raw-comparison.png" alt="Mean Big Five raw scores for LessWrong, the online test's assumed distribution, and a representative U.S. sample" width="1000" />
    </figure>
</div>

<p>As you can see, the online test assumed higher average Openness and lower average Conscientiousness and Agreeableness than the U.S. sample, so it made LessWrong’s differences on those traits look smaller. And while the online test made rats look less Neurotic than average, their mean is almost exactly the U.S. mean.</p>

<p>So, using the U.S. means and SDs, the median LessWrong score on each trait comes out to roughly:<sup id="fnref:normal-percentiles" role="doc-noteref"><a href="#fn:normal-percentiles" class="footnote" rel="footnote">3</a></sup></p>

<ul>
  <li>88th percentile (+1.2 SD) on Openness</li>
  <li>10th percentile (−1.3 SD) on Conscientiousness</li>
  <li>27th percentile (−0.6 SD) on Extraversion</li>
  <li>28th percentile (−0.6 SD) on Agreeableness</li>
  <li>50th percentile (0 SD) on Neuroticism</li>
</ul>

<div style="text-align: center;">
    <figure>
        <img src="/assets/images/rat-personality/lesswrong-2012-big-five-us-z-histograms-normal-overlay-1.png" alt="Histograms of reconstructed LessWrong Big Five scores standardized to U.S. norms, with median and interquartile range marked for each trait" width="1000" />
    </figure>
</div>

<p>The revised scores now match my intuitive sense of LessWrong users. LessWrong covers ideas outside the norm, and the Openness score now reflects that. Additionally, many users reported suffering from <a href="https://www.lesswrong.com/w/akrasia">akrasia</a>, which is also now reflected in the revised Conscientiousness score.</p>

<p>What lesson should we draw from this? Perhaps, when an online test gives you a percentile, ask: compared with whom? If the comparison group is unusual (which it usually is), that number will give you a misleading picture of where you stand in relation to the the general population.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:out-of-service" role="doc-endnote">
      <p>Yes, that’s the name of the website. Interestingly enough, it’s still <em>in</em> service years later. <a href="#fnref:out-of-service" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:score-reconstruction" role="doc-endnote">
      <p>VincentYu reported the test’s assumed means (SDs) on the 1–5 scale: O 3.85 (0.65), C 3.40 (0.76), E 3.30 (0.88), A 3.66 (0.70), and N 3.15 (0.85). I inverted the Gaussian percentile mapping over the BFI’s possible scores, allowing for rounding in the published parameters. Of 2,112 nonmissing trait responses, 1,982 (94%) had a unique match; 95 were ambiguous and 35 unmatched. The charts omit the latter two groups. <a href="#fnref:score-reconstruction" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:normal-percentiles" role="doc-endnote">
      <p>Each z-score is the median reconstructed LessWrong raw score minus the U.S. mean, divided by the U.S. SD. The percentiles assume a normal U.S. score distribution; they are not empirical percentiles from Bogg &amp; Vo’s sample. The opening table uses the survey’s reported median percentiles, while these medians use only scores that could be uniquely reconstructed from the public data. <a href="#fnref:normal-percentiles" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name></name></author><category term="rationalism-ea" /><summary type="html"><![CDATA[If you ask an LLM what the personality profile of rationalists is, they’ll probably direct you to the 2012 LW Survey results, which state:]]></summary></entry><entry><title type="html">LiveBench Mostly Measures One General Ability</title><link href="https://stanichor.net/livebench-factors/" rel="alternate" type="text/html" title="LiveBench Mostly Measures One General Ability" /><published>2026-10-03T00:00:00+00:00</published><updated>2026-10-03T00:00:00+00:00</updated><id>https://stanichor.net/livebench-factors</id><content type="html" xml:base="https://stanichor.net/livebench-factors/"><![CDATA[<p><a href="https://livebench.ai/#/">LiveBench</a> is an LLM benchmark with tasks grouped into categories. I’ll analyze its April 7, 2025 public release, which included 18 tasks across six categories. Item-level data were available for seven tasks across three categories: Coding (LCB Generation and Coding Completion), Language (Connections, Plot Unscrambling, and Typos), and Instruction Following (Paraphrase and Story Generation). Unfortunately, I can’t analyze the other tasks because LiveBench didn’t release their item-level data, which is essential for this kind of analysis.</p>

<p>Here’s what each task involves:</p>

<style>
  .display-note {
    margin: 0.35rem 0 0.9rem !important;
    color: #667085;
    font-size: 0.82em;
    line-height: 1.45;
  }
  .task-description { margin: 0 !important; border-top: 1px solid #d0d7de; }
  .task-description:last-of-type {
    border-bottom: 1px solid #d0d7de;
    margin-bottom: 1.5rem !important;
  }
  .task-description summary {
    display: flex;
    align-items: center;
    justify-content: space-between;
    padding: 0.8rem 0.25rem;
    cursor: pointer;
    list-style: none;
  }
  .task-description summary::-webkit-details-marker { display: none; }
  .task-description summary h3 { margin: 0 !important; font-size: 1.25em; }
  .task-description summary::after {
    content: "+";
    margin-left: 1rem;
    color: #57606a;
    font-size: 1.35rem;
    font-weight: 400;
    line-height: 1;
  }
  .task-description[open] summary::after { content: "-"; }
  .task-description > :not(summary) { margin-right: 1rem; margin-left: 1rem; }
  .task-description > :last-child { margin-bottom: 1.5rem; }
</style>

<details class="task-description">
  <summary><h3 id="lcb-generation">LCB Generation (Category: Coding)</h3></summary>

  <p>The model receives a programming problem (typically from LiveCodeBench) and must write a complete solution which is then run against test cases. The model receives a score of <code class="language-plaintext highlighter-rouge">1</code> if it passes and <code class="language-plaintext highlighter-rouge">0</code> if it fails; there is no partial credit.</p>

</details>

<details class="task-description">
  <summary><h3 id="coding-completion">Coding Completion (Category: Coding)</h3></summary>

  <p>The model receives a programming problem and a fragment of a correct solution, which it must complete. Scoring is the same as for LCB Generation.</p>

</details>

<details class="task-description">
  <summary><h3 id="connections">Connections (Category: Language)</h3></summary>

  <p>The Connections task works much like the NYT game of the same name. The model receives a shuffled list of words and must sort them into groups of four, with each group sharing a theme. For example:</p>

  <ul>
    <li><code class="language-plaintext highlighter-rouge">bass, cod, salmon, trout</code> → fish</li>
    <li><code class="language-plaintext highlighter-rouge">apple, banana, pear, peach</code> → fruit</li>
  </ul>

  <p>LiveBench uses 8 words (2 groups), 12 words (3 groups), or 16 words (4 groups). The score is the fraction of complete groups identified correctly.</p>

</details>

<details class="task-description">
  <summary><h3 id="plot-unscrambling">Plot Unscrambling (Category: Language)</h3></summary>

  <p>The model receives the sentences of a recent movie synopsis in random order and must reconstruct the original narrative. The evaluator fuzzy-matches the response to the original sentences, allowing minor transcription changes, then measures the edit distance between the proposed and correct orders:</p>

\[\text{score} = 1 - \frac{d}{n}\]

  <p>where $d$ is the ordering distance and $n$ is the number of sentences.</p>

</details>

<details class="task-description">
  <summary><h3 id="typos">Typos (Category: Language)</h3></summary>

  <p>The model receives text (usually based on a recent arXiv abstract) with synthetic spelling errors inserted. It must correct the misspellings while leaving everything else unchanged. That means it shouldn’t rewrite the text, change punctuation, change US spelling to UK spelling or vice versa, or add stylistic “improvements”. The scorer gives <code class="language-plaintext highlighter-rouge">1</code> if the ground-truth text appears anywhere in the output and <code class="language-plaintext highlighter-rouge">0</code> otherwise. Thus, despite the instruction, extra surrounding text does not necessarily cause a failure.</p>

</details>

<details class="task-description">
  <summary><h3 id="paraphrase">Paraphrase (Category: Instruction Following)</h3></summary>

  <p>The model receives the beginning of a recent Guardian article and is asked to paraphrase it while following several mechanically verifiable instructions. For example:</p>

  <blockquote>
    <p>Paraphrase this article.
Include a title. Use the words “course,” “media,” and “sun.” Write exactly three paragraphs. Begin the first paragraph with “hand.”</p>
  </blockquote>

  <p>LiveBench scores compliance with the explicit instructions, not the quality of the paraphrase. In fact, it doesn’t care about the actual paraphrase at all; models can receive full credit without ever attempting to paraphrase the article. It averages two components:</p>

  <ul>
    <li>Prompt-level accuracy: <code class="language-plaintext highlighter-rouge">1</code> only if every instruction was followed; otherwise <code class="language-plaintext highlighter-rouge">0</code>.</li>
    <li>Instruction-level accuracy: the fraction of individual instructions followed.</li>
  </ul>

</details>

<details class="task-description">
  <summary><h3 id="story-generation">Story Generation (Category: Instruction Following)</h3></summary>

  <p>The model receives a recent news article and is asked to generate a story based on it while following mechanically verifiable instructions. LiveBench uses the same two-component scoring method as Paraphrase.</p>

</details>

<p>The seven tasks report different kinds of item scores. Here is how LiveBench’s published scores relate to the responses I use in this analysis:</p>

<table>
  <thead>
    <tr>
      <th>Task</th>
      <th>LiveBench’s published item score</th>
      <th>Response used in this analysis</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>LCB Generation</td>
      <td>Pass/fail: 0 or 1</td>
      <td>Binary, unchanged</td>
    </tr>
    <tr>
      <td>Coding Completion</td>
      <td>Pass/fail: 0 or 1</td>
      <td>Binary, unchanged</td>
    </tr>
    <tr>
      <td>Connections</td>
      <td>Fraction of complete four-word groups identified correctly; an item has two, three, or four groups</td>
      <td>Ordered partial credit, unchanged</td>
    </tr>
    <tr>
      <td>Typos</td>
      <td>Exact-match pass/fail: 0 or 1</td>
      <td>Binary, unchanged</td>
    </tr>
    <tr>
      <td>Plot Unscrambling</td>
      <td>One minus ordering edit distance divided by the number of sentences; bounded between 0 and 1</td>
      <td>Count-adjusted logit of the score, treated as continuous</td>
    </tr>
    <tr>
      <td>Paraphrase</td>
      <td>Average of all-instructions-correct accuracy and the fraction of individual instructions followed</td>
      <td>Instruction-level fraction only, modeled as ordered partial credit</td>
    </tr>
    <tr>
      <td>Story Generation</td>
      <td>Same two-component score as Paraphrase</td>
      <td>Instruction-level fraction only, modeled as ordered partial credit</td>
    </tr>
  </tbody>
</table>

<p>I don’t like how Paraphrase and Story Generation are currently graded. Their published scores average instruction-level accuracy with an all-or-nothing prompt-level component, so missing just one instruction costs more than half the grade. I therefore use instruction-level accuracy alone.</p>

<p>For Plot Unscrambling, I logit-transform the score using this count-adjusted formula:</p>

\[\text{transformed score} = \log\left(\frac{n - D + \frac{1}{2}}{D + \frac{1}{2}}\right)\]

<p>where $D$ is the edit distance and $n$ is the number of sentences.</p>

<h2 id="determining-the-number-of-factors">Determining the Number of Factors</h2>

<p>Naturally, I used parallel analysis to determine the number of factors. It yielded 19, which is a lot (see the <a href="#classical-parallel-analysis">plot</a>).</p>

<p>However, some item pairs have no models in common, and others have only a few, making their correlations unavailable or imprecise. So I modeled the correlation matrix Bayesianly and repeated parallel analysis across posterior draws to see how much the recommended factor count varies. Each of the first 12 factors exceeds the chance threshold in at least 95% of posterior draws, and the 90% interval for the number retained is 12–13.</p>

<div style="text-align: center;">
    <figure>
        <img src="/assets/images/livebench-factor-analysis/bayesian_mixed_correlation_parallel_analysis.png" width="1000" />
    </figure>
</div>

<p>Even 12–13 factors is a lot for seven tasks. I would have expected something closer to seven, so let’s look at the tasks individually to see where the extra dimensions might be coming from.</p>

<p class="display-note">Select a task to see its Bayesian parallel analysis. The badge counts leading factors above the chance threshold in at least 95% of posterior draws.</p>

<style>
  .task-parallel-analysis { margin: 0 !important; border-top: 1px solid #d0d7de; }
  .task-parallel-analysis:last-of-type {
    border-bottom: 1px solid #d0d7de;
    margin-bottom: 1.5rem !important;
  }
  .task-parallel-analysis summary {
    display: flex;
    align-items: center;
    gap: 0.75rem;
    padding: 0.75rem 0.25rem;
    cursor: pointer;
    list-style: none;
  }
  .task-parallel-analysis summary::-webkit-details-marker { display: none; }
  .task-parallel-analysis summary .task-name { font-weight: 600; }
  .task-parallel-analysis summary .factor-count {
    padding: 0.1rem 0.55rem;
    border-radius: 1rem;
    background: #eef2f6;
    color: #465568;
    font-size: 0.78em;
    white-space: nowrap;
  }
  .task-parallel-analysis summary::after {
    content: "+";
    margin-left: auto;
    color: #57606a;
    font-size: 1.35rem;
    line-height: 1;
  }
  .task-parallel-analysis[open] summary::after { content: "−"; }
  .task-parallel-analysis figure { margin: 0.25rem 0.25rem 1.25rem; }
  .task-parallel-analysis img { display: block; width: 100%; height: auto; }
</style>

<details class="task-parallel-analysis" name="task-parallel-analysis">
<summary><span class="task-name">LCB Generation</span><span class="factor-count">2 factors</span></summary>
<figure><img src="/assets/images/livebench-factor-analysis/task-parallel-analysis/LCB_generation/bayesian_mixed_correlation_parallel_analysis.png" alt="Bayesian parallel analysis for LCB Generation" loading="lazy" /></figure>
</details>

<details class="task-parallel-analysis" name="task-parallel-analysis">
<summary><span class="task-name">Coding Completion</span><span class="factor-count">2 factors</span></summary>
<figure><img src="/assets/images/livebench-factor-analysis/task-parallel-analysis/coding_completion/bayesian_mixed_correlation_parallel_analysis.png" alt="Bayesian parallel analysis for Coding Completion" loading="lazy" /></figure>
</details>

<details class="task-parallel-analysis" name="task-parallel-analysis">
<summary><span class="task-name">Connections</span><span class="factor-count">3 factors</span></summary>
<figure><img src="/assets/images/livebench-factor-analysis/task-parallel-analysis/connections/bayesian_mixed_correlation_parallel_analysis.png" alt="Bayesian parallel analysis for Connections" loading="lazy" /></figure>
</details>

<details class="task-parallel-analysis" name="task-parallel-analysis">
<summary><span class="task-name">Plot Unscrambling</span><span class="factor-count">2 factors</span></summary>
<figure><img src="/assets/images/livebench-factor-analysis/task-parallel-analysis/plot_unscrambling/bayesian_mixed_correlation_parallel_analysis.png" alt="Bayesian parallel analysis for Plot Unscrambling" loading="lazy" /></figure>
</details>

<details class="task-parallel-analysis" name="task-parallel-analysis">
<summary><span class="task-name">Typos</span><span class="factor-count">7 factors</span></summary>
<figure><img src="/assets/images/livebench-factor-analysis/task-parallel-analysis/typos/bayesian_mixed_correlation_parallel_analysis.png" alt="Bayesian parallel analysis for Typos" loading="lazy" /></figure>
</details>

<details class="task-parallel-analysis" name="task-parallel-analysis">
<summary><span class="task-name">Paraphrase</span><span class="factor-count">2 factors</span></summary>
<figure><img src="/assets/images/livebench-factor-analysis/task-parallel-analysis/paraphrase/bayesian_mixed_correlation_parallel_analysis.png" alt="Bayesian parallel analysis for Paraphrase" loading="lazy" /></figure>
</details>

<details class="task-parallel-analysis" name="task-parallel-analysis">
<summary><span class="task-name">Story Generation</span><span class="factor-count">2 factors</span></summary>
<figure><img src="/assets/images/livebench-factor-analysis/task-parallel-analysis/story_generation/bayesian_mixed_correlation_parallel_analysis.png" alt="Bayesian parallel analysis for Story Generation" loading="lazy" /></figure>
</details>

<p>Although the Bayesian parallel analyses support more than one factor for every task, the first dimension dominates in six of them. The ratio of the first two (posterior-median) eigenvalues ranges from 3.1 to 10.6 for those tasks, compared with just 1.5 for Typos. The within-task item-correlation matrices offer another way to see this:</p>

<p class="display-note">Items are ordered by median task-factor loading. Grey means unavailable, not zero. Scroll to compare tasks; select a matrix for full size.</p>

<style>
  .task-matrix-gallery {
    display: flex;
    gap: 1rem;
    overflow-x: auto;
    scroll-snap-type: x proximity;
    -webkit-overflow-scrolling: touch;
    padding: 0.25rem 0.25rem 1rem;
    margin: 0.75rem 0 1.5rem;
  }
  .task-matrix-gallery .task-matrix-card {
    flex: 0 0 min(78vw, 24rem);
    scroll-snap-align: start;
    margin: 0 !important;
    border: 1px solid #d0d7de;
    border-radius: 0.5rem;
    overflow: hidden;
    background: #fff;
  }
  .task-matrix-card a { display: block; }
  .task-matrix-card img { display: block; width: 100%; height: auto; }
  .task-matrix-card figcaption {
    padding: 0.6rem 0.8rem;
    border-top: 1px solid #d0d7de;
    font-weight: 600;
  }
</style>

<div class="task-matrix-gallery" role="region" aria-label="Task item-correlation matrices; scroll horizontally" tabindex="0">
  <figure class="task-matrix-card">
    <a href="/assets/images/livebench-factor-analysis/task-correlation-matrices/typos/item_correlation_matrix.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/task-correlation-matrices/typos/item_correlation_matrix.png" alt="Typos item-correlation matrix, ordered by task loading" loading="lazy" /></a>
    <figcaption>Typos</figcaption>
  </figure>
  <figure class="task-matrix-card">
    <a href="/assets/images/livebench-factor-analysis/task-correlation-matrices/LCB_generation/item_correlation_matrix.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/task-correlation-matrices/LCB_generation/item_correlation_matrix.png" alt="LCB Generation item-correlation matrix, ordered by task loading" loading="lazy" /></a>
    <figcaption>LCB Generation</figcaption>
  </figure>
  <figure class="task-matrix-card">
    <a href="/assets/images/livebench-factor-analysis/task-correlation-matrices/coding_completion/item_correlation_matrix.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/task-correlation-matrices/coding_completion/item_correlation_matrix.png" alt="Coding Completion item-correlation matrix, ordered by task loading" loading="lazy" /></a>
    <figcaption>Coding Completion</figcaption>
  </figure>
  <figure class="task-matrix-card">
    <a href="/assets/images/livebench-factor-analysis/task-correlation-matrices/connections/item_correlation_matrix.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/task-correlation-matrices/connections/item_correlation_matrix.png" alt="Connections item-correlation matrix, ordered by task loading" loading="lazy" /></a>
    <figcaption>Connections</figcaption>
  </figure>
  <figure class="task-matrix-card">
    <a href="/assets/images/livebench-factor-analysis/task-correlation-matrices/plot_unscrambling/item_correlation_matrix.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/task-correlation-matrices/plot_unscrambling/item_correlation_matrix.png" alt="Plot Unscrambling item-correlation matrix, ordered by task loading" loading="lazy" /></a>
    <figcaption>Plot Unscrambling</figcaption>
  </figure>
  <figure class="task-matrix-card">
    <a href="/assets/images/livebench-factor-analysis/task-correlation-matrices/paraphrase/item_correlation_matrix.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/task-correlation-matrices/paraphrase/item_correlation_matrix.png" alt="Paraphrase item-correlation matrix, ordered by task loading" loading="lazy" /></a>
    <figcaption>Paraphrase</figcaption>
  </figure>
  <figure class="task-matrix-card">
    <a href="/assets/images/livebench-factor-analysis/task-correlation-matrices/story_generation/item_correlation_matrix.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/task-correlation-matrices/story_generation/item_correlation_matrix.png" alt="Story Generation item-correlation matrix, ordered by task loading" loading="lazy" /></a>
    <figcaption>Story Generation</figcaption>
  </figure>
</div>

<p>Most of the matrices suggest a clear positive manifold. Typos is the exception: it still looks quite ugly after the items are sorted by loading. That makes me want to check whether the Typos items themselves are sound. Auditing every item would take too long, but CTT/IRT measures can flag suspicious ones for us to focus on…</p>

<h2 id="item-flags">Item Flags</h2>

<p>In <a href="/benchmark-flags/">my post about flagging suspicious questions in AI benchmarks</a>, I discussed flags based on item discrimination and distractor behavior. None of the items here is multiple choice, so there are no distractors to examine. That leaves item discrimination, which I measure using the corrected item–task score correlation: the correlation between an item and its task score calculated <em>without</em> that item.</p>

<p class="display-note">Red: below 0 · Yellow: 0–0.2 · Green: above 0.2 · Grey: undefined. Choose a task, then select its histogram for full size.</p>

<style>
  .item-flag-viewer { margin: 0.75rem 0 1.25rem; }
  .item-flag-viewer input[type="radio"] {
    position: absolute;
    width: 1px;
    height: 1px;
    opacity: 0;
  }
  .item-flag-options {
    display: flex;
    flex-wrap: wrap;
    gap: 0.4rem;
    margin-bottom: 0.65rem;
  }
  .item-flag-options label {
    padding: 0.3rem 0.65rem;
    border: 1px solid #d0d7de;
    border-radius: 1rem;
    cursor: pointer;
    font-size: 0.85em;
    line-height: 1.2;
  }
  .item-flag-viewer input[type="radio"]:focus-visible ~ .item-flag-options {
    outline: 2px solid #2563eb;
    outline-offset: 3px;
  }
  #item-flag-typos:checked ~ .item-flag-options label[for="item-flag-typos"],
  #item-flag-lcb:checked ~ .item-flag-options label[for="item-flag-lcb"],
  #item-flag-coding:checked ~ .item-flag-options label[for="item-flag-coding"],
  #item-flag-connections:checked ~ .item-flag-options label[for="item-flag-connections"],
  #item-flag-plot:checked ~ .item-flag-options label[for="item-flag-plot"],
  #item-flag-paraphrase:checked ~ .item-flag-options label[for="item-flag-paraphrase"],
  #item-flag-story:checked ~ .item-flag-options label[for="item-flag-story"] {
    background: #28384d;
    border-color: #28384d;
    color: #fff;
  }
  .item-flag-panels figure {
    display: none;
    margin: 0 !important;
    border: 1px solid #d0d7de;
    border-radius: 0.5rem;
    overflow: hidden;
    background: #fff;
  }
  #item-flag-typos:checked ~ .item-flag-panels #item-flag-panel-typos,
  #item-flag-lcb:checked ~ .item-flag-panels #item-flag-panel-lcb,
  #item-flag-coding:checked ~ .item-flag-panels #item-flag-panel-coding,
  #item-flag-connections:checked ~ .item-flag-panels #item-flag-panel-connections,
  #item-flag-plot:checked ~ .item-flag-panels #item-flag-panel-plot,
  #item-flag-paraphrase:checked ~ .item-flag-panels #item-flag-panel-paraphrase,
  #item-flag-story:checked ~ .item-flag-panels #item-flag-panel-story {
    display: block;
  }
  .item-flag-panels a { display: block; }
  .item-flag-panels img { display: block; width: 100%; max-height: 26rem; object-fit: contain; }
  .item-flag-table { overflow-x: auto; margin: 0.75rem 0 1.5rem; }
  .item-flag-table table { width: 100%; border-collapse: collapse; font-size: 0.85em; }
  .item-flag-table th, .item-flag-table td {
    padding: 0.3rem 0.5rem;
    border: 1px solid #d0d7de;
    white-space: nowrap;
  }
  .item-flag-table td { text-align: right; }
  .item-flag-table tbody th { text-align: left; font-weight: 500; }
  .item-flag-table .item-flag-total { font-weight: 700; background: #f6f8fa; }
  .item-flag-table .flag-dot {
    display: inline-block;
    width: 0.7em;
    height: 0.7em;
    margin-right: 0.3em;
    border-radius: 50%;
  }
  .item-flag-table .flag-red { background: #e52421; }
  .item-flag-table .flag-yellow { background: #f0b900; }
  .item-flag-table .flag-green { background: #16a34a; }
  .item-flag-table .flag-grey { background: #94a3b8; }
</style>

<div class="item-flag-viewer" role="group" aria-label="Item correlation histograms by task">
  <input type="radio" name="item-flag-task" id="item-flag-typos" checked="" />
  <input type="radio" name="item-flag-task" id="item-flag-lcb" />
  <input type="radio" name="item-flag-task" id="item-flag-coding" />
  <input type="radio" name="item-flag-task" id="item-flag-connections" />
  <input type="radio" name="item-flag-task" id="item-flag-plot" />
  <input type="radio" name="item-flag-task" id="item-flag-paraphrase" />
  <input type="radio" name="item-flag-task" id="item-flag-story" />
  <div class="item-flag-options">
    <label for="item-flag-typos">Typos</label>
    <label for="item-flag-lcb">LCB Generation</label>
    <label for="item-flag-coding">Coding Completion</label>
    <label for="item-flag-connections">Connections</label>
    <label for="item-flag-plot">Plot Unscrambling</label>
    <label for="item-flag-paraphrase">Paraphrase</label>
    <label for="item-flag-story">Story Generation</label>
  </div>
  <div class="item-flag-panels">
    <figure id="item-flag-panel-typos"><a href="/assets/images/livebench-factor-analysis/item-task-correlations/typos.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/item-task-correlations/typos.png" alt="Histogram of corrected item–task correlations for Typos" loading="lazy" /></a></figure>
    <figure id="item-flag-panel-lcb"><a href="/assets/images/livebench-factor-analysis/item-task-correlations/LCB_generation.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/item-task-correlations/LCB_generation.png" alt="Histogram of corrected item–task correlations for LCB Generation" loading="lazy" /></a></figure>
    <figure id="item-flag-panel-coding"><a href="/assets/images/livebench-factor-analysis/item-task-correlations/coding_completion.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/item-task-correlations/coding_completion.png" alt="Histogram of corrected item–task correlations for Coding Completion" loading="lazy" /></a></figure>
    <figure id="item-flag-panel-connections"><a href="/assets/images/livebench-factor-analysis/item-task-correlations/connections.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/item-task-correlations/connections.png" alt="Histogram of corrected item–task correlations for Connections" loading="lazy" /></a></figure>
    <figure id="item-flag-panel-plot"><a href="/assets/images/livebench-factor-analysis/item-task-correlations/plot_unscrambling.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/item-task-correlations/plot_unscrambling.png" alt="Histogram of corrected item–task correlations for Plot Unscrambling" loading="lazy" /></a></figure>
    <figure id="item-flag-panel-paraphrase"><a href="/assets/images/livebench-factor-analysis/item-task-correlations/paraphrase.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/item-task-correlations/paraphrase.png" alt="Histogram of corrected item–task correlations for Paraphrase" loading="lazy" /></a></figure>
    <figure id="item-flag-panel-story"><a href="/assets/images/livebench-factor-analysis/item-task-correlations/story_generation.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/item-task-correlations/story_generation.png" alt="Histogram of corrected item–task correlations for Story Generation" loading="lazy" /></a></figure>
  </div>
</div>

<p class="display-note">All 494 items are counted. Percentages use each task's total, including undefined correlations, and are rounded to whole numbers; rows may not sum to 100%.</p>

<div class="item-flag-table" role="region" aria-label="Item correlation flag counts by task" tabindex="0">
  <table>
    <thead><tr><th scope="col">Task</th><th scope="col">Items</th><th scope="col"><span class="flag-dot flag-red"></span>Red (&lt; 0)</th><th scope="col"><span class="flag-dot flag-yellow"></span>Yellow (0–0.2)</th><th scope="col"><span class="flag-dot flag-green"></span>Green (&gt; 0.2)</th><th scope="col"><span class="flag-dot flag-grey"></span>Undefined</th></tr></thead>
    <tbody>
      <tr><th scope="row">Typos</th><td>100</td><td>3 (3%)</td><td>15 (15%)</td><td>80 (80%)</td><td>2 (2%)</td></tr>
      <tr><th scope="row">LCB Generation</th><td>78</td><td>0 (0%)</td><td>1 (1%)</td><td>72 (92%)</td><td>5 (6%)</td></tr>
      <tr><th scope="row">Coding Completion</th><td>50</td><td>0 (0%)</td><td>2 (4%)</td><td>48 (96%)</td><td>0 (0%)</td></tr>
      <tr><th scope="row">Connections</th><td>100</td><td>0 (0%)</td><td>2 (2%)</td><td>98 (98%)</td><td>0 (0%)</td></tr>
      <tr><th scope="row">Plot Unscrambling</th><td>90</td><td>0 (0%)</td><td>0 (0%)</td><td>90 (100%)</td><td>0 (0%)</td></tr>
      <tr><th scope="row">Paraphrase</th><td>50</td><td>0 (0%)</td><td>3 (6%)</td><td>47 (94%)</td><td>0 (0%)</td></tr>
      <tr><th scope="row">Story Generation</th><td>26</td><td>1 (4%)</td><td>3 (12%)</td><td>22 (85%)</td><td>0 (0%)</td></tr>
      <tr class="item-flag-total"><th scope="row">All tasks</th><td>494</td><td>4 (1%)</td><td>26 (5%)</td><td>457 (93%)</td><td>7 (1%)</td></tr>
    </tbody>
  </table>
</div>

<p>The flags uncovered two Typos items with valid alternative answers, three LCB Generation items with grading problems, and three Story Generation items with prompt or checker problems. The <a href="#flagged-item-audit">item-by-item audit</a> gives the archived-answer counts and rescoring checks. I could inspect only a subset of items and models, so other problems may remain.</p>

<h2 id="factor-structure">Factor Structure</h2>

<p>Time to look at the factor structure. I’ll exclude the items I found problems with: <code class="language-plaintext highlighter-rouge">832610e9</code> and <code class="language-plaintext highlighter-rouge">0becbf34</code> from Typos; “Wrong Answer,” “Takahashi Quest,” and “Bad Juice” from LCB Generation; and <code class="language-plaintext highlighter-rouge">0d828b10</code>, <code class="language-plaintext highlighter-rouge">6c5eb0ac</code>, and <code class="language-plaintext highlighter-rouge">230fffb5</code> from Story Generation. Other problematic items may remain because I couldn’t inspect them. As noted above, every task except Typos is strongly unidimensional. A separate factor analysis of Typos produced factors that were hard to interpret, so for simplicity I’ll model each task, including Typos, with a single factor.</p>

<figure>
  <img src="/assets/images/livebench-factor-analysis/correlated_tasks_factor_structure.svg" alt="Seven task factors each load onto their own item responses, and a heptagram-like network connects every pair of task factors with one of 21 freely estimated correlations." loading="lazy" />
</figure>

<p>The item loadings in the correlated-task model look like this:</p>

<figure>
  <a href="/assets/images/livebench-factor-analysis/correlated_tasks_item_loadings.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/correlated_tasks_item_loadings.png" alt="Item loading distributions by task in the correlated-task model; points show posterior medians and vertical bars show 90% intervals, with zero-crossing intervals distinguished from positive intervals." loading="lazy" /></a>
</figure>

<p>Fourteen items have 90% loading intervals that include zero. That doesn’t mean they’re flawed, but it does make them worth a closer look. I’ve put my notes in a collapsible section so they don’t interrupt the main discussion.</p>

<style>
  .loading-audit { margin: 1rem 0 1.5rem; border-top: 1px solid #d0d7de; border-bottom: 1px solid #d0d7de; }
  .loading-audit summary { display: flex; align-items: center; padding: 0.65rem 0.25rem; cursor: pointer; list-style: none; font-weight: 600; }
  .loading-audit summary::-webkit-details-marker { display: none; }
  .loading-audit summary::after { content: "+"; margin-left: auto; color: #57606a; font-size: 1.25rem; font-weight: 400; line-height: 1; }
  .loading-audit[open] summary::after { content: "−"; }
  .loading-audit > :not(summary) { margin-right: 1rem; margin-left: 1rem; }
  .loading-audit > :last-child { margin-bottom: 1rem; }
</style>

<details class="loading-audit">
  <summary>Items whose 90% loading intervals include zero (14)</summary>

  <p><strong>Typos</strong></p>

  <ul>
    <li>For the item with ID prefix <code class="language-plaintext highlighter-rouge">d889972c</code>, the corrupted <code class="language-plaintext highlighter-rouge">vectorfiel-based</code> is keyed as <code class="language-plaintext highlighter-rouge">vector field-based</code>, but some models correct it to <code class="language-plaintext highlighter-rouge">vector-field-based</code>, which seems like a valid alternative. Of the 77 model answers I could check, 51 were marked wrong; 3 of those otherwise match the key exactly and differ only in hyphenation. Since the prompt asks models to preserve stylistic choices, I’d call this ambiguous rather than a definite scoring error.</li>
    <li>For the item with ID prefix <code class="language-plaintext highlighter-rouge">2b05709f</code>, <code class="language-plaintext highlighter-rouge">anbdhten</code> is keyed as <code class="language-plaintext highlighter-rouge">and the</code>, matching the original abstract. But <code class="language-plaintext highlighter-rouge">and then</code> is also a plausible correction in context. Of the 77 model answers I could check, 62 were marked wrong; 28 of those otherwise match the key exactly and differ only in using <code class="language-plaintext highlighter-rouge">and then</code>.</li>
    <li>For <code class="language-plaintext highlighter-rouge">c705c2cb</code>, I found no key or scoring problem among the 77 archived answers I could check.</li>
    <li>For the other eight Typos items, the public data provides neither prompts and keys nor archived answers, so I could not audit them.</li>
  </ul>

  <p><strong>Paraphrase</strong></p>

  <ul>
    <li>The sole Paraphrase item has a coherent prompt and key, but no archived answers, so I cannot determine whether any scores were wrong.</li>
  </ul>

  <p><strong>Story Generation</strong></p>

  <ul>
    <li>For both the items, the archived scores agree with the prompt.</li>
  </ul>

</details>

<p>The task-factor correlations look like this:</p>

<figure>
  <a href="/assets/images/livebench-factor-analysis/correlated_tasks_factor_correlations.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/correlated_tasks_factor_correlations.png" alt="Posterior correlations among the seven task factors; each cell shows the median correlation and its 90% interval on a vivid red-to-green scale." loading="lazy" /></a>
</figure>

<p>There’s a clear positive manifold, and parallel analysis of the <em>task-factor correlation matrix</em> supports a single factor. This would imply a hierarchical model in which a higher-order factor explains the correlation among tasks. However, since LiveBench groups tasks into categories, we might instead add Coding, Language, and Instruction Following domain factors. Those domains could correlate freely or load on an even higher-order factor.</p>

<figure>
  <a href="/assets/images/livebench-factor-analysis/higher_order_model_structures.svg" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/higher_order_model_structures.svg" alt="Three proposed factor structures: one general factor above all seven tasks; three freely correlated domain factors above their respective tasks; and one grand factor above the three domains, which in turn explain their respective tasks." loading="lazy" /></a>
</figure>

<p>Unfortunately, each of these models fits worse than the correlated-task model (see the <a href="#initial-model-comparison">initial comparison</a>).</p>

<p>The domain estimates help explain why. In the correlated-domains model, the domain correlations are quite high:</p>

<figure>
  <a href="/assets/images/livebench-factor-analysis/domain_correlations.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/domain_correlations.png" alt="Posterior correlations among the Coding, Language, and Instruction Following domains, with medians and 90% intervals." loading="lazy" /></a>
</figure>

<p>Even so, they can’t account for some of the task-factor correlations, as we’ll see below. In the grand-factor model, all three domain loadings are <em>very</em> close to 1; the lowest loading is <em>0.9993</em>. That leaves little domain-specific variance, so modeling the categories doesn’t seem to add much.</p>

<p>That brings us back to the hierarchical model. It also fits worse than the correlated-task model, but I find it more plausible a priori. The positive manifold is what I’d expect from LLMs, and parallel analysis of the task-factor correlations supports one common factor. Its poorer fit suggests that the general factor alone misses some relationships between tasks. We can allow for those relationships by adding correlated residuals, so that selected task factors can correlate more than the general factor predicts.</p>

<p>To see which links might be worth including, I fit an exploratory hierarchical model with positive-only shrinkage priors on all 21 task-residual correlations:</p>

<figure>
  <a href="/assets/images/livebench-factor-analysis/exploratory_residual_correlation_matrix.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/exploratory_residual_correlation_matrix.png" alt="Positive-only shrinkage fit: residual correlations among all seven task factors, with posterior medians and 90% intervals." loading="lazy" /></a>
</figure>

<p>The largest estimated residual correlations are Plot Unscrambling–Typos (+.52), LCB Generation–Coding Completion (+.29), and Connections–Story Generation (+.27). But Coding Completion’s loading on the general factor is almost one in this exploratory fit, leaving virtually no task-specific variance. Its +.29 residual correlation therefore adds only about +.002 to the implied correlation between the two tasks. I’ll include Plot–Typos and Connections–Story, but not LCB–Coding. Because the exploratory prior rules out negative residual correlations, an interval above zero is not, by itself, a reason to include a link.</p>

<p>In the final hierarchical model, only those two residual correlations are estimated, with priors that allow either sign. All other residual correlations are fixed at zero. The posterior estimates are:</p>

<table>
  <thead>
    <tr>
      <th>Task pair</th>
      <th style="text-align: right">Median residual correlation</th>
      <th style="text-align: right">90% interval</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Plot Unscrambling–Typos</td>
      <td style="text-align: right">+.57</td>
      <td style="text-align: right">[+.52, +.61]</td>
    </tr>
    <tr>
      <td>Connections–Story Generation</td>
      <td style="text-align: right">+.35</td>
      <td style="text-align: right">[+.19, +.48]</td>
    </tr>
  </tbody>
</table>

<figure>
  <a href="/assets/images/livebench-factor-analysis/hierarchical_selected_residuals.svg" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/hierarchical_selected_residuals.svg" alt="One general factor loads on seven task factors; dashed arcs show the only two additional task-residual correlations, Connections–Story Generation and Plot Unscrambling–Typos." loading="lazy" /></a>
</figure>

<p>Comparing the new model with the previous ones:</p>

<figure>
  <a href="/assets/images/livebench-factor-analysis/waic_elpd_vs_correlated_tasks.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/waic_elpd_vs_correlated_tasks.png" alt="WAIC expected log predictive density differences from correlated task factors, including the hierarchy with Plot–Typos and Connections–Story residual links; bars show one paired pointwise standard error." loading="lazy" /></a>
</figure>

<p>The two residual correlations improve the hierarchical model’s fit.</p>

<p>The loadings of the seven tasks on the general factor in this fit are:</p>

<table>
  <thead>
    <tr>
      <th>Task factor</th>
      <th style="text-align: right">Median loading</th>
      <th style="text-align: right">90% interval</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>LCB Generation</td>
      <td style="text-align: right">.91</td>
      <td style="text-align: right">[.90, .92]</td>
    </tr>
    <tr>
      <td>Coding Completion</td>
      <td style="text-align: right">1.00<sup id="fnref:coding-loading-rounding" role="doc-noteref"><a href="#fn:coding-loading-rounding" class="footnote" rel="footnote">1</a></sup></td>
      <td style="text-align: right">[1.00, 1.00]</td>
    </tr>
    <tr>
      <td>Connections</td>
      <td style="text-align: right">.82</td>
      <td style="text-align: right">[.80, .83]</td>
    </tr>
    <tr>
      <td>Plot Unscrambling</td>
      <td style="text-align: right">.70</td>
      <td style="text-align: right">[.69, .71]</td>
    </tr>
    <tr>
      <td>Typos</td>
      <td style="text-align: right">.79</td>
      <td style="text-align: right">[.77, .81]</td>
    </tr>
    <tr>
      <td>Paraphrase</td>
      <td style="text-align: right">.77</td>
      <td style="text-align: right">[.74, .80]</td>
    </tr>
    <tr>
      <td>Story Generation</td>
      <td style="text-align: right">.86</td>
      <td style="text-align: right">[.82, .89]</td>
    </tr>
  </tbody>
</table>

<p>We can also ask how much of each task’s total-score variance is attributable to the general factor, its task-specific factor, or item-specific variation. This decomposition is for an equal-weighted sum of <em>underlying</em> item responses, not the observed mixed-format LiveBench score:</p>

<figure>
  <a href="/assets/images/livebench-factor-analysis/task_variance_decomposition.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/task_variance_decomposition.png" alt="Stacked bars for each task showing the posterior mean percentages of latent total-score variance attributable to the general factor, task-specific factor, and item-specific variation." loading="lazy" /></a>
</figure>

<h3 id="correlations-with-the-eci">Correlations with the ECI</h3>

<p>This compares Epoch’s ECI with posterior-mean factor scores from the selected hierarchical fit.<sup id="fnref:eci-one-version-match" role="doc-noteref"><a href="#fn:eci-one-version-match" class="footnote" rel="footnote">2</a></sup> The general factor correlates strongly with ECI. The task factors do too, though much of that correlation appears to come from their shared general component. Once that component is removed, only the Plot Unscrambling and Paraphrase residuals have 90% intervals entirely above zero.</p>

<figure>
  <a href="/assets/images/livebench-factor-analysis/eci_task_factor_correlations.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/eci_task_factor_correlations.png" alt="Pearson correlations with ECI for the general factor and, for each of seven tasks, the full task factor and its task-specific residual, with 90% bootstrap intervals and matched-model counts." loading="lazy" /></a>
</figure>

<p class="display-note">Blue: full task factor · Orange: task residual after removing the general factor · General factor at left. Bars show 90% bootstrap intervals; counts include direct task responses only.</p>

<h2 id="takeaways">Takeaways</h2>

<p>Once again, benchmark item flags proved useful for finding problems. The task factors display a positive manifold, as expected. What’s more notable is the lack of clear domain factors beyond the general factor. Human cognitive ability is well modeled by g, but not perfectly: someone may be better at spatial tasks, and someone else better at verbal tasks, than their levels of g would predict. They could have the same g, yet if you needed to navigate an unfamiliar city or write an essay, you might have a clear choice between them.</p>

<p>That distinction is much less apparent for the AI models and tasks tested here. LiveBench divides its tasks into categories, but I find little evidence that these categories capture distinct abilities. It doesn’t seem especially useful to say “use Model A for Coding and Model B for Language” when performance across those domains is so closely tied to general performance. Individual tasks can still differ: Plot Unscrambling and Typos, for example, are more closely related than the general factor alone predicts. But LCB Generation and Coding Completion do not show much extra association, despite both involving coding. The distinctions worth paying attention to seem to lie with particular tasks, not the broad category labels.</p>

<h2 id="appendix">Appendix</h2>

<h3 id="classical-parallel-analysis">Classical Parallel Analysis</h3>

<div style="text-align: center;">
    <figure>
        <img src="/assets/images/livebench-factor-analysis/classical_mixed_correlation_parallel_analysis.png" width="1000" />
    </figure>
</div>

<h3 id="flagged-item-audit">Flagged-item audit</h3>

<p>Looking at the flagged items, along with items that have constant scores across models:</p>

<p><strong>Typos</strong></p>

<ul>
  <li>For the item with ID prefix <code class="language-plaintext highlighter-rouge">832610e9</code>, the scoring key says “algebraical” should be changed to “algebraic”. However, “algebraical” is a valid word listed in the <a href="https://www.oed.com/dictionary/algebraical_adj?tl=true">Oxford English Dictionary</a>. Of the 77 archived model answers I can check, 45 were scored wrong, and 19 of those retained “algebraical”. Among these models, crediting those otherwise-valid answers raises the item–Typos correlation from +.07 to +.52.</li>
  <li>For the item with ID prefix <code class="language-plaintext highlighter-rouge">0becbf34</code>, “behavour” can be corrected to either the US “behavior” or the UK “behaviour”, but only the US spelling was accepted. Of the 77 archived answers I can check, 74 were scored wrong, and 3 of those used the UK spelling. Among these models, crediting those answers raises the item–Typos correlation from +.19 to +.35.</li>
  <li>The other six Typos items I was able to check showed no issues.</li>
  <li>I can’t audit 12 flagged Typos items. The public release has their scores but neither their prompts and keys nor archived answers.</li>
</ul>

<p><strong>LCB Generation</strong></p>

<ul>
  <li>The “<a href="https://atcoder.jp/contests/abc343/tasks/abc343_a?lang=en">Wrong Answer</a>” item is simple: given $A$ and $B$, the program should print any digit from 0 to 9 except $A+B$. For the input <code class="language-plaintext highlighter-rouge">2 5</code>, both <code class="language-plaintext highlighter-rouge">0</code> and <code class="language-plaintext highlighter-rouge">2</code> are valid, but the stored test expects one specific output, such as <code class="language-plaintext highlighter-rouge">2</code>. A valid program that prints <code class="language-plaintext highlighter-rouge">0</code> therefore fails. Despite the item’s simplicity, all 159 judged models received zero. Among the 76 with archived answers, 34 have at least one well-formatted program that prints a digit other than the sum. The current item–task correlation is undefined; crediting these 34 apparently valid archived programs results in a correlation of +.56.</li>
  <li>“<a href="https://atcoder.jp/contests/abc333/tasks/abc333_e?lang=en">Takahashi Quest</a>” has another fixed-output grading issue. All 178 judged models scored zero, making the current item–task correlation undefined. Among models with archived answers, crediting the 34 apparently valid programs results in a correlation of +.24.</li>
  <li>“<a href="https://atcoder.jp/contests/abc337/tasks/abc337_e?lang=en">Bad Juice</a>” is an interactive problem, but it was not evaluated with an interactive judge. A correct program first prints how it will distribute the bottles among the minimum number of friends, reads the judge’s reply about which friends became sick, and then prints the spoiled bottle. LiveBench instead supplies one static input string, <code class="language-plaintext highlighter-rouge">3 1\n</code>, and expects one fixed output transcript. It doesn’t provide replies tailored to each program’s printed groups. All 159 judged models scored zero. Under my provisional re-scoring of the archived answers, the item–task correlation becomes +.31, but I still cannot determine how many programs would pass a real interactive judge.</li>
  <li>The other three LCB Generation items I was able to check showed no issues.</li>
</ul>

<p><strong>Coding Completion</strong></p>

<ul>
  <li>The two flagged items had correct answer keys and scoring.</li>
</ul>

<p><strong>Connections</strong></p>

<ul>
  <li>The two flagged items had correct answer keys and scoring.</li>
</ul>

<p><strong>Plot Unscrambling</strong> (No flagged items)</p>

<p><strong>Paraphrase</strong></p>

<ul>
  <li>The three flagged items had correct answer keys and scoring.</li>
</ul>

<p><strong>Story Generation</strong></p>

<ul>
  <li>For the items with ID prefixes <code class="language-plaintext highlighter-rouge">0d828b10</code> and <code class="language-plaintext highlighter-rouge">6c5eb0ac</code>, the prompt is somewhat contradictory. Models were given a news story and told to “Please generate a story based on the sentences provided. Answer with one of the following options: (‘My answer is yes.’, ‘My answer is no.’, ‘My answer is maybe.’)”. A typical prompt instead adds instructions such as “Entire output should be wrapped in JSON format. You can use markdown ticks such as ```.”, which modify the story’s format or content. Nonetheless, 91/97 models received full credit for <code class="language-plaintext highlighter-rouge">0d828b10</code>, and 93/97 received full credit for <code class="language-plaintext highlighter-rouge">6c5eb0ac</code>. This concerns me: among 77 archived answers per item, 34 and 31, respectively, consisted solely of a permitted phrase. Many models passed the mechanical check, but their scores did not reflect the story-writing request.</li>
  <li>For the item with ID prefix <code class="language-plaintext highlighter-rouge">230fffb5</code>, models were instructed to generate a story with fewer than 241 words and a <code class="language-plaintext highlighter-rouge">P.P.S</code> postscript at the end. The checker counted <code class="language-plaintext highlighter-rouge">\w+</code> tokens, which can split hyphenated expressions and <code class="language-plaintext highlighter-rouge">P.P.S</code> into multiple “words” even when whitespace-based counting would count each as one. Of the 77 models with archived answers, 32 did not receive full credit, and 11 of those were affected by this word-count issue. The postscript check had a separate problem: three models put <code class="language-plaintext highlighter-rouge">P.P.S</code> at the <em>beginning</em> of their answers but still received credit for that instruction; two received full item credit. Among models with archived answers, counting words by whitespace and requiring <code class="language-plaintext highlighter-rouge">P.P.S</code> at the end raises the item–task correlation from +.08 to +.29.</li>
  <li>The other flagged item had no scoring issue.</li>
</ul>

<p>I can inspect archived answers for only a subset of items and models, so I cannot determine the full scope of these problems. Nonetheless, I’ve tried my best doing what I can do.</p>

<h3 id="initial-model-comparison">Initial Model Comparison</h3>

<figure>
  <a href="/assets/images/livebench-factor-analysis/waic_elpd_vs_correlated_tasks_without_residual_links.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/waic_elpd_vs_correlated_tasks_without_residual_links.png" alt="WAIC expected log predictive density differences from the correlated-task reference for the single higher-order factor, correlated domains, and grand-factor-over-domains models, with bars showing one paired standard error." loading="lazy" /></a>
</figure>

<h3 id="loadings-vs-difficulties">Loadings vs Difficulties</h3>

<p>These plots place each item’s task-factor loading against its estimated difficulty in the selected hierarchical fit.</p>

<p class="display-note">Choose a task, then select its plot for full size and 90% intervals. Difficulty uses a task-specific response scale, so compare horizontal positions only within a task.</p>

<style>
  .item-difficulty-viewer { margin: 0.75rem 0 1.25rem; }
  .item-difficulty-viewer input[type="radio"] {
    position: absolute;
    width: 1px;
    height: 1px;
    opacity: 0;
  }
  .item-difficulty-options {
    display: flex;
    flex-wrap: wrap;
    gap: 0.4rem;
    margin-bottom: 0.65rem;
  }
  .item-difficulty-options label {
    padding: 0.3rem 0.65rem;
    border: 1px solid #d0d7de;
    border-radius: 1rem;
    cursor: pointer;
    font-size: 0.85em;
    line-height: 1.2;
  }
  .item-difficulty-viewer input[type="radio"]:focus-visible ~ .item-difficulty-options {
    outline: 2px solid #2563eb;
    outline-offset: 3px;
  }
  #item-difficulty-lcb:checked ~ .item-difficulty-options label[for="item-difficulty-lcb"],
  #item-difficulty-coding:checked ~ .item-difficulty-options label[for="item-difficulty-coding"],
  #item-difficulty-connections:checked ~ .item-difficulty-options label[for="item-difficulty-connections"],
  #item-difficulty-plot:checked ~ .item-difficulty-options label[for="item-difficulty-plot"],
  #item-difficulty-typos:checked ~ .item-difficulty-options label[for="item-difficulty-typos"],
  #item-difficulty-paraphrase:checked ~ .item-difficulty-options label[for="item-difficulty-paraphrase"],
  #item-difficulty-story:checked ~ .item-difficulty-options label[for="item-difficulty-story"] {
    background: #28384d;
    border-color: #28384d;
    color: #fff;
  }
  .item-difficulty-panels figure {
    display: none;
    margin: 0 !important;
    border: 1px solid #d0d7de;
    border-radius: 0.5rem;
    overflow: hidden;
    background: #fff;
  }
  #item-difficulty-lcb:checked ~ .item-difficulty-panels #item-difficulty-panel-lcb,
  #item-difficulty-coding:checked ~ .item-difficulty-panels #item-difficulty-panel-coding,
  #item-difficulty-connections:checked ~ .item-difficulty-panels #item-difficulty-panel-connections,
  #item-difficulty-plot:checked ~ .item-difficulty-panels #item-difficulty-panel-plot,
  #item-difficulty-typos:checked ~ .item-difficulty-panels #item-difficulty-panel-typos,
  #item-difficulty-paraphrase:checked ~ .item-difficulty-panels #item-difficulty-panel-paraphrase,
  #item-difficulty-story:checked ~ .item-difficulty-panels #item-difficulty-panel-story {
    display: block;
  }
  .item-difficulty-panels a { display: block; }
  .item-difficulty-panels img { display: block; width: 100%; max-height: 26rem; object-fit: contain; }
</style>

<div class="item-difficulty-viewer" role="group" aria-label="Item loadings versus difficulty by task">
  <input type="radio" name="item-difficulty-task" id="item-difficulty-lcb" checked="" />
  <input type="radio" name="item-difficulty-task" id="item-difficulty-coding" />
  <input type="radio" name="item-difficulty-task" id="item-difficulty-connections" />
  <input type="radio" name="item-difficulty-task" id="item-difficulty-plot" />
  <input type="radio" name="item-difficulty-task" id="item-difficulty-typos" />
  <input type="radio" name="item-difficulty-task" id="item-difficulty-paraphrase" />
  <input type="radio" name="item-difficulty-task" id="item-difficulty-story" />
  <div class="item-difficulty-options">
    <label for="item-difficulty-lcb">LCB Generation</label>
    <label for="item-difficulty-coding">Coding Completion</label>
    <label for="item-difficulty-connections">Connections</label>
    <label for="item-difficulty-plot">Plot Unscrambling</label>
    <label for="item-difficulty-typos">Typos</label>
    <label for="item-difficulty-paraphrase">Paraphrase</label>
    <label for="item-difficulty-story">Story Generation</label>
  </div>
  <div class="item-difficulty-panels">
    <figure id="item-difficulty-panel-lcb"><a href="/assets/images/livebench-factor-analysis/item-loadings-by-difficulty/LCB_generation.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/item-loadings-by-difficulty/LCB_generation.png" alt="LCB Generation item loadings versus difficulty" loading="lazy" /></a></figure>
    <figure id="item-difficulty-panel-coding"><a href="/assets/images/livebench-factor-analysis/item-loadings-by-difficulty/coding_completion.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/item-loadings-by-difficulty/coding_completion.png" alt="Coding Completion item loadings versus difficulty" loading="lazy" /></a></figure>
    <figure id="item-difficulty-panel-connections"><a href="/assets/images/livebench-factor-analysis/item-loadings-by-difficulty/connections.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/item-loadings-by-difficulty/connections.png" alt="Connections item loadings versus difficulty" loading="lazy" /></a></figure>
    <figure id="item-difficulty-panel-plot"><a href="/assets/images/livebench-factor-analysis/item-loadings-by-difficulty/plot_unscrambling.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/item-loadings-by-difficulty/plot_unscrambling.png" alt="Plot Unscrambling item loadings versus difficulty" loading="lazy" /></a></figure>
    <figure id="item-difficulty-panel-typos"><a href="/assets/images/livebench-factor-analysis/item-loadings-by-difficulty/typos.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/item-loadings-by-difficulty/typos.png" alt="Typos item loadings versus difficulty" loading="lazy" /></a></figure>
    <figure id="item-difficulty-panel-paraphrase"><a href="/assets/images/livebench-factor-analysis/item-loadings-by-difficulty/paraphrase.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/item-loadings-by-difficulty/paraphrase.png" alt="Paraphrase item loadings versus difficulty" loading="lazy" /></a></figure>
    <figure id="item-difficulty-panel-story"><a href="/assets/images/livebench-factor-analysis/item-loadings-by-difficulty/story_generation.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/item-loadings-by-difficulty/story_generation.png" alt="Story Generation item loadings versus difficulty" loading="lazy" /></a></figure>
  </div>
</div>

<h3 id="task-information-curves">Task Information Curves</h3>

<p>The seven task-factor information curves are overlaid on the same axes, so their heights are directly comparable.</p>

<p class="display-note">Select a task or its curve to highlight it. Hover, focus, or tap a model marker for its name, median task score, 90% interval, and information at that score. The full-size plots also show individual item curves.</p>

<style>
  .task-information-interactive {
    margin: 0.8rem 0 0.65rem;
    border: 1px solid #d0d7de;
    border-radius: 0.55rem;
    background: #fff;
    overflow: hidden;
  }
  .task-information-viewport { overflow-x: auto; }
  .task-information-viewport > a { display: block; }
  .task-information-viewport img { display: block; width: 100%; height: auto; }
  .task-information-svg { display: block; min-width: 730px; width: 100%; height: auto; }
  .task-information-controls {
    display: flex;
    flex-wrap: wrap;
    gap: 0.4rem;
    padding: 0.7rem;
    border-bottom: 1px solid #d0d7de;
    background: #f6f8fa;
  }
  .task-information-task-button {
    border: 1px solid #cbd5e1;
    border-left: 4px solid var(--task-color);
    border-radius: 0.35rem;
    background: #fff;
    color: #334155;
    cursor: pointer;
    font: inherit;
    font-size: 0.82em;
    line-height: 1.25;
    padding: 0.34rem 0.5rem;
  }
  .task-information-task-button:hover,
  .task-information-task-button:focus-visible { border-color: var(--task-color); }
  .task-information-task-button[aria-pressed="true"] {
    border-color: var(--task-color);
    background: #eff6ff;
    color: #172554;
    font-weight: 700;
    box-shadow: inset 0 0 0 1px var(--task-color);
  }
  .task-information-grid { stroke: #e9edf3; stroke-width: 1; }
  .task-information-grid-zero { stroke: #cbd5e1; stroke-width: 1; }
  .task-information-fill { opacity: 0.08; }
  .task-information-curve { fill: none; stroke-width: 2; opacity: 0.38; }
  .task-information-curve.is-active { stroke-width: 4; opacity: 1; }
  .task-information-curve-hit { fill: none; stroke: transparent; stroke-width: 13; cursor: pointer; }
  .task-information-curve-hit:focus-visible { stroke: #1e293b; stroke-width: 1; stroke-dasharray: 3 3; }
  .task-information-axis-label, .task-information-axis-title { fill: #475569; font-size: 12px; }
  .task-information-strip-label { fill: #475569; font-size: 12px; font-weight: 600; }
  .task-information-strip-baseline { stroke: #cbd5e1; stroke-width: 1; }
  .task-information-interval { stroke: #1e3a8a; stroke-width: 2.5; opacity: 0.42; }
  .task-information-tick { stroke: #1e3a8a; stroke-width: 2.5; }
  .task-information-marker { cursor: pointer; }
  .task-information-hitbox { fill: transparent; stroke: transparent; stroke-width: 2; }
  .task-information-marker:hover .task-information-tick,
  .task-information-marker:focus .task-information-tick,
  .task-information-marker.is-selected .task-information-tick { stroke: #d97706; stroke-width: 4; }
  .task-information-marker:focus-visible .task-information-hitbox { stroke: #d97706; }
  .task-information-selected-guide { stroke: #d97706; stroke-width: 1.6; stroke-dasharray: 3 3; }
  .task-information-selected-point { fill: #d97706; stroke: #fff; stroke-width: 1.5; }
  .task-information-detail {
    display: flex;
    flex-wrap: wrap;
    align-items: baseline;
    gap: 0.2rem 0.6rem;
    min-height: 2.8rem;
    padding: 0.55rem 0.8rem;
    border-top: 1px solid #d0d7de;
    background: #f6f8fa;
    color: #24292f;
    font-size: 0.88em;
    line-height: 1.4;
  }
  .task-information-detail strong { color: #1d4ed8; }
  .task-information-model-id { color: #64748b; font-size: 0.82em; }
  .task-information-links {
    display: flex;
    flex-wrap: wrap;
    gap: 0.3rem 0.9rem;
    margin: 0 0 1.5rem;
    font-size: 0.9em;
  }
  .task-information-links a { text-decoration: underline; text-underline-offset: 2px; }
</style>

<div class="task-information-interactive" data-livebench-task-information="" data-src="/assets/jsons/livebench_task_information.json">
  <div class="task-information-viewport"><a href="/assets/images/livebench-factor-analysis/task-information-curves/overview.png" target="_blank" rel="noopener"><img src="/assets/images/livebench-factor-analysis/task-information-curves/overview.png" alt="Seven overlaid task-factor information curves on shared score and information axes" loading="lazy" /></a></div>
  <div class="task-information-detail" aria-live="polite">Loading interactive model markers…</div>
</div>
<script src="/assets/js/livebench-task-information.js" defer=""></script>

<nav class="task-information-links" aria-label="Full-size task information plots">
  <a href="/assets/images/livebench-factor-analysis/task-information-curves/LCB_generation.png" target="_blank" rel="noopener">LCB Generation</a>
  <a href="/assets/images/livebench-factor-analysis/task-information-curves/coding_completion.png" target="_blank" rel="noopener">Coding Completion</a>
  <a href="/assets/images/livebench-factor-analysis/task-information-curves/connections.png" target="_blank" rel="noopener">Connections</a>
  <a href="/assets/images/livebench-factor-analysis/task-information-curves/plot_unscrambling.png" target="_blank" rel="noopener">Plot Unscrambling</a>
  <a href="/assets/images/livebench-factor-analysis/task-information-curves/typos.png" target="_blank" rel="noopener">Typos</a>
  <a href="/assets/images/livebench-factor-analysis/task-information-curves/paraphrase.png" target="_blank" rel="noopener">Paraphrase</a>
  <a href="/assets/images/livebench-factor-analysis/task-information-curves/story_generation.png" target="_blank" rel="noopener">Story Generation</a>
</nav>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:coding-loading-rounding" role="doc-endnote">
      <p>All values in the task-loading table are rounded to two decimal places. Coding Completion’s displayed 1.00 values are slightly below 1 before rounding. <a href="#fnref:coding-loading-rounding" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:eci-one-version-match" role="doc-endnote">
      <p>For this figure, I include ECI models with exactly one distinct model version in their benchmark records and ECI scores dated no later than April 7, 2025. I match version names to LiveBench after ignoring case and punctuation, but not version numbers or words. If several LiveBench runs match the same ECI model, I use the one with the most item responses. This leaves 38 ECI models before task-specific response requirements. <a href="#fnref:eci-one-version-match" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name></name></author><category term="ai-benchmarks" /><summary type="html"><![CDATA[LiveBench is an LLM benchmark with tasks grouped into categories. I’ll analyze its April 7, 2025 public release, which included 18 tasks across six categories. Item-level data were available for seven tasks across three categories: Coding (LCB Generation and Coding Completion), Language (Connections, Plot Unscrambling, and Typos), and Instruction Following (Paraphrase and Story Generation). Unfortunately, I can’t analyze the other tasks because LiveBench didn’t release their item-level data, which is essential for this kind of analysis.]]></summary></entry><entry><title type="html">People Like AI Slop</title><link href="https://stanichor.net/ai-slop/" rel="alternate" type="text/html" title="People Like AI Slop" /><published>2026-09-29T00:00:00+00:00</published><updated>2026-09-29T00:00:00+00:00</updated><id>https://stanichor.net/ai-slop</id><content type="html" xml:base="https://stanichor.net/ai-slop/"><![CDATA[<p>I often hear people talk about how someone’s decision to use AI-generated content is going to backfire because AI-generated content, as we all know, is slop, and people will be disgusted and turn away.</p>

<p>This is wrong. People dislike the <em>idea</em> of AI-generated content. But when you don’t tell them it’s AI-generated, people love AI slop. Sure, maybe you and your friends hate AI-generated content<sup id="fnref:tells" role="doc-noteref"><a href="#fn:tells" class="footnote" rel="footnote">1</a></sup>. However, you<sup id="fnref:you" role="doc-noteref"><a href="#fn:you" class="footnote" rel="footnote">2</a></sup> and your circle are quite unrepresentative.</p>

<p>In studies where participants aren’t told whether something was made by a human or an AI, people generally rate the AI-generated version as about as good as or better than the human-generated version. This has been found for <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC11564748/">poetry</a>, <a href="https://journals.plos.org/plosone/article?id=10.1371%2Fjournal.pone.0305364">jokes</a>, <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC11065427/">short stories</a>, and <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC11750838/">visual art</a>.</p>

<p>The poetry results are particularly illustrative. Participants couldn’t reliably distinguish AI-generated poems from poems by famous poets, and rated the AI poems higher in overall quality. Every one of the five AI poems was rated more highly than every one of the five human poems. In another study, ChatGPT-3.5’s jokes were rated as funnier than jokes written by ordinary people, while its satirical headlines were rated about as funny as headlines from <em>The Onion</em>. For short stories, researchers found no significant difference in human-rated creativity between human- and AI-written stories. And in one visual-art experiment, participants significantly preferred DALL·E 2 images to human-made artworks.</p>

<p>What seems to hurt AI-generated content isn’t the quality of the content itself, but people knowing or believing that it’s AI-generated. In the poetry study, the exact same poems received lower ratings when participants were told they were AI-generated rather than human-written. A 2024 study of visual art found something similar: participants generally preferred the AI-generated images on their actual merits, but rated images less favorably the more strongly they believed those images had been generated by AI.</p>

<p>And some of these studies are from up to <em>two</em> years ago, ancient by AI standards. We should expect current models to perform much better.</p>

<h2 id="solved-hyperstimuli">Solved hyperstimuli</h2>

<p>A framework for thinking about “taste” that I think is apt is laid out in <a href="https://cameronharwick.com/writing/high-culture-and-hyperstimulus/">“High Culture and Hyperstimulus”</a>: what separates low culture from high culture is not that low culture consists of hyperstimuli, since the same is true of high culture, but that low culture consists of <em>solved</em> hyperstimuli, while high culture consists of <em>unsolved</em> hyperstimuli.</p>

<p>A Playboy centerfold is an obvious example of tasteless low culture, seemingly because it presents the hyperstimulus of a naked body. But Renaissance nudes <em>also</em> contain naked bodies and are considered <em>high</em> culture. Why are alcoholics considered distasteful while sommeliers are considered sophisticated? Why is K-pop disposable mass culture while symphonies are high culture? Each of these cases involves roughly the same kind of underlying reward, so what makes one vulgar while the other is respectable?</p>

<p>The answer, on this framework, is whether the hyperstimulus is solved, whether we know how to produce it reliably on demand.</p>

<p>The alcoholic is indiscriminating in the alcohol he consumes, provided it gets him drunk, while the sommelier focuses on subtle flavors such as the notes that (<a href="https://www.cremieux.xyz/p/the-myth-of-the-sommelier">supposedly</a>) tell you where the grapes grew and under what conditions. It’s harder to mass-produce these subtle characteristics, or so the sommeliers say, than to make some kind of alcohol that can get any fool drunk. Anyone can go on the internet and look at naked bodies, but to view a Renaissance nude in a museum, you have to pretend to know something about various historical artists and movements. K-pop is famous for having a solved formula that allows producers to manufacture predictable hits.</p>

<p>This also explains the barber pole of culture. A particular form of art starts out as ineffable, with no one quite sure what the secret is behind it and only a blessed few able to pull it off. But as people figure out the formula, it becomes solved and, inevitably, low culture.</p>

<p>You can see how this poses a problem for AI-generated content: it is the epitome of low culture. Any kind of art it can produce has, almost by definition, been solved so completely that a machine can churn out effectively infinite content on demand.</p>

<p>Perhaps if the user is heavily involved in the creation process, then that particular piece of art might not be so low culture. But I expect that, in practice, we’ll always be a little suspicious that the “artist” was not as involved as they claim, like the person who says they wrote most of their 10,000-word essay when in reality they wrote three bullet points and asked Claude to expand on their thoughts.</p>

<h2 id="revealed-taste">Revealed taste</h2>

<p>I also think some of this is a mixture of preference falsification and people being bad at knowing their own preferences.</p>

<p>The high-status thing to say in these circles is that, of course, while AI may be better than most creators in a technical sense, it’s missing that <em>spark</em> that is the reason we consume the work of the top 1% of creators. Or that we value human provenance<sup id="fnref:provenance" role="doc-noteref"><a href="#fn:provenance" class="footnote" rel="footnote">3</a></sup> so much that we’d actually prefer to consume genuinely kinda terrible human-made content over better AI-generated content.</p>

<p>I think this is mostly posturing and self-deception.</p>

<p>When people know they’re looking at AI-generated work, they may dislike what it represents, dislike the production process, dislike the social implications, or simply believe that liking it reflects badly on their taste. But when you strip away the label and ask them to judge the thing itself, the evidence so far suggests that their preferences lean the opposite way.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:tells" role="doc-endnote">
      <p>At least, you say you hate the content you can <em>tell</em> is AI-generated. By the way, how good are you actually at telling when content is AI-generated? Sure, you can spot the obvious cases, but what about the more subtle ones? <a href="#fnref:tells" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:you" role="doc-endnote">
      <p>I feel comfortable making some assumptions about the kind of person likely to read this, such as that you’re much more knowledgeable about and exposed to AI than the general population. <a href="#fnref:you" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:provenance" role="doc-endnote">
      <p>Did that word make you fear that an AI wrote this? <a href="#fnref:provenance" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name></name></author><category term="ai-benchmarks" /><summary type="html"><![CDATA[I often hear people talk about how someone’s decision to use AI-generated content is going to backfire because AI-generated content, as we all know, is slop, and people will be disgusted and turn away.]]></summary></entry><entry><title type="html">#AlwaysLookAtTheItems</title><link href="https://stanichor.net/look-at-items/" rel="alternate" type="text/html" title="#AlwaysLookAtTheItems" /><published>2026-09-29T00:00:00+00:00</published><updated>2026-09-29T00:00:00+00:00</updated><id>https://stanichor.net/look-at-items</id><content type="html" xml:base="https://stanichor.net/look-at-items/"><![CDATA[<p>“<a href="https://www.psypost.org/people-with-higher-cognitive-ability-have-weaker-moral-foundations-new-study-finds/">People with higher cognitive ability have weaker moral foundations</a>”</p>

<p>Taken at face value, this is true. The study used a scale called the Moral Foundations Questionnaire 2 (MFQ-2), which says it measures moral foundations, and cognitive ability was negatively correlated with <em>all</em> of its sub-scales, which happen to be labelled “Care”, “Equality”, “Proportionality”, “Loyalty”, “Authority”, and “Purity”.</p>

<p>But looking at just the names psychologists give to the traits they measure is a good way to be misled about reality. Looking at how they <em>define</em> those traits can also be a good way to be misled. If you really want to understand what a psychological scale is measuring, you should look at how it actually measures it: the questions people are being asked.</p>

<p>For example, with the <a href="https://stanichor.net/mfq-2/">MFQ-2</a>, names like “Equality” and “Loyalty” make you think the sub-scales are measuring broad traits like preferences for equal treatment and rights, or how much one values attachment and obligations to one’s friends, family, and other groups. But if you look at the items, you’ll see that the Equality sub-scale largely measures support for equalizing incomes and resources (e.g., “I believe it would be ideal if everyone in society wound up with roughly the same amount of money”), while the Loyalty sub-scale largely measures patriotism and national identity (e.g., “I think children should be taught to be loyal to their country”).</p>

<p>So, after looking at the items, we realize “people with higher cognitive ability have weaker moral foundations” is rather misleading. It isn’t that people with higher cognitive ability are less morally egalitarian in some broad sense, or that they’re particularly disloyal to their family and friends. What the study actually establishes is quite different: people with higher cognitive ability show less support equalizing income and resources, and have less patriotism and national identity.</p>

<p>The <a href="https://stanichor.net/systemizing/">Systemizing Quotient</a> (SQ) gives another example of the same problem. The SQ is supposed to measure a trait called “systemizing”, defined as “the drive to analyze or construct systems”, and men <em>do</em> score higher on the SQ than women do. If you stop there, the natural conclusion is that men have a substantially stronger drive to analyze and construct systems.</p>

<p>But when we actually look at the items, we find that, while many do seem related to thinking analytically, they’re unusually focused on male-typical interests like cars, computers, sports, stocks, technology, etc. When we instead look at factors made up of more gender-neutral content, such as nature, language, or attention to detail, the gender difference decreases dramatically. So at least part of the famous sex difference in “systemizing” is really a sex difference in the particular <em>kinds of systems the questionnaire asks about</em>.</p>

<p>Looking at the items themselves gets us much closer to understanding what a scale measures, but we can go another step. Not every item contributes equally to a scale or factor. We should also look at the item <em>loadings</em>: how strongly each item is related to the latent construct the scale is supposed to be measuring.</p>

<p>For example, <a href="https://stanichor.net/targeted-personality/">taking a look at Tailcalled’s Targeted Personality Test</a>, one might assume that a factor named “Charisma” measures how charismatic the respondent actually is. And, to be fair, there are some items dealing with relatively concrete social behaviors (e.g., “In conversations, I jump straight to the point rather than doing smalltalk and other irrelevant things”). But the highest-loading items are much more directly about how charismatic or socially skilled the respondent <em>perceives themself to be</em>.</p>

<p>This matters if, for example, one wants to look at correlations between Charisma and other variables like self-esteem. If the Charisma score is heavily determined by questions asking people whether they think they’re charismatic or socially competent, then a correlation between Charisma and self-esteem is partly a correlation between two kinds of positive self-evaluation. That’s a rather different result from showing that people who are <em>actually perceived by others as charismatic</em> have higher self-esteem.</p>

<p>There are related issues whenever the method of measurement introduces something important into the construct. If you’re reading a study (or Substack post) that’s looking at the relationship between attractiveness and other variables, but attractiveness is measured not through external raters but through self-reports, then a substantial component of the “attractiveness” variable is going to be how attractive people <em>think</em> they are. So the people who score as most “attractive” won’t necessarily just be the best-looking people; they’ll also tend to be the people with the highest opinions of their own appearance.</p>

<p>There was <a href="https://x.com/davidshor/status/2097856198544617801">a recent analysis</a> that found that people who read horoscopes every day report personalities more similar to what we’d expect from stereotypes about their Zodiac sign, while the effect among people who never read horoscopes was approximately zero. But remember how we usually measure personality: through self-reports. One possibility is that there’s a sort of placebo affect where astrology changes people’s actual personalities if people believe in astrology. Another, much more likely possibility is that people who believe in astrology incorporate the stereotypes associated with their Zodiac sign into their <em>self-perception</em> and therefore into how they answer personality questionnaires. If we want to investigate the effect of reading horoscopes on someone’s <em>actual</em> personality, we’d have to look at other-ratings or behavioral measures rather than solely self-ratings.</p>

<p>The broader lesson is this: <strong>look at the items</strong>. The name of a scale is just the researcher’s summary of what they think the scale is measuring (sometimes not even that!), and researchers are often <em>wrong</em>. The items, on the other hand, are the essence of the scale. So, always look at the items.</p>]]></content><author><name></name></author><category term="measurement" /><summary type="html"><![CDATA[“People with higher cognitive ability have weaker moral foundations”]]></summary></entry><entry><title type="html">How to Spot Suspicious Questions in AI Benchmarks</title><link href="https://stanichor.net/benchmark-flags/" rel="alternate" type="text/html" title="How to Spot Suspicious Questions in AI Benchmarks" /><published>2026-09-27T00:00:00+00:00</published><updated>2026-09-27T00:00:00+00:00</updated><id>https://stanichor.net/benchmark-flags</id><content type="html" xml:base="https://stanichor.net/benchmark-flags/"><![CDATA[<style>
  .pf-review {
    margin: 1rem 0;
    border: 1px solid #d0d7de;
    border-radius: .5rem;
  }
  .pf-review summary {
    cursor: pointer;
    font-weight: 600;
    list-style: none;
    padding: .75rem 1rem;
    background: #f6f8fa;
    border-radius: .5rem;
  }
  .pf-review summary::-webkit-details-marker {
    display: none;
  }
  .pf-review summary::before {
    content: "▸";
    display: inline-block;
    margin-right: .55rem;
    transition: transform .15s ease;
  }
  .pf-review[open] summary::before {
    transform: rotate(90deg);
  }
  .pf-review > :not(summary) {
    margin-left: 1rem;
    margin-right: 1rem;
  }
  .pf-theory {
    max-width: 48rem;
    margin: 1.75rem 0 2rem 1.25rem;
    padding: 1rem 1.25rem;
    border-left: 4px solid #6f9fc2;
    border-radius: 0 .4rem .4rem 0;
    background: #f3f7fb;
  }
  .pf-theory summary {
    position: relative;
    padding-right: 1.5rem;
    cursor: pointer;
    list-style: none;
  }
  .pf-theory summary::-webkit-details-marker {
    display: none;
  }
  .pf-theory summary::before {
    content: "BACKGROUND";
    display: block;
    margin-bottom: .3rem;
    color: #315b76;
    font-size: .7rem;
    font-weight: 700;
    letter-spacing: .1em;
  }
  .pf-theory summary h2 {
    margin: 0 !important;
    padding: 0;
    border: 0;
    color: #193b55;
    font-size: 1.1em;
    line-height: 1.3;
  }
  .pf-theory summary::after {
    content: "▸";
    position: absolute;
    right: 0;
    top: 50%;
    transform: translateY(-50%);
    color: #315b76;
  }
  .pf-theory details[open] summary::after {
    content: "▾";
  }
  @media (max-width: 600px) {
    .pf-theory {
      margin-left: 0;
    }
  }
</style>

<p>AI benchmarks have a problem: lots of the questions are scored incorrectly. Sometimes the answer key selects the wrong answer as correct, sometimes the question is literally impossible and there is no correct answer, and sometimes there’s more than one correct answer. Obviously, this makes measuring AI progress harder. We might think models are plateauing when really they’ve hit the benchmark’s ceiling and many of the remaining items are incorrectly scored, ambiguous, or otherwise flawed.</p>

<p>For example, in September 2026, GPT-5.6-Sol scored a 47.3% mean@4 on the Physics section of Humanity’s Last Exam. However, <a href="https://arxiv.org/abs/2609.13009">when researchers audited questions models repeatedly got wrong</a>, they discovered that many of the questions were incorrect, ambiguous, or otherwise flawed. In many cases, the problem was with the question, not the model. After removing or repairing the flawed questions, GPT-5.6-Sol’s measured mean@4 increased from 47.3% to 78.7%.</p>

<p>How are we to avoid problems like these? Well, the basic idea behind benchmarks is to measure how capable AI models are and how much they know. As it turns out, there’s a field already dedicated to figuring out how to measure abilities and knowledge accurately in humans: psychometrics.</p>

<p>Psychometrics is especially pertinent because cognitive and achievement tests can be <em>very</em> important: they can determine what college one goes to, whether one is able to enter a profession they’ve spent years and hundreds of thousands of dollars preparing for, or inform education policies affecting millions of students. In short, these tests have an enormous effect on the lives of millions.</p>

<p>As such, it’s very important that these tests have correct items, measure the intended trait accurately, and are free of bias. Items undergo rigorous screening before they appear on an exam. I’m going to go through a few of the flags<sup id="fnref:other-flags" role="doc-noteref"><a href="#fn:other-flags" class="footnote" rel="footnote">1</a></sup> that psychometricians look at when deciding whether an item is suspicious and in need of a closer look: item discrimination, item difficulty, and distractor behavior. I’ll be looking at MMLU-Pro<sup id="fnref:mmlu-pro" role="doc-noteref"><a href="#fn:mmlu-pro" class="footnote" rel="footnote">2</a></sup>. For each flag, I’ll look<sup id="fnref:looking" role="doc-noteref"><a href="#fn:looking" class="footnote" rel="footnote">3</a></sup> at some of the most extreme items it identifies, so we can see whether they really are bad items.</p>

<aside class="pf-theory">
  <details>
    <summary><h2 id="a-brief-and-probably-inadequate-intro-to-item-response-theory">A Brief (and Inadequate) Intro to Item Response Theory</h2></summary>

    <p>Most high-stakes assessments make extensive use of <a href="https://en.wikipedia.org/wiki/Item_response_theory">item response theory</a> (IRT). IRT, at least the way it’s typically used, makes three assumptions:</p>

    <ul>
      <li>There is a unidimensional trait, $\theta$, that determines how a respondent answers items.</li>
      <li>The items are <a href="https://en.wikipedia.org/wiki/Local_independence">locally independent</a>, in that relationships between item responses are mediated through the aforementioned trait $\theta$.</li>
      <li>A respondent’s response to an item can be modeled by an item response function (IRF), which gives the probability that a respondent with a given ability level, $\theta$, will answer correctly.</li>
    </ul>

    <p>Usually, the item response function is based on the logistic function, $\sigma(x)$:</p>

\[\sigma(x) = \frac{1}{1 + e^{-x}}\]

    <p>There’s the two-parameter logistic (2PL) model:</p>

\[p_i(\theta) = \sigma(a_i(\theta - b_i))\]

    <p>which has parameters $a_i$, which denotes the item’s discrimination, and $b_i$, which denotes the item’s difficulty.</p>

    <p>There’s also the three-parameter logistic (3PL) model, which adds a pseudo-guessing parameter, $c_i$:</p>

\[p_i(\theta) = c_i + \frac{1-c_i}{1+\exp(-a_i(\theta - b_i))}\]

    <p>While most uses of IRT involve the logistic function, the normal ogive (the CDF of the standard normal distribution) is also popular. A scaling factor of roughly 1.7 is often used to make the logistic function closely approximate the normal ogive, and so you’ll sometimes see, say, a 2PL item response function written as</p>

\[p_i(\theta) = \frac{1}{1+\exp(-Da_i(\theta - b_i))}\]

    <p>where $D$ is usually set to 1.7.</p>

  </details>
</aside>

<h2 id="discrimination">Discrimination</h2>

<p>An item’s discrimination is the degree to which the item is able to distinguish between high-ability and low-ability respondents. There are a number of different measures we could use: the $a_i$ parameter from our fitted IRF, the point-biserial correlation between item correctness and total score, or the discrimination index, which is the difference in the proportion answering correctly between the top 27% and the bottom 27%.</p>

<p>Any of these being negative for an item is a bad sign: it means that less capable respondents are <em>more</em> likely to answer that specific item correctly than more capable respondents. That should make us wonder whether the indicated correct answer is actually correct, the question is ambiguous, or something else has gone wrong with the item.</p>

<p>But even positive, near-zero discriminations are suspect. For example, the <a href="https://research.collegeboard.org/media/pdf/Digital%20SAT%20Suite%20of%20Assessments%20Technical%20Manual-FINAL.pdf">SAT</a> flags pretest items for review when their item-total score correlations are below 0.20, while the <a href="https://www.act.org/content/dam/act/unsecured/documents/ACT_Technical_Manual.pdf">ACT</a> generally expects items to have item-total correlations of at least 0.20. <a href="https://www.oecd.org/content/dam/oecd/en/publications/reports/2024/03/pisa-2022-technical-report_599753f0/01820d6d-en.pdf">PISA</a> flags items when their IRT discrimination parameter is below 0.1, while the <a href="https://thebarexaminer.ncbex.org/article/september-2015/the-testing-column-equating-the-mbe">MBE</a> says that, for item selection, the discrimination index should at least be positive and preferably greater than 0.20. Such low discriminations indicate that item correctness barely changes with ability, which is still a problem. Is an item really measuring ability if more capable respondents are barely more likely to answer it correctly? At the very least, something about the item warrants investigation. And even if nothing is technically <em>wrong</em> with it, there’s not much <em>point</em> in including an item that contributes almost nothing to distinguishing between more and less capable respondents. We’re just wasting time (and tokens).</p>

<p>When looking at discrimination in the IRT fit, instead of using a logistic model, I’ll be using a probit model and reporting the standardized loading. A common rule of thumb treats loadings with magnitudes below 0.3 as weak, so I’ll use loadings below 0.3 as a flag for review. Negative loadings are even more suspect: they mean that respondents become <em>less</em> likely to answer the item correctly as their ability increases.</p>

<div style="text-align: center;">
    <figure>
        <img src="/assets/images/psychometric-flags/item_metric_scatterplot_matrix.png" width="1000" />
    </figure>
</div>

<p>It looks like it doesn’t much matter <em>which</em> measure we use, since they’re all very correlated with each other. They also all agree that ~35% of the items should be flagged for review, indicated by the red and yellow, with ~15% having discriminations of the wrong sign, indicated by the red, and another ~20% merely having very weak discriminations, indicated by the yellow.</p>

<p>Now, let’s look at some of the items with the worst discriminations. We’ll use the item-total score correlation, since it has the highest average correlation with the other measures:</p>

<details class="pf-review">
  <summary>Question 8006</summary>

  <blockquote>
    <p>In one study half of a class were instructed to watch exactly 1 hour of television per day, the other half were told to watch 5 hours per day, and then their class grades were compared. In a second study students in a class responded to a questionnaire asking about their television usage and their class grades.</p>
  </blockquote>

  <ul>
    <li>A. The first study was an experiment without a control group, while the second was an observational study. <strong>[recorded key]</strong></li>
    <li>B. Both studies were observational studies.</li>
    <li>C. The first study was a controlled experiment, while the second was an observational study.</li>
    <li>D. The first study was an observational study, while the second was an experiment without a control group.</li>
    <li>E. The first study was an experiment with a control group, while the second was an observational study.</li>
    <li>F. Both studies were controlled experiments.</li>
    <li>G. Both studies were experiments without control groups.</li>
    <li>H. The first study was an observational study, the second was neither an observational study nor a controlled experiment.</li>
    <li>I. The first study was a controlled experiment, while the second was neither an observational study nor a controlled experiment.</li>
    <li>J. The first study was an observational study, while the second was a controlled experiment.</li>
  </ul>

  <p>While the fact that the second study is observational seems unambiguous, leaving only A, B, C, and E as our remaining options, what to call the first study is a bit trickier. The first study is obviously an experiment, eliminating B, but is it an experiment with or without a control group? In one sense, the first study does not have a control group, since both groups are assigned some amount of TV exposure. In another sense, the 1-hour-of-TV-per-day group serves as an active control, allowing us to estimate the effect of an additional 4 hours of TV exposure per day. So both A and E could be said to be either correct or incorrect depending on what one means by a “control group”. Meanwhile, C is unambiguously a controlled experiment. The item hinges on ambiguous terminology rather than a substantive statistical distinction, and so this item would not pass review.</p>

</details>

<details class="pf-review">
  <summary>Question 7977</summary>

  <blockquote>
    <p>Find the smallest positive integer that leaves a remainder of 2 when divided by 3, a remainder of 3 when divided by 5, and a remainder of 1 when divided by 7.</p>
  </blockquote>

  <ul>
    <li>A. 8 <strong>[recorded key]</strong></li>
    <li>B. 31</li>
    <li>C. 10</li>
    <li>D. 37</li>
    <li>E. 23</li>
    <li>F. 12</li>
    <li>G. 14</li>
    <li>H. 52</li>
    <li>I. 15</li>
    <li>J. 26</li>
  </ul>

  <p>So, 8 is clearly the correct answer, and when I checked there didn’t seem to be an issue like the models reporting “8” instead of “A”. The most popular model choices are 31 and 37, with 37 being more popular among the stronger models. The item passes review, though the question of what’s going on with the models remains.</p>

</details>

<details class="pf-review">
  <summary>Question 6931</summary>

  <blockquote>
    <p>Define Gross National Product (GNP).</p>
  </blockquote>

  <ul>
    <li>A. The total income earned by a nation’s residents in a year</li>
    <li>B. The total amount of money in circulation in an economy in a year</li>
    <li>C. The total market value of all final goods and services produced in the economy in one year <strong>[recorded key]</strong></li>
    <li>D. The market value of all goods and services produced abroad by the residents of a nation in a year</li>
    <li>E. The total cost of all goods and services purchased in a year</li>
    <li>F. The total savings rate of a nation’s residents plus the value of imports minus the value of exports in a year</li>
    <li>G. The aggregate of all wages paid to employees, plus profits of businesses and taxes, minus any subsidies</li>
    <li>H. The sum of all financial transactions within a country’s borders in a year</li>
    <li>I. The total value of all consumer spending, government spending, investments, and net exports in a year</li>
    <li>J. The total value of all goods and services produced by a nation’s residents, regardless of the location</li>
  </ul>

  <p>So, the US Bureau of Economic Analysis says the <a href="https://www.bea.gov/help/glossary/gross-national-product-gnp">gross national product</a> is “[t]he market value of goods and services produced by labor and property supplied by U.S. residents, regardless of where they are located”. The <a href="https://www.bea.gov/help/glossary/gross-domestic-product-gdp">gross domestic product</a> (GDP), on the other hand, is a measure of “the value of final goods and services produced within the United States”. So C is describing GDP, not GNP. J is actually the correct answer. This item would not provide review.</p>

</details>

<details class="pf-review">
  <summary>Question 8473</summary>

  <blockquote>
    <p>Let V and W be 4-dimensional subspaces of a 7-dimensional vector space X. Which of the following CANNOT be the dimension of the subspace V intersect W?</p>
  </blockquote>

  <ul>
    <li>A. 0 <strong>[recorded key]</strong></li>
    <li>B. 1</li>
    <li>C. 9</li>
    <li>D. 8</li>
    <li>E. 2</li>
    <li>F. 5</li>
    <li>G. 4</li>
    <li>H. 7</li>
    <li>I. 3</li>
    <li>J. 6</li>
  </ul>

  <p>While it’s correct that 0 cannot be the dimension of the subspace V intersect W, it’s also the case that 5, 6, 7, 8, and 9 cannot be the dimensions either. This means options C, D, F, H, and J are correct, in addition to the recorded key, A. So there’s more than one correct answer. This item would not pass review.</p>

</details>

<details class="pf-review">
  <summary>Question 10750</summary>

  <blockquote>
    <p>Which is the largest asymptotically?</p>
  </blockquote>

  <ul>
    <li>A. O(n^2) <strong>[recorded key]</strong></li>
    <li>B. O(n^3)</li>
    <li>C. O(sqrt(n))</li>
    <li>D. O(2^n)</li>
    <li>E. O(log n)</li>
    <li>F. O(n log n)</li>
    <li>G. O(log log n)</li>
    <li>H. O(1)</li>
    <li>I. O(n)</li>
  </ul>

  <p>O(n^3) is obviously larger than O(n^2), but O(2^n) is much larger than <em>both</em> of them. So this is another case of an incorrect key: D, not A, is the intended correct answer. This item would not pass review.</p>

</details>

<p>4 out of the 5 items we looked at would not pass review. This suggests that using the discrimination flag would be quite helpful for identifying defective or anomalous items.</p>

<h2 id="difficulty">Difficulty</h2>

<p>An item’s difficulty is, uh, how difficult it is. One way to measure it is to look at what proportion of respondents answer the item correctly. Another way is to look at the $b_i$ parameter of the fitted IRF, where $b_i$ indicates the ability level at which a respondent has a 50% chance of answering the item correctly.</p>

<p>In typical high-stakes assessments, test-makers want items to be neither too difficult nor too easy. The SAT flags items for review when the proportion answering correctly exceeds 0.90 or falls below 0.20. The ACT flags items for review when the proportion answering correctly falls outside the range 0.100–0.899 and, additionally, tries to ensure a balanced distribution of items across defined difficulty bands. PISA flags items where $|b_i| &gt; 5$.</p>

<p>While it makes sense for benchmark creators to want to exclude items that are too easy, the case for excluding very difficult items is trickier. If we’re interested in measuring progress in AI, we may actually want to retain some easy items so that the benchmark remains informative for weaker models. Likewise, if you ensure that none of your items are too difficult, then your benchmark will quickly saturate and become obsolete. But the problem with items that no model can answer correctly is that, a lot of the time, the reason no model can answer them correctly is that the item is faulty, as in the case of Humanity’s Last Exam mentioned above.</p>

<p>Extreme difficulties also pose problems for estimating other item parameters, such as discrimination. It’s difficult to determine how well an item discriminates when nearly every model gives the same response, either because almost all models answer it correctly or almost all answer it incorrectly, regardless of ability. That said, let’s look at the most difficult items, as measured by the percentage of models answering correctly, and see what happens:</p>

<details class="pf-review">
  <summary>Question 10053</summary>

  <blockquote>
    <p>A block is dragged along a table and experiences a frictional force, f, that opposes its movement. The force exerted on the block by the table is</p>
  </blockquote>

  <p>Options (models selecting each letter):</p>

  <ul>
    <li>A. parallel to the table — 49/500</li>
    <li>B. equal to the frictional force — 57/500</li>
    <li>C. in the opposite direction of movement — 193/500</li>
    <li>D. perpendicular to the table — 146/500</li>
    <li>E. in the direction of movement — 27/500</li>
    <li>F. zero — 10/500</li>
    <li>G. equal to the gravitational force — 3/500</li>
    <li>H. neither parallel nor perpendicular to the table — 1/500 <strong>[scored key]</strong></li>
    <li>I. always greater than the frictional force — 8/500</li>
    <li>J. always less than the frictional force — 6/500</li>
  </ul>

  <p>There are two forces exerted on the block by the table: the normal force, which is perpendicular to the table, and the frictional force, which is parallel to the table and opposite the direction of motion. The net contact force exerted by the table is the vector sum of these two forces, so it is neither parallel nor perpendicular to the table. Thus, H is in fact the correct option. This item would pass review.</p>

</details>

<details class="pf-review">
  <summary>Question 6998</summary>

  <blockquote>
    <p>Suppose a bank has $250,000 in deposits, and $10,000 in ex-cess reserves. If the required reserve ratio is 20%, what are the bank's actual reserves?</p>
  </blockquote>

  <p>Options (models selecting each letter):</p>

  <ul>
    <li>A. $30,000 — 119/500</li>
    <li>B. $50,000 — 136/500</li>
    <li>C. $20,000 — 68/500</li>
    <li>D. $110,000 — 107/500</li>
    <li>E. $80,000 — 26/500</li>
    <li>F. $40,000 — 16/500</li>
    <li>G. $100,000 — 16/500</li>
    <li>H. $60,000 — 1/500 <strong>[scored key]</strong></li>
    <li>I. $90,000 — 7/500</li>
    <li>J. $70,000 — 4/500</li>
  </ul>

  <p>I would assume excess reserves refers to reserves the bank holds in excess of the required amount, so the bank’s actual reserves should be $(0.20 \times $250,000) + $10,000 = $60,000$. Thus, H is the correct option. This item would pass review.</p>

</details>

<details class="pf-review">
  <summary>Question 5401</summary>

  <blockquote>
    <p>By what nickname is the Federal National Mortgage Association known?</p>
  </blockquote>

  <p>Options (models selecting each letter):</p>

  <ul>
    <li>A. Feddie Mac — 26/500</li>
    <li>B. Fannie Mac — 245/500</li>
    <li>C. FedNat — 6/500</li>
    <li>D. Federal Mae — 5/500</li>
    <li>E. FEMA — 1/500 <strong>[scored key]</strong></li>
    <li>F. FedMort — 7/500</li>
    <li>G. Frankie Mae — 1/500</li>
    <li>H. Freddie Mac — 203/500</li>
    <li>I. Morty — 6/500</li>
  </ul>

  <p>A quick Wikipedia search tells us that the <a href="https://en.wikipedia.org/wiki/Fannie_Mae">Federal National Mortgage Association</a> is commonly known as Fannie Mae. Interestingly, that’s not an option here. Its abbreviation, FNMA, isn’t an option either. So, not only is the scored key incorrect, but there isn’t even a correct option. This item would not pass review.</p>

</details>

<details class="pf-review">
  <summary>Question 7918</summary>

  <blockquote>
    <p>Compute 22 / 2 + 9.</p>
  </blockquote>

  <p>Options (models selecting each letter):</p>

  <ul>
    <li>A. 24 — 158/500</li>
    <li>B. 23 — 196/500</li>
    <li>C. 21 — 23/500</li>
    <li>D. 22 — 55/500</li>
    <li>E. 10 — 10/500</li>
    <li>F. 25 — 27/500</li>
    <li>G. 2 — 3/500</li>
    <li>H. 20 — 1/500 <strong>[scored key]</strong></li>
    <li>I. 19 — 11/500</li>
    <li>J. 11 — 16/500</li>
  </ul>

  <p>Using the order of operations, $22 / 2 + 9 = 11 + 9 = 20$, so H is the correct option. In fact, I’m not sure how one arrives at most of the popular answers here. Even evaluating the expression strictly from left to right gives 20. You could get 2 only by incorrectly treating the expression as $22 / (2 + 9)$. So I am confused why models prefer options such as A or B. Anyway, the item would pass review, though I do wonder what’s going on with the models.</p>

</details>

<details class="pf-review">
  <summary>Question 8280</summary>

  <blockquote>
    <p>John is playing a game in which he tries to obtain the highest number possible. He must put the symbols +, $\times$, and - (plus, times, and minus) in the following blanks, using each symbol exactly once:[2 \underline{\hphantom{8}} 4 \underline{\hphantom{8}} 6 \underline{\hphantom{8}} 8.] John cannot use parentheses or rearrange the numbers. What is the highest possible number that John could obtain?</p>
  </blockquote>

  <p>Options (models selecting each letter):</p>

  <ul>
    <li>A. 22 — 36/500</li>
    <li>B. 90 — 119/500</li>
    <li>C. 100 — 108/500</li>
    <li>D. 78 — 13/500</li>
    <li>E. 99 — 204/500</li>
    <li>F. 46 — 2/500 <strong>[scored key]</strong></li>
    <li>G. 56 — 5/500</li>
    <li>H. 50 — 2/500</li>
    <li>I. 66 — 7/500</li>
    <li>J. 38 — 4/500</li>
  </ul>

  <p>It does seem like the highest possible number one could obtain is 46, from $2 - 4 + 6 \times 8 = 46$, so F is the correct option. In fact, the only values obtainable from the six possible arrangements of the three operations are 18, -42, 6, 10, 46, and -14, so I’m confused why a model would choose any other option. Regardless, this item would pass review.</p>

</details>

<p>1 out of the 5 items we looked at would not pass review. In fact, the fact that so many models aren’t able to answer some of these items correctly confuses me a bit. Interestingly, 47 of the 102 (46%) most difficult items had H<sup id="fnref:h" role="doc-noteref"><a href="#fn:h" class="footnote" rel="footnote">4</a></sup> as the scored key, compared to 1,100 of 11,836 (9%) items overall. Regardless, the difficulty flag seems less useful than the discrimination flag for identifying defective items.</p>

<h2 id="distractor-behavior">Distractor Behavior</h2>

<p>Distractors are the incorrect answer choices in multiple-choice questions. In the same way that we expect the probability of selecting the correct option to increase with capability, we expect the probability of selecting an incorrect option to decrease with capability. If, instead, the probability of selecting a particular distractor increases with capability, we have to wonder what’s going on. Perhaps the indicated correct answer is not actually correct, perhaps there’s more than one defensible answer, perhaps the question itself is ambiguous, or perhaps something else has gone wrong. This is why the SAT flags items for review when “a distractor has a correlation with the total score greater than 0.05.”</p>

<p>So, I’ll take a look at the items with the highest distractor-total score correlations. However, because I’m taking the highest distractor-total correlation for each item and I’m <em>also</em> looking at 11,832 items, some correlations may appear quite high purely by chance. So, I use p-values<sup id="fnref:bayesian" role="doc-noteref"><a href="#fn:bayesian" class="footnote" rel="footnote">5</a></sup> and adjust them for the false discovery rate using the <a href="https://en.wikipedia.org/wiki/False_discovery_rate#Benjamini%E2%80%93Hochberg_procedure">Benjamini-Hochberg procedure</a>.</p>

<div style="text-align: center;">
    <figure>
        <img src="/assets/images/psychometric-flags/distractor_permutation_calibration.png" width="1000" />
    </figure>
</div>

<p>So, lots of items were flagged: more than half. This is another indicator of potential problems. Items with concerning distractor behavior also tend to have concerning discriminations.</p>

<div style="text-align: center;">
    <figure>
        <img src="/assets/images/psychometric-flags/item_metric_scatterplot_matrix_spearman_with_distractors.png" width="1000" />
    </figure>
</div>

<p>Let’s look at the worst offenders:</p>

<details class="pf-review">
  <summary>Question 9811</summary>

  <blockquote>
    <p>Two identical containers are filled with different gases. Container 1 is filled with hydrogen and container 2 is filled with nitrogen. Each container is set on a lab table and allowed to come to thermal equilibrium with the room. Which of the following correctly compares the properties of the two gases?</p>
  </blockquote>

  <ul>
    <li>A. The pressures of the gases cannot be compared without knowing the number of molecules in each container. <strong>[scored key]</strong></li>
    <li>B. The pressures of the gases cannot be compared without knowing the temperature of the room.</li>
    <li>C. The thermal conductivity of the hydrogen gas is less than the nitrogen gas.</li>
    <li>D. The average force exerted on the container by the hydrogen gas is greater than the nitrogen gas.</li>
    <li>E. The viscosity of the hydrogen gas is greater than the nitrogen gas.</li>
    <li>F. The diffusion rate of the hydrogen gas is less than the nitrogen gas.</li>
    <li>G. The average speed of the hydrogen gas molecules is less than the nitrogen gas molecules.</li>
    <li>H. The density of the hydrogen gas is less than the nitrogen gas. <strong>[highest-r distractor]</strong></li>
    <li>I. The average kinetic energy of the hydrogen gas is greater than the nitrogen gas.</li>
    <li>J. The average kinetic energy of the nitrogen gas is greater than the hydrogen gas.</li>
  </ul>

  <p>Luckily, $PV = nRT$, the ideal gas law, has been burned into my brain. Volume is equal between the gases since the containers are identical, and temperature is also equal since they’ve both reached thermal equilibrium with the room. So, the only two variables left free to vary are $P$, the pressure, and $n$, the number of moles, and we’re not given any information about either. As such, A is the correct answer, so the key is correct. H would be true <em>if</em> we were told that the two containers held equal numbers of molecules, but we weren’t, so it’s incorrect, though it’s understandable how a model could arrive at that answer. All that said, though the item would be flagged for review, it would pass, since upon inspection I see no problems.</p>

</details>

<details class="pf-review">
  <summary>Question 5160</summary>

  <blockquote>
    <p>As of 2019, which of the following had the lowest life expectancy?</p>
  </blockquote>

  <ul>
    <li>A. Australia</li>
    <li>B. Japan</li>
    <li>C. United States</li>
    <li>D. Canada</li>
    <li>E. Iran</li>
    <li>F. Germany</li>
    <li>G. Brazil</li>
    <li>H. Russia <strong>[highest-r distractor]</strong></li>
    <li>I. China</li>
    <li>J. Mexico <strong>[scored key]</strong></li>
  </ul>

  <p>According to <a href="https://ourworldindata.org/grapher/life-expectancy-unwpp?tab=line&amp;country=MEX~RUS">Our World in Data</a>, in 2019 Mexico had a life expectancy of 74.5 years, while Russia had a life expectancy of 73.1 years. So Russia, not Mexico, is the correct option. This item would not pass review.</p>

</details>

<details class="pf-review">
  <summary>Question 9205</summary>

  <blockquote>
    <p>An object of mass 2 kg is acted upon by three external forces, each of magnitude 4 N. Which of the following could NOT be the resulting acceleration of the object?</p>
  </blockquote>

  <ul>
    <li>A. 0 m/s^2</li>
    <li>B. 12 m/s^2</li>
    <li>C. 18 m/s^2 <strong>[highest-r distractor]</strong></li>
    <li>D. 4 m/s^2</li>
    <li>E. 10 m/s^2</li>
    <li>F. 14 m/s^2</li>
    <li>G. 16 m/s^2</li>
    <li>H. 8 m/s^2 <strong>[scored key]</strong></li>
    <li>I. 2 m/s^2</li>
    <li>J. 6 m/s^2</li>
  </ul>

  <p>Since each force has magnitude 4 N, the greatest possible net force occurs when they’re all pointing in the same direction, giving a total force of 12 N. Since the mass is 2 kg, the maximum possible acceleration is $6\ \mathrm{m/s^2}$. As such, while H is a correct answer, so are B, C, E, F, and G. This item would not pass review.</p>

</details>

<details class="pf-review">
  <summary>Question 10478</summary>

  <p>Question:</p>

  <blockquote>
    <p>The procedure below is intended to display the index in a list of unique names (nameList) where a particular name (targetName) is found. lf targetName is not found in nameList, the code should display 0.
 PROCEDURE FindName (nameList, targetName)
 {
  index ← 0
  FOR EACH name IN nameList
  {
   index ← index + 1
   IF (name = targetName)
   {
   foundIndex ← index
   }
   ELSE
   {
   foundIndex ← 0
   }
  }
  DISPLAY (foundIndex)
 }
 Which of the following procedure calls can be used to demonstrate that the procedure does NOT Work as intended?</p>
  </blockquote>

  <ul>
    <li>A. FindName ([“Andrea”, “Ben”, “Chris”, “Diane”, “Eva”, “Frank”, “Grace”, “Hannah”, “Igor”], “Igor” )</li>
    <li>B. FindName ([“Andrea”, “Ben”, “Chris”, “Diane”], “Diane” )</li>
    <li>C. FindName ([“Andrea”, “Ben”, “Chris”, “Diane”, “Eva”, “Frank”], “Frank” )</li>
    <li>D. FindName ([“Andrea”, “Ben”], “Ben” )</li>
    <li>E. FindName ([“Andrea”, “Chris”, “Diane”], “Ben”) <strong>[highest-r distractor]</strong></li>
    <li>F. FindName ([“Andrea”, “Ben” ], “Diane” )</li>
    <li>G. FindName ([“Andrea”, “Ben”, “Chris”], “Ben”) <strong>[scored key]</strong></li>
    <li>H. FindName ([“Andrea”, “Ben”, “Chris”, “Diane”, “Eva”], “Eva” )</li>
    <li>I. FindName ([“Andrea”, “Ben”, “Chris”, “Diane”, “Eva”, “Frank”, “Grace”, “Hannah”], “Hannah” )</li>
    <li>J. FindName ([“Andrea”, “Ben”, “Chris”, “Diane”, “Eva”, “Frank”, “Grace”], “Grace” )</li>
  </ul>

  <p>G does seem to be the correct answer. The problem with the procedure is that, while it sets <code class="language-plaintext highlighter-rouge">foundIndex</code> to index when it finds <code class="language-plaintext highlighter-rouge">targetName</code>, it resets <code class="language-plaintext highlighter-rouge">foundIndex</code> to 0 if it processes any nonmatching names afterward. G exposes this problem: it finds “Ben” at index 2, then resets <code class="language-plaintext highlighter-rouge">foundIndex</code> to 0 when it processes “Chris.” E is notable, along with F, in that <code class="language-plaintext highlighter-rouge">nameList</code> does not contain <code class="language-plaintext highlighter-rouge">targetName</code>, but that doesn’t expose a problem with the procedure, since returning 0 is exactly what it’s supposed to do in that case. The item seems correct and would pass review.</p>

</details>

<details class="pf-review">
  <summary>Question 10666</summary>

  <blockquote>
    <table>
      <tbody>
        <tr>
          <td>Statement 1</td>
          <td>Overfitting is more likely when the set of training data is small. Statement 2</td>
          <td>Overfitting is more likely when the hypothesis space is small.</td>
        </tr>
      </tbody>
    </table>
  </blockquote>

  <ul>
    <li>A. False, False, False</li>
    <li>B. True, True</li>
    <li>C. False, False</li>
    <li>D. False, True <strong>[scored key]</strong></li>
    <li>E. False, Not enough information</li>
    <li>F. True, False <strong>[highest-r distractor]</strong></li>
    <li>G. Not enough information, True</li>
    <li>H. True, Not enough information</li>
    <li>I. Not enough information, Not enough information</li>
    <li>J. Not enough information, False</li>
  </ul>

  <p>Statement 1 is generally true: holding other things equal, overfitting is more likely with a smaller training set. So D, the scored key, is already incorrect. Statement 2 also seems false: if anything, overfitting is generally more likely when the hypothesis space is <em>larger</em>, since a more flexible hypothesis class can fit idiosyncrasies in the training data more easily. So F seems to be the correct option, not D. This item would not pass review.</p>

</details>

<p>3 out of the 5 items we looked at would not pass review. So, the distractor flag also seems helpful for identifying defective or anomalous items.</p>

<h2 id="takeaways">Takeaways</h2>

<p>As you can see, all of the above criteria are <em>flags</em> for further review. They’re not automatic exclusion criteria that mean an item should immediately be eliminated. It’s also the case that the discrimination and distractor flags seem much more informative than the difficulty flags, which makes sense: they more directly indicate that something may be wrong with the item itself.</p>

<p>This is also much faster than going through every item one by one or randomly selecting items to inspect. And it’s better than simply looking at the items that many models got wrong, since, as we saw, difficulty by itself isn’t nearly as informative as discrimination or distractor behavior.</p>

<p>Psychometric flags can’t automatically tell us which benchmarks are faulty. But they can tell us where to look. Instead of waiting until models appear to hit a mysterious performance ceiling and only then auditing hundreds or thousands of items, we can use the same kinds of statistical diagnostics that psychometricians already use to identify suspicious items for closer review. If we’re going to use benchmarks to measure increasingly capable AI systems, the current state of affairs is abysmal, and we need to start holding benchmarks to higher standards.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:other-flags" role="doc-endnote">
      <p>There’s other flags like item response time (how quickly respondents spend on an item), item omission (what proportion of items attempt to answer the item), and differential item functioning (DIF) / measurement invariance (whether the item functions differently between groups). Flags like reponse time and item omission aren’t as applicable to AI models as they would be in humans. DIF, on the other hand, deserves its own post. <a href="#fnref:other-flags" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:mmlu-pro" role="doc-endnote">
      <p>Because I <em>was</em> going to use the MMLU-Pro for a different post idea, before realizing how utterly terrible it is. <a href="#fnref:mmlu-pro" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:looking" role="doc-endnote">
      <p>Well, since I’m not a polymath<sup id="fnref:polymath" role="doc-noteref"><a href="#fn:polymath" class="footnote" rel="footnote">6</a></sup>, I’ll be looking through the most flagrant violators that <em>I</em> can actualy verify the veracity of. I <em>could</em> just ask an LLM whether the answer is scored correctly. But that seems akin to asking the students you’re testing whether their answers are correct. If they’re going to be more correct than you, then what’s the point in testing them? And if you think there is a point in testing them, then you necessarily don’t think you can trust their judgement of what is correct. <a href="#fnref:looking" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:h" role="doc-endnote">
      <p>H has become a most ominous letter <a href="#fnref:h" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:bayesian" role="doc-endnote">
      <p>This is unfortunately quite Frequentist of me. <a href="#fnref:bayesian" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:polymath" role="doc-endnote">
      <p>Looking through the items, I seem to know less than a high schooler. 😔 <a href="#fnref:polymath" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name></name></author><category term="ai-benchmarks" /><summary type="html"><![CDATA[]]></summary></entry><entry><title type="html">Factors of Attraction Towards Men: A Replication</title><link href="https://stanichor.net/attraction-to-men/" rel="alternate" type="text/html" title="Factors of Attraction Towards Men: A Replication" /><published>2026-09-25T00:00:00+00:00</published><updated>2026-09-25T00:00:00+00:00</updated><id>https://stanichor.net/attraction-to-men</id><content type="html" xml:base="https://stanichor.net/attraction-to-men/"><![CDATA[<p>In <a href="https://thingstoread.substack.com/p/factors-of-attraction-toward-men">Factors of Attraction Toward Men</a>, Apple Pie conducted surveys on romantic preferences, asking participants to rate several traits based on how attractive they are. The following is a replication of the analysis using confirmatory factor analysis (CFA) and item response theory (IRT), rather than using principal component analysis (PCA) as Apple Pie did. PCA forces factors to be orthogonal, while CFA allows factors to correlate. IRT also provides more item-level information.</p>

<h2 id="factor-analysis-results">Factor Analysis Results</h2>

<p>Parallel analysis suggested nine factors. Expand each factor below for its interpretation and loadings.</p>

<style>
  .factor-result {
    margin: 0 !important;
    border-top: 1px solid #d0d7de;
  }

  .factor-result:last-of-type {
    border-bottom: 1px solid #d0d7de;
  }

  .factor-result summary {
    display: flex;
    align-items: center;
    justify-content: space-between;
    padding: 0.8rem 0.25rem;
    cursor: pointer;
    list-style: none;
  }

  .factor-result summary::-webkit-details-marker {
    display: none;
  }

  .factor-result summary h3 {
    margin: 0 !important;
    font-size: 1.25em;
  }

  .factor-result summary::after {
    content: "+";
    margin-left: 1rem;
    color: #57606a;
    font-size: 1.35rem;
    font-weight: 400;
    line-height: 1;
  }

  .factor-result[open] summary::after {
    content: "-";
  }

  .factor-result > :not(summary) {
    margin-right: 1rem;
    margin-left: 1rem;
  }

  .factor-result > :last-child {
    margin-bottom: 1.5rem;
  }
</style>

<details class="factor-result">
  <summary><h3 id="ruggedness">Ruggedness</h3></summary>

  <div style="text-align: center;">
    <figure>
        <img src="/assets/images/female-attraction/visualized-ruggedness.png" alt="Ruggedness factor illustration" width="600" />
    </figure>
</div>

  <p>This factor indicates a preference for roughness and danger. The items dealing with bodily preferences (‘Smooth Skin’ (negative loading), ‘Scars’, ‘Rough Skin’, ‘Sweat’) indicate a preference for a sort of ruggedness: a weathered, physically tough appearance rather than a smooth or polished one. The items dealing with personality (‘Dangerous’, ‘Cocky Attitude’, ‘Adventurous’) indicate the behavioral archetype of the “bad boy”: someone daring, cocky, and somewhat risk-seeking. Other items, such as ‘Guns’, ‘Motorcycles’, and ‘Clunky Old Cars’, fit into the same general aesthetic of roughness, danger, and disregard for polish. Taken together, the factor seems to capture a preference for a rough, dangerous form of masculinity, encompassing both rugged appearance and roguish behavior.</p>

  <p><strong>Factor Loadings</strong></p>

  <table>
    <thead>
      <tr>
        <th>Trait</th>
        <th style="text-align: right">Loading</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Scars</td>
        <td style="text-align: right">0.64</td>
      </tr>
      <tr>
        <td>Smooth Skin</td>
        <td style="text-align: right">-0.63</td>
      </tr>
      <tr>
        <td>Rough Skin</td>
        <td style="text-align: right">0.61</td>
      </tr>
      <tr>
        <td>Dangerous</td>
        <td style="text-align: right">0.54</td>
      </tr>
      <tr>
        <td>Sweat</td>
        <td style="text-align: right">0.49</td>
      </tr>
      <tr>
        <td>Pretty Faces</td>
        <td style="text-align: right">-0.36</td>
      </tr>
      <tr>
        <td>Clunky Old Cars</td>
        <td style="text-align: right">0.36</td>
      </tr>
      <tr>
        <td>Motorcycles</td>
        <td style="text-align: right">0.35</td>
      </tr>
      <tr>
        <td>Guns</td>
        <td style="text-align: right">0.34</td>
      </tr>
      <tr>
        <td>Cocky Attitude</td>
        <td style="text-align: right">0.34</td>
      </tr>
      <tr>
        <td>Adventurous</td>
        <td style="text-align: right">0.30</td>
      </tr>
      <tr>
        <td>Rugged Looks</td>
        <td style="text-align: right">0.24</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result">
  <summary><h3 id="companionability">Companionability</h3></summary>

  <div style="text-align: center;">
    <figure>
        <img src="/assets/images/female-attraction/visualized-companionability.png" alt="Companionability factor illustration" width="600" />
    </figure>
</div>

  <p>This factor represents a preference for a warm, lively, enjoyable personality. This is obvious from the items that deal directly with personality, such as ‘Enthusiastic’, ‘Silly’, ‘Humorous’, and ‘Adventurous’, but we can learn more by looking at the other items. This is a relatively embodied personality, as evidenced by items such as ‘A Good Dancer’, ‘A Good Cook’, and ‘The Life of the Party’: someone who expresses their personality through doing things rather than merely possessing abstract personality traits. There’s also an emphasis on the man being “safe”, as you can see from items such as ‘Kind’, ‘Sympathetic’, and ‘Good with Children’, as well as the negative loading of ‘Unfaithful’. Taken together, this seems to describe someone energetic, fun, warm, and socially engaging.</p>

  <p><strong>Factor Loadings</strong></p>

  <table>
    <thead>
      <tr>
        <th>Trait</th>
        <th style="text-align: right">Loading</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Enthusiastic</td>
        <td style="text-align: right">0.54</td>
      </tr>
      <tr>
        <td>A Good Dancer</td>
        <td style="text-align: right">0.53</td>
      </tr>
      <tr>
        <td>Kind</td>
        <td style="text-align: right">0.52</td>
      </tr>
      <tr>
        <td>Silly</td>
        <td style="text-align: right">0.45</td>
      </tr>
      <tr>
        <td>Sympathetic</td>
        <td style="text-align: right">0.40</td>
      </tr>
      <tr>
        <td>A Good Cook</td>
        <td style="text-align: right">0.40</td>
      </tr>
      <tr>
        <td>The Life of the Party</td>
        <td style="text-align: right">0.38</td>
      </tr>
      <tr>
        <td>Humorous</td>
        <td style="text-align: right">0.36</td>
      </tr>
      <tr>
        <td>Good with Children</td>
        <td style="text-align: right">0.36</td>
      </tr>
      <tr>
        <td>Adventurous</td>
        <td style="text-align: right">0.35</td>
      </tr>
      <tr>
        <td>Unfaithful</td>
        <td style="text-align: right">-0.28</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result">
  <summary><h3 id="affluence">Affluence</h3></summary>

  <div style="text-align: center;">
    <figure>
        <img src="/assets/images/female-attraction/visualized-affluence.png" alt="Affluence factor illustration" width="600" />
    </figure>
</div>

  <p>Honestly, there’s not much to say here: this factor is obvious and only has a measly three items. It represents a preference for rich men. ‘Money’ and ‘Wealthy’ are obvious. ‘Fast Cars’ loads positively, but weakly, probably because fast cars are a visible symbol of wealth. Notably, this factor isn’t about a preference for high-status men <em>per se</em>. Education, intelligence, suits, dominance, and other potential indicators of status don’t load on this factor. It really does seem to be about money specifically.</p>

  <p><strong>Factor Loadings</strong></p>

  <table>
    <thead>
      <tr>
        <th>Trait</th>
        <th style="text-align: right">Loading</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Money</td>
        <td style="text-align: right">0.64</td>
      </tr>
      <tr>
        <td>Wealthy</td>
        <td style="text-align: right">0.60</td>
      </tr>
      <tr>
        <td>Fast Cars</td>
        <td style="text-align: right">0.29</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result">
  <summary><h3 id="boyishness">Boyishness</h3></summary>

  <div style="text-align: center;">
    <figure>
        <img src="/assets/images/female-attraction/visualized-boyishness.png" alt="Boyishness factor illustration" width="600" />
    </figure>
</div>

  <p>This factor indicates a preference for youthful prettiness. The age component is obvious: ‘Young’ and ‘Teenagers’ load positively. There’s also a facial and bodily component: ‘Pretty in the Face’, ‘Pretty Faces’, and ‘Slender’ all load positively. Evidently, hair is not seen as youthful, as shown by the negative loadings of ‘Chest Hair’, ‘Facial Hair’, and ‘Receding Hairlines’.</p>

  <p><strong>Factor Loadings</strong></p>

  <table>
    <thead>
      <tr>
        <th>Trait</th>
        <th style="text-align: right">Loading</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Pretty in the Face</td>
        <td style="text-align: right">0.63</td>
      </tr>
      <tr>
        <td>Chest Hair</td>
        <td style="text-align: right">-0.61</td>
      </tr>
      <tr>
        <td>Receding Hairlines</td>
        <td style="text-align: right">-0.53</td>
      </tr>
      <tr>
        <td>Facial Hair</td>
        <td style="text-align: right">-0.48</td>
      </tr>
      <tr>
        <td>Pretty Faces</td>
        <td style="text-align: right">0.48</td>
      </tr>
      <tr>
        <td>Young</td>
        <td style="text-align: right">0.46</td>
      </tr>
      <tr>
        <td>Slender</td>
        <td style="text-align: right">0.44</td>
      </tr>
      <tr>
        <td>Short Necks</td>
        <td style="text-align: right">-0.37</td>
      </tr>
      <tr>
        <td>Teenagers</td>
        <td style="text-align: right">0.33</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result">
  <summary><h3 id="intellect">Intellect</h3></summary>

  <div style="text-align: center;">
    <figure>
        <img src="/assets/images/female-attraction/visualized-intellect.png" alt="Intellect factor illustration" width="600" />
    </figure>
</div>

  <p>This factor indicates a preference for intelligence in all its forms: education, brilliance, maturity, independent thought, and wit. This is obvious when we look at item pairs such as ‘Educated’ (positive loading) vs. ‘Uneducated’ (negative loading), or ‘Brilliant’ (positive loading) vs. ‘Not so bright’ (negative loading). Beyond book smarts, there’s also a focus on verbal intelligence, as indicated by items such as ‘Witty’ and ‘Simple Spoken’ (negative loading). Interestingly, ‘Humorous’ does <em>not</em> load on this factor, despite the loading of ‘Witty’, suggesting that the relevant distinction is verbal cleverness rather than simply being funny.</p>

  <p><strong>Factor Loadings</strong></p>

  <table>
    <thead>
      <tr>
        <th>Trait</th>
        <th style="text-align: right">Loading</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Uneducated</td>
        <td style="text-align: right">-0.71</td>
      </tr>
      <tr>
        <td>Educated</td>
        <td style="text-align: right">0.71</td>
      </tr>
      <tr>
        <td>Not so bright</td>
        <td style="text-align: right">-0.60</td>
      </tr>
      <tr>
        <td>Brilliant</td>
        <td style="text-align: right">0.48</td>
      </tr>
      <tr>
        <td>Witty</td>
        <td style="text-align: right">0.39</td>
      </tr>
      <tr>
        <td>Free Thinking</td>
        <td style="text-align: right">0.39</td>
      </tr>
      <tr>
        <td>Simple Spoken</td>
        <td style="text-align: right">-0.38</td>
      </tr>
      <tr>
        <td>Mature</td>
        <td style="text-align: right">0.38</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result">
  <summary><h3 id="pigmentation">Pigmentation</h3></summary>

  <div style="text-align: center;">
    <figure>
        <img src="/assets/images/female-attraction/visualized-pigmentation.png" alt="Pigmentation factor illustration" width="600" />
    </figure>
</div>

  <p>This is a factor that I was not expecting, yet it’s fairly straightforward: it represents a preference for darker coloration over lighter coloration. ‘Dark Skin’, ‘Dark Eyes’, and ‘Black Hair’ all load positively, while ‘Fair Skin’, ‘Light Hair’, and ‘Light Eyes’ load negatively. In other words, all the items dealing with light/dark coloration are present and have exactly the loading signs you would expect. The only slight oddity is that ‘Black Hair’ loads considerably more weakly than the skin, hair-lightness, and eye-color items. I’m not sure why. Other than that, this seems to be an unusually clean preference for darker pigmentation across skin, hair, and eyes.</p>

  <p><strong>Factor Loadings</strong></p>

  <table>
    <thead>
      <tr>
        <th>Trait</th>
        <th style="text-align: right">Loading</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Fair Skin</td>
        <td style="text-align: right">-0.60</td>
      </tr>
      <tr>
        <td>Dark Skin</td>
        <td style="text-align: right">0.53</td>
      </tr>
      <tr>
        <td>Light Hair</td>
        <td style="text-align: right">-0.49</td>
      </tr>
      <tr>
        <td>Dark Eyes</td>
        <td style="text-align: right">0.41</td>
      </tr>
      <tr>
        <td>Light Eyes</td>
        <td style="text-align: right">-0.37</td>
      </tr>
      <tr>
        <td>Black Hair</td>
        <td style="text-align: right">0.22</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result">
  <summary><h3 id="adiposity">Adiposity</h3></summary>

  <div style="text-align: center;">
    <figure>
        <img src="/assets/images/female-attraction/visualized-adiposity.png" alt="Adiposity factor illustration" width="600" />
    </figure>
</div>

  <p>This is another straightforward factor: it represents a preference for fat men, as evidenced by the positive loadings of ‘Comfortably Overweight’ and ‘Heavyset’ and the negative loading of ‘Slender’. Importantly, items such as ‘Athletic’ and ‘Bulging Muscles’ don’t load on this factor, suggesting that this isn’t simply a preference for larger or more physically substantial men in general. Rather, it seems specifically to capture a preference for greater body fat and a heavier build.</p>

  <p><strong>Factor Loadings</strong></p>

  <table>
    <thead>
      <tr>
        <th>Trait</th>
        <th style="text-align: right">Loading</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Comfortably Overweight</td>
        <td style="text-align: right">0.58</td>
      </tr>
      <tr>
        <td>Heavyset</td>
        <td style="text-align: right">0.58</td>
      </tr>
      <tr>
        <td>Slender</td>
        <td style="text-align: right">-0.26</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result">
  <summary><h3 id="bohemianism">Bohemianism</h3></summary>

  <div style="text-align: center;">
    <figure>
        <img src="/assets/images/female-attraction/visualized-bohemianism.png" alt="Bohemianism factor illustration" width="600" />
    </figure>
</div>

  <p>This factor indicates a preference for an artistic, expressive, unconventional presentation, or a sort of bohemian type in that it is both artistic and socially unconventional. The artistic side of the factor can be seen in items such as ‘Artistic’ and ‘Musically Talented’ while ‘Long Hair’ fits naturally with the same alternative aesthetic. At the other end are items associated with more conventional forms of masculinity and presentation: ‘Men in Uniform’, ‘Jocks’, ‘Suits and Ties’, ‘Cologne’, and ‘Guns’ all load negatively. Taken together, the factor seems to contrast an artistic, expressive, alternative masculinity with a more conventional, institutional, and stereotypically masculine presentation.</p>

  <p><strong>Factor Loadings</strong></p>

  <table>
    <thead>
      <tr>
        <th>Trait</th>
        <th style="text-align: right">Loading</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Men in Uniform</td>
        <td style="text-align: right">-0.54</td>
      </tr>
      <tr>
        <td>Long Hair</td>
        <td style="text-align: right">0.50</td>
      </tr>
      <tr>
        <td>Guns</td>
        <td style="text-align: right">-0.47</td>
      </tr>
      <tr>
        <td>Jocks</td>
        <td style="text-align: right">-0.45</td>
      </tr>
      <tr>
        <td>Artistic</td>
        <td style="text-align: right">0.42</td>
      </tr>
      <tr>
        <td>Musically Talented</td>
        <td style="text-align: right">0.33</td>
      </tr>
      <tr>
        <td>Cologne</td>
        <td style="text-align: right">-0.32</td>
      </tr>
      <tr>
        <td>Cocky Attitude</td>
        <td style="text-align: right">-0.32</td>
      </tr>
      <tr>
        <td>Suits and Ties</td>
        <td style="text-align: right">-0.31</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result">
  <summary><h3 id="virility">Virility</h3></summary>

  <div style="text-align: center;">
    <figure>
        <img src="/assets/images/female-attraction/visualized-virility.png" alt="Virility factor illustration" width="600" />
    </figure>
</div>

  <p>This factor indicates a preference for sexually dimorphic masculinity. Most of the items relate to the physical side of things: broad rather than narrow shoulders, broad rather than weak jaws, deep voices, greater musculature, tall rather than short stature, and larger rather than smaller genitals. The items relating to personality point in the same direction: ‘Dominant’ loads positively while ‘Shy’ loads negatively, suggesting a preference for a stereotypically masculine personality alongside the stereotypically masculine body.</p>

  <p><strong>Factor Loadings</strong></p>

  <table>
    <thead>
      <tr>
        <th>Trait</th>
        <th style="text-align: right">Loading</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Narrow Shoulders</td>
        <td style="text-align: right">-0.69</td>
      </tr>
      <tr>
        <td>Broad Shoulders</td>
        <td style="text-align: right">0.63</td>
      </tr>
      <tr>
        <td>Broad Jaws</td>
        <td style="text-align: right">0.61</td>
      </tr>
      <tr>
        <td>Deep Voices</td>
        <td style="text-align: right">0.58</td>
      </tr>
      <tr>
        <td>Weak Chins</td>
        <td style="text-align: right">-0.55</td>
      </tr>
      <tr>
        <td>Short</td>
        <td style="text-align: right">-0.55</td>
      </tr>
      <tr>
        <td>Bulging Muscles</td>
        <td style="text-align: right">0.53</td>
      </tr>
      <tr>
        <td>Small Genitals</td>
        <td style="text-align: right">-0.49</td>
      </tr>
      <tr>
        <td>Dominant</td>
        <td style="text-align: right">0.48</td>
      </tr>
      <tr>
        <td>Shy</td>
        <td style="text-align: right">-0.47</td>
      </tr>
      <tr>
        <td>Virile</td>
        <td style="text-align: right">0.46</td>
      </tr>
      <tr>
        <td>Extremely Tall</td>
        <td style="text-align: right">0.39</td>
      </tr>
      <tr>
        <td>Very Large Genitals</td>
        <td style="text-align: right">0.35</td>
      </tr>
      <tr>
        <td>Rugged Looks</td>
        <td style="text-align: right">0.31</td>
      </tr>
    </tbody>
  </table>

</details>

<h3 id="comparison-with-apple-pie">Comparison with Apple Pie</h3>

<p>Apple Pie’s factors map most closely onto the factors found here as follows:</p>

<table>
  <thead>
    <tr>
      <th>Apple Pie factor</th>
      <th>Closest factor(s) here (sign relative to first-named pole)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Artists vs Heroes</td>
      <td>Bohemianism (+); Virility (-), Affluence (-)</td>
    </tr>
    <tr>
      <td>Pretty Boys vs Masculine Men</td>
      <td>Boyishness (+); Virility (-), Ruggedness (-)</td>
    </tr>
    <tr>
      <td>Romantic Comedy</td>
      <td>Companionability (+)</td>
    </tr>
    <tr>
      <td>Kind vs Dangerous</td>
      <td>Companionability (+); Ruggedness (-)</td>
    </tr>
    <tr>
      <td>Ditsy Teens vs Sophisticated Gentlemen</td>
      <td>Boyishness (+); Intellect (-)</td>
    </tr>
  </tbody>
</table>

<h2 id="another-aside-about-acquiescence">An(other) Aside About Acquiescence</h2>

<p>At this point, I should mention <a href="/acquiescence/">acquiescence bias</a>. Acquiescence bias is the tendency to agree with statements in a questionnaire regardless of what those statements actually assert. For example, in this dataset, there are traits that are natural opposites: Tall vs. Short, Fair skin vs. Dark skin, Young vs Old. A respondent high in acquiescence would say that they strongly prefer both tall men <em>and</em> short men, young men <em>and</em> old men, and so on. Here’s a visualization:</p>

<div style="text-align: center;">
    <figure>
        <img src="/assets/images/female-attraction/acquiescence-opposite-pairs.png" width="600" />
    </figure>
</div>

<p>Now, it’s perfectly fine for someone to prefer many traits, even if they are ‘opposites’, but this doesn’t help us determine the factors, that is, the substantive sources of covariation among preferences. In fact, acquiescence will obscure those factors, because a general tendency to endorse items will make all indicators <em>more</em> positively correlated with one another, regardless of their content.</p>

<p>For this reason, I’ve modeled an acquiescence factor that affects all items equally, allowing this general endorsement tendency to be separated from the substantive preference factors. This seems to have been warranted, seeing as the common loading on the acquiescence factor was about 0.20 (90% posterior HDI: [0.18, 0.21]).</p>

<h2 id="factor-correlation-matrix">Factor Correlation Matrix</h2>

<div style="text-align: center;">
    <figure>
        <img src="/assets/images/female-attraction/factor-correlation.png" alt="Correlations among the nine preference factors" width="800" />
    </figure>
</div>

<p>The correlations between preference factors make intuitive sense. Virility is a preference for a stereotypically masculine man, and so naturally has a negative correlation with Bohemianism, which is a preference for an <em>unconventional</em> man. Preferences for youthful prettiness (Boyishness) are negatively correlated with preferences for roughness and danger (Ruggedness), as well as with the aforementioned stereotypical manliness (Virility). The negative correlation between preferences for Affluence and Bohemianism makes sense if one keeps in mind the stereotype of the starving artist. As for positive correlations, the one between preferences for Ruggedness and Virility makes sense, since both are, in their own way, preferences for a particular kind of masculinity.</p>

<h2 id="takeaways">Takeaways</h2>

<p>To be honest, I’m not sure what to put here. I think I’ve said everything that needs to be said. I have no further opinions. Thanks for reading.</p>

<p><em>Thanks to <a href="https://thingstoread.substack.com/">Apple Pie</a> for sharing the data!</em></p>

<h2 id="appendix">Appendix</h2>

<h3 id="correlation-fishing">Correlation Fishing</h3>

<p>Let’s take a look at some of the correlations<sup id="fnref:correlation" role="doc-noteref"><a href="#fn:correlation" class="footnote" rel="footnote">1</a></sup> between preferences and other traits. First, preferences and personality.</p>

<div style="text-align: center;">
    <figure>
        <img src="/assets/images/female-attraction/preference-personality.png" alt="Preferences and personality factor-score regressions" width="600" />
    </figure>
</div>

<p>There isn’t much going on here, so let’s move on to the correlations between preferences and politics.</p>

<p>The political items used were:</p>

<ul>
  <li>Hierarchy: It is important for society to have a hierarchy.</li>
  <li>Religion: Religion is very important in my life.</li>
  <li>Euthanasia: People suffering from incurable diseases should have the right to be put painlessly to death.</li>
  <li>Equality: We would have fewer problems if we treated people more equally.</li>
  <li>Free speech: Free speech is important, and should be protected even if some people’s feelings are hurt.</li>
  <li>Compulsory schooling: Teenagers should be legally required to go to school.</li>
</ul>

<div style="text-align: center;">
    <figure>
        <img src="/assets/images/female-attraction/preference-politics.png" alt="Preferences and politics factor-score regressions" width="600" />
    </figure>
</div>

<p>Responses to the hierarchy item were positively correlated with the Virility factor and negatively correlated with the Bohemianism factor (as well as with the Companionability and Pigmentation factors). It seems like women who think societal hierarchies are important have stronger preferences for conventionally masculine men, along with stronger preferences for fairer-skinned men.</p>

<p>Responses to the equality item were positively correlated with the Bohemianism factor (and the Pigmentation factor) and negatively correlated with the Affluence factor. All in all, these correlations seem to be broadly the reverse of the pattern for the hierarchy item.</p>

<p>Other correlations were weaker, though one possibly interesting finding is the positive correlation between endorsing euthanasia and preference for Intellect.</p>

<p>Finally, social desirability:</p>

<div style="text-align: center;">
    <figure>
        <img src="/assets/images/female-attraction/social-desirability.png" alt="Social desirability factor-score regressions" width="600" />
    </figure>
</div>

<p>It seems that the most socially desirable preference is for a man who would make a good lead in a romantic comedy?</p>

<style>
  .measurement-model { margin: 0.8rem 0 2.5rem; }
  .measurement-model .model-figure { margin: 0 0 1.5rem; overflow-x: auto; }
  .measurement-model .model-figure img { display: block; width: 100%; min-width: 720px; max-width: 900px; height: auto; margin: 0 auto; }
  .measurement-model h4 { margin: 1.3rem 0 0.55rem; font-size: 1.08em; }
  .measurement-model .model-factor-grid { display: grid; grid-template-columns: repeat(2, minmax(0, 1fr)); column-gap: 1.5rem; border-top: 1px solid #d0d7de; }
  .measurement-model .model-factor { min-width: 0; padding: 0.75rem 0 0.65rem; border-bottom: 1px solid #d0d7de; }
  .measurement-model .model-factor h5 { margin: 0 0 0.35rem; font-size: 1em; }
  .measurement-model.politics-model .model-factor-grid { grid-template-columns: minmax(0, 1fr); }
  .measurement-model.politics-model .model-factor { display: grid; grid-template-columns: 11rem minmax(0, 1fr); gap: 1rem; }
  .measurement-model.politics-model .model-factor ul { display: grid; grid-template-columns: repeat(2, minmax(0, 1fr)); gap: 0 1.5rem; }
  .measurement-model ul { margin: 0; padding: 0; list-style: none; }
  .measurement-model li { margin: 0.25rem 0; line-height: 1.5; }
  .measurement-model .model-estimate { font-weight: 600; font-variant-numeric: tabular-nums; white-space: nowrap; }
  .measurement-model .model-hdi { color: #57606a; font-size: 0.88em; white-space: nowrap; }
  .measurement-model .model-shared { margin: 1rem 0 1.25rem; padding-left: 0.75rem; border-left: 3px solid #3069a5; }
  .measurement-model .model-method-list { display: grid; grid-template-columns: repeat(2, minmax(0, 1fr)); gap: 0.15rem 1.5rem; padding-top: 0.4rem; border-top: 1px solid #d0d7de; }
  .measurement-model .model-correlations { margin: 0; border-top: 1px solid #d0d7de; }
  .measurement-model .model-correlations > div { display: flex; align-items: center; justify-content: space-between; gap: 1rem; padding: 0.45rem 0; border-bottom: 1px solid #d0d7de; }
  .measurement-model .model-correlations dt { margin: 0; font-weight: 400; font-style: normal; }
  .measurement-model .model-correlations dd { margin: 0; white-space: nowrap; }
  @media (max-width: 680px) {
    .measurement-model .model-factor-grid, .measurement-model .model-method-list { grid-template-columns: minmax(0, 1fr); }
    .measurement-model .model-correlations > div { display: block; }
    .measurement-model.politics-model .model-factor { display: block; }
    .measurement-model.politics-model .model-factor ul { grid-template-columns: minmax(0, 1fr); }
  }
</style>

<h3 id="factor-model-of-personality">Factor Model of Personality</h3>

<p>I modelled the HEXACO items as resulting from 8 orthogonal factors: the 6 HEXACO factors, an acquiescence factor that all items loaded on equally, and a social desirability factor.</p>

<div class="measurement-model">
  <figure class="model-figure">
    <img src="/assets/images/female-attraction/hexaco-factor-model.svg" alt="Six orthogonal HEXACO factors, orthogonal acquiescence, and orthogonal social desirability. Each HEXACO factor loads on two signed items; acquiescence has one shared loading and social desirability has a signed loading per item." />
  </figure>
  <h4>HEXACO item loadings</h4>
  <div class="model-factor-grid"><section class="model-factor"><h5>Honesty-Humility</h5><ul><li><span class="model-item">Honest, Honorable:</span> <span class="model-estimate">0.20</span> <span class="model-hdi">(90% HDI: [0.09, 0.29])</span></li><li><span class="model-item">Amoral, Carefree:</span> <span class="model-estimate">-0.31</span> <span class="model-hdi">(90% HDI: [-0.42, -0.20])</span></li></ul></section><section class="model-factor"><h5>Emotionality</h5><ul><li><span class="model-item">Sentimental, Soft:</span> <span class="model-estimate">0.42</span> <span class="model-hdi">(90% HDI: [0.33, 0.51])</span></li><li><span class="model-item">Rugged, Unemotional:</span> <span class="model-estimate">-0.44</span> <span class="model-hdi">(90% HDI: [-0.53, -0.34])</span></li></ul></section><section class="model-factor"><h5>Extraversion</h5><ul><li><span class="model-item">Active, Talkative:</span> <span class="model-estimate">0.62</span> <span class="model-hdi">(90% HDI: [0.55, 0.69])</span></li><li><span class="model-item">Introverted, Withdrawn:</span> <span class="model-estimate">-0.63</span> <span class="model-hdi">(90% HDI: [-0.70, -0.56])</span></li></ul></section><section class="model-factor"><h5>Agreeableness</h5><ul><li><span class="model-item">Easygoing, Calm:</span> <span class="model-estimate">0.49</span> <span class="model-hdi">(90% HDI: [0.39, 0.58])</span></li><li><span class="model-item">Tense, Hot-Tempered:</span> <span class="model-estimate">-0.49</span> <span class="model-hdi">(90% HDI: [-0.59, -0.40])</span></li></ul></section><section class="model-factor"><h5>Conscientiousness</h5><ul><li><span class="model-item">Organized, Thorough:</span> <span class="model-estimate">0.40</span> <span class="model-hdi">(90% HDI: [0.30, 0.49])</span></li><li><span class="model-item">Disorganized, Careless:</span> <span class="model-estimate">-0.37</span> <span class="model-hdi">(90% HDI: [-0.46, -0.27])</span></li></ul></section><section class="model-factor"><h5>Openness</h5><ul><li><span class="model-item">Interested in Art, Deep:</span> <span class="model-estimate">0.44</span> <span class="model-hdi">(90% HDI: [0.34, 0.54])</span></li><li><span class="model-item">Down-to-Earth, Unimaginative:</span> <span class="model-estimate">-0.44</span> <span class="model-hdi">(90% HDI: [-0.56, -0.35])</span></li></ul></section></div>
  <p class="model-shared"><strong>Acquiescence:</strong> one shared loading of 0.035 <span class="model-hdi">(90% HDI: [0.013, 0.059])</span>.</p>
  <h4>Social desirability item loadings</h4>
  <ul class="model-method-list">
    <li><span class="model-item">Honest, Honorable:</span> <span class="model-estimate">0.55</span> <span class="model-hdi">(90% HDI: [0.47, 0.64])</span></li>
      <li><span class="model-item">Disorganized, Careless:</span> <span class="model-estimate">-0.48</span> <span class="model-hdi">(90% HDI: [-0.57, -0.39])</span></li>
      <li><span class="model-item">Organized, Thorough:</span> <span class="model-estimate">0.43</span> <span class="model-hdi">(90% HDI: [0.34, 0.52])</span></li>
      <li><span class="model-item">Active, Talkative:</span> <span class="model-estimate">0.20</span> <span class="model-hdi">(90% HDI: [0.11, 0.28])</span></li>
      <li><span class="model-item">Amoral, Carefree:</span> <span class="model-estimate">-0.19</span> <span class="model-hdi">(90% HDI: [-0.31, -0.08])</span></li>
      <li><span class="model-item">Introverted, Withdrawn:</span> <span class="model-estimate">-0.16</span> <span class="model-hdi">(90% HDI: [-0.24, -0.07])</span></li>
      <li><span class="model-item">Sentimental, Soft:</span> <span class="model-estimate">0.16</span> <span class="model-hdi">(90% HDI: [0.05, 0.25])</span></li>
      <li><span class="model-item">Easygoing, Calm:</span> <span class="model-estimate">0.15</span> <span class="model-hdi">(90% HDI: [0.05, 0.25])</span></li>
      <li><span class="model-item">Tense, Hot-Tempered:</span> <span class="model-estimate">-0.12</span> <span class="model-hdi">(90% HDI: [-0.22, -0.03])</span></li>
      <li><span class="model-item">Interested in Art, Deep:</span> <span class="model-estimate">0.11</span> <span class="model-hdi">(90% HDI: [0.03, 0.21])</span></li>
      <li><span class="model-item">Rugged, Unemotional:</span> <span class="model-estimate">-0.08</span> <span class="model-hdi">(90% HDI: [-0.16, -0.01])</span></li>
      <li><span class="model-item">Down-to-Earth, Unimaginative:</span> <span class="model-estimate">-0.06</span> <span class="model-hdi">(90% HDI: [-0.14, -0.002])</span></li>
  </ul>
</div>

<p>The loadings on the HEXACO factors are moderate, though some, such as Honesty-Humility, are lower than I’d like. Overall, though, the model looks decent, and there doesn’t seem to be much acquiescence occurring.</p>

<h3 id="factor-model-of-politics">Factor Model of Politics</h3>

<p>Apple Pie chose the six politics items to represent the three factors of political views that they found in their earlier research: Conservatism, Tough-mindedness, and Libertarianism. Apple Pie says the factors are orthogonal, but I thought I’d test that assumption anyway by allowing the factors to correlate. I also included an orthogonal acquiescence factor.</p>

<div class="measurement-model politics-model">
  <figure class="model-figure">
    <img src="/assets/images/female-attraction/politics-factor-model.svg" alt="Three correlated political factors each load on two signed items. Independent acquiescence has one shared positive loading on all six items." />
  </figure>
  <h4>Political item loadings</h4>
  <div class="model-factor-grid"><section class="model-factor"><h5>Conservatism</h5><ul><li><span class="model-item">It is important for society to have a hierarchy:</span> <span class="model-estimate">0.28</span> <span class="model-hdi">(90% HDI: [0.17, 0.39])</span></li><li><span class="model-item">Religion is very important in my life:</span> <span class="model-estimate">0.29</span> <span class="model-hdi">(90% HDI: [0.18, 0.40])</span></li></ul></section><section class="model-factor"><h5>Tough-mindedness</h5><ul><li><span class="model-item">People suffering from incurable diseases should have the right to be put painlessly to death:</span> <span class="model-estimate">0.11</span> <span class="model-hdi">(90% HDI: [0.03, 0.21])</span></li><li><span class="model-item">We would have fewer problems if we treated people more equally:</span> <span class="model-estimate">-0.11</span> <span class="model-hdi">(90% HDI: [-0.20, -0.01])</span></li></ul></section><section class="model-factor"><h5>Libertarianism</h5><ul><li><span class="model-item">Free speech is important, and should be protected even if some people&#39;s feelings are hurt:</span> <span class="model-estimate">0.23</span> <span class="model-hdi">(90% HDI: [0.12, 0.36])</span></li><li><span class="model-item">Teenagers should be legally required to go to school:</span> <span class="model-estimate">-0.21</span> <span class="model-hdi">(90% HDI: [-0.31, -0.10])</span></li></ul></section></div>
  <p class="model-shared"><strong>Acquiescence:</strong> one shared loading of 0.053 <span class="model-hdi">(90% HDI: [0.015, 0.089])</span>.</p>
  <h4>Factor correlations</h4>
  <dl class="model-correlations">
    <div><dt>Conservatism / Tough-mindedness</dt><dd><span class="model-estimate">0.001</span> <span class="model-hdi">(90% HDI: [-0.123, 0.130])</span></dd></div>
      <div><dt>Conservatism / Libertarianism</dt><dd><span class="model-estimate">0.015</span> <span class="model-hdi">(90% HDI: [-0.101, 0.132])</span></dd></div>
      <div><dt>Tough-mindedness / Libertarianism</dt><dd><span class="model-estimate">0.015</span> <span class="model-hdi">(90% HDI: [-0.100, 0.130])</span></dd></div>
  </dl>
</div>

<p>Unfortunately, it’s not a good model: the factor loadings are very weak. Perhaps this is because there are substantial cross-loadings; for example, we might expect the euthanasia item to also load on Libertarianism. I’d also expect the hierarchy and equality items to be more negatively correlated than would be implied by the factor correlations and loadings alone. Regardless, modeling the factors seems like a bad idea here, so I look at correlations using the raw items instead.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:correlation" role="doc-endnote">
      <p>Really, these are standardized regression coefficients between factor scores. Because the scores are calculated rather than jointly modeled, I expect the associations to be attenuated to some extent. <a href="#fnref:correlation" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name></name></author><category term="human-traits" /><summary type="html"><![CDATA[In Factors of Attraction Toward Men, Apple Pie conducted surveys on romantic preferences, asking participants to rate several traits based on how attractive they are. The following is a replication of the analysis using confirmatory factor analysis (CFA) and item response theory (IRT), rather than using principal component analysis (PCA) as Apple Pie did. PCA forces factors to be orthogonal, while CFA allows factors to correlate. IRT also provides more item-level information.]]></summary></entry><entry><title type="html">EA Is Pretty Left-Wing, Actually</title><link href="https://stanichor.net/ea-left-wing/" rel="alternate" type="text/html" title="EA Is Pretty Left-Wing, Actually" /><published>2026-09-21T00:00:00+00:00</published><updated>2026-09-21T00:00:00+00:00</updated><id>https://stanichor.net/ea-left-wing</id><content type="html" xml:base="https://stanichor.net/ea-left-wing/"><![CDATA[<blockquote>
  <p>“You called yourself luck,” Akua said, “but that is a lie, Intercessor. You are not a blind roll of the dice. You take sides.”</p>

  <p>“I’ve helped both sides of the Game,” the Intercessor dismissed, “I-“</p>

  <p>“You help Good,” Akua said. “When you have the choice, that is the truth of you.”</p>
</blockquote>

<p>— ErraticErrata, A Practical Guide to Evil (Book 7: <a href="https://practicalguidetoevil.wordpress.com/2022/02/18/chapter-68-hallow-hollow/">Chapter 68: Hallow; Hollow</a>)</p>

<p>Now, the title of this post shouldn’t be a surprise to anyone, but ever since the Trump administration began promoting “<a href="https://x.com/DoWCTO/status/2099536442594582922">Americanism, not Effective Altruism</a>” and thrust EA even further into the political spotlight, I’ve seen EA-aligned people claiming that EA isn’t inherently left-wing or right-wing. I’ll get into the “inherently” part later, but right now, I want to argue that, in practice, EA is pretty obviously left-wing<sup id="fnref:america" role="doc-noteref"><a href="#fn:america" class="footnote" rel="footnote">1</a></sup>.</p>

<h2 id="political-leanings-of-effective-altruists">Political Leanings of Effective Altruists</h2>

<p>An obvious place to look is the political leanings of effective altruists themselves. Looking at the (semi-)yearly EA Surveys, <a href="https://stanichor.net/rat-demographics/#effective-altruism-3">it has never been the case that more than 5% of effective altruists identify as either right or center-right</a>. Meanwhile, the proportion of effective altruists identifying as left, excluding for the moment center-left, varies between 25% and 40%. Including the center-left brings that figure to between 70% and 85%. So left-leaning effective altruists outnumber right-leaning effective altruists by, at minimum, about 14 to 1. That’s <em>very</em> lopsided<sup id="fnref:libertarian" role="doc-noteref"><a href="#fn:libertarian" class="footnote" rel="footnote">2</a></sup>.</p>

<div style="text-align: center;">
    <figure>
        <img src="/assets/images/rat-demographics/ea-politics.png" width="700" alt="Stacked time-series chart showing the political identification of Effective Altruism survey respondents by year." />
    </figure>
</div>

<h2 id="political-leanings-of-ea-cause-areas">Political Leanings of EA Cause Areas</h2>

<p>What about the cause areas EA prioritizes? What partisan leanings do those issues have? The <a href="https://forum.effectivealtruism.org/posts/CKwDgZGLipchAoxtN/ea-survey-2024-cause-prioritization">major causes</a> EAs seem to care about are AI risk, global poverty and health, biosecurity, and animal welfare. Looking at <a href="https://forum.effectivealtruism.org/posts/brAkfA5AKTpuER7wv/pulse-wave-2-cause-prioritization-of-the-us-public">Rethink Priorities’ Pulse Wave 2</a>, Democrats give higher importance ratings than Republicans to every cause area surveyed, and also show higher support for donating to every cause area. The largest gaps are for climate change, civil rights, global health and development (GHD), and pandemic preparedness<sup id="fnref:covid" role="doc-noteref"><a href="#fn:covid" class="footnote" rel="footnote">3</a></sup>. The gaps are smallest for AI risk, cancer research, and nuclear weapons.</p>

<p>The partisan leaning of AI risk is currently… fluid, so I’ll set it aside for now.</p>

<p>Looking at other sources, a <a href="https://www.pewresearch.org/wp-content/uploads/sites/20/2025/04/pg_2025.05.01_us-engagement_report.pdf">2025 Pew Research poll</a> shows a 14-point partisan gap on whether the US should provide foreign aid in the form of medicine and medical supplies to developing countries, and a 34-point gap on foreign aid aimed at supporting economic development in developing countries. In <a href="https://www.pewresearch.org/politics/2018/11/29/conflicting-partisan-priorities-for-u-s-foreign-policy/">a 2018 Pew poll</a>, 32% of Democrats, compared to 12% of Republicans, said “helping improve living standards in developing nations” should be a top foreign-policy priority.</p>

<p>On pandemic preparedness, a <a href="https://yougov.com/en-us/articles/51911-five-years-after-what-americans-think-about-covid-19-pandemic-poll">2025 YouGov poll</a> found that 71% of Democrats thought the US needed to do more to prepare for future pandemics, compared to 31% of Republicans. In <a href="https://www.pewresearch.org/global/2024/04/23/what-are-americans-top-foreign-policy-priorities/">a 2024 Pew Research poll</a>, 63% of Democrats versus 41% of Republicans said “reducing the spread of infectious diseases” should be a top priority in US long-range foreign policy.</p>

<p>A <a href="https://faunalytics.org/public-acceptability-of-standard-u-s-animal-agriculture-practices/">2025 Faunalytics survey</a> found that Democrats, on average, rated standard animal-agriculture practices as more unacceptable than Republicans did. A <a href="https://www.aspca.org/sites/default/files/2023_industrial_ag_survey_results_report_052523_1.pdf">2023 ASPCA/Ipsos national survey</a> found that Democrats consistently showed higher support for humane-treatment policies.</p>

<p>So, all in all, looking at the partisan skew of major EA cause areas, Democrats consistently show more support for them.</p>

<h2 id="political-donations-and-candidacies">Political Donations and Candidacies</h2>

<p>Another way to get at the political leanings of a group is to look at where its political donations end up. Looking at some of the most prominent EA-associated donors:</p>

<ul>
  <li>In 2020, Dustin Moskovitz and Cari Tuna donated ~\$50 million to Democratic candidates and liberal groups, and \$0 to Republican candidates and conservative groups.</li>
  <li>In 2024, Dustin Moskovitz donated ~\$50 million to Democratic candidates and liberal groups, and \$0 to Republican candidates and conservative groups.</li>
  <li>In 2022, Protect Our Future, an SBF-funded PAC <a href="https://www.politico.com/minutes/congress/01-27-2022/new-pac-launches/">organized around pandemic preparedness</a>, spent \$21 million in independent expenditures, exclusively in Democratic House primaries.</li>
  <li>In 2022, <a href="https://www.citizensforethics.org/wp-content/uploads/2022/12/SBF-FEC-Complaint-FINAL.pdf">SBF publicly disclosed ~\$37 million in contributions to Democratic candidates and only $320,000 in contributions to Republicans</a>. However, he also later said that he had secretly been the source of approximately \$37 million, and potentially as much as \$80 million, in contributions to Republicans.</li>
</ul>

<p>What about actual EAs running for office? In 2022, Carrick Flynn, whom <em>The New Yorker</em> described as <a href="https://www.newyorker.com/magazine/2022/08/15/the-reluctant-prophet-of-effective-altruism?utm_source=chatgpt.com">the first explicitly EA-affiliated congressional candidate</a>, ran as a Democrat.</p>

<p>Admittedly, this evidence is weaker, since I’ve focused on a handful of major donors and the one explicitly EA-affiliated congressional candidate. But, given everything else we’ve seen, it is at least suggestive of the same overall pattern.</p>

<h2 id="is-ea-inherently-left-wing">Is EA <em>Inherently</em> Left-Wing?</h2>

<p>So, all in all, it seems to me that EA is, in practice, <em>very</em> left-leaning. But is it <em>inherently</em> left-wing?</p>

<p>I’m not a fan of theories that try to reduce the two American political sides to a single simple principle<sup id="fnref:thrive-survive" role="doc-noteref"><a href="#fn:thrive-survive" class="footnote" rel="footnote">4</a></sup>. They seem to overlook the fact that, in the American system, there aren’t two inherent political sides. There are two major parties, for historical and structural reasons, but those parties are better thought of as coalitions of various groups and interests. Additionally, the partisan alignment of particular issues can vary for somewhat contingent reasons<sup id="fnref:pandemics" role="doc-noteref"><a href="#fn:pandemics" class="footnote" rel="footnote">5</a></sup>.</p>

<p>So I wouldn’t say EA is <em>inherently</em> anything. But, given the political environment we actually live in, it does seem likely that EA was always going to end up left-leaning. Focusing heavily on the well-being of foreigners and animals is already politically left-leaning in the contemporary US. Pandemic preparedness seems like a more unfortunate case of contingent polarization.</p>

<p>None of this is meant as a criticism of EA. It doesn’t seem to me that EA has been deliberately trying to become overtly partisan. The partisan leanings of its cause areas and of the people who make up the movement are what they are. What, if anything, can be done about that? I don’t know.</p>

<p>But I do think it’s important not to deceive ourselves about the situation, especially as EA is thrust further into the political spotlight. If EA looks overwhelmingly left-leaning from the outside, that <em>matters</em> in an environment where people are increasingly suspicious that institutions associated with the political out-group are really vehicles for the out-group’s interests. Whatever EA does in response, it seems better to start from an accurate picture of how the movement is actually positioned.</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:america" role="doc-endnote">
      <p>In this post, I’ll be referring to left and right in the American sense. I’m not well-versed enough in the politics of other countries to make the same claims about them. The United States also seems like the most important country to focus on here, given its outsized role in many of the areas EA cares about. <a href="#fnref:america" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:libertarian" role="doc-endnote">
      <p>What about the proportion of effective altruists who identify as libertarian? It varies between 5% and 15%, which seems to roughly match the proportion of Americans who identify as libertarian. That said, EA libertarians seem more likely to be “liberaltarians” (socially liberal with free-market preferences) than, say, right-libertarians. <a href="#fnref:libertarian" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:covid" role="doc-endnote">
      <p>This is probably partly downstream of the broader partisan polarization surrounding COVID. <a href="#fnref:covid" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:thrive-survive" role="doc-endnote">
      <p>E.g., <a href="https://slatestarcodex.com/2013/03/04/a-thrivesurvive-theory-of-the-political-spectrum/">Scott Alexander’s thrive/survive theory</a>. <a href="#fnref:thrive-survive" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:pandemics" role="doc-endnote">
      <p>For example, during the 2014 Ebola scare, Republicans were more supportive than Democrats of quarantining travelers returning from affected countries. During COVID, however, Democrats were generally more supportive of restrictive public-health measures than Republicans. That said, the policies proposed during Ebola and COVID differed in <em>who</em> they targeted: travelers and foreigners in the former case, ordinary Americans in the latter. <a href="#fnref:pandemics" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name></name></author><category term="rationalism-ea" /><summary type="html"><![CDATA[“You called yourself luck,” Akua said, “but that is a lie, Intercessor. You are not a blind roll of the dice. You take sides.” “I’ve helped both sides of the Game,” the Intercessor dismissed, “I-“ “You help Good,” Akua said. “When you have the choice, that is the truth of you.”]]></summary></entry><entry><title type="html">Foundations of Mogging: Chapter 3</title><link href="https://stanichor.net/mog/" rel="alternate" type="text/html" title="Foundations of Mogging: Chapter 3" /><published>2026-09-21T00:00:00+00:00</published><updated>2026-09-21T00:00:00+00:00</updated><id>https://stanichor.net/mog</id><content type="html" xml:base="https://stanichor.net/mog/"><![CDATA[<p>Audience is a reward for mogging. Necessarily so. To see this, let’s take a look at the etymology of ‘mog’. It derives from the noun ‘AMOG’, or Alpha Male Of the Group. The Alpha Male’s social dominance causes most interaction within a group to be focused on him. The group is, in a sense, his audience.</p>

<p>To understand mogging, we must first grasp that it is a transfer of status. That is, most forms of mogging attempt to raise or lower the status of individuals via game-like structures, with defined roles and a <em>structurally</em> predictable script. There is always a mogger, a moggee, and, crucially, an <em>audience</em> (which can coincide with the moggee, by design or accident). The moggee may or may not be present. So, there are at least three roles in a mog, of which the role of audience may be played by a group. This gives us three basic forms of mogging.</p>

<h3 id="autist-two-person-mogging">Autist (Two-Person) Mogging</h3>

<p>Two-person mogging is Autist mogging. If you attempt a mog with just one other person present, and you are only capable of gaining status if the other person loses status (as is necessary, since status is zero-sum), you get a terminally stupid situation that only Autists will attempt to enact. Note that two-person moggings with an absent moggee are really three-role moggings, so they don’t count.</p>

<p>In a two-person situation, you either get non-adversarial self-mogs (which reinforce existing status), or an adversarial mog. Since social proof works by majority vote, two-person adversarial mogs cannot work unless the moggee laughs at himself, accepting a loss of status. If the moggee fights back, with no neutral audience to determine status changes, you get a pointless game of one-upmanship. This is why the mog is autist. It boils down to he-said-she-said. Without an audience, there is no mogging, and there is no meaningful transfer of AMOG status.</p>

<h3 id="sociopath-one-person-mogging">Sociopath (One-Person) Mogging</h3>

<p>One-person mogging is Sociopath mogging, and is psychologically more complex. It can only happen when the mogger and <em>audience</em> are the same person (which replaces social proof with individual judgment), and everybody else present is a moggee, often unaware that they are being mogged. Sociopath mogging often possesses push-button cruelty, in the sense of a cat-and-mouse “pushing buttons” exercise in viciousness for private pleasure, and is “objective” mogging in that sense, since the moggee’s loss of status is reflective of an “objective” fact about the world.</p>

<p>But mogging gets <em>really</em> interesting when there are more than two active participants. That gets you to Normie mogging, an engine of social status transfers.</p>

<h3 id="normie-group-mogging">Normie (Group) Mogging</h3>

<p>In Normie mogging, as with Autist mogging, the status transfer comes via the judgment of somebody other than the mogger (audience != mogger != moggee). But unlike the two-person stalemate that is the norm in Autist mogging, Normie mogging usually creates clear outcomes because democratic social proof can work. The smallest meaningful Normie group is three people (including some special cases where the moggee is absent, and both mogger and audience accept the proposed status transfer, providing a 2/3 social-proof majority).</p>

<p>In general, the transfer of status depends <em>entirely</em> on the reactions of the audience. What breaks the status stalemate in groups of three is that meaningful status movements can occur. Alphas can become Betas, and Betas can become Alphas. Due to additive effects, bigger groups are even better: if three people, who collectively are all Betas, are present, then we can add up all the micro-shifts in status and get a macro-shift in status.</p>

<p>How does mogging create status transfers? Consider a simple three-person situation. A attempts a mog of B. Without C present, you’d get Autist dynamics. But if C is present and deems the mog successful, he bonds with the mogger and transfers status to A, in the sense that B (and C, to an extent) are now envious of A. If he frowns or otherwise indicates that the mog was in bad taste, he bonds with the moggee and transfers status in the other direction (A is now envious of B). And crucially, if he does not react, no status is transferred.</p>

<p>Laugh/frown votes are a powerful weapon for the passive members of any situational group. In the most extreme situation – the smallest possible group of three people – there is enormous power wielded by just one person.</p>

<p><em>Apologies to Venkat</em></p>]]></content><author><name></name></author><category term="uncategorized" /><summary type="html"><![CDATA[Audience is a reward for mogging. Necessarily so. To see this, let’s take a look at the etymology of ‘mog’. It derives from the noun ‘AMOG’, or Alpha Male Of the Group. The Alpha Male’s social dominance causes most interaction within a group to be focused on him. The group is, in a sense, his audience.]]></summary></entry><entry><title type="html">Factor Analysis of Segregation Measures: A Replication</title><link href="https://stanichor.net/segregation/" rel="alternate" type="text/html" title="Factor Analysis of Segregation Measures: A Replication" /><published>2026-09-19T00:00:00+00:00</published><updated>2026-09-19T00:00:00+00:00</updated><id>https://stanichor.net/segregation</id><content type="html" xml:base="https://stanichor.net/segregation/"><![CDATA[<p>While reading <a href="https://www.tandfonline.com/doi/abs/10.1080/00222500601188486"><em>A Wealth and Status-Based Model of Residential Segregation</em></a>, I noticed its discussion of Massey and Denton (1988), who reviewed 20 segregation indices and applied them to 1980 census data from 60 metropolitan statistical areas (MSAs). Using factor analysis, Massey and Denton classified these indices into five underlying dimensions: evenness, exposure, clustering, centralization, and concentration.</p>

<p>However, they appear to have chosen five factors <em>a priori</em>, based on their review of the literature, rather than using the data to determine how many factors to retain. So, I decided to replicate their analysis using parallel analysis to determine the number of factors empirically.</p>

<p>I also broadened the analysis in two ways. Instead of restricting the sample to 60 MSAs, I used every contemporary metropolitan area with at least 1,000 members of the focal racial group and at least 1,000 non-Hispanic White-alone residents<sup id="fnref:threshold" role="doc-noteref"><a href="#fn:threshold" class="footnote" rel="footnote">1</a></sup>. And instead of examining only 1980, I repeated the analysis for 1990, 2000, 2010, and 2020 to see how stable the findings are over four decades. The national-sample result is remarkably consistent: parallel analysis retains three factors in every census year. The fixed panel corresponding to the original 60 metropolitan identities produces a more conservative result, as discussed below.</p>

<p>Massey and Denton’s five dimensions describe the different ways in which residential segregation might occur:</p>

<ul>
  <li>Evenness: How evenly distributed is a group across residential areas?</li>
  <li>Exposure: How much potential contact do members of one group have with members of another group?</li>
  <li>Concentration: How much physical space does a group occupy?</li>
  <li>Centralization: How concentrated is a group in the city center?</li>
  <li>Clustering: To what degree does a group disproportionately live in contiguous areas?</li>
</ul>

<h2 id="how-many-factors-are-there">How many factors are there?</h2>

<p>The national sample consistently yields three factors, not five.</p>

<figure>
    <img src="/assets/images/segregation/parallel_analysis_overview__national__r2km__extended22.png" width="1000" alt="Five-panel parallel-analysis chart for the national metropolitan sample in 1980, 1990, 2000, 2010, and 2020 using a 2-kilometer spatial radius. In every year, the first three observed eigenvalues exceed the 95th-percentile null eigenvalues, so three components are retained." />
</figure>

<p>I’ve shown the findings for when the local-environment radius for the spatial measures is set at 2km, but the same result happens when we use 0.5, 1, or 4 km (see the Appendix).</p>

<h3 id="does-the-result-depend-on-the-sample">Does the result depend on the sample?</h3>

<p>As a robustness check, I repeated the analysis using a fixed panel corresponding to Massey and Denton’s original 60 metropolitan identities, following the same metro names while using each census year’s contemporary boundaries. The number of retained factors is the same at every spatial radius within a given sample and year:</p>

<table>
  <thead>
    <tr>
      <th>Sample</th>
      <th style="text-align: right">1980</th>
      <th style="text-align: right">1990</th>
      <th style="text-align: right">2000</th>
      <th style="text-align: right">2010</th>
      <th style="text-align: right">2020</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>National sample</td>
      <td style="text-align: right">3</td>
      <td style="text-align: right">3</td>
      <td style="text-align: right">3</td>
      <td style="text-align: right">3</td>
      <td style="text-align: right">3</td>
    </tr>
    <tr>
      <td>Original-metro panel</td>
      <td style="text-align: right">2</td>
      <td style="text-align: right">2</td>
      <td style="text-align: right">2</td>
      <td style="text-align: right">2</td>
      <td style="text-align: right">3</td>
    </tr>
  </tbody>
</table>

<p>The fixed panel has a much smaller weighted effective sample size (between 69 and 92 observations, compared with 230 to 472 in the national sample) so parallel analysis is more conservative. Still, we can see that the data does not support five factors.</p>

<h2 id="what-do-the-factors-measure">What do the factors measure?</h2>

<p>The three factors are Evenness, Isolation, and Concentration. Expand each factor below for its interpretation and loadings.</p>

<style>
  .factor-result {
    margin: 0 !important;
    border-top: 1px solid #d0d7de;
  }

  .factor-result.factor-result-last {
    border-bottom: 1px solid #d0d7de;
  }

  .factor-result summary {
    display: flex;
    align-items: center;
    justify-content: space-between;
    padding: 0.8rem 0.25rem;
    cursor: pointer;
    list-style: none;
  }

  .factor-result summary::-webkit-details-marker {
    display: none;
  }

  .factor-result summary h3,
  .factor-result summary h5 {
    margin: 0 !important;
    font-size: 1.25em;
  }

  .factor-result summary::after {
    content: "+";
    margin-left: 1rem;
    color: #57606a;
    font-size: 1.35rem;
    font-weight: 400;
    line-height: 1;
  }

  .factor-result[open] summary::after {
    content: "-";
  }

  .factor-result > :not(summary) {
    margin-right: 1rem;
    margin-left: 1rem;
  }

  .factor-result > :last-child {
    margin-bottom: 1.5rem;
  }
</style>

<!-- factor-loading-tables:2km:start -->

<details class="factor-result">
  <summary><h3 id="evenness">Evenness</h3></summary>

  <p>This is pretty similar to the original Evenness dimension that Massey and Denton found; it measures how evenly distributed the minority and majority groups are.</p>

  <p><strong>Factor loadings</strong></p>

  <table>
    <thead>
      <tr>
        <th>Index</th>
        <th style="text-align: right">1980</th>
        <th style="text-align: right">1990</th>
        <th style="text-align: right">2000</th>
        <th style="text-align: right">2010</th>
        <th style="text-align: right">2020</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Atkinson, $b = 0.5$</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">0.99</td>
      </tr>
      <tr>
        <td>Atkinson, $b = 0.9$</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">0.99</td>
      </tr>
      <tr>
        <td>Information theory ($H$)</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">0.99</td>
      </tr>
      <tr>
        <td>Spatial information theory ($H_s$)</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.96</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">1.00</td>
      </tr>
      <tr>
        <td>Spatial dissimilarity ($D_s$)</td>
        <td style="text-align: right">0.96</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
      </tr>
      <tr>
        <td>Gini ($G$)</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.98</td>
      </tr>
      <tr>
        <td>Dissimilarity ($D$)</td>
        <td style="text-align: right">0.96</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.99</td>
      </tr>
      <tr>
        <td>Atkinson, $b = 0.1$</td>
        <td style="text-align: right">0.88</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.96</td>
        <td style="text-align: right">0.96</td>
        <td style="text-align: right">0.96</td>
      </tr>
      <tr>
        <td>Correlation ratio ($\eta^2$)</td>
        <td style="text-align: right">0.94</td>
        <td style="text-align: right">0.95</td>
        <td style="text-align: right">0.90</td>
        <td style="text-align: right">0.94</td>
        <td style="text-align: right">0.94</td>
      </tr>
      <tr>
        <td>Spatial proximity ($\mathrm{SP}$)</td>
        <td style="text-align: right">0.82</td>
        <td style="text-align: right">0.92</td>
        <td style="text-align: right">0.90</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.91</td>
      </tr>
      <tr>
        <td>Divergence index ($\mathrm{DIV}$)</td>
        <td style="text-align: right">0.81</td>
        <td style="text-align: right">0.81</td>
        <td style="text-align: right">0.77</td>
        <td style="text-align: right">0.85</td>
        <td style="text-align: right">0.86</td>
      </tr>
      <tr>
        <td>Relative clustering ($\mathrm{RCL}$)</td>
        <td style="text-align: right">0.60</td>
        <td style="text-align: right">0.69</td>
        <td style="text-align: right">0.72</td>
        <td style="text-align: right">0.59</td>
        <td style="text-align: right">0.55</td>
      </tr>
      <tr>
        <td>Relative centralization ($\mathrm{RCE}$)</td>
        <td style="text-align: right">0.22</td>
        <td style="text-align: right">0.46</td>
        <td style="text-align: right">0.64</td>
        <td style="text-align: right">0.45</td>
        <td style="text-align: right">0.46</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result">
  <summary><h3 id="isolation-factor">Isolation</h3></summary>

  <p>Broadly speaking, these indices measure how much contact a random member of the focal group is expected to have with other members of the same group.</p>

  <p><strong>Factor loadings</strong></p>

  <table>
    <thead>
      <tr>
        <th>Index</th>
        <th style="text-align: right">1980</th>
        <th style="text-align: right">1990</th>
        <th style="text-align: right">2000</th>
        <th style="text-align: right">2010</th>
        <th style="text-align: right">2020</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Distance-decay isolation ($\mathrm{DP}_{xx}$)</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.95</td>
        <td style="text-align: right">0.96</td>
      </tr>
      <tr>
        <td>Spatial isolation (${}_xP_x^{(s)}$)</td>
        <td style="text-align: right">0.89</td>
        <td style="text-align: right">0.90</td>
        <td style="text-align: right">0.94</td>
        <td style="text-align: right">0.92</td>
        <td style="text-align: right">0.93</td>
      </tr>
      <tr>
        <td>Isolation (${}_xP_x$)</td>
        <td style="text-align: right">0.85</td>
        <td style="text-align: right">0.87</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.91</td>
        <td style="text-align: right">0.93</td>
      </tr>
      <tr>
        <td>Absolute concentration ($\mathrm{ACO}$)</td>
        <td style="text-align: right">-0.83</td>
        <td style="text-align: right">-0.87</td>
        <td style="text-align: right">-0.31</td>
        <td style="text-align: right">-0.90</td>
        <td style="text-align: right">-0.92</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result factor-result-last">
  <summary><h3 id="concentration-factor">Concentration</h3></summary>

  <p>This combines indices that measure how concentrated a group is in the city center with indices that measure how concentrated a group is in general.</p>

  <p><strong>Factor loadings</strong></p>

  <table>
    <thead>
      <tr>
        <th>Index</th>
        <th style="text-align: right">1980</th>
        <th style="text-align: right">1990</th>
        <th style="text-align: right">2000</th>
        <th style="text-align: right">2010</th>
        <th style="text-align: right">2020</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Absolute centralization ($\mathrm{ACE}$)</td>
        <td style="text-align: right">0.84</td>
        <td style="text-align: right">0.87</td>
        <td style="text-align: right">0.62</td>
        <td style="text-align: right">0.90</td>
        <td style="text-align: right">0.90</td>
      </tr>
      <tr>
        <td>Delta ($\mathrm{DEL}$)</td>
        <td style="text-align: right">0.84</td>
        <td style="text-align: right">0.82</td>
        <td style="text-align: right">0.68</td>
        <td style="text-align: right">0.83</td>
        <td style="text-align: right">0.86</td>
      </tr>
      <tr>
        <td>Proportion central city ($\mathrm{PCC}$)</td>
        <td style="text-align: right">0.82</td>
        <td style="text-align: right">0.83</td>
        <td style="text-align: right">0.61</td>
        <td style="text-align: right">0.77</td>
        <td style="text-align: right">0.77</td>
      </tr>
      <tr>
        <td>Relative concentration ($\mathrm{RCO}$)</td>
        <td style="text-align: right">0.82</td>
        <td style="text-align: right">0.74</td>
        <td style="text-align: right">0.92</td>
        <td style="text-align: right">0.50</td>
        <td style="text-align: right">0.55</td>
      </tr>
      <tr>
        <td>Relative centralization ($\mathrm{RCE}$)</td>
        <td style="text-align: right">0.69</td>
        <td style="text-align: right">0.46</td>
        <td style="text-align: right">0.28</td>
        <td style="text-align: right">0.40</td>
        <td style="text-align: right">0.39</td>
      </tr>
    </tbody>
  </table>

</details>

<!-- factor-loading-tables:2km:end -->

<h2 id="why-are-there-fewer-than-five-factors">Why are there fewer than five factors?</h2>

<p>So, are there really only three dimensions of segregation? The national sample says three, while the fixed panel generally says two. But neither supports five. I’d argue that there are three <em>relevant</em> dimensions, even if the smaller fixed panel does not always distinguish all three empirically.</p>

<p>A problem that immediately jumped out at me when reading about the five dimensions was that the Centralization dimension makes an assumption not about <em>how</em> a group is distributed, but about <em>where</em> the group is distributed. This assumption makes me uneasy. I don’t like my mathematical measures to bake in such messy, <em>empirical</em> assumptions. Now, it does make sense that, <em>given</em> the fact that minority groups tend to be located near the center of urban areas, we’d expect centralization to be correlated with concentration. But I’d prefer to use concentration measures and then, once we determine that certain groups tend to end up concentrated, ask where those groups are concentrated. If that turns out to be the city center, then great.</p>

<p>What about the two dimensions that don’t show up in my factor analysis, Exposure and Clustering? Well, conceptually, clustering seems related to evenness: if a group is all clustered together, then it isn’t evenly distributed. As for exposure, that seems related to isolation. If a group is isolated, that is, its members are particularly likely to live near one another, then they’re not very exposed to members of other groups.</p>

<p>So, it doesn’t actually seem like we needed five dimensions in the first place.</p>

<h2 id="which-measures-should-we-use">Which measures should we use?</h2>

<p>Which measures should be used for each factor? For the Evenness factor, there are seven indices with consistently near-perfect loadings, so really, we could choose any of them and we’d be fine. But the dissimilarity index ($D$) seems to be the most popular, so we’ll go with that.</p>

<p>For the Isolation factor, the three highest-loading indices are, well, isolation indices. Of these three, distance-decay isolation and spatial isolation seem the most conceptually appropriate: we should expect neighboring tracts to matter when determining isolation. Of these, distance-decay isolation consistently loads higher, so let’s go with that.</p>

<p>For the Concentration factor, I’d prefer to avoid any measures that explicitly refer to or rely on the idea of a city center, which leaves delta ($\text{DEL}$) as our best option.</p>

<p>So, if you want to measure segregation, I would distinguish three dimensions: evenness, isolation, and concentration. If you’re interested in evenness, use the dissimilarity index. If you’re interested in isolation, use distance-decay isolation. If you’re interested in concentration, use the delta index.</p>

<h2 id="appendix">Appendix</h2>

<h3 id="measures-of-segregation">Measures of Segregation</h3>

<p>The analysis uses 22 nonredundant indices: 18 from the original Massey-Denton battery and four later extensions. Expand any measure below for its interpretation and formula.</p>

<p>Population and area notation:</p>

<ul>
  <li>$x_i$ and $y_i$ are the focal-group and non-Hispanic-White populations of tract $i$.</li>
  <li>$t_i=x_i+y_i$ is the tract’s two-group population.</li>
  <li>$X$, $Y$, and $T$ are the corresponding metropolitan totals.</li>
  <li>$p_i=x_i/t_i$ is the focal-group share of tract $i$, and $P=X/T$ is the focal-group share of the metropolitan population.</li>
  <li>$a_i$ is tract land area (in square miles).</li>
</ul>

<p>For the original proximity measures:</p>

<ul>
  <li>$d_{ij}$ is the distance in miles between the representative points of tracts $i$ and $j$.</li>
  <li>$z_{ij}=e^{-d_{ij}}$ is the distance-decay weight.</li>
  <li>Within a tract, self-distance is approximated as $d_{ii}=\sqrt{0.6\,a_i}$. This is to avoid treating every resident of a tract as if they’re in the exact same place.</li>
</ul>

<p>For the added local-environment measures:</p>

<ul>
  <li>$r_{ij}$ is the distance in kilometers between tracts $i$ and $j$.</li>
  <li>$h$ is the local-environment radius.</li>
</ul>

<p>Nearby tracts are weighted using the biweight kernel:</p>

\[w_{ij}(h)=
\begin{cases}
[1-(r_{ij}/h)^2]^2, &amp; r_{ij}&lt;h,\\
0, &amp; r_{ij}\ge h.
\end{cases}\]

<p>The resulting focal-group share of tract $i$’s local environment is:</p>

\[q_i=\frac{\sum_j w_{ij}(h)x_j}{\sum_j w_{ij}(h)t_j}.\]

<p><strong>Indices of Segregation</strong></p>

<!-- measure-sections:start -->

<details>
  <summary id="dissimilarity">Dissimilarity ($D$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Evenness</p>

    <p>Dissimilarity is the share of either group that would have to move to a different tract for every tract to have the same two-group composition as the metropolitan area. It ranges from zero under identical tract distributions to one under complete separation.</p>

\[D=\frac{1}{2}\sum_i\left|\frac{x_i}{X}-\frac{y_i}{Y}\right|\]

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="gini">Gini ($G$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Evenness</p>

    <p>The segregation Gini compares the focal-group proportions of every pair of tracts, weighting pairs by their two-group populations. Like dissimilarity, it is zero when tract compositions are identical and one under complete separation, but it uses the entire segregation curve rather than a single cutoff.</p>

\[G=\frac{\sum_i\sum_j t_i t_j|p_i-p_j|}{2T^2P(1-P)}\]

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="information-theory">Information Theory ($H$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Evenness</p>

    <p>Information theory measures the proportional reduction in racial/ethnic entropy obtained by knowing a person’s tract. It is zero when every tract reproduces the metropolitan composition and approaches one as tracts become internally homogeneous.</p>

    <p>Writing $E(q)=-q\log q-(1-q)\log(1-q)$,</p>

\[H=1-\frac{\sum_i t_iE(p_i)}{TE(P)}\]

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="atkinson-01">Atkinson, $b = 0.1$</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Evenness</p>

    <p>This is the low-parameter member of the Atkinson segregation family. The parameter changes which parts of the tract-composition distribution receive the most weight; using several values tests whether the result depends on that normative weighting. Higher values indicate greater unevenness.</p>

\[A_b=1-\frac{P}{1-P}\left[\frac{\sum_i t_i p_i^b(1-p_i)^{1-b}}{PT}\right]^{1/(1-b)},\qquad b=0.1\]

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="atkinson-05">Atkinson, $b = 0.5$</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Evenness</p>

    <p>This is the midpoint specification of the Atkinson index. It gives a comparatively balanced weighting to focal- and reference-group representation across tracts. Higher values indicate greater unevenness.</p>

\[A_b=1-\frac{P}{1-P}\left[\frac{\sum_i t_i p_i^b(1-p_i)^{1-b}}{PT}\right]^{1/(1-b)},\qquad b=0.5\]

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="atkinson-09">Atkinson, $b = 0.9$</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Evenness</p>

    <p>This is the high-parameter member of the Atkinson segregation family. Together with the 0.1 and 0.5 versions, it checks the sensitivity of measured unevenness to the Atkinson weighting parameter. Higher values indicate greater unevenness.</p>

\[A_b=1-\frac{P}{1-P}\left[\frac{\sum_i t_i p_i^b(1-p_i)^{1-b}}{PT}\right]^{1/(1-b)},\qquad b=0.9\]

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="divergence">Divergence Index ($\mathrm{DIV}$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Not included in their review</p>

    <p>The divergence index is the population-weighted Kullback-Leibler divergence between each tract’s two-group composition and the metropolitan composition. It is zero when all tracts match the metropolitan area and increases as their compositions diverge. Unlike many normalized indices, it is not constrained to a zero-to-one scale.</p>

\[\mathrm{DIV}=\sum_i\frac{t_i}{T}\left[p_i\log\frac{p_i}{P}+(1-p_i)\log\frac{1-p_i}{1-P}\right]\]

    <p>Formula source: <a href="https://arxiv.org/abs/1508.01167">Roberto (2015), <em>The Divergence Index</em></a>.</p>
  </div>
</details>

<details>
  <summary id="spatial-information-theory">Spatial Information Theory ($H_s$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Not included in their review</p>

    <p>This replaces each tract’s composition in the information-theory index with the composition of its kernel-weighted local environment.</p>

    <p>If $q_i$ is the focal-group proportion in tract $i$’s local environment,</p>

\[H_s=1-\frac{\sum_i t_iE(q_i)}{TE(P)}\]

    <p>Formula framework: <a href="https://doi.org/10.1111/j.0081-1750.2004.00150.x">Reardon and O’Sullivan (2004), <em>Measures of Spatial Segregation</em></a>. The displayed equation is the tract-centroid, biweight-kernel implementation used in this analysis.</p>
  </div>
</details>

<details>
  <summary id="spatial-dissimilarity">Spatial Dissimilarity ($D_s$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Not included in their review</p>

    <p>Spatial dissimilarity ($D_s$) measures the population-weighted absolute difference between each tract’s local-environment composition and the metropolitan composition.</p>

\[D_s=\frac{\sum_i t_i|q_i-P|}{2TP(1-P)}\]

    <p>Formula framework: <a href="https://doi.org/10.1111/j.0081-1750.2004.00150.x">Reardon and O’Sullivan (2004), <em>Measures of Spatial Segregation</em></a>. The displayed equation is the tract-centroid, biweight-kernel implementation used in this analysis.</p>
  </div>
</details>

<details>
  <summary id="isolation">Isolation (${}_xP_x$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Exposure</p>

    <p>Isolation is the expected focal-group share of the tract occupied by a randomly selected focal-group member. It therefore combines segregation with the focal group’s overall metropolitan prevalence: even under equal tract composition, its baseline is $P$ rather than zero.</p>

\[xP_x=\sum_i\frac{x_i}{X}p_i\]

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="correlation-ratio">Correlation Ratio ($\eta^2$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Exposure</p>

    <p>The correlation ratio rescales isolation relative to the focal group’s metropolitan share. It can be interpreted as the proportion of variance in individual group membership associated with tract membership: zero indicates no tract differentiation and one indicates complete separation.</p>

\[\eta^2=\frac{xP_x-P}{1-P}\]

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="distance-decay-isolation">Distance-Decay Isolation ($\mathrm{DP}_{xx}$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Clustering</p>

    <p>Distance-decay isolation replaces same-tract contact with potential contact across all tract pairs. Pairwise influence declines exponentially with distance, while a tract’s self-distance is approximated from its area. Higher values mean that the average focal-group resident’s spatial surroundings contain a larger focal-group share.</p>

\[\mathrm{DP}_{xx}=\sum_i\frac{x_i}{X}\left(\frac{\sum_j z_{ij}x_j}{\sum_j z_{ij}t_j}\right)\]

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="spatial-isolation">Spatial Isolation (${}_xP_x^{(s)}$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Not included in their review</p>

    <p>Spatial isolation (${}_xP_x^{(s)}$) is the expected focal-group proportion in the kernel-weighted local environment of a randomly selected focal-group resident. Unlike distance-decay isolation, its influence has a finite outer radius and follows a biweight kernel.</p>

\[xP_x^{(s)}=\sum_i\frac{x_i}{X}q_i\]

    <p>Formula framework: <a href="https://doi.org/10.1111/j.0081-1750.2004.00150.x">Reardon and O’Sullivan (2004), <em>Measures of Spatial Segregation</em></a>. The displayed equation is the tract-centroid, biweight-kernel implementation used in this analysis.</p>
  </div>
</details>

<details>
  <summary id="delta">Delta ($\mathrm{DEL}$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Concentration</p>

    <p>Delta compares the focal group’s distribution across tracts with the distribution of metropolitan land area. It is the share of the focal population that would have to relocate for its tract distribution to match the distribution of land area.</p>

    <p>If $a_i$ is tract land area and $A=\sum_i a_i$,</p>

\[\mathrm{DEL}=\frac{1}{2}\sum_i\left|\frac{x_i}{X}-\frac{a_i}{A}\right|\]

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="absolute-concentration">Absolute Concentration ($\mathrm{ACO}$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Concentration</p>

    <p>Absolute concentration compares the average tract area occupied by focal-group members with the smallest and largest areas that could contain the same population under the observed tract structure. Let $\bar a_x=\sum_i x_i a_i/X$. Let $\bar a_{\min}$ and $\bar a_{\max}$ be the population-weighted mean areas obtained by filling the smallest and largest tracts respectively until the accumulated two-group population reaches $X$. Then</p>

\[\mathrm{ACO}=1-\frac{\bar a_x-\bar a_{\min}}{\bar a_{\max}-\bar a_{\min}}.\]

    <p>Higher values indicate that the focal group occupies relatively little physical space. The published normalization can leave its nominal range in unusual metros where the designated focal group is overwhelmingly dominant; those values were retained rather than clipped.</p>

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="relative-concentration">Relative Concentration ($\mathrm{RCO}$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Concentration</p>

    <p>Relative concentration compares the focal group’s population-weighted mean tract area with that of the reference group, normalized by the metropolitan area’s feasible area limits. With $\bar a_x=\sum_i x_i a_i/X$, $\bar a_y=\sum_i y_i a_i/Y$, and the same $\bar a_{\min}$ and $\bar a_{\max}$ used for $\mathrm{ACO}$,</p>

\[\mathrm{RCO}=\frac{\bar a_x/\bar a_y-1}{\bar a_{\min}/\bar a_{\max}-1}.\]

    <p>Positive values indicate that the focal group occupies less space than the reference group; negative values indicate the reverse.</p>

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="proportion-central-city">Proportion Central City ($\mathrm{PCC}$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Centralization</p>

    <p>$\mathrm{PCC}$ is simply the proportion of the focal metropolitan population living within the year’s official central or principal city boundaries. When a metropolitan area has multiple official central/principal cities, all of them count.</p>

\[\mathrm{PCC}=\frac{\sum_{i\in\text{central city}}x_i}{X}\]

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="absolute-centralization">Absolute Centralization ($\mathrm{ACE}$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Centralization</p>

    <p>After ordering tracts from nearest to farthest from the metropolitan center, define $C_i^x=\sum_{j\le i}x_j/X$ and $C_i^a=\sum_{j\le i}a_j/\sum_j a_j$. Absolute centralization is</p>

\[\mathrm{ACE}=\sum_{i=1}^{n-1}\left(C_i^xC_{i+1}^a-C_{i+1}^xC_i^a\right).\]

    <p>Positive values indicate centralization, zero indicates no systematic central tendency, and negative values indicate decentralization toward the periphery.</p>

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="relative-centralization">Relative Centralization ($\mathrm{RCE}$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Centralization</p>

    <p>After ordering tracts from nearest to farthest from the center, define $C_i^x=\sum_{j\le i}x_j/X$ and $C_i^y=\sum_{j\le i}y_j/Y$. Relative centralization is</p>

\[\mathrm{RCE}=\sum_{i=1}^{n-1}\left(C_i^xC_{i+1}^y-C_{i+1}^xC_i^y\right).\]

    <p>Positive values mean the focal group is more centralized than non-Hispanic White residents; negative values mean it is less centralized.</p>

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="absolute-clustering">Absolute Clustering ($\mathrm{ACL}$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Clustering</p>

    <p>Absolute clustering compares the distance-weighted proximity of focal-group residents to one another with the proximity expected under a uniform spatial distribution, normalized by the proximity of the total two-group population. Define</p>

\[C_x=\sum_i\frac{x_i}{X}\sum_jz_{ij}x_j,\qquad C_t=\sum_i\frac{x_i}{X}\sum_jz_{ij}t_j,\qquad U=\frac{X}{n^2}\sum_i\sum_jz_{ij}.\]

    <p>Then</p>

\[\mathrm{ACL}=\frac{C_x-U}{C_t-U}.\]

    <p>Higher values indicate a more tightly clustered focal population.</p>

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="spatial-proximity">Spatial Proximity ($\mathrm{SP}$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Clustering</p>

    <p>Define focal-group, reference-group, and total-population proximity as</p>

\[P_{xx}=\frac{\sum_i\sum_jx_ix_jz_{ij}}{X^2},\qquad P_{yy}=\frac{\sum_i\sum_jy_iy_jz_{ij}}{Y^2},\qquad P_{tt}=\frac{\sum_i\sum_jt_it_jz_{ij}}{T^2}.\]

    <p>Spatial proximity is</p>

\[\mathrm{SP}=\frac{XP_{xx}+YP_{yy}}{TP_{tt}}.\]

    <p>A value of one indicates no excess same-group proximity; values above one indicate that members of the two groups tend to live nearer members of their own group.</p>

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<details>
  <summary id="relative-clustering">Relative Clustering ($\mathrm{RCL}$)</summary>
<div style="margin-left: 15px; border: 1px solid #ccc; padding: 10px; background-color: #f9f9f9;">

    <p><strong>Massey and Denton (1988) dimension:</strong> Clustering</p>

    <p>Using $P_{xx}$ and $P_{yy}$ as defined for $\mathrm{SP}$, relative clustering is</p>

\[\mathrm{RCL}=\frac{P_{xx}}{P_{yy}}-1.\]

    <p>Zero means equal clustering, positive values mean the focal group is more clustered, and negative values mean the reference group is more clustered.</p>

    <p>Formula sources: <a href="https://doi.org/10.1093/sf/67.2.281">Massey and Denton (1988)</a> and <a href="https://www.census.gov/topics/housing/housing-patterns/guidance/appendix-b.html">U.S. Census Bureau, <em>Housing Patterns: Appendix B</em></a>.</p>
  </div>
</details>

<!-- measure-sections:end -->

<h3 id="parallel-analysis-for-other-distances">Parallel Analysis for Other Distances</h3>

<figure>
    <img src="/assets/images/segregation/parallel_analysis_overview__national__r0.5km__extended22.png" width="1000" alt="Five-panel parallel-analysis chart for the national metropolitan sample from 1980 through 2020 using a 0.5-kilometer spatial radius. Three components are retained in every year." />
    <figcaption>National-sample parallel analysis using a 0.5-km spatial radius. Three components are retained in every census year.</figcaption>
</figure>

<figure>
    <img src="/assets/images/segregation/parallel_analysis_overview__national__r1km__extended22.png" width="1000" alt="Five-panel parallel-analysis chart for the national metropolitan sample from 1980 through 2020 using a 1-kilometer spatial radius. Three components are retained in every year." />
    <figcaption>National-sample parallel analysis using a 1-km spatial radius. Three components are retained in every census year.</figcaption>
</figure>

<figure>
    <img src="/assets/images/segregation/parallel_analysis_overview__national__r4km__extended22.png" width="1000" alt="Five-panel parallel-analysis chart for the national metropolitan sample from 1980 through 2020 using a 4-kilometer spatial radius. Three components are retained in every year." />
    <figcaption>National-sample parallel analysis using a 4-km spatial radius. Three components are retained in every census year.</figcaption>
</figure>

<!-- factor-loading-tables:other-radii:start -->

<h3 id="factor-loadings-at-other-distances">Factor Loadings at Other Distances</h3>

<p>The factors are matched and sign-aligned to each radius’s 1980 solution. As in the primary 2-km analysis, each table includes only indices whose mean absolute loading on that factor across the five censuses is at least .40, ordered by that mean.</p>

<h4 id="05-km-radius">0.5-km radius</h4>

<details class="factor-result">
  <summary><h5 id="r05-evenness">Factor 1: Evenness</h5></summary>

  <table>
    <thead>
      <tr>
        <th>Index</th>
        <th style="text-align: right">1980</th>
        <th style="text-align: right">1990</th>
        <th style="text-align: right">2000</th>
        <th style="text-align: right">2010</th>
        <th style="text-align: right">2020</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Atkinson, $b = 0.5$</td>
        <td style="text-align: right">1.01</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">1.00</td>
      </tr>
      <tr>
        <td>Atkinson, $b = 0.9$</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">1.00</td>
      </tr>
      <tr>
        <td>Spatial information theory ($H_s$)</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">1.00</td>
      </tr>
      <tr>
        <td>Information theory ($H$)</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">1.00</td>
      </tr>
      <tr>
        <td>Gini ($G$)</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.98</td>
      </tr>
      <tr>
        <td>Dissimilarity ($D$)</td>
        <td style="text-align: right">0.96</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.99</td>
      </tr>
      <tr>
        <td>Spatial dissimilarity ($D_s$)</td>
        <td style="text-align: right">0.96</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.99</td>
      </tr>
      <tr>
        <td>Atkinson, $b = 0.1$</td>
        <td style="text-align: right">0.89</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.96</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.96</td>
      </tr>
      <tr>
        <td>Correlation ratio ($\eta^2$)</td>
        <td style="text-align: right">0.94</td>
        <td style="text-align: right">0.95</td>
        <td style="text-align: right">0.90</td>
        <td style="text-align: right">0.94</td>
        <td style="text-align: right">0.94</td>
      </tr>
      <tr>
        <td>Spatial proximity ($\mathrm{SP}$)</td>
        <td style="text-align: right">0.82</td>
        <td style="text-align: right">0.92</td>
        <td style="text-align: right">0.90</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.91</td>
      </tr>
      <tr>
        <td>Divergence index ($\mathrm{DIV}$)</td>
        <td style="text-align: right">0.81</td>
        <td style="text-align: right">0.81</td>
        <td style="text-align: right">0.77</td>
        <td style="text-align: right">0.85</td>
        <td style="text-align: right">0.87</td>
      </tr>
      <tr>
        <td>Relative clustering ($\mathrm{RCL}$)</td>
        <td style="text-align: right">0.60</td>
        <td style="text-align: right">0.69</td>
        <td style="text-align: right">0.72</td>
        <td style="text-align: right">0.59</td>
        <td style="text-align: right">0.55</td>
      </tr>
      <tr>
        <td>Relative centralization ($\mathrm{RCE}$)</td>
        <td style="text-align: right">0.22</td>
        <td style="text-align: right">0.46</td>
        <td style="text-align: right">0.64</td>
        <td style="text-align: right">0.45</td>
        <td style="text-align: right">0.46</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result">
  <summary><h5 id="r05-isolation">Factor 2: Isolation</h5></summary>

  <table>
    <thead>
      <tr>
        <th>Index</th>
        <th style="text-align: right">1980</th>
        <th style="text-align: right">1990</th>
        <th style="text-align: right">2000</th>
        <th style="text-align: right">2010</th>
        <th style="text-align: right">2020</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Distance-decay isolation ($\mathrm{DP}_{xx}$)</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.95</td>
        <td style="text-align: right">0.96</td>
      </tr>
      <tr>
        <td>Spatial isolation (${}_xP_x^{(s)}$)</td>
        <td style="text-align: right">0.85</td>
        <td style="text-align: right">0.88</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.91</td>
        <td style="text-align: right">0.92</td>
      </tr>
      <tr>
        <td>Isolation (${}_xP_x$)</td>
        <td style="text-align: right">0.85</td>
        <td style="text-align: right">0.88</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.91</td>
        <td style="text-align: right">0.92</td>
      </tr>
      <tr>
        <td>Absolute concentration ($\mathrm{ACO}$)</td>
        <td style="text-align: right">-0.83</td>
        <td style="text-align: right">-0.87</td>
        <td style="text-align: right">-0.30</td>
        <td style="text-align: right">-0.90</td>
        <td style="text-align: right">-0.92</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result factor-result-last">
  <summary><h5 id="r05-concentration">Factor 3: Concentration</h5></summary>

  <table>
    <thead>
      <tr>
        <th>Index</th>
        <th style="text-align: right">1980</th>
        <th style="text-align: right">1990</th>
        <th style="text-align: right">2000</th>
        <th style="text-align: right">2010</th>
        <th style="text-align: right">2020</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Absolute centralization ($\mathrm{ACE}$)</td>
        <td style="text-align: right">0.85</td>
        <td style="text-align: right">0.87</td>
        <td style="text-align: right">0.62</td>
        <td style="text-align: right">0.90</td>
        <td style="text-align: right">0.90</td>
      </tr>
      <tr>
        <td>Delta ($\mathrm{DEL}$)</td>
        <td style="text-align: right">0.84</td>
        <td style="text-align: right">0.82</td>
        <td style="text-align: right">0.68</td>
        <td style="text-align: right">0.83</td>
        <td style="text-align: right">0.86</td>
      </tr>
      <tr>
        <td>Proportion central city ($\mathrm{PCC}$)</td>
        <td style="text-align: right">0.81</td>
        <td style="text-align: right">0.83</td>
        <td style="text-align: right">0.61</td>
        <td style="text-align: right">0.77</td>
        <td style="text-align: right">0.77</td>
      </tr>
      <tr>
        <td>Relative concentration ($\mathrm{RCO}$)</td>
        <td style="text-align: right">0.82</td>
        <td style="text-align: right">0.73</td>
        <td style="text-align: right">0.92</td>
        <td style="text-align: right">0.50</td>
        <td style="text-align: right">0.54</td>
      </tr>
      <tr>
        <td>Relative centralization ($\mathrm{RCE}$)</td>
        <td style="text-align: right">0.68</td>
        <td style="text-align: right">0.45</td>
        <td style="text-align: right">0.28</td>
        <td style="text-align: right">0.40</td>
        <td style="text-align: right">0.39</td>
      </tr>
    </tbody>
  </table>

</details>

<h4 id="1-km-radius">1-km radius</h4>

<details class="factor-result">
  <summary><h5 id="r1-evenness">Factor 1: Evenness</h5></summary>

  <table>
    <thead>
      <tr>
        <th>Index</th>
        <th style="text-align: right">1980</th>
        <th style="text-align: right">1990</th>
        <th style="text-align: right">2000</th>
        <th style="text-align: right">2010</th>
        <th style="text-align: right">2020</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Atkinson, $b = 0.5$</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">1.00</td>
      </tr>
      <tr>
        <td>Atkinson, $b = 0.9$</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">1.00</td>
      </tr>
      <tr>
        <td>Spatial information theory ($H_s$)</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">1.00</td>
      </tr>
      <tr>
        <td>Information theory ($H$)</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">1.00</td>
      </tr>
      <tr>
        <td>Gini ($G$)</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.98</td>
      </tr>
      <tr>
        <td>Spatial dissimilarity ($D_s$)</td>
        <td style="text-align: right">0.96</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.99</td>
      </tr>
      <tr>
        <td>Dissimilarity ($D$)</td>
        <td style="text-align: right">0.96</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.99</td>
      </tr>
      <tr>
        <td>Atkinson, $b = 0.1$</td>
        <td style="text-align: right">0.89</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.96</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.96</td>
      </tr>
      <tr>
        <td>Correlation ratio ($\eta^2$)</td>
        <td style="text-align: right">0.94</td>
        <td style="text-align: right">0.95</td>
        <td style="text-align: right">0.90</td>
        <td style="text-align: right">0.94</td>
        <td style="text-align: right">0.94</td>
      </tr>
      <tr>
        <td>Spatial proximity ($\mathrm{SP}$)</td>
        <td style="text-align: right">0.82</td>
        <td style="text-align: right">0.92</td>
        <td style="text-align: right">0.90</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.91</td>
      </tr>
      <tr>
        <td>Divergence index ($\mathrm{DIV}$)</td>
        <td style="text-align: right">0.80</td>
        <td style="text-align: right">0.81</td>
        <td style="text-align: right">0.77</td>
        <td style="text-align: right">0.85</td>
        <td style="text-align: right">0.86</td>
      </tr>
      <tr>
        <td>Relative clustering ($\mathrm{RCL}$)</td>
        <td style="text-align: right">0.60</td>
        <td style="text-align: right">0.69</td>
        <td style="text-align: right">0.72</td>
        <td style="text-align: right">0.59</td>
        <td style="text-align: right">0.55</td>
      </tr>
      <tr>
        <td>Relative centralization ($\mathrm{RCE}$)</td>
        <td style="text-align: right">0.22</td>
        <td style="text-align: right">0.46</td>
        <td style="text-align: right">0.64</td>
        <td style="text-align: right">0.45</td>
        <td style="text-align: right">0.46</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result">
  <summary><h5 id="r1-isolation">Factor 2: Isolation</h5></summary>

  <table>
    <thead>
      <tr>
        <th>Index</th>
        <th style="text-align: right">1980</th>
        <th style="text-align: right">1990</th>
        <th style="text-align: right">2000</th>
        <th style="text-align: right">2010</th>
        <th style="text-align: right">2020</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Distance-decay isolation ($\mathrm{DP}_{xx}$)</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.95</td>
        <td style="text-align: right">0.96</td>
      </tr>
      <tr>
        <td>Spatial isolation (${}_xP_x^{(s)}$)</td>
        <td style="text-align: right">0.86</td>
        <td style="text-align: right">0.88</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.91</td>
        <td style="text-align: right">0.93</td>
      </tr>
      <tr>
        <td>Isolation (${}_xP_x$)</td>
        <td style="text-align: right">0.85</td>
        <td style="text-align: right">0.88</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.91</td>
        <td style="text-align: right">0.92</td>
      </tr>
      <tr>
        <td>Absolute concentration ($\mathrm{ACO}$)</td>
        <td style="text-align: right">-0.83</td>
        <td style="text-align: right">-0.87</td>
        <td style="text-align: right">-0.31</td>
        <td style="text-align: right">-0.90</td>
        <td style="text-align: right">-0.92</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result factor-result-last">
  <summary><h5 id="r1-concentration">Factor 3: Concentration</h5></summary>

  <table>
    <thead>
      <tr>
        <th>Index</th>
        <th style="text-align: right">1980</th>
        <th style="text-align: right">1990</th>
        <th style="text-align: right">2000</th>
        <th style="text-align: right">2010</th>
        <th style="text-align: right">2020</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Absolute centralization ($\mathrm{ACE}$)</td>
        <td style="text-align: right">0.85</td>
        <td style="text-align: right">0.87</td>
        <td style="text-align: right">0.62</td>
        <td style="text-align: right">0.90</td>
        <td style="text-align: right">0.90</td>
      </tr>
      <tr>
        <td>Delta ($\mathrm{DEL}$)</td>
        <td style="text-align: right">0.84</td>
        <td style="text-align: right">0.82</td>
        <td style="text-align: right">0.68</td>
        <td style="text-align: right">0.83</td>
        <td style="text-align: right">0.86</td>
      </tr>
      <tr>
        <td>Proportion central city ($\mathrm{PCC}$)</td>
        <td style="text-align: right">0.81</td>
        <td style="text-align: right">0.83</td>
        <td style="text-align: right">0.61</td>
        <td style="text-align: right">0.77</td>
        <td style="text-align: right">0.77</td>
      </tr>
      <tr>
        <td>Relative concentration ($\mathrm{RCO}$)</td>
        <td style="text-align: right">0.82</td>
        <td style="text-align: right">0.73</td>
        <td style="text-align: right">0.92</td>
        <td style="text-align: right">0.50</td>
        <td style="text-align: right">0.54</td>
      </tr>
      <tr>
        <td>Relative centralization ($\mathrm{RCE}$)</td>
        <td style="text-align: right">0.68</td>
        <td style="text-align: right">0.46</td>
        <td style="text-align: right">0.28</td>
        <td style="text-align: right">0.40</td>
        <td style="text-align: right">0.39</td>
      </tr>
    </tbody>
  </table>

</details>

<h4 id="4-km-radius">4-km radius</h4>

<details class="factor-result">
  <summary><h5 id="r4-evenness">Factor 1: Evenness</h5></summary>

  <table>
    <thead>
      <tr>
        <th>Index</th>
        <th style="text-align: right">1980</th>
        <th style="text-align: right">1990</th>
        <th style="text-align: right">2000</th>
        <th style="text-align: right">2010</th>
        <th style="text-align: right">2020</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Atkinson, $b = 0.5$</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">0.99</td>
      </tr>
      <tr>
        <td>Atkinson, $b = 0.9$</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
      </tr>
      <tr>
        <td>Spatial dissimilarity ($D_s$)</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
      </tr>
      <tr>
        <td>Information theory ($H$)</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.99</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">1.00</td>
        <td style="text-align: right">0.99</td>
      </tr>
      <tr>
        <td>Gini ($G$)</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.96</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.98</td>
      </tr>
      <tr>
        <td>Dissimilarity ($D$)</td>
        <td style="text-align: right">0.95</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.98</td>
      </tr>
      <tr>
        <td>Spatial information theory ($H_s$)</td>
        <td style="text-align: right">0.92</td>
        <td style="text-align: right">0.95</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.97</td>
      </tr>
      <tr>
        <td>Atkinson, $b = 0.1$</td>
        <td style="text-align: right">0.87</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.95</td>
        <td style="text-align: right">0.96</td>
        <td style="text-align: right">0.95</td>
      </tr>
      <tr>
        <td>Correlation ratio ($\eta^2$)</td>
        <td style="text-align: right">0.94</td>
        <td style="text-align: right">0.95</td>
        <td style="text-align: right">0.90</td>
        <td style="text-align: right">0.94</td>
        <td style="text-align: right">0.94</td>
      </tr>
      <tr>
        <td>Spatial proximity ($\mathrm{SP}$)</td>
        <td style="text-align: right">0.84</td>
        <td style="text-align: right">0.92</td>
        <td style="text-align: right">0.90</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.92</td>
      </tr>
      <tr>
        <td>Divergence index ($\mathrm{DIV}$)</td>
        <td style="text-align: right">0.81</td>
        <td style="text-align: right">0.81</td>
        <td style="text-align: right">0.78</td>
        <td style="text-align: right">0.85</td>
        <td style="text-align: right">0.86</td>
      </tr>
      <tr>
        <td>Relative clustering ($\mathrm{RCL}$)</td>
        <td style="text-align: right">0.61</td>
        <td style="text-align: right">0.69</td>
        <td style="text-align: right">0.72</td>
        <td style="text-align: right">0.60</td>
        <td style="text-align: right">0.56</td>
      </tr>
      <tr>
        <td>Relative centralization ($\mathrm{RCE}$)</td>
        <td style="text-align: right">0.22</td>
        <td style="text-align: right">0.46</td>
        <td style="text-align: right">0.63</td>
        <td style="text-align: right">0.46</td>
        <td style="text-align: right">0.46</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result">
  <summary><h5 id="r4-isolation">Factor 2: Isolation</h5></summary>

  <table>
    <thead>
      <tr>
        <th>Index</th>
        <th style="text-align: right">1980</th>
        <th style="text-align: right">1990</th>
        <th style="text-align: right">2000</th>
        <th style="text-align: right">2010</th>
        <th style="text-align: right">2020</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Distance-decay isolation ($\mathrm{DP}_{xx}$)</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.97</td>
        <td style="text-align: right">0.98</td>
        <td style="text-align: right">0.95</td>
        <td style="text-align: right">0.96</td>
      </tr>
      <tr>
        <td>Spatial isolation (${}_xP_x^{(s)}$)</td>
        <td style="text-align: right">0.94</td>
        <td style="text-align: right">0.94</td>
        <td style="text-align: right">0.96</td>
        <td style="text-align: right">0.93</td>
        <td style="text-align: right">0.95</td>
      </tr>
      <tr>
        <td>Isolation (${}_xP_x$)</td>
        <td style="text-align: right">0.84</td>
        <td style="text-align: right">0.87</td>
        <td style="text-align: right">0.92</td>
        <td style="text-align: right">0.91</td>
        <td style="text-align: right">0.92</td>
      </tr>
      <tr>
        <td>Absolute concentration ($\mathrm{ACO}$)</td>
        <td style="text-align: right">-0.82</td>
        <td style="text-align: right">-0.86</td>
        <td style="text-align: right">-0.31</td>
        <td style="text-align: right">-0.90</td>
        <td style="text-align: right">-0.92</td>
      </tr>
    </tbody>
  </table>

</details>

<details class="factor-result factor-result-last">
  <summary><h5 id="r4-concentration">Factor 3: Concentration</h5></summary>

  <table>
    <thead>
      <tr>
        <th>Index</th>
        <th style="text-align: right">1980</th>
        <th style="text-align: right">1990</th>
        <th style="text-align: right">2000</th>
        <th style="text-align: right">2010</th>
        <th style="text-align: right">2020</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>Absolute centralization ($\mathrm{ACE}$)</td>
        <td style="text-align: right">0.84</td>
        <td style="text-align: right">0.87</td>
        <td style="text-align: right">0.62</td>
        <td style="text-align: right">0.90</td>
        <td style="text-align: right">0.90</td>
      </tr>
      <tr>
        <td>Delta ($\mathrm{DEL}$)</td>
        <td style="text-align: right">0.84</td>
        <td style="text-align: right">0.81</td>
        <td style="text-align: right">0.68</td>
        <td style="text-align: right">0.83</td>
        <td style="text-align: right">0.86</td>
      </tr>
      <tr>
        <td>Proportion central city ($\mathrm{PCC}$)</td>
        <td style="text-align: right">0.83</td>
        <td style="text-align: right">0.83</td>
        <td style="text-align: right">0.61</td>
        <td style="text-align: right">0.77</td>
        <td style="text-align: right">0.77</td>
      </tr>
      <tr>
        <td>Relative concentration ($\mathrm{RCO}$)</td>
        <td style="text-align: right">0.82</td>
        <td style="text-align: right">0.74</td>
        <td style="text-align: right">0.92</td>
        <td style="text-align: right">0.50</td>
        <td style="text-align: right">0.55</td>
      </tr>
      <tr>
        <td>Relative centralization ($\mathrm{RCE}$)</td>
        <td style="text-align: right">0.68</td>
        <td style="text-align: right">0.46</td>
        <td style="text-align: right">0.28</td>
        <td style="text-align: right">0.39</td>
        <td style="text-align: right">0.39</td>
      </tr>
    </tbody>
  </table>

</details>

<!-- factor-loading-tables:other-radii:end -->

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:threshold" role="doc-endnote">
      <p>The threshold is applied separately for each racial group and census year. <a href="#fnref:threshold" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name></name></author><category term="uncategorized" /><summary type="html"><![CDATA[While reading A Wealth and Status-Based Model of Residential Segregation, I noticed its discussion of Massey and Denton (1988), who reviewed 20 segregation indices and applied them to 1980 census data from 60 metropolitan statistical areas (MSAs). Using factor analysis, Massey and Denton classified these indices into five underlying dimensions: evenness, exposure, clustering, centralization, and concentration.]]></summary></entry><entry><title type="html">Do AI Benchmarks Measure the Same Thing Over Time?</title><link href="https://stanichor.net/benchmark-dif/" rel="alternate" type="text/html" title="Do AI Benchmarks Measure the Same Thing Over Time?" /><published>2026-09-15T00:00:00+00:00</published><updated>2026-09-15T00:00:00+00:00</updated><id>https://stanichor.net/benchmark-dif</id><content type="html" xml:base="https://stanichor.net/benchmark-dif/"><![CDATA[<link rel="stylesheet" href="/assets/css/benchmark-dif.css" />

<p>AI’s moving pretty fast these days. How fast? Really fast. You want a quantitative answer? That’s what benchmarks are for. Just have the AI models answer questions and complete tasks, then score the results. You can use the scores to compare models and chart AI progress. But benchmarks have a problem. They saturate too fast. AI progress is so fast that benchmarks can become useless within a couple of years, sometimes much less.</p>

<p>So, you might think to solve this problem by ‘linking’ benchmarks together. If you know how much more difficult one benchmark is than another, you can sort of combine them into a mega-benchmark that works even as model capabilities saturate easier benchmarks. This is how Epoch’s ECI works.</p>

<p>It borrows from item response theory, specifically the 2PL model. The 2PL model is so called because it uses the logistic function and two parameters, the benchmark discrimination ($\alpha_b$) and difficulty ($D_b$), in the following way (for Epoch’s ECI):</p>

\[\mu_{mb} = \sigma(\alpha_b[C_m-D_b])\]

<p>where:</p>

\[\sigma(x) = \frac{1}{1+e^{-x}}\]

<p>where:</p>

<ul>
  <li>$\mu_{mb}$ is the predicted performance of a model, $m$, on a benchmark $b$.</li>
  <li>$\alpha_b$ is the benchmark’s discrimination. It tells us how good the benchmark is at ‘discriminating’ between high- and low-capability models. On a benchmark with high discrimination, small differences in capability produce large differences in performance, while on a benchmark with low discrimination, performance changes more gradually with capability.</li>
  <li>$D_b$ is the benchmark’s difficulty. It tells us, well, how difficult the benchmark is.</li>
  <li>$C_m$ is the model’s capability. Models with high capabilities will be able to do well even on difficult benchmarks, while those with low capabilities will struggle with even easy benchmarks.</li>
</ul>

<p>The ECI assumes these parameters are constant across time, but that’s not a safe assumption to make. Whenever we’re comparing a latent construct across groups, we need to ensure that our measurement instrument is measuring the same construct in the same way. In psychometrics, this is called <a href="https://www.the100.ci/2024/01/10/a-casual-but-causal-take-on-measurement-invariance/"><em>measurement invariance</em></a>. Otherwise, we could have biased items that make one group artificially score higher than another, giving us misleading results.</p>

<p>For example, vocabulary is known to be very <em>g</em>-loaded, and so one might decide to use a vocabulary test as a proxy for cognitive ability, like WORDSUM in the General Social Survey. However, there are words that men are more likely to know than women and vice versa<sup id="fnref:vocab" role="doc-noteref"><a href="#fn:vocab" class="footnote" rel="footnote">1</a></sup>.</p>

<div class="vocabulary-gap-table" role="region" aria-label="Vocabulary knowledge differences by gender">
<table>
  <thead>
    <tr>
      <th colspan="2" class="vocabulary-gap-men">Men more likely to know</th>
      <th colspan="2" class="vocabulary-gap-women">Women more likely to know</th>
    </tr>
    <tr>
      <th>Word</th>
      <th>Advantage</th>
      <th>Word</th>
      <th>Advantage</th>
    </tr>
  </thead>
  <tbody>
    <tr><td>howitzer</td><td>+31 pp</td><td>peplum</td><td>+51 pp</td></tr>
    <tr><td>thermistor</td><td>+31 pp</td><td>tulle</td><td>+50 pp</td></tr>
    <tr><td>azimuth</td><td>+31 pp</td><td>chignon</td><td>+48 pp</td></tr>
    <tr><td>femtosecond</td><td>+32 pp</td><td>bandeau</td><td>+46 pp</td></tr>
    <tr><td>milliamp</td><td>+32 pp</td><td>freesia</td><td>+45 pp</td></tr>
    <tr><td>aileron</td><td>+33 pp</td><td>chenille</td><td>+42 pp</td></tr>
    <tr><td>servo</td><td>+33 pp</td><td>kohl</td><td>+41 pp</td></tr>
    <tr><td>degauss</td><td>+33 pp</td><td>verbena</td><td>+40 pp</td></tr>
    <tr><td>boson</td><td>+32 pp</td><td>doula</td><td>+38 pp</td></tr>
    <tr><td>checksum</td><td>+33 pp</td><td>ruche</td><td>+37 pp</td></tr>
  </tbody>
</table>
<p>“Advantage” is the difference in the proportion of men and women who knew the word, in percentage points.</p>
</div>

<p>If we have a test that consists of many of these words, it’ll be biased: one gender will have a higher chance of getting items correct than the other, even holding cognitive ability constant.</p>

<p>Going back to AI models, we might decide to group models by <em>when</em> they were released. In this case, for example, a benchmark might become popular, so much so that labs start explicitly optimizing for performance on it. We might then see models’ performance on that benchmark increase very rapidly, much faster than their general capability. In that case, making use of the benchmark without accounting for this shift would cause us to overestimate AI progress. And so, we must check whether the parameters are actually static across time.</p>

<p>For this analysis, I use the public ECI data downloaded on September 15, 2026, containing 2,745 reported scores from 264 models across 58 benchmarks. I measure time using each model’s release month. Thus, a change over time means that models released in different months have different expected performance on a benchmark after accounting for their estimated general capability.</p>

<p>First, what happens when we allow the benchmark discriminations to vary depending on model release date?</p>

<section class="benchmark-dif-interactive" data-benchmark-dif-chart="" data-kind="discrimination" data-default-benchmark="b47" data-data-url="/assets/jsons/benchmark_dif.json" aria-labelledby="benchmark-discrimination-title">
  <div class="benchmark-dif-heading">
    <div>
      <h2 id="benchmark-discrimination-title" data-role="title">Benchmark discrimination over time</h2>
      <p>Absolute posterior mean with a 90% credible band for temporal shape</p>
    </div>
    <label class="benchmark-dif-control">
      <span>Benchmark</span>
      <select data-role="benchmark-select" disabled="">
        <option>Loading benchmarks…</option>
      </select>
    </label>
  </div>
  <div class="benchmark-dif-marker-key" aria-label="Month marker legend">
    <span><i class="is-shape-band" aria-hidden="true"></i>90% temporal-shape band</span>
    <span><i aria-hidden="true"></i>Month with observations</span>
    <span><i class="is-interpolated" aria-hidden="true"></i>Interpolated month</span>
    <span><i class="is-epoch" aria-hidden="true"></i>Epoch static estimate</span>
  </div>
  <div class="benchmark-dif-plot" data-role="plot"></div>
  <p class="benchmark-dif-status" data-role="status" aria-live="polite">Loading discrimination estimates…</p>
  <p class="benchmark-dif-footnote">The band removes uncertainty in the curve’s observation-weighted overall level, isolating uncertainty in its temporal shape. Hover over, tap, or use the arrow keys for the ordinary absolute interval. Interpolated months receive zero weight; blank periods fall outside the observed range.</p>
  <p class="sr-only" data-role="readout" aria-live="polite"></p>
</section>

<p>There are two benchmarks that display a unique, peculiar U-shape: VPCT and GSM8K. I’m not quite sure why this is. Other than those two, most benchmarks have either relatively static discriminations or decreasing discriminations over time, the most prominent of which are ARC-AGI-2, DeepSWE, GeoBench, and GDPval. There aren’t actually any benchmarks I can confidently say have had increasing discriminations, rather than the weird U-shape. This might be because model capabilities eventually progress past the point at which benchmarks have peak discriminative ability and into a region where they become increasingly poor at discriminating between models, because they don’t have enough items of the requisite difficulty.</p>

<p>What happens when we allow a benchmark’s difficulty to vary depending on model release date?</p>

<section class="benchmark-dif-interactive" data-benchmark-dif-chart="" data-kind="difficulty" data-default-benchmark="b7" data-data-url="/assets/jsons/benchmark_dif.json" aria-labelledby="benchmark-difficulty-title">
  <div class="benchmark-dif-heading">
    <div>
      <h2 id="benchmark-difficulty-title" data-role="title">Benchmark difficulty over time</h2>
      <p>Absolute posterior mean EDI with a 90% credible band for temporal shape</p>
    </div>
    <label class="benchmark-dif-control">
      <span>Benchmark</span>
      <select data-role="benchmark-select" disabled="">
        <option>Loading benchmarks…</option>
      </select>
    </label>
  </div>
  <div class="benchmark-dif-marker-key" aria-label="Month marker legend">
    <span><i class="is-shape-band" aria-hidden="true"></i>90% temporal-shape band</span>
    <span><i aria-hidden="true"></i>Month with observations</span>
    <span><i class="is-interpolated" aria-hidden="true"></i>Interpolated month</span>
    <span><i class="is-epoch" aria-hidden="true"></i>Epoch static estimate</span>
  </div>
  <div class="benchmark-dif-plot" data-role="plot"></div>
  <p class="benchmark-dif-status" data-role="status" aria-live="polite">Loading difficulty estimates…</p>
  <p class="benchmark-dif-footnote">The band removes uncertainty in the curve’s observation-weighted overall level, isolating uncertainty in its temporal shape. Hover over, tap, or use the arrow keys for the ordinary absolute interval. Interpolated months receive zero weight; blank periods fall outside the observed range.</p>
  <p class="sr-only" data-role="readout" aria-live="polite"></p>
</section>

<p>Benchmarks with increasing difficulty include Winogrande, Fiction.LiveBench, DeepResearch Bench, and ARC AI2, which seem to share a focus on language tasks. Benchmarks with decreasing difficulty include DeepSWE, GSM8K, MATH Level 5, and OSWorld. DeepSWE is the benchmark that’s exhibited the biggest decrease in difficulty, so much so that it’s an outlier. It’s also focused on long-horizon software engineering, which seems to be a popular focus of labs as of late, a fact which I don’t think is a coincidence. The other fast-decreasing benchmarks are focused on math and software ability, which have also been a focus of labs recently.</p>

<p>Now, I don’t think it makes sense to think of benchmarks as literally getting more or less difficult. Ideally, they should be the same difficulty across time. So I think it’s best to interpret increasing benchmark “difficulty” as models improving on that benchmark more slowly than they’re improving in general capability, and decreasing benchmark “difficulty” as models improving on that benchmark faster than they’re improving in general capability. Under that interpretation, the pattern makes sense. We see models increasing their math and coding capabilities faster than their general ability, while improving at language tasks more slowly than their general ability. This makes sense. Labs are focusing heavily on math and coding through post-training regimes such as RLVR, while comparatively less attention is being paid to language tasks.</p>

<p>Is there a correlation between average discrimination changes and average difficulty changes?</p>

<section class="benchmark-dif-interactive benchmark-dif-scatter-card" data-benchmark-dif-scatter="" data-data-url="/assets/jsons/benchmark_dif.json" aria-labelledby="benchmark-change-correlation-title">
  <div class="benchmark-dif-heading">
    <div>
      <h2 id="benchmark-change-correlation-title">Average monthly changes</h2>
      <p data-role="subtitle">Each point shows one benchmark’s posterior-mean change.</p>
    </div>
  </div>
  <div class="benchmark-dif-correlation-summary" data-role="correlation-summary" hidden="">
    <div class="benchmark-dif-correlation-row">
      <span class="benchmark-dif-correlation-label">All benchmarks</span>
      <strong data-role="correlation-all-estimate"></strong>
      <span class="benchmark-dif-correlation-interval" data-role="correlation-all-interval"></span>
    </div>
    <div class="benchmark-dif-correlation-row">
      <span class="benchmark-dif-correlation-label">Without DeepSWE and VPCT</span>
      <strong data-role="correlation-restricted-estimate"></strong>
      <span class="benchmark-dif-correlation-interval" data-role="correlation-restricted-interval"></span>
    </div>
  </div>
  <div class="benchmark-dif-scatter" data-role="plot"></div>
  <p class="benchmark-dif-status" data-role="status" aria-live="polite">Loading benchmark changes…</p>
  <p class="benchmark-dif-footnote">The correlation is recalculated for every paired posterior draw; its interval therefore includes uncertainty in both temporal parameters. Hover over or tap a point to see the benchmark and exact posterior-mean changes. Selecting a point also updates both trend charts.</p>
  <p class="sr-only" data-role="readout" aria-live="polite"></p>
</section>

<p>If we ignore the outliers that are DeepSWE and VPCT, our mean estimate is that the correlation is exactly… 0. So, it doesn’t seem like there’s anything interesting there.</p>

<p>So, should we throw out the ECI because it uses static parameters and try something new? No. It turns out that predicted model capabilities taking changing parameters into account correlate at 0.995 with Epoch’s predicted model capabilities, making them nearly identical. The same holds for benchmark difficulties, which correlate at 0.98, though at the extremes my estimates suggest that the hardest benchmarks aren’t quite as hard as Epoch estimates and the easiest benchmarks aren’t quite as easy. Benchmark discriminations are less correlated, at only 0.85, but there’s no discernible systematic pattern to the differences.</p>

<p>Still, it does mean we need to be careful when treating AI capability as a unitary construct, as the ECI does. Which capabilities are most relevant changes over time, partly because labs shift their focus. It used to be reading and understanding language, but as models became proficient at that, attention shifted toward coding and mathematics. A chart that represents AI progress with a single capability score is therefore somewhat misleading: what counts as capability changes over time. Progress in currently relevant capabilities, such as mathematics and coding, may be underestimated by combining them with benchmarks that emphasize domains in which progress is now slower, such as general language understanding. Those benchmarks pull the construct the ECI is measuring towards slower-moving domains.</p>

<p>In conclusion, the ECI remains a (very) useful summary of average benchmark performance, but it should not be mistaken for a measure of a single, unchanging construct. We should take care to think about <em>which</em> capabilities we think are useful to measure.</p>

<h2 id="appendix">Appendix</h2>

<h3 id="the-real-treasure-was-the-models-we-made-along-the-way">The Real Treasure Was The Models We Made Along The Way</h3>

<p>I didn’t begin with the final two-stage model. I started by reproducing Epoch’s estimator as literally as possible, then changed one assumption at a time, to ensure I wasn’t making any silly mistakes. Expand the steps below to follow the progression from the public ECI implementation to the two-stage model used in this post. (You <em>could</em> always just skip to the final model, but I think it’s easier to go step-by-step.)</p>

<p>The code, frozen data, and compact results for all six models are available in the <a href="https://github.com/stanichor/eci-temporal-invariance">accompanying GitHub repository</a>.</p>

<details class="model-step">
  <summary><h4 id="bayesianizing-eci">1. Bayesianizing the ECI</h4></summary>

  <p>The core of Epoch’s model is, as described above, the logistic function (<a href="https://epoch.ai/data/eci-documentation/methodology">source</a>) (<a href="https://github.com/epoch-research/eci-public/blob/main/src/eci/fitting.py#L242">line 242</a> of the ECI implementation):</p>

\[\mu_{mb} = \sigma(\alpha_b[C_m-D_b])\]

  <p>where:</p>

\[\sigma(x) = \frac{1}{1+e^{-x}}\]

  <p><strong>Free parameter vector</strong></p>

  <p>There are $M$ models and $B$ benchmarks. Epoch’s implementation estimates:</p>

\[\phi = (C_1, \dots, C_M,D_1,\dots,D_B,\alpha_1,\dots,\alpha_{B-1})\]

  <p>There is one benchmark discrimination excluded from the free parameter vector: Winogrande’s discrimination, which is fixed at 1. The number of free parameters is therefore:</p>

\[K = M + 2B - 1\]

  <p>The raw parameter bounds are (<a href="https://github.com/epoch-research/eci-public/blob/main/src/eci/fitting.py#L257-L266">lines 257-266</a> of the ECI implementation):</p>

\[-10 \leq C_m \leq 10\]

\[-10 \leq D_b \leq 10\]

\[0.1 \leq \alpha_b \leq 10\]

  <p>Observed scores are clipped to $[0.001,0.999]$ before fitting (<a href="https://github.com/epoch-research/eci-public/blob/main/src/eci/fitting.py#L196-L197">lines 196-197</a> of the ECI implementation).</p>

  <p><strong>Implemented residual vector</strong></p>

  <p>For every observed score, the code (<a href="https://github.com/epoch-research/eci-public/blob/main/src/eci/fitting.py#L243">line 243</a> of the ECI implementation) supplies SciPy with its residual:</p>

\[r_{mb}(\phi) = \mu_{mb}(\phi) - s_{mb}\]

  <p>It also appends a residual used for regularization (<a href="https://github.com/epoch-research/eci-public/blob/main/src/eci/fitting.py#L245-L246">lines 245-246</a> of the ECI implementation):</p>

\[r_\text{reg}(\phi) = \sqrt{\lambda \frac{1}{K} \sum_{k=1}^{K}\phi_k^2}\]

  <p>where the default regularization strength, $\lambda$ is set to 0.1 (<a href="https://github.com/epoch-research/eci-public/blob/main/src/eci/fitting.py#L138">line 138</a> of the ECI implementation).</p>

  <p>Scipy’s <code class="language-plaintext highlighter-rouge">least_squares</code> minimizes one-half of the sum of squared residuals, so Epoch’s actual implemented objective is:</p>

\[\begin{aligned}
\mathcal{L}(\phi) &amp;= \frac{1}{2}\left(\left(\sum_{(m,b) \in \mathcal{O}} r_{mb}(\phi)^2\right) + r_\text{reg}(\phi)^2\right) \\
&amp;= \frac{1}{2}\left(\left(\sum_{(m,b) \in \mathcal{O}} [\mu_{mb}(\phi) - s_{mb}]^2\right) + \lambda\frac{1}{K}\sum_{k=1}^{K}\phi_k^2\right) \\
&amp;= \frac{1}{2}\sum_{(m,b) \in \mathcal{O}}[\mu_{mb}(\phi) - s_{mb}]^2 + \frac{1}{2}\frac{\lambda}{K}\sum_{k=1}^{K}\phi_k^2
\end{aligned}\]

  <p><strong>Public scale</strong></p>

  <p>After fitting, Epoch converts the raw capabilities to the official ECI scale using an affine transformation while setting Claude 3.5 Sonnet at 130 and GPT-5 at 150.</p>

  <p><strong>Exact MAP-equivalent model</strong></p>

  <p>A model’s performance on a benchmark is modeled as:</p>

\[s_{mb}\mid\phi
\sim
\mathcal N\left(\mu_{mb}(\phi),\sigma_\varepsilon^2\right).\]

  <p>where $\sigma_\varepsilon$ is a fixed residual standard deviation.</p>

  <p>Ignoring constants, the negative log-likelihood is:</p>

\[-\log p(s\mid\phi)
=
\frac{1}{2\sigma_\varepsilon^2}
\sum_{(m,b)\in\mathcal O}
\left[s_{mb}-\mu_{mb}(\phi)\right]^2.\]

  <p>We place independent Gaussian priors with a common standard deviation $\tau$ on the free parameters, subject to the same bounds used by Epoch:</p>

\[C_m\sim\operatorname{TruncatedNormal}(0,\tau^2;-10,10),\]

\[D_b\sim\operatorname{TruncatedNormal}(0,\tau^2;-10,10),\]

  <p>and:</p>

\[\alpha_b
\sim
\operatorname{TruncatedNormal}(0,\tau^2;0.1,10)\]

  <p>for each non-Winogrande benchmark. The Winogrande slope remains fixed:</p>

\[\alpha_{\text{Winogrande}}=1.\]

  <p>Away from the bounds, the negative log-prior is:</p>

\[-\log p(\phi)
=
\frac{1}{2\tau^2}
\sum_{k=1}^{K}\phi_k^2
+\text{constant}\]

  <p>So, the negative log-posterior is</p>

\[-\log p(s\mid\phi) - \log p(\phi) = \frac{1}{2\sigma_\varepsilon^2}
\sum_{(m,b)\in\mathcal O}
\left[s_{mb}-\mu_{mb}(\phi)\right]^2 + \frac{1}{2\tau^2}
\sum_{k=1}^{K}\phi_k^2\]

  <p>If we multiply the negative log-posterior by $\sigma_\epsilon^2$, we get:</p>

\[\frac{1}{2}
\sum_{(m,b)\in\mathcal O}
\left[s_{mb}-\mu_{mb}(\phi)\right]^2
+
\frac{1}{2}
\frac{\sigma_\varepsilon^2}{\tau^2}
\sum_{k=1}^{K}\phi_k^2\]

  <p>This matches Epoch’s objective when:</p>

\[\boxed{
\frac{\lambda}{K}
=
\frac{\sigma_\varepsilon^2}{\tau^2}
}\]

  <p>or equivalently:</p>

\[\boxed{
\tau
=
\sigma_\varepsilon\sqrt{\frac{K}{\lambda}}
}.\]

  <p>It’s important to note that there’s no unique pair $(\sigma_\epsilon, \tau)$ implied by the objective. Only the ratio is implied. We could set $\sigma_\epsilon = 1$, but, since the scores lie in $[0.001, 0.999]$, a <em>residual</em> standard deviation of 1 would be <em>absurdly</em> large on the observed scale.</p>

  <p>What’s more, this also affects $\tau$. For example, if $K = 379$ and $\lambda = 0.1$, then setting $\epsilon_e = 1$ implies</p>

\[\tau = \sqrt{379/0.1} \approx 61.6\]

  <p>This is nearly flat over the capability and difficulty bounds $[-10, 10]$ and the discrimination bounds $[0.1, 10]$. As such, the penalty isn’t really doing much.</p>

  <p>Unfortunately, I first ended up doing taking the ‘convenient’ approach of setting $\sigma_\epsilon = 1$, and so while the MAP estimates match exactly, the posterior means are all over the place.</p>

  <div class="model-comparison-figures">
  <figure>
    <img src="/assets/images/benchmark-dif/models/model-1-map.png" alt="Epoch parameters compared with Model 1 Bayesian MAP estimates" />
  </figure>
  <figure>
    <img src="/assets/images/benchmark-dif/models/model-1-mcmc-mean.png" alt="Epoch parameters compared with Model 1 Bayesian posterior means" />
  </figure>
</div>

</details>

<details class="model-step">
  <summary><h4 id="learning-the-scale">2. Learning the Residual Scale</h4></summary>

  <p>The first model set $\sigma_\varepsilon$ to 1 because that was convenient. However, a residual standard deviation of 1 is absurdly large for scores bounded between 0 and 1 (well, technically between 0.001 and 0.999). Given that only the ratio $\frac{\sigma_\varepsilon^2}{\tau^2}=\frac{\lambda}{K}$ matters, it might be better to let the model fit the residual standard deviation instead.</p>

  <p>So, the second model I fit assigns</p>

\[\log\sigma_\varepsilon\sim\mathcal N(\log 0.1,1)\]

  <p>and deterministically sets</p>

\[\tau=\sigma_\varepsilon\sqrt{\frac{K}{\lambda}}.\]

  <p>For any value of $\sigma_\varepsilon$, the model has the same variance-to-penalty ratio as Epoch, but now it can learn the appropriate residual variance from the data. Everything else remains unchanged: Winogrande’s discrimination is fixed at 1, capabilities and difficulties remain bounded to $[-10,10]$, and the other discriminations remain bounded to $[0.1,10]$. As you can see, the MAP estimates still match Epoch’s exactly, while the posterior mean estimates are much more reasonable, though there’s a bend in the curve starting below an ECI of ~120 such that the Bayesian estimates are higher than Epoch’s estimates.</p>

  <div class="model-comparison-figures">
  <figure>
    <img src="/assets/images/benchmark-dif/models/model-2-map.png" alt="Epoch parameters compared with Model 2 Bayesian MAP estimates" />
  </figure>
  <figure>
    <img src="/assets/images/benchmark-dif/models/model-2-mcmc-mean.png" alt="Epoch parameters compared with Model 2 Bayesian posterior means" />
  </figure>
</div>

</details>

<details class="model-step">
  <summary><h4 id="unfixing-winogrande">3. Unfixing Winogrande</h4></summary>

  <p>The next model estimates Winogrande’s discrimination along with every other benchmark discrimination. There are now</p>

\[K=M+2B\]

  <p>free ECI parameters, and every discrimination receives the same bounded Gaussian prior:</p>

\[\alpha_b\sim\operatorname{TruncatedNormal}(0,\tau^2;0.1,10).\]

  <p>Removing the Winogrande anchor creates a scale identifiability issue. For any $a&gt;0$,</p>

\[C_m'=aC_m,\qquad D_b'=aD_b,\qquad \alpha_b'=\frac{\alpha_b}{a}\]

  <p>produces exactly the same predictions. The likelihood is also unchanged when the same constant is added to every capability and difficulty. The bounded proper priors make the posterior proper and select a particular origin and unit, but I still don’t like the model.</p>

  <p>Nevertheless, it’s still a useful intermediate model. If we’re to model changes in discriminations, we can’t also fix the discrimination of one of the benchmarks. The model produces MAP estimates, that no longer exactly match those of Epoch, but are still extremely close. The posterior mean estimates have also gotten closer to Epoch’s estimates.</p>

  <div class="model-comparison-figures">
  <figure>
    <img src="/assets/images/benchmark-dif/models/model-3-map.png" alt="Epoch parameters compared with Model 3 Bayesian MAP estimates" />
  </figure>
  <figure>
    <img src="/assets/images/benchmark-dif/models/model-3-mcmc-mean.png" alt="Epoch parameters compared with Model 3 Bayesian posterior means" />
  </figure>
</div>

</details>

<details class="model-step">
  <summary><h4 id="symmetric-identification">4. Replacing the Anchor and Bounds with Symmetric Identification</h4></summary>

  <p>Our next model removes the hard bounds and resolves the identifiability issue without privileging a particular model or benchmark.</p>

  <p>Capabilities and difficulties are combined into one vector and assigned a zero-sum Gaussian prior:</p>

\[(C_1,\ldots,C_M,D_1,\ldots,D_B)
\sim\operatorname{ZeroSumNormal}(\tau),\]

  <p>which enforces</p>

\[\sum_m C_m+\sum_bD_b=0.\]

  <p>This identifies the origin of the latent scale. Discriminations are modeled on a logarithmic scale:</p>

\[s_\alpha\sim\operatorname{HalfNormal}(0.5),\]

\[(\log\alpha_1,\ldots,\log\alpha_B)
\sim\operatorname{ZeroSumNormal}(s_\alpha).\]

  <p>Therefore,</p>

\[\sum_b\log\alpha_b=0\]

  <p>and the geometric mean of the discriminations is one:</p>

\[\left(\prod_b\alpha_b\right)^{1/B}=1.\]

  <p>This identifies the multiplicative scale. This static model becomes the foundation for the parameter-varying models. Also, our posterior mean estimates are finally lining up with Epoch’s estimates, which is nice.</p>

  <div class="model-comparison-figures">
  <figure>
    <img src="/assets/images/benchmark-dif/models/model-4-map.png" alt="Epoch parameters compared with Model 4 Bayesian MAP estimates" />
  </figure>
  <figure>
    <img src="/assets/images/benchmark-dif/models/model-4-mcmc-mean.png" alt="Epoch parameters compared with Model 4 Bayesian posterior means" />
  </figure>
</div>

</details>

<details class="model-step">
  <summary><h4 id="temporal-discriminations">5. Allowing Discriminations to Change Over Time</h4></summary>

  <p>The first parameter-varying model keeps capabilities and difficulties static but allows benchmark discriminations to depend on model-release month. Let $t_{0b}$ be benchmark $b$’s earliest observed month. Its initial log discrimination receives the prior described above:</p>

\[s_\alpha\sim\operatorname{HalfNormal}(0.5),\]

\[(a_{1,0},\ldots,a_{B,0})
\sim\operatorname{ZeroSumNormal}(s_\alpha).\]

  <p>Temporal change follows a stationary RBF Gaussian process. All benchmarks share the same length scale,</p>

\[\ell_\alpha\sim\operatorname{Uniform}(2,36),\]

  <p>which is measured in months. The covariance function is given by</p>

\[K_{\alpha,tt'}
=
\exp\left[-\frac{(t-t')^2}{2\ell_\alpha^2}\right]+10^{-4}I.\]

  <p>Each benchmark has its own Gaussian process trajectory $f_b$:</p>

\[f_b\sim\mathcal N(\mathbf 0,K_\alpha),\]

\[\kappa_b\sim\operatorname{HalfNormal}(0.5).\]

  <p>A benchmark’s monthly discrimination is given by</p>

\[\log\alpha_{b,t}
=
a_{b,0}+\kappa_b\left(f_{b,t}-f_{b,t_{0b}}\right).\]

  <p>The likelihood for observed scores is</p>

\[s_{mb}
\sim
\mathcal N\left(
\sigma\left[\alpha_{b,t_m}(C_m-D_b)\right],
\sigma_\varepsilon^2
\right).\]

  <p>When a single discrimination is needed for a model (such as in the following graphs), I use the geometric mean across a benchmark’s supported months. This is the first stage of the final analysis.</p>

  <div class="model-comparison-figures">
  <figure>
    <img src="/assets/images/benchmark-dif/models/model-5-map.png" alt="Epoch parameters compared with Model 5 Bayesian MAP estimates" />
  </figure>
  <figure>
    <img src="/assets/images/benchmark-dif/models/model-5-mcmc-mean.png" alt="Epoch parameters compared with Model 5 Bayesian posterior means" />
  </figure>
</div>

</details>

<details class="model-step">
  <summary><h4 id="two-stage-temporal-model">6. The Final Two-Stage Model</h4></summary>

  <p>The analysis in this post uses a two-stage model. Stage 1 was the model described I just described above. For Stage 2, rather than passing only Stage 1’s posterior means, I use discrimination draws</p>

\[\boldsymbol\alpha^{(q)}
=
\{\alpha^{(q)}_{b,t}:b=1,\ldots,B;\ t=1,\ldots,T\}.\]

  <p>This allows us to use and model the posterior distribution of Stage 1’s results, rather than treating the posterior mean as the only possible set of discriminations.</p>

  <p>Each draw is normalized such that its geometric mean of the discriminations is equal to 1. If $\mathcal T_b$ is benchmark $b$’s supported calendar range, define the average log discrimination, $g^{(q)}$ as</p>

\[g^{(q)}
=
\frac{1}{B}
\sum_b
\left[
\frac{1}{|\mathcal T_b|}
\sum_{t\in\mathcal T_b}
\log\alpha^{(q)}_{b,t}
\right].\]

  <p>The draw we end up using is</p>

\[\widetilde\alpha^{(q)}_{b,t}
=
\exp\left[\log\alpha^{(q)}_{b,t}-g^{(q)}\right].\]

  <p>This normalization does not alter relative differences between benchmarks or temporal changes within a benchmark; it only fixes the otherwise arbitrary unit of the latent scale.</p>

  <p>For every propagated draw, I fit a model in which the difficulties can vary across time. The set of discriminations is fixed within that fit, while capabilities and benchmark difficulties are re-estimated. Difficulty deviations follow another shared-length-scale RBF process:</p>

\[\ell_D\sim\operatorname{Uniform}(2,36),\]

\[h_b\sim\mathcal N(\mathbf 0,K_D),\]

\[\omega_b\sim\operatorname{HalfNormal}(0.5).\]

  <p>Let the raw temporal-difficulty be</p>

\[r_{b,t}=\omega_bh_{b,t}.\]

  <p>The raw temporal-difficulties could absorb both benchmark averages and a movement shared by all benchmarks in a calendar month. To prevent that, we impose the constraint</p>

\[\sum_{t\in\mathcal T_b}\delta_{b,t}=0\]

  <p>for every benchmark. Thus, the deviations average to zero across each benchmark’s supported months and do not change its overall difficulty. We also impose</p>

\[\sum_{b:(b,t)\text{ supported}}\delta_{b,t}=0\]

  <p>for every month. Thus, within each month, the deviations average to zero across the benchmarks used in that month. So deviations measure how benchmark difficulties change <em>relative</em> to one another. Monthly difficulty is then</p>

\[D_{b,t}=\bar D_b+\delta_{b,t},\]

  <p>and the conditional likelihood is</p>

\[s_{mb}
\sim
\mathcal N\left(
\sigma\left[
\widetilde\alpha^{(q)}_{b,t_m}
(C_m-D_{b,t_m})
\right],
\sigma_\varepsilon^2
\right).\]

  <p>I repeat this fit for complete surfaces drawn from Stage 1 and pool the conditional posteriors:</p>

\[p_{\mathrm{mod}}(\Theta_D\mid s)
\approx
\frac{1}{Q}
\sum_{q=1}^{Q}
p_2\left(
\Theta_D\mid s,\widetilde{\boldsymbol\alpha}^{(q)}
\right).\]

  <p>This results in our final estimates, which, when we ignore changes in discriminations and difficulties over time, match Epoch’s estimates pretty well.</p>

  <div class="model-comparison-figures">
  <figure>
    <img src="/assets/images/benchmark-dif/models/model-6-map.png" alt="Epoch parameters compared with Model 6 Bayesian MAP estimates" />
  </figure>
  <figure>
    <img src="/assets/images/benchmark-dif/models/model-6-mcmc-mean.png" alt="Epoch parameters compared with Model 6 Bayesian posterior means" />
  </figure>
</div>

</details>

<script src="/assets/js/benchmark-dif.js"></script>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:vocab" role="doc-endnote">
      <p>Marc Brysbaert, Paweł Mandera, Samantha F. McCormick, and Emmanuel Keuleers, “Word Prevalence Norms for 62,000 English Lemmas,” Behavior Research Methods 51 (2019): 467–479, https://doi.org/10.3758/s13428-018-1077-9 <a href="#fnref:vocab" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name></name></author><category term="ai-benchmarks" /><summary type="html"><![CDATA[]]></summary></entry></feed>