<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://davecturner.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://davecturner.github.io/" rel="alternate" type="text/html" /><updated>2026-03-20T16:53:31+00:00</updated><id>https://davecturner.github.io/feed.xml</id><title type="html">David Turner says…</title><subtitle>Github Pages</subtitle><author><name></name></author><entry><title type="html">TCP backpressure</title><link href="https://davecturner.github.io/2025/12/21/tcp-backpressure.html" rel="alternate" type="text/html" title="TCP backpressure" /><published>2025-12-21T00:00:00+00:00</published><updated>2025-12-21T00:00:00+00:00</updated><id>https://davecturner.github.io/2025/12/21/tcp-backpressure</id><content type="html" xml:base="https://davecturner.github.io/2025/12/21/tcp-backpressure.html"><![CDATA[<p>Backpressure in a distributed system allows receiving nodes to notify sending
nodes that they temporarily lack the capacity to handle further requests.
Without backpressure the receiving node must attempt to handle every request
and, if overloaded, must shed load by returning error responses, often causing
more work as the sender retries the failed requests.</p>

<p>A very useful (if somewhat misunderstood) feature of TCP is its built-in
support for backpressure. Every TCP packet includes a 16-bit <em>window</em> field in
its header which indicates how many more bytes the sender can accept. Typically
this value is shifted by some number of bits known as the <em>window scale</em> set
with a TCP option during the initial handshake (and potentially modified later)
as defined in <a href="https://datatracker.ietf.org/doc/html/rfc7323#section-2">RFC
7323</a>. A node,
receiving a stream of data faster than it can handle it, applies backpressure
by reducing the window size in its TCP acknowledgements. Crucially, the node
may reduce this window size all the way down to zero, indicating that it can
handle no more data. This backpressure mechanism is built into the operating
system’s TCP implementation and the only thing that an application needs to do
to access it is to stop reading data from the socket in question while it is
overloaded. Once the overload has passed, the application reads from the socket
again which causes the kernel to transmit a packet advertising that the window
is open again, inviting the sender to continue sending data.</p>

<h2 id="infinite-patience">Infinite patience</h2>

<p>One of the misunderstandings about this mechanism is the belief that there must
be some kind of timeout after which a connection in the zero-window
backpressure state indicates an error condition and should be closed. <a href="https://datatracker.ietf.org/doc/html/rfc1122#page-92">RFC
1122</a> says that this is
not the case:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>A TCP MAY keep its offered receive window closed
indefinitely.  As long as the receiving TCP continues to
send acknowledgments in response to the probe segments, the
sending TCP MUST allow the connection to stay open.
</code></pre></div></div>

<p>There is no need for any timeout here because the zero-window state is actively
maintained by the two endpoints: the prospective sender repeatedly sends
so-called zero-window probes to which the receiver responds, indicating that
both ends remain alive but that the backpressure situation persists. This is
essential because when the backpressure is released the receiver sends a single
window-open advertisement but this packet may be lost. As long as the sender
eventually sends another probe it will eventually discover that the
backpressure has been released.</p>

<p>These repeated zero-window probes also ensure that the sender eventually
detects a network partition by watching for a sufficiently long sequence of
consecutive probes to which it has not received a response, as if it were
sending TCP keepalives.</p>

<h2 id="probe-timings">Probe timings</h2>

<p><a href="https://datatracker.ietf.org/doc/html/rfc1122#page-92">RFC 1122</a> has some
recommendations about the exact timings of the zero-window probes:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The transmitting host SHOULD send the first zero-window
probe when a zero window has existed for the retransmission
timeout period (see Section 4.2.2.15), and SHOULD increase
exponentially the interval between successive probes.

[...]                   Exponential backoff is
recommended, possibly with some maximum interval not
specified here.
</code></pre></div></div>

<p>In practice a maximum is essential, or else the sender may wait for
unreasonably long before discovering that the window has reopened. For example,
if it backed off by a factor of 2 on every probe with no maximum and the
window-opening packet went undelivered then the sender would effectively wait
for twice the length of the backpressure period before discovering that the
window has reopened.</p>

<p>But how exactly are these timings calculated?</p>

<p>In Linux by default the zero-window probes are scheduled similarly to regular
<a href="/2025/12/02/tcp-retries2.html">retransmissions</a>, starting at
<code class="language-plaintext highlighter-rouge">RTO_MIN</code> (200ms) and backing off repeatedly by a factor of 2 up to a maximum
of <code class="language-plaintext highlighter-rouge">RTO_MAX</code> (2 minutes), which it reaches after the backpressure has lasted
for a little under 3½ minutes:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: right">Probe</th>
      <th style="text-align: right">Start/mm:ss</th>
      <th style="text-align: right">Timeout</th>
      <th style="text-align: right">Timeout/s</th>
      <th style="text-align: right">End/mm:ss.s</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: right">0</td>
      <td style="text-align: right">0:00.0</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MIN</code></td>
      <td style="text-align: right">0.2</td>
      <td style="text-align: right">0:00.2</td>
    </tr>
    <tr>
      <td style="text-align: right">1</td>
      <td style="text-align: right">0:00.2</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">2 × RTO_MIN</code></td>
      <td style="text-align: right">0.4</td>
      <td style="text-align: right">0:00.6</td>
    </tr>
    <tr>
      <td style="text-align: right">2</td>
      <td style="text-align: right">0:00.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">4 × RTO_MIN</code></td>
      <td style="text-align: right">0.8</td>
      <td style="text-align: right">0:01.4</td>
    </tr>
    <tr>
      <td style="text-align: right">3</td>
      <td style="text-align: right">0:01.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">8 × RTO_MIN</code></td>
      <td style="text-align: right">1.6</td>
      <td style="text-align: right">0:03.0</td>
    </tr>
    <tr>
      <td style="text-align: right">4</td>
      <td style="text-align: right">0:03.0</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">16 × RTO_MIN</code></td>
      <td style="text-align: right">3.2</td>
      <td style="text-align: right">0:06.2</td>
    </tr>
    <tr>
      <td style="text-align: right">5</td>
      <td style="text-align: right">0:06.2</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">32 × RTO_MIN</code></td>
      <td style="text-align: right">6.4</td>
      <td style="text-align: right">0:12.6</td>
    </tr>
    <tr>
      <td style="text-align: right">6</td>
      <td style="text-align: right">0:12.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">64 × RTO_MIN</code></td>
      <td style="text-align: right">12.8</td>
      <td style="text-align: right">0:25.4</td>
    </tr>
    <tr>
      <td style="text-align: right">7</td>
      <td style="text-align: right">0:25.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">128 × RTO_MIN</code></td>
      <td style="text-align: right">25.6</td>
      <td style="text-align: right">0:51.0</td>
    </tr>
    <tr>
      <td style="text-align: right">8</td>
      <td style="text-align: right">0:51.0</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">258 × RTO_MIN</code></td>
      <td style="text-align: right">51.2</td>
      <td style="text-align: right">1:42.2</td>
    </tr>
    <tr>
      <td style="text-align: right">9</td>
      <td style="text-align: right">1:42.2</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">512 × RTO_MIN</code></td>
      <td style="text-align: right">102.4</td>
      <td style="text-align: right">3:24.6</td>
    </tr>
    <tr>
      <td style="text-align: right">10</td>
      <td style="text-align: right">3:24.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">120.0</td>
      <td style="text-align: right">5:24.6</td>
    </tr>
    <tr>
      <td style="text-align: right">11</td>
      <td style="text-align: right">5:24.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">120.0</td>
      <td style="text-align: right">7:24.6</td>
    </tr>
    <tr>
      <td style="text-align: right">12</td>
      <td style="text-align: right">7:24.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">120.0</td>
      <td style="text-align: right">9:24.6</td>
    </tr>
    <tr>
      <td style="text-align: right">13</td>
      <td style="text-align: right">9:24.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">120.0</td>
      <td style="text-align: right">11:24.6</td>
    </tr>
    <tr>
      <td style="text-align: right">14</td>
      <td style="text-align: right">11:24.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">120.0</td>
      <td style="text-align: right">13:24.6</td>
    </tr>
    <tr>
      <td style="text-align: right">⋮</td>
      <td style="text-align: right">⋮</td>
      <td style="text-align: right">⋮</td>
      <td style="text-align: right">⋮</td>
      <td style="text-align: right">⋮</td>
    </tr>
  </tbody>
</table>

<p>Thus the default behaviour is that the resolution of a backpressure situation
which lasted for just a few minutes might not be noticed for a further two
minutes (<code class="language-plaintext highlighter-rouge">RTO_MAX</code>) in the unfortunate, but not uncommon, event that the single
window-opening packet goes undelivered. Occasionally the first probes after the
window opens may also go unacknowledged, each time adding another two minutes
to any backpressure-related delays. That seems awfully long to me.</p>

<p>This also raises the question of how the system deals with unacknowledged
probes, such as would happen if the network were partitioned or the receiving
process were no longer running.</p>

<p>The answer is that the <code class="language-plaintext highlighter-rouge">tcp_retries2</code> sysctl works similarly on zero-window
probes to how it works with regular retransmissions: if more than
<code class="language-plaintext highlighter-rouge">tcp_retries2</code> consecutive zero-window probes go unacknowledged then the
connection fails. With the default value of <code class="language-plaintext highlighter-rouge">15</code>, this means that a connection
across a network partition might not be considered faulty for a whopping 30
minutes (<code class="language-plaintext highlighter-rouge">15 × RTO_MAX</code>) after the start of the partition. Yikes!</p>

<h2 id="less-is-more">Less is more</h2>

<p>The best way to reduce the time it takes to detect a network partition in a
backpressure situation is to reduce <code class="language-plaintext highlighter-rouge">tcp_retries2</code> to a <a href="/2025/12/02/tcp-retries2.html">more reasonable
value</a>, just as in the non-backpressure
case. By setting <code class="language-plaintext highlighter-rouge">tcp_retries2</code> to <code class="language-plaintext highlighter-rouge">5</code> the system will close the connection and
report the failure to the application after just 5 unacknowledged probes in a
row.</p>

<p>If the interval between the zero-window probes were allowed to grow up to
<code class="language-plaintext highlighter-rouge">RTO_MAX</code> then waiting for 5 of them to go unacknowledged would still take a
pretty dreadful 10 minutes. However, the <code class="language-plaintext highlighter-rouge">tcp_retries2</code> sysctl also limits the
time between the probes. I couldn’t find this mentioned in any documentation
but <a href="https://github.com/torvalds/linux/blob/9094662f6707d1d4b53d18baba459604e8bb0783/net/ipv4/tcp_output.c#L4583-L4584">this is where it’s implemented in
<code class="language-plaintext highlighter-rouge">tcp_send_probe0()</code></a>:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>if (icsk-&gt;icsk_backoff &lt; READ_ONCE(net-&gt;ipv4.sysctl_tcp_retries2))
    icsk-&gt;icsk_backoff++;
</code></pre></div></div>

<p>Here <code class="language-plaintext highlighter-rouge">icsk-&gt;icsk_backoff</code> is the backoff counter, visible using tools such as
<code class="language-plaintext highlighter-rouge">ss -tonie</code>, and from which the re-probe interval is computed. The effect of
this code is to stop increasing the backoff counter, and thus the re-probe
interval, once it reaches <code class="language-plaintext highlighter-rouge">tcp_retries2</code>. By default this allows the re-probe
interval to increase all the way to <code class="language-plaintext highlighter-rouge">RTO_MAX</code>, but if <code class="language-plaintext highlighter-rouge">tcp_retries2</code> is <code class="language-plaintext highlighter-rouge">5</code>
then the interval between zero-window probes will not increase beyond <code class="language-plaintext highlighter-rouge">32 ×
RTO_MIN</code> which is a little over 6 seconds:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: right">Probe</th>
      <th style="text-align: right">Start/mm:ss</th>
      <th style="text-align: right">Timeout</th>
      <th style="text-align: right">Timeout/s</th>
      <th style="text-align: right">End/mm:ss.s</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: right">0</td>
      <td style="text-align: right">0:00.0</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MIN</code></td>
      <td style="text-align: right">0.2</td>
      <td style="text-align: right">0:00.2</td>
    </tr>
    <tr>
      <td style="text-align: right">1</td>
      <td style="text-align: right">0:00.2</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">2 × RTO_MIN</code></td>
      <td style="text-align: right">0.4</td>
      <td style="text-align: right">0:00.6</td>
    </tr>
    <tr>
      <td style="text-align: right">2</td>
      <td style="text-align: right">0:00.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">4 × RTO_MIN</code></td>
      <td style="text-align: right">0.8</td>
      <td style="text-align: right">0:01.4</td>
    </tr>
    <tr>
      <td style="text-align: right">3</td>
      <td style="text-align: right">0:01.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">8 × RTO_MIN</code></td>
      <td style="text-align: right">1.6</td>
      <td style="text-align: right">0:03.0</td>
    </tr>
    <tr>
      <td style="text-align: right">4</td>
      <td style="text-align: right">0:03.0</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">16 × RTO_MIN</code></td>
      <td style="text-align: right">3.2</td>
      <td style="text-align: right">0:06.2</td>
    </tr>
    <tr>
      <td style="text-align: right">5</td>
      <td style="text-align: right">0:06.2</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">32 × RTO_MIN</code></td>
      <td style="text-align: right">6.4</td>
      <td style="text-align: right">0:12.6</td>
    </tr>
    <tr>
      <td style="text-align: right">6</td>
      <td style="text-align: right">0:12.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">32 × RTO_MIN</code></td>
      <td style="text-align: right">6.4</td>
      <td style="text-align: right">0:19.0</td>
    </tr>
    <tr>
      <td style="text-align: right">7</td>
      <td style="text-align: right">0:19.0</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">32 × RTO_MIN</code></td>
      <td style="text-align: right">6.4</td>
      <td style="text-align: right">0:25.4</td>
    </tr>
    <tr>
      <td style="text-align: right">8</td>
      <td style="text-align: right">0:25.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">32 × RTO_MIN</code></td>
      <td style="text-align: right">6.4</td>
      <td style="text-align: right">0:31.8</td>
    </tr>
    <tr>
      <td style="text-align: right">9</td>
      <td style="text-align: right">0:31.8</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">32 × RTO_MIN</code></td>
      <td style="text-align: right">6.4</td>
      <td style="text-align: right">0:38.2</td>
    </tr>
    <tr>
      <td style="text-align: right">10</td>
      <td style="text-align: right">0:38.2</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">32 × RTO_MIN</code></td>
      <td style="text-align: right">6.4</td>
      <td style="text-align: right">0:44.6</td>
    </tr>
    <tr>
      <td style="text-align: right">⋮</td>
      <td style="text-align: right">⋮</td>
      <td style="text-align: right">⋮</td>
      <td style="text-align: right">⋮</td>
      <td style="text-align: right">⋮</td>
    </tr>
  </tbody>
</table>

<p>This means that a prospective sender will be able to pick up the open window
within a few seconds even if the window-opening packet goes undelivered, losing
only a few more seconds on each undelivered probe, and a network partition will
be detected in <code class="language-plaintext highlighter-rouge">5 × 32 × RTO_MIN</code> which is a little over 30 seconds, surely
vastly preferable to the 30-minute default.</p>

<h2 id="user-timeouts">User timeouts</h2>

<p>Cloudflare has a <a href="https://blog.cloudflare.com/when-tcp-sockets-refuse-to-die/">blog post about detecting dead TCP
connections</a> which
concludes that typical applications sending data to the internet should set
Linux’s <code class="language-plaintext highlighter-rouge">TCP_USER_TIMEOUT</code> socket option to be equal to the overall TCP
keepalive timeout (i.e. <code class="language-plaintext highlighter-rouge">TCP_KEEPIDLE + TCP_KEEPINTVL * TCP_KEEPCNT</code>) so that
sockets with nonempty send buffers can still detect network partitions in as
timely a fashion as ones with empty send buffers.</p>

<p>Linux’s <code class="language-plaintext highlighter-rouge">TCP_USER_TIMEOUT</code> socket option has the following meaning according to
<a href="https://man7.org/linux/man-pages/man7/tcp.7.html"><code class="language-plaintext highlighter-rouge">man tcp</code></a>:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>TCP_USER_TIMEOUT (since Linux 2.6.37)
    This option takes an unsigned int as an argument.  When the
    value is greater than 0, it specifies the maximum amount of
    time in milliseconds that transmitted data may remain
    unacknowledged, or buffered data may remain untransmitted
    (due to zero window size) before TCP will forcibly close
    the corresponding connection and return ETIMEDOUT to the
    application.  If the option value is specified as 0, TCP
    will use the system default.

    [...]

    Further details on the user timeout feature can be found in
    RFC 793 and RFC 5482 ("TCP User Timeout Option").
</code></pre></div></div>

<p>However, <a href="https://datatracker.ietf.org/doc/html/rfc5482">RFC 5482</a> describes
a subtly different timeout:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The Transmission Control Protocol (TCP) specification [RFC0793]
defines a local, per-connection "user timeout" parameter that
specifies the maximum amount of time that transmitted data may remain
unacknowledged before TCP will forcefully close the corresponding
connection.
</code></pre></div></div>

<p>This RFC specifies a TCP option allowing endpoints to communicate such a
timeout to each other, but as far as I can tell Linux doesn’t make use of this
facility even if <code class="language-plaintext highlighter-rouge">TCP_USER_TIMEOUT</code> is set.</p>

<p>Confusingly <a href="https://datatracker.ietf.org/doc/html/rfc0793">RFC 793</a> describes
such a “user timeout” with different semantics again:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The timeout, if present, permits the caller to set up a timeout
for all data submitted to TCP.  If data is not successfully
delivered to the destination within the timeout period, the TCP
will abort the connection.  The present global default is five
minutes.
</code></pre></div></div>

<p>This is the timeout specified in Linux by the <code class="language-plaintext highlighter-rouge">SO_SNDTIMEO</code> socket option, not
<code class="language-plaintext highlighter-rouge">TCP_USER_TIMEOUT</code>.</p>

<p>The difference between the timeout described in RFC 5482 and the implementation
of the <code class="language-plaintext highlighter-rouge">TCP_USER_TIMEOUT</code> option in Linux is subtle but vitally important when
considering TCP backpressure. The RFC 5482 timeout only considers
<em>unacknowledged</em> data, but a TCP connection in a zero-window state has no
unacknowledged data and thus this timeout should have no effect. In contrast,
the <code class="language-plaintext highlighter-rouge">TCP_USER_TIMEOUT</code> socket option also considers <em>untransmitted</em> data and
thus imposes a time limit on any backpressure situation after which the
connection is closed, violating <a href="https://datatracker.ietf.org/doc/html/rfc1122#page-92">RFC
1122</a>:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>A TCP MAY keep its offered receive window closed
indefinitely.  As long as the receiving TCP continues to
send acknowledgments in response to the probe segments, the
sending TCP MUST allow the connection to stay open.
</code></pre></div></div>

<p>Unfortunately this makes this feature useless, indeed harmful, in a system that
relies on TCP backpressure. I can imagine ways that it might be appropriate to
use in Cloudflare’s particular situation, but it does not apply more generally.</p>

<p>Although the <code class="language-plaintext highlighter-rouge">tcp_retries2</code> option does get a brief mention in Cloudflare’s
post, the author fails to explore the consequences of reducing this from its
unreasonably large default of <code class="language-plaintext highlighter-rouge">15</code> down to something more sensible. Had they
done so, they might have concluded that this is a more effective solution to
the problems they were describing than the backpressure-incompatible
<code class="language-plaintext highlighter-rouge">TCP_USER_TIMEOUT</code> option.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[Backpressure in a distributed system allows receiving nodes to notify sending nodes that they temporarily lack the capacity to handle further requests. Without backpressure the receiving node must attempt to handle every request and, if overloaded, must shed load by returning error responses, often causing more work as the sender retries the failed requests.]]></summary></entry><entry><title type="html">TCP retransmissions</title><link href="https://davecturner.github.io/2025/12/02/tcp-retries2.html" rel="alternate" type="text/html" title="TCP retransmissions" /><published>2025-12-02T00:00:00+00:00</published><updated>2025-12-02T00:00:00+00:00</updated><id>https://davecturner.github.io/2025/12/02/tcp-retries2</id><content type="html" xml:base="https://davecturner.github.io/2025/12/02/tcp-retries2.html"><![CDATA[<p>Distributed systems must always be able to <em>correctly</em> deal with the
non-delivery (or
<a href="https://en.wikipedia.org/wiki/Two_Generals%27_Problem">non-acknowledgement</a>)
of a message, but in any reasonable network environment one can assume that
message non-delivery is fairly rare. This means we need not behave <em>optimally</em>
in the case of a message non-delivery. In principle the system should still
work correctly even if retransmissions were totally disabled, but in practice
TCP uses packet loss to signal network congestion and apply backpressure so
some small number of retransmissions are to be expected and need to be handled
gracefully.</p>

<p>The question is therefore how tenaciously to retry a delivery before eventually
giving up and performing some corrective action. Of course this depends on the
exact use-case but typically a few seconds of retries is going to be more than
enough. Assuming for instance that network-congestion-signalling packet loss
affects even a whopping 1% of packets, and that those packets are chosen at
random, the probability of losing a few retransmissions in a row from the same
connection quickly becomes <em>vanishingly</em> small.</p>

<p>The latest standard on the subject of retransmissions is <a href="https://datatracker.ietf.org/doc/html/rfc1122#page-101">RFC
1122</a>, which uses <code class="language-plaintext highlighter-rouge">R2</code>
to denote the overall timeout on message deliveries and specifies the
following:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The value of R2 SHOULD correspond to at least 100 seconds.
</code></pre></div></div>

<p>Note however that the scope of this RFC is rather narrow:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The requirements spelled out in this document are designed
for a full-function Internet host, capable of full
interoperation over an arbitrary Internet path.
</code></pre></div></div>

<p>Most nodes in a modern distributed system are not a <code class="language-plaintext highlighter-rouge">full-function Internet
host</code> in this sense. Aside from nodes at the edge, very little traffic in such
a system is sent over an <code class="language-plaintext highlighter-rouge">arbitrary Internet path</code>, and none goes over paths
that behave like the internet did in October 1989 when this RFC was written.
Instead, the nodes in the interior of the system will be communicating over a
much more reliable and performant network and different configuration choices
are appropriate for this kind of network environment.</p>

<p>See also this less-opinionated <a href="https://pracucci.com/linux-tcp-rto-min-max-and-tcp-retries2.html">blog post by Marco
Pracucci</a> on
the same subject.</p>

<h2 id="linux">Linux</h2>

<p>In Linux, the retransmission behaviour is controlled by <a href="https://github.com/torvalds/linux/blob/4a26e7032d7d57c998598c08a034872d6f0d3945/Documentation/networking/ip-sysctl.rst#L799-L814">the
<code class="language-plaintext highlighter-rouge">net.ipv4.tcp_retries2</code>
sysctl</a>:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>tcp_retries2 - INTEGER
    This value influences the timeout of an alive TCP connection,
    when RTO retransmissions remain unacknowledged.
    Given a value of N, a hypothetical TCP connection following
    exponential backoff with an initial RTO of TCP_RTO_MIN would
    retransmit N times before killing the connection at the (N+1)th RTO.

    The default value of 15 yields a hypothetical timeout of 924.6
    seconds and is a lower bound for the effective timeout.
    TCP will effectively time out at the first RTO which exceeds the
    hypothetical timeout.
</code></pre></div></div>

<p>The mentioned “exponential backoff” uses a factor of two, and is also bounded
above by <code class="language-plaintext highlighter-rouge">TCP_RTO_MAX</code>, where the constants are <a href="https://github.com/torvalds/linux/blob/4a26e7032d7d57c998598c08a034872d6f0d3945/include/net/tcp.h#L161-L163">defined as the following
numbers of
jiffies</a>:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>#define TCP_RTO_MAX_SEC 120
#define TCP_RTO_MAX ((unsigned)(TCP_RTO_MAX_SEC * HZ))
#define TCP_RTO_MIN ((unsigned)(HZ / 5))
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">HZ</code> is a compile-time-configurable constant representing the number of jiffies
per second, so these values work out to 200ms and 120s respectively. This is
where the documented “hypothetical timeout of 924.6 seconds” comes from:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: right">Attempt</th>
      <th style="text-align: right">Start/mm:ss</th>
      <th style="text-align: right">Timeout</th>
      <th style="text-align: right">Timeout/s</th>
      <th style="text-align: right">End/mm:ss.s</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: right">0</td>
      <td style="text-align: right">0:00.0</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MIN</code></td>
      <td style="text-align: right">0.2</td>
      <td style="text-align: right">0:00.2</td>
    </tr>
    <tr>
      <td style="text-align: right">1</td>
      <td style="text-align: right">0:00.2</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">2 × RTO_MIN</code></td>
      <td style="text-align: right">0.4</td>
      <td style="text-align: right">0:00.6</td>
    </tr>
    <tr>
      <td style="text-align: right">2</td>
      <td style="text-align: right">0:00.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">4 × RTO_MIN</code></td>
      <td style="text-align: right">0.8</td>
      <td style="text-align: right">0:01.4</td>
    </tr>
    <tr>
      <td style="text-align: right">3</td>
      <td style="text-align: right">0:01.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">8 × RTO_MIN</code></td>
      <td style="text-align: right">1.6</td>
      <td style="text-align: right">0:03.0</td>
    </tr>
    <tr>
      <td style="text-align: right">4</td>
      <td style="text-align: right">0:03.0</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">16 × RTO_MIN</code></td>
      <td style="text-align: right">3.2</td>
      <td style="text-align: right">0:06.2</td>
    </tr>
    <tr>
      <td style="text-align: right">5</td>
      <td style="text-align: right">0:06.2</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">32 × RTO_MIN</code></td>
      <td style="text-align: right">6.4</td>
      <td style="text-align: right">0:12.6</td>
    </tr>
    <tr>
      <td style="text-align: right">6</td>
      <td style="text-align: right">0:12.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">64 × RTO_MIN</code></td>
      <td style="text-align: right">12.8</td>
      <td style="text-align: right">0:25.4</td>
    </tr>
    <tr>
      <td style="text-align: right">7</td>
      <td style="text-align: right">0:25.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">128 × RTO_MIN</code></td>
      <td style="text-align: right">25.6</td>
      <td style="text-align: right">0:51.0</td>
    </tr>
    <tr>
      <td style="text-align: right">8</td>
      <td style="text-align: right">0:51.0</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">258 × RTO_MIN</code></td>
      <td style="text-align: right">51.2</td>
      <td style="text-align: right">1:42.2</td>
    </tr>
    <tr>
      <td style="text-align: right">9</td>
      <td style="text-align: right">1:42.2</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">512 × RTO_MIN</code></td>
      <td style="text-align: right">102.4</td>
      <td style="text-align: right">3:24.6</td>
    </tr>
    <tr>
      <td style="text-align: right">10</td>
      <td style="text-align: right">3:24.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">120.0</td>
      <td style="text-align: right">5:24.6</td>
    </tr>
    <tr>
      <td style="text-align: right">11</td>
      <td style="text-align: right">5:24.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">120.0</td>
      <td style="text-align: right">7:24.6</td>
    </tr>
    <tr>
      <td style="text-align: right">12</td>
      <td style="text-align: right">7:24.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">120.0</td>
      <td style="text-align: right">9:24.6</td>
    </tr>
    <tr>
      <td style="text-align: right">13</td>
      <td style="text-align: right">9:24.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">120.0</td>
      <td style="text-align: right">11:24.6</td>
    </tr>
    <tr>
      <td style="text-align: right">14</td>
      <td style="text-align: right">11:24.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">120.0</td>
      <td style="text-align: right">13:24.6</td>
    </tr>
    <tr>
      <td style="text-align: right">15</td>
      <td style="text-align: right">13:24.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">120.0</td>
      <td style="text-align: right">15:24.6</td>
    </tr>
  </tbody>
</table>

<p>I don’t think I’ve ever encountered a situation where waiting for ≥15 minutes
in case a message is finally delivered and acknowledged is the right thing to
do, and even waiting the RFC-1122-specified ≥100 seconds for 8 retransmissions
is almost always excessive.</p>

<p>If you instead set <code class="language-plaintext highlighter-rouge">tcp_retries2=5</code> then the system will report failure to
deliver a message after a little under 13s, allowing for much more prompt
corrective action. This will only happen if the connection in question failed
to deliver 6 packets in a row which is incredibly unlikely to happen naturally.
Assuming random and independent packet loss due to congestion etc. of 1%, the
loss of 6 packets in a row would have probability 0.0000000001%. Put
differently, if you find you are getting connection timeouts with
<code class="language-plaintext highlighter-rouge">tcp_retries2=5</code> then the packet loss is almost certainly non-random, and
therefore something deserving of investigation and a remedy.</p>

<h3 id="linux-615">Linux ≥6.15</h3>

<p>Starting in Linux 6.15 the <code class="language-plaintext highlighter-rouge">TCP_RTO_MAX</code> value can be adjusted at runtime via
the <code class="language-plaintext highlighter-rouge">net.ipv4.tcp_rto_max_ms</code> sysctl, affecting all connections on the system,
and applications can also override the system-wide maximum on each socket via
the <code class="language-plaintext highlighter-rouge">TCP_RTO_MAX_MS</code> socket option. For instance, setting the shortest
permitted <code class="language-plaintext highlighter-rouge">TCP_RTO_MAX_MS=1000</code> yields the following retransmission schedule:</p>

<table>
  <thead>
    <tr>
      <th style="text-align: right">Attempt</th>
      <th style="text-align: right">Start/mm:ss</th>
      <th style="text-align: right">Timeout</th>
      <th style="text-align: right">Timeout/s</th>
      <th style="text-align: right">End/mm:ss.s</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: right">0</td>
      <td style="text-align: right">0:00.0</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MIN</code></td>
      <td style="text-align: right">0.2</td>
      <td style="text-align: right">0:00.2</td>
    </tr>
    <tr>
      <td style="text-align: right">1</td>
      <td style="text-align: right">0:00.2</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">2 × RTO_MIN</code></td>
      <td style="text-align: right">0.4</td>
      <td style="text-align: right">0:00.6</td>
    </tr>
    <tr>
      <td style="text-align: right">2</td>
      <td style="text-align: right">0:00.6</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">4 × RTO_MIN</code></td>
      <td style="text-align: right">0.8</td>
      <td style="text-align: right">0:01.4</td>
    </tr>
    <tr>
      <td style="text-align: right">3</td>
      <td style="text-align: right">0:01.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">1.0</td>
      <td style="text-align: right">0:02.4</td>
    </tr>
    <tr>
      <td style="text-align: right">4</td>
      <td style="text-align: right">0:02.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">1.0</td>
      <td style="text-align: right">0:03.4</td>
    </tr>
    <tr>
      <td style="text-align: right">5</td>
      <td style="text-align: right">0:03.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">1.0</td>
      <td style="text-align: right">0:04.4</td>
    </tr>
    <tr>
      <td style="text-align: right">6</td>
      <td style="text-align: right">0:04.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">1.0</td>
      <td style="text-align: right">0:05.4</td>
    </tr>
    <tr>
      <td style="text-align: right">7</td>
      <td style="text-align: right">0:05.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">1.0</td>
      <td style="text-align: right">0:06.4</td>
    </tr>
    <tr>
      <td style="text-align: right">8</td>
      <td style="text-align: right">0:06.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">1.0</td>
      <td style="text-align: right">0:07.4</td>
    </tr>
    <tr>
      <td style="text-align: right">9</td>
      <td style="text-align: right">0:07.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">1.0</td>
      <td style="text-align: right">0:08.4</td>
    </tr>
    <tr>
      <td style="text-align: right">10</td>
      <td style="text-align: right">0:08.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">1.0</td>
      <td style="text-align: right">0:09.4</td>
    </tr>
    <tr>
      <td style="text-align: right">11</td>
      <td style="text-align: right">0:09.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">1.0</td>
      <td style="text-align: right">0:10.4</td>
    </tr>
    <tr>
      <td style="text-align: right">12</td>
      <td style="text-align: right">0:10.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">1.0</td>
      <td style="text-align: right">0:11.4</td>
    </tr>
    <tr>
      <td style="text-align: right">13</td>
      <td style="text-align: right">0:11.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">1.0</td>
      <td style="text-align: right">0:12.4</td>
    </tr>
    <tr>
      <td style="text-align: right">14</td>
      <td style="text-align: right">0:12.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">1.0</td>
      <td style="text-align: right">0:13.4</td>
    </tr>
    <tr>
      <td style="text-align: right">15</td>
      <td style="text-align: right">0:13.4</td>
      <td style="text-align: right"><code class="language-plaintext highlighter-rouge">RTO_MAX</code></td>
      <td style="text-align: right">1.0</td>
      <td style="text-align: right">0:14.4</td>
    </tr>
  </tbody>
</table>

<p>A timeout in under 15 seconds is vastly preferable to waiting over 15 minutes.
It remains to be seen whether such frequent retransmissions can have negative
consequences, perhaps even causing additional network congestion and further
packet loss, but I expect this will behave better in very many situations.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[Distributed systems must always be able to correctly deal with the non-delivery (or non-acknowledgement) of a message, but in any reasonable network environment one can assume that message non-delivery is fairly rare. This means we need not behave optimally in the case of a message non-delivery. In principle the system should still work correctly even if retransmissions were totally disabled, but in practice TCP uses packet loss to signal network congestion and apply backpressure so some small number of retransmissions are to be expected and need to be handled gracefully.]]></summary></entry><entry><title type="html">Unicode glyphs for Mac OS keyboard modifiers</title><link href="https://davecturner.github.io/2025/09/03/mac-os-keyboard-glyphs.html" rel="alternate" type="text/html" title="Unicode glyphs for Mac OS keyboard modifiers" /><published>2025-09-03T00:00:00+00:00</published><updated>2025-09-03T00:00:00+00:00</updated><id>https://davecturner.github.io/2025/09/03/mac-os-keyboard-glyphs</id><content type="html" xml:base="https://davecturner.github.io/2025/09/03/mac-os-keyboard-glyphs.html"><![CDATA[<p>Sometimes it’s useful to represent all the odd keys on a Mac OS keyboard with
their proper glyphs in a textual conversation. There’s various ways to do that
but honestly the easiest for me is just to search for a web page that contains
the right glyphs and copy-paste them into the chat. Recently I’ve found that
becoming less effective, requiring several different searches before finally
finding a good page, so I’m recording them here where I can find them more
easily.</p>

<table>
  <thead>
    <tr>
      <th>Glyph</th>
      <th>Description</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>⌘</td>
      <td>Command</td>
    </tr>
    <tr>
      <td>⌥</td>
      <td>Option</td>
    </tr>
    <tr>
      <td>⌃</td>
      <td>Control</td>
    </tr>
    <tr>
      <td>⇧</td>
      <td>Shift</td>
    </tr>
    <tr>
      <td>⇪</td>
      <td>Caps lock</td>
    </tr>
    <tr>
      <td>␣</td>
      <td>Space</td>
    </tr>
    <tr>
      <td>⇥</td>
      <td>Tab</td>
    </tr>
    <tr>
      <td>↑</td>
      <td>Up arrow</td>
    </tr>
    <tr>
      <td>↓</td>
      <td>Down arrow</td>
    </tr>
    <tr>
      <td>←</td>
      <td>Left arrow</td>
    </tr>
    <tr>
      <td>→</td>
      <td>Right arrow</td>
    </tr>
    <tr>
      <td>⏎</td>
      <td>Enter</td>
    </tr>
    <tr>
      <td>⏏</td>
      <td>Eject</td>
    </tr>
    <tr>
      <td>⌫</td>
      <td>Delete</td>
    </tr>
  </tbody>
</table>]]></content><author><name></name></author><summary type="html"><![CDATA[Sometimes it’s useful to represent all the odd keys on a Mac OS keyboard with their proper glyphs in a textual conversation. There’s various ways to do that but honestly the easiest for me is just to search for a web page that contains the right glyphs and copy-paste them into the chat. Recently I’ve found that becoming less effective, requiring several different searches before finally finding a good page, so I’m recording them here where I can find them more easily.]]></summary></entry><entry><title type="html">Bayesian bisection</title><link href="https://davecturner.github.io/2024/11/11/bayesian-bisection.html" rel="alternate" type="text/html" title="Bayesian bisection" /><published>2024-11-11T00:00:00+00:00</published><updated>2024-11-11T00:00:00+00:00</updated><id>https://davecturner.github.io/2024/11/11/bayesian-bisection</id><content type="html" xml:base="https://davecturner.github.io/2024/11/11/bayesian-bisection.html"><![CDATA[<p><a href="/2016/01/31/bisection-in-dcvss.html">Bisection</a> is a mightily
effective technique for debugging those tricky issues that were quietly
introduced into a codebase and not noticed for an extended period of time. You
can find the first bad commit from a range of thousands (or more) in a
logarithmic number of steps via a binary search, halving the range of possible
bad commits on each step.</p>

<p>At least, you can do this as long as you have a way to say for each commit
whether it is definitely good or definitely bad. But sometimes you might have a
test case that started failing extremely rarely, and this is where traditional
binary search struggles. If you can reproduce the failure on a commit then that
commit is definitely bad, but if the failure doesn’t reproduce on some other
commit then it could be that that commit is good or it could be that you just
didn’t get lucky this time. So you run it again. And again. But how many good
runs in a row do you need before you can consider the commit to be good?</p>

<p>To give some figures from a recent example, I was recently chasing down a test
that started to fail approximately once in every ~2700 runs. Each run of the
test took around 6 seconds so you could expect a bad commit to fail in around
4-and-a-half hours. The offending commit was somewhere in a range of about a
thousand commits, so in principle should be findable with a binary search of 10
steps.</p>

<p>One possible strategy would be to consider a commit that gets 2700 successful
runs in a row to be good, and carry on with the normal binary search process.
This would take a couple of days to complete the 10 steps needed to narrow the
search down to a single commit. But consider what happens with this strategy if
we were just unlucky and a bad commit took a little longer than expected to
fail. In this case, we discard the half of the commit range that actually
contains the first bad commit, and the binary search process will terminate
pointing to an earlier, and actually good, commit.</p>

<p>We can improve on this by looking for longer sequences of successful runs
before considering a commit to be good. For instance, we could wait until we’ve
seen twice as many successes in a row. But this means that on a good commit
we’ll have to wait nine hours before moving to the next step. With ten steps to
run that’d be around four days of work, and then we’d still have the concern
that at least one of those steps was unlucky and didn’t capture the failure,
leading to low confidence that the binary search really found the first bad
commit. We could increase that confidence by running the tests for much longer
on the (claimed) last good commit, but if any of the preceding runs was unlucky
then we’d find the (claimed) last good commit actually to be bad, and from
there we have limited options apart from restarting the binary search on the
smaller range, because all those commits we thought to be good are now under
suspicion again.</p>

<p>It’d be awfully nice if we could use the output of the preceding hours of test
runs to influence the restarted search somehow.</p>

<h3 id="bayesian-inference">Bayesian inference</h3>

<p>Bayesian inference allows us to build a probabilistic model of our beliefs
about the quality of each commit, and to update this model as we learn more
about the system by running its tests on the different commits. The resulting
process generalizes a traditional binary search into one which takes account of
the probabilistic nature of the test results, iteratively updating its beliefs
about which commit may have introduced the flakiness based on the history of
observed failures and successes.</p>

<p>Rather than treating the outcome of a test run as a binary pass-or-fail thing,
we model the system in terms of probabilities as follows. Consider a sequence
of commits C<sub>i</sub> for i ∈ {0..N-1} together with some k ∈ {1..N-1} such
that the initial subsequence C<sub>i&lt;k</sub> are all <em>good</em>, reliably
passing some test, and the remainder C<sub>i≥k</sub> are all <em>bad</em>,
occasionally failing the test with flakiness probability <em>p</em>. Assume also that
the failures of distinct test runs are independent. Our goal is to identify the
first bad commit C<sub>k</sub>, which we will do by computing a discrete
probability distribution P<sub>i</sub> = P(C<sub>i</sub> is the first bad
commit) based on a history of successes and failures on the various commits in
the sequence, and then selecting further tests to run so as to concentrate the
probability mass on a single commit in the sequence as efficiently as possible.</p>

<p>The computation of the probability distribution works iteratively: given any
prior distribution P<sub>i</sub> and another test result at commit
C<sub>j</sub>, we can compute a posterior distribution P’<sub>i</sub> which
incorporates the evidence from the new test result into the distribution using
<a href="https://en.wikipedia.org/wiki/Bayes%27_theorem">Bayes’ theorem</a>:</p>

<p>P’<sub>i</sub> = P(C<sub>i</sub> is the first bad commit | test result at C<sub>j</sub>)<br />
    = P(test result at C<sub>j</sub> | C<sub>i</sub> is the first bad commit)<br />
          × P(C<sub>i</sub> is the first bad commit)<br />
          ÷ P(test result at commit C<sub>j</sub>)<br />
    ∝ P(test result at C<sub>j</sub> | C<sub>i</sub> is the first bad commit) × P<sub>i</sub></p>

<p>… noting that the denominator P(test result at commit C<sub>j</sub>) is
independent of i and is really more like a normalization factor that scales all
the numbers so that they sum to 1 again. Thus we can compute the posterior
distribution by multiplying each probability in the prior distribution by a
factor which depends on the probability <em>p</em> of a test failing on a bad commit:</p>

<ul>
  <li>P(test passes at C<sub>j</sub> | C<sub>i</sub> is the first bad commit) = (1-<em>p</em>) if i ≤ j else 1</li>
  <li>P(test fails at C<sub>j</sub> | C<sub>i</sub> is the first bad commit) = <em>p</em> if i ≤ j else 0</li>
</ul>

<p>In words: given a prior distribution of probabilities that each commit in a
sequence is the first bad commit, and another test result on a commit
C<sub>j</sub>, compute the posterior distribution of those probabilities as
follows. If the test passed at C<sub>j</sub> then multiply by (1-<em>p</em>) the
probabilities of all C<sub>i</sub> for i ≤ j to reflect that it has now become
slightly less likely that any of these commits is the first bad one. In
contrast, if the test failed on a commit C<sub>j</sub> then it has now become
impossible for any earlier commit to be the first bad one, so set to zero the
probabilities of all commits C<sub>i</sub> for j &lt; i. In the latter case there
is in fact no need to multiply the remaining probabilities by <em>p</em>, as the above
formula suggests, because we are only working up to proportionality and we must
renormalize everything to sum to 1 to compute the true probabilities anyway.</p>

<p>Doing this repeatedly for a sequence of test results yields a posterior
probability distribution which reflects all the knowledge gathered from those
test runs. The initial prior distribution is not particularly important, but a
uniform distribution is easiest to implement so I suggest to use that.</p>

<h3 id="testing-the-median">Testing the median</h3>

<p>Given such a probability distribution, it makes the most sense to do the next
test run on the <em>median</em> commit, because the outcome of this run (pass or fail)
most evenly divides the remaining probability space, meaning it provides the
maximum possible expected reduction in uncertainty about the location of the
first bad commit.</p>

<p>Repeatedly testing the median commit according to the posterior probability
distribution has an intuitively useful behaviour: as tests pass the median will
slowly creep upwards towards the first known-bad commit, and then whenever a
test fails it will jump downwards again. In effect, it re-tests some commits
that we were previously believed to be good, gathering more information about
their quality and possibly discovering a failure on an earlier commit than the
previous best. If we’ve seen a failure on the actual first-bad commit then the
median will creep upwards until it is between the first-bad and last-good
commits, where it will stay.</p>

<p>Note that the speed at which the median creeps towards the first known-bad
commit is determined by <em>p</em>. The rarer the test failures on bad commits, the
slower the median commit moves, so that the algorithm spends more time testing
commits further from the known-bad ones. In contrast, if test failures are
relatively likely then the median commits moves more quickly. In the limit, if
the test is not flaky at all (so <em>p</em> is 1) and we start from a uniform prior
distribution then testing the median will do a standard binary search exactly
like regular bisection does.</p>

<p>Since this is a discrete probability distribution there typically is no commit
which is the exact median. Instead, one must choose between the <em>submedian</em>
commit, i.e. the last commit C<sub>j</sub> such that ∑<sub>i≤j</sub>
P<sub>i</sub> ≤ ½, and the <em>supermedian</em> commit, i.e. the first commit
C<sub>j</sub> such that ∑<sub>i≤j</sub> P<sub>i</sub> ≥ ½. When the algorithm
is just starting out the difference is fairly unimportant, but as things
converge it turns out that neither of these choices alone behaves quite as
desired: we must repeatedly test the believed-last-good commit in case it turns
out to be bad, but always choosing the submedian can end up focussing on the
commit just before the believed-last-good commit, whereas always choosing the
supermedian will end up focussing on the believed-first-bad commit instead.</p>

<p>My preferred solution is to choose whichever of the submedian or supermedian
commits has had fewer test runs so far. In the endgame, this scheme means the
algorithm alternates between the last-good and first-bad commits, which has
some useful consequences. Running further tests on the first-bad commit helps
to refine our estimate of the flakiness probability <em>p</em> (see below), while
repeatedly checking the believed-last-good commit will eventually see a failure
if this commit is in fact bad. Additionally, this scheme produces a sequence of
test runs on a bad commit, some of which failed, and a similar length of
sequence of test runs on the previous commit, all of which passed, which forms
an intuitively compelling argument that we’ve really found the first bad commit
even without having to resort to any probability calculations.</p>

<p>As a slight further refinement, I instead choose the commit with the smaller
value of ⌊runs÷10⌋, preferring the submedian in case of a tie. Rather than
strictly alternating between the two commits, using ⌊runs÷10⌋ runs ten tests in
a row on each commit before switching to the other one. Doing several runs on
each commit amortises the overheads associated with switching commits, such as
having to recompile the system under test. You may find a different number
works better in other cases according to your commit-switching overheads.</p>

<h3 id="estimating-the-flakiness-probability-p">Estimating the flakiness probability <em>p</em></h3>

<p>The process described above assumes we know roughly how flaky the test is, i.e.
we have a good estimate for <em>p</em>. In fact whatever value we use for <em>p</em> will
still yield the right quantitative behaviour and find the right commit in the
end, it just might move the median around at a suboptimal speed.</p>

<p>I find it works well to estimate <em>p</em> simply as the number of observed failures
divided by the total number of test runs on all known-bad commits. When a
failure is observed on a commit that was previously believed to be good, this
estimate for <em>p</em> may drop significantly. The more successful runs there have
been on the commits now known to be bad, the greater will be the reduction in
the estimate for <em>p</em>, and therefore the earlier in the commit history will be
the new median commit.</p>

<p>At the very start of the process, when no test failures have been observed,
this estimate for <em>p</em> is not well-defined. Instead, seed the process by running
the test repeatedly on the last commit in the sequence C<sub>N-1</sub> until it
fails. For instance, if the test succeeds for the first nine runs on
C<sub>N-1</sub> and fails on the tenth then the initial estimate for <em>p</em> will
be ⅒.</p>

<h3 id="algorithm-outline">Algorithm outline</h3>

<p>A conceptual outline of the full Bayesian bisection process:</p>

<ol>
  <li><strong>Initialize</strong>:
    <ul>
      <li><strong>Define commit range</strong>: Establish the range of commits to investigate,
identifying a known good commit at one end and a known bad commit at
the other.</li>
      <li><strong>Initialize probability distribution</strong>: Assign a distribution
P<sub>i</sub> to represent the probability of each commit C<sub>i</sub>
being the first bad one, setting all these probabilities to be
initially equal.</li>
      <li><strong>Find the first failure</strong>: Repeatedly run the test on the last commit
C<sub>N-1</sub> until it fails for the first time.</li>
    </ul>
  </li>
  <li><strong>Iterate</strong>:
    <ul>
      <li><strong>Estimate flakiness probability <em>p</em></strong>: Recompute the estimate of the
flakiness probability <em>p</em> as the number of failures of commits known to
be bad divided by the total number of runs on those commits.</li>
      <li><strong>Recompute posterior distribution</strong>: Given the new estimate for <em>p</em>
and the history of other test results, compute the current probability
distribution P<sub>i</sub>.</li>
      <li><strong>Select next commit</strong>: Identify the submedian and supermedian commits
according to the current probability distribution P<sub>i</sub>. Choose
whichever has smaller value of ⌊runs÷10⌋, preferring the submedian in
case of a tie, and call it C<sub>j</sub>.</li>
      <li><strong>Run test</strong>: Execute the flaky test on C<sub>j</sub>.</li>
    </ul>
  </li>
  <li><strong>Conclude</strong>: Once the probability mass is concentrated as desired on a
single commit, and we have a sufficiently long run of successes on the
previous commit, declare that commit as the most likely first bad commit.</li>
</ol>]]></content><author><name></name></author><summary type="html"><![CDATA[Bisection is a mightily effective technique for debugging those tricky issues that were quietly introduced into a codebase and not noticed for an extended period of time. You can find the first bad commit from a range of thousands (or more) in a logarithmic number of steps via a binary search, halving the range of possible bad commits on each step.]]></summary></entry><entry><title type="html">The Doomsday rule: mental day-of-week calculations</title><link href="https://davecturner.github.io/2021/12/27/doomsday-rule.html" rel="alternate" type="text/html" title="The Doomsday rule: mental day-of-week calculations" /><published>2021-12-27T00:00:00+00:00</published><updated>2021-12-27T00:00:00+00:00</updated><id>https://davecturner.github.io/2021/12/27/doomsday-rule</id><content type="html" xml:base="https://davecturner.github.io/2021/12/27/doomsday-rule.html"><![CDATA[<p>The <a href="https://en.wikipedia.org/wiki/Doomsday_rule">Doomsday rule</a> is an
algorithm for working out the day of the week of a given date. It’s based on
John Conway’s observation that certain memorable dates called <em>doomsdays</em> (4/4,
6/6, 8/8, 10/10, 12/12, 9/5, 5/9, 7/11, 11/7, …) always occur on the same day
of the week in any given year. This day is known as the year’s <em>anchor day</em>. To
compute an arbitary day of week you work out the anchor day for the given year
and the offset between the target date and an appropriate doomsday. It’s
intended to be simple enough that you can do the computation in your head with
a bit of practice.</p>

<p>A couple of recent Youtube videos have raised its profile, first one by <a href="https://www.youtube.com/watch?v=z2x3SSBVGJU">James
Grimes on Numberphile</a> and then a
second by <a href="https://www.youtube.com/watch?v=eSpW4I5moiA">Mike Boyd</a>. In both
videos they show that it’s possible to do the mental calculations quite
quickly. I tried learning it some time ago and struggled to achieve that kind
of speed on the algorithm as described in the literature. The first part of the
computation involves counting leap years (i.e. dividing by 4) and the second
part involves subtracting numbers mod 7 which means you have to get the sign
right and sometimes have to deal with negative numbers. All definitely
possible, but tricky to do quickly in your head.</p>

<p>I found that the effects of practice was basically to inline various bits of
the strategy and commit the inlined version to memory, but the inlining wasn’t
consistently applicable which meant that some dates ended up harder to compute
than others. Since the point of practice seemed to be to commit various things
to memory, I figured it’d be simpler to rework the algorithm to make use of
memorisation directly.</p>

<p>There’s not actually that many things to remember, it only took me a couple of
hours, and once I’d done the memorisation work I found I can run the algorithm
pretty quickly even if distracted by other things such as continuing a
conversation. I wonder if when other folks say they’re doing Conway’s method
they’re really doing something more like this, or whether it’s just me that
finds recalling memorised facts much easier than doing actual computations.</p>

<p>One handy feature of the method presented here is that you combine all the
sub-computations together wih addition (mod 7) which is commutative and
associative so you can process the components of the date in any order and
reduce them to a single value at each step which saves on working memory slots.
Here in the UK most folks say the date in day-month-year order; in the US
month-day-year seems preferred, but really any order is reasonable.</p>

<h3 id="overview">Overview</h3>

<p>Process the day-of-month, the month, the century and the year-of-century in any
order you choose. Each component yields a value which you add to a running
total, computed mod 7. The final running total after all four values are added
(occasionally with a simple leap-year correction) indicates the day of the week
of the target date, with 0 meaning Sunday, 1 meaning Monday and so on.</p>

<h3 id="day-of-month">Day-of-month</h3>

<p>When you hear the day of the month, add it to the running total, working mod 7.</p>

<h3 id="month">Month</h3>

<p>When you hear the month, convert it to a value according to the following table
and add the value to your running total, working mod 7.</p>

<table>
  <thead>
    <tr>
      <th>Month</th>
      <th>Value</th>
      <th>Month</th>
      <th>Value</th>
      <th>Month</th>
      <th>Value</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>January</td>
      <td>4</td>
      <td>May</td>
      <td>5</td>
      <td>September</td>
      <td>2</td>
    </tr>
    <tr>
      <td>February</td>
      <td>0</td>
      <td>June</td>
      <td>1</td>
      <td>October</td>
      <td>4</td>
    </tr>
    <tr>
      <td>March</td>
      <td>0</td>
      <td>July</td>
      <td>3</td>
      <td>November</td>
      <td>0</td>
    </tr>
    <tr>
      <td>April</td>
      <td>3</td>
      <td>August</td>
      <td>6</td>
      <td>December</td>
      <td>2</td>
    </tr>
  </tbody>
</table>

<p>Just learn this mapping. There’s kind of a pattern but learning the values
doesn’t take long and means you can process the month very quickly while
listening for the rest of the date.</p>

<h3 id="century">Century</h3>

<p>When you hear the century—the first two digits of the year—convert it to a
value according to the following table and add that value to your running
total, working mod 7.</p>

<table>
  <thead>
    <tr>
      <th>Century</th>
      <th>Value</th>
      <th>Century</th>
      <th>Value</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>17xx</td>
      <td>0</td>
      <td>21xx</td>
      <td>0</td>
    </tr>
    <tr>
      <td>18xx</td>
      <td>5</td>
      <td>22xx</td>
      <td>5</td>
    </tr>
    <tr>
      <td>19xx</td>
      <td>3</td>
      <td>23xx</td>
      <td>3</td>
    </tr>
    <tr>
      <td>20xx</td>
      <td>2</td>
      <td>24xx</td>
      <td>2</td>
    </tr>
  </tbody>
</table>

<p>Again, just learn this mapping. There is a repeating pattern because the
Gregorian calendar repeats every 400 years, but in practice you’ll mostly get
one of these centuries so it’s simpler to remember them. You can reasonably
reject years before 1700 because the UK and US didn’t start using the Gregorian
calendar until 1752, and if you get a year after 2499 then it’s usually ok to
take a bit longer to work out to which of 20xx, 21xx, 22xx or 23xx it’s
equivalent.</p>

<h3 id="year-of-century">Year-of-century</h3>

<p>The year of the century—the last two digits of the year—is the trickiest bit to
deal with. If it’s a multiple of 4 then it maps to a remainder mod 7 as follows
which should be added to the running total:</p>

<table>
  <thead>
    <tr>
      <th>Year</th>
      <th>Value</th>
      <th>Year</th>
      <th>Value</th>
      <th>Year</th>
      <th>Value</th>
      <th>Year</th>
      <th>Value</th>
      <th>Year</th>
      <th>Value</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>xx00</td>
      <td>0</td>
      <td>xx20</td>
      <td>4</td>
      <td>xx40</td>
      <td>1</td>
      <td>xx60</td>
      <td>5</td>
      <td>xx80</td>
      <td>2</td>
    </tr>
    <tr>
      <td>xx04</td>
      <td>5</td>
      <td>xx24</td>
      <td>2</td>
      <td>xx44</td>
      <td>6</td>
      <td>xx64</td>
      <td>3</td>
      <td>xx84</td>
      <td>0</td>
    </tr>
    <tr>
      <td>xx08</td>
      <td>3</td>
      <td>xx28</td>
      <td>0</td>
      <td>xx48</td>
      <td>4</td>
      <td>xx68</td>
      <td>1</td>
      <td>xx88</td>
      <td>5</td>
    </tr>
    <tr>
      <td>xx12</td>
      <td>1</td>
      <td>xx32</td>
      <td>5</td>
      <td>xx52</td>
      <td>2</td>
      <td>xx72</td>
      <td>6</td>
      <td>xx92</td>
      <td>3</td>
    </tr>
    <tr>
      <td>xx16</td>
      <td>6</td>
      <td>xx36</td>
      <td>3</td>
      <td>xx56</td>
      <td>0</td>
      <td>xx76</td>
      <td>4</td>
      <td>xx96</td>
      <td>1</td>
    </tr>
  </tbody>
</table>

<p>The values for the nine multiples of 12 are just the result of dividing by 12.
The other 16 are all ±4 away from a multiple of twelve so you can in principle
compute the nearest multiple of 12 and then add or subtract 2 according to the
pattern, but it’s only 16 values so still not that hard to just memorise the
mapping.</p>

<p>If the year isn’t a multiple of 4 then add its remainder mod 4 to your running
total, then round the year down to the nearest multiple of 4, grab the value
from the table above, and add that to the running total as well. There’s
various other ways you could achieve this, for instance you could memorise all
100 values, but representing the year as <code class="language-plaintext highlighter-rouge">4k+r</code> and then processing <code class="language-plaintext highlighter-rouge">4k</code> and
<code class="language-plaintext highlighter-rouge">r</code> separately works best for me.</p>

<h3 id="leap-year-correction">Leap year correction</h3>

<p>If the target is in January or February of a leap year then decrement the
running total by one. This is kind of irritating to remember. However note that
leap years are slightly simpler to compute since they’re always a multiple of 4
so there’s no remainder mod 4 to deal with at the year-of-century step.</p>

<h3 id="finishing-up">Finishing up</h3>

<p>The final total represents the day of the week of the chosen date, with 0
meaning Sunday, 1 meaning Monday, and so on.</p>

<h2 id="worked-examples">Worked examples</h2>

<h3 id="19-july-1989">19 July 1989</h3>

<p>The day-of-month 19 becomes 5 which is the initial running total. July has
value 3, and adding this to the running total gives 1. The century 19xx is also
3, which adds to the running total to give 4. The year-of-century xx89 is 1
greater than a multiple of 4, namely xx88, so we add the 1 to the running total
to give 5. Finally xx88 has value 5 which is added to the running total to give
3, which means that 19 July 1989 was a Wednesday.</p>

<h3 id="february-27-2204">February 27, 2204</h3>

<p>February has value 0 so the running total starts out as zero. The day-of-month
27 becomes 6 which replaces the running total of zero.  The century 22xx has
value 5 which adds to the running total to give 4 (recalling that adding 6 is
the same as subtracting 1). The year-of-century xx04 is a multiple of 4 with
value 5 which adds to the running total to give 2. But 2204 is a leap year and
the date is in February so we must subtract 1, giving a final total of 1 which
means that 2 February 2204 will be a Monday.</p>

<h3 id="1900-01-15">1900-01-15</h3>

<p>The century 19xx has value 3 so that’s the initial running total.  The
year-of-century xx00 is a multiple of 4 with value 0 so the running total is
unchanged.  January has value 4 which is added to the running total to give 0.
The day-of-month 15 becomes 1 mod 7 so this is the new running total.  The
month is January and the year is a multiple of 4 but note that 1900 was <em>not</em> a
leap year so no further correction is needed. 1900-01-15 was therefore a
Monday.</p>

<h2 id="how-it-works">How it works</h2>

<p>The day and month values added together (minus one for January and February in
leap years) tells you the offset in days from the year’s anchor day to the
target date. You can check that the memorable doomsdays all add up to a
multiple of seven: 4 April = 4 + 3 = 7; 9 May = 9 + 5 = 14; …</p>

<p>The century and year-of-century values added together tells you the anchor day
for the year. The year-of-century table adds the number of years to the number
of earlier leap years in the century, i.e. it computes <code class="language-plaintext highlighter-rouge">5n/4 mod 7</code> for
multiples of 4.</p>

<h2 id="speed-tips">Speed tips</h2>

<p>Recalling a value from memory is enormously faster than working it out. I’d
rather remember a table of 25 things instead of adding another computation
step.</p>

<p>I spent some time actively memorising the remainders mod 7 of the numbers from
1 to 31, even though I can compute them fairly well, because you need to do
this step quickly whilst distracted by listening for the rest of the date.</p>

<p>I also found it worth practicing doing sums of remainders mod 7. One of my very
early maths teachers made us do speed exercises for adding small numbers, much
like multiplication tables, which everyone hated at the time but I’ve since
realised how powerful it is to be able to do these sums <em>instantly</em> rather than
merely quickly. There’s not actually that many to learn in this case: the sums
that don’t exceed 7 are just the regular ones; the sums that equal 7 are easy
to map to zero. Sums involving zero are also easy, and adding 6 is easier to
treat as subtracting 1. That just leaves 3+5, 4+4, 4+5 and 5+5 to learn.</p>

<p>I learned the year-of-century values in three batches: first the multiples of
12, then 04, 08, 16, 28, 32, 40, 76, 80, and finally 20, 44, 52, 56, 64, 68,
88, 92. Learning all 16 values at once didn’t work well for me, but working 8
at a time was very effective. The batches were pretty much randomly chosen
except I made sure the split of values was as even as possible. There’s not
much science behind that but it seemed to work.</p>

<p>I recommend learning the values for any years that you will use a lot, such as
the current year and a couple of nearby ones. 2021’s is 0 and 2022’s will be 1.</p>

<p>If you’re coming at this cold then the only things you don’t already know are
the values for the 12 months, 8 centuries, and the 16 years-of-century that
aren’t multiples of 12, which is 36 individual items. To get faster you
probably also need to practice the 31 remainders for days-of-the-month, 9
remainders for the multiple-of-12 years, and the 21 sums mod 7. That is 97
things in total, plus maybe the value for the current year ±1 to make it up to
an even hundred. There’s loads of spaced-repetition flashcard apps out there
that should help with this kind of thing. With a couple hours of up-front
investment you can save literally seconds per month of looking up dates in a
calendar.</p>

<p><a href="https://xkcd.com/1205/"><img src="https://imgs.xkcd.com/comics/is_it_worth_the_time.png" alt="Obligatory XKCD" /></a></p>]]></content><author><name></name></author><summary type="html"><![CDATA[The Doomsday rule is an algorithm for working out the day of the week of a given date. It’s based on John Conway’s observation that certain memorable dates called doomsdays (4/4, 6/6, 8/8, 10/10, 12/12, 9/5, 5/9, 7/11, 11/7, …) always occur on the same day of the week in any given year. This day is known as the year’s anchor day. To compute an arbitary day of week you work out the anchor day for the given year and the offset between the target date and an appropriate doomsday. It’s intended to be simple enough that you can do the computation in your head with a bit of practice.]]></summary></entry><entry><title type="html">Tracking down a seven-year-old segfault</title><link href="https://davecturner.github.io/2021/08/30/seven-year-old-segfault.html" rel="alternate" type="text/html" title="Tracking down a seven-year-old segfault" /><published>2021-08-30T00:00:00+00:00</published><updated>2021-08-30T00:00:00+00:00</updated><id>https://davecturner.github.io/2021/08/30/seven-year-old-segfault</id><content type="html" xml:base="https://davecturner.github.io/2021/08/30/seven-year-old-segfault.html"><![CDATA[<p>Back in August 2014 a user reported they <a href="https://discuss.elastic.co/t/segfault-in-ffi-prep-closure-loc/19227">couldn’t run
Elasticsearch</a>:
it would immediately crash with a segmentation fault. Elasticsearch is almost
entirely written in Java which as a managed language is supposed to protect us
from low-level issues like segmentation faults. The “almost” in the previous
sentence is the problem: Elasticsearch calls out to native code in a few
places, and it was one of these places that was triggering the crash.</p>

<p>Elasticsearch makes its native calls using
<a href="https://github.com/java-native-access/jna">JNA</a> which is a deeply magical
library that makes it easy for Java code to call into libraries written in C.
There was definitely a suspicion that this was something to do with JNA, but
the investigation fizzled out before it got very far.</p>

<p>Then in May 2016 Github user <a href="https://github.com/fxh"><strong>@fxh</strong></a> reported <a href="https://github.com/elastic/elasticsearch/issues/18272">the
same thing on Github</a>.
The investigation got further this time: the issue seemed to manifest only when
SELinux was enabled, and seemed to be related to whether temporary files were
permitted to contain executable code. Forbidding temporary executables is a
security measure: anyone can write to <code class="language-plaintext highlighter-rouge">/tmp</code> so it’s possible to attack a
system by writing a nefarious executable to <code class="language-plaintext highlighter-rouge">/tmp</code> and then tricking someone
more privileged into running it.</p>

<p>In this context “executable” means more than just programs that you can run
from the command line. Modern processors <a href="https://en.wikipedia.org/wiki/NX_bit">distinguish code from
data</a> at a very low level, with a flag on
each page of memory that determines whether the data it contains can ever be
interpreted as instructions that the CPU will execute. If <code class="language-plaintext highlighter-rouge">/tmp</code> is mounted
with the <code class="language-plaintext highlighter-rouge">noexec</code> option then every page that’s associated with a file under
<code class="language-plaintext highlighter-rouge">/tmp</code> will have the no-execute bit set. This forbids fully-fledged programs
and dynamically-linked libraries and also, crucially, any other kind of
memory-mapped executable page that is backed by a file in <code class="language-plaintext highlighter-rouge">/tmp</code>.</p>

<p>Some of JNA’s deep magic works by dynamically generating a temporary library
(wrapping around the C library) into which the Java code can call. It’s
sometimes possible to generate code dynamically in pages that aren’t backed by
a file, but for security reasons the system might also proscribe pages from
being both writeable and executable, and obviously we need write access to the
memory in which we’re generating the code. The usual solution seems to be to
write the code into a file and then use something like <code class="language-plaintext highlighter-rouge">mmap()</code> to load it
again into read-only-but-executable pages. In order to do this we need some
temporary space that isn’t mounted <code class="language-plaintext highlighter-rouge">noexec</code>.</p>

<p>However the crash wasn’t <em>just</em> caused by having <code class="language-plaintext highlighter-rouge">/tmp</code> mounted with the
<code class="language-plaintext highlighter-rouge">noexec</code> option: if you do that then you get a <a href="https://github.com/elastic/elasticsearch/issues/18272#issuecomment-224140253">different
error</a>
and not a segmentation fault. And anyway you can tell JNA to create its
temporary library in a different location by setting the <code class="language-plaintext highlighter-rouge">java.io.tmpdir</code> or
<code class="language-plaintext highlighter-rouge">jna.tmpdir</code> system properties, which is a sensible workaround for when
executables are forbidden in the default temporary directory. This doesn’t
always fix the problem.</p>

<p>At this point it became hard to make further progress: there was no
sufficiently well-locked-down SELinux system on which to analyse things further
and it’s not a configuration that gets a lot of testing. We verified that the
problem really wasn’t in Elasticsearch code itself and
<a href="https://github.com/elastic/elasticsearch/issues/18272#issuecomment-234687922">concluded</a>
that it must be a problematic and untested interaction between SELinux and JNA
in the hope that a future version of SELinux and/or JNA would fix it.</p>

<p>A few months later user <a href="https://github.com/vineet01"><strong>@vineet01</strong></a> reported
that they were having the same problem and that they <a href="https://github.com/elastic/elasticsearch/issues/18272#issuecomment-294321118">fixed
it</a>
by creating a home directory for the user as which Elasticsearch was running.
They hypothesised that this was because the JVM wanted to create a
usage-tracking file in the home directory. More recently user
<a href="https://github.com/cyamal1b4"><strong>@cyamal1b4</strong></a> reported the <a href="https://github.com/elastic/elasticsearch/issues/73309">same
fix</a> and blamed the same
usage-tracking file, although neither user gave an explanation of how a failure
to write this file might lead to a segmentation fault in JNA-related code.</p>

<p>Over the years many users have reported this same crash. It still definitely
exists on very locked-down systems. It always appears to be related to
temporary executable files and is generally fixed by fiddling with environment
variables/system properties/permissions until the segfault goes away. The
trouble is that the users that have these very locked-down systems are also the
users that can provide the least amount of debugging context, and that struggle
to make changes to permissions or other environmental settings. It’s often
quite a long process to find the right combination of settings that fix it, and
all these cases consume much time from many engineers. At least the crash
reliably happens at every startup. It’d be much worse if it were
nondeterministic or took a long time to manifest, but Elasticsearch really
should not be failing with a segmentation fault like this, and we really
shouldn’t be spending this much time helping users solve the same issue over
and over again.</p>

<p>Recently I decided it was worth taking another look to see if I could work out
what was at the bottom of it.</p>

<h2 id="jvm-error-log-analysis">JVM error log analysis</h2>

<p><strong>@fxh</strong> included the JVM error log file in the <a href="https://github.com/elastic/elasticsearch/issues/18272">original Github
issue</a> which contains a
fantastic amount of detail about the state of the JVM at the time of the crash.
Here’s an excerpt of the important bits:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>#
# A fatal error has been detected by the Java Runtime Environment:
#
#  SIGSEGV (0xb) at pc=0x00007f424226f40a, pid=28216, tid=139922878629632
#
# JRE version: Java(TM) SE Runtime Environment (7.0_75-b13) (build 1.7.0_75-b13)
# Java VM: Java HotSpot(TM) 64-Bit Server VM (24.75-b04 mixed mode linux-amd64 compressed oops)
# Problematic frame:
# C  [jna4948368637624641726.tmp+0x1240a]  ffi_prep_closure_loc+0x1a
#

Registers:
RAX=0x00007f424226f8c2, RBX=0x00007f425579ed48, RCX=0x00007f424c44a7b0, RDX=0x00007f4242264590
RSP=0x00007f425579eae0, RBP=0x00007f425579eae0, RSI=0x00007f424c44a7d0, RDI=0x0000000000000000
R8 =0x00007f425007ae43, R9 =0x0000000000000002, R10=0x00007f425579e870, R11=0x00007f424226f3f0
R12=0x0000000000000000, R13=0x0000000000000008, R14=0x00007f424c44a7b0, R15=0x0000000000000004
RIP=0x00007f424226f40a, EFLAGS=0x0000000000010246, CSGSFS=0x0000000000000033, ERR=0x0000000000000006
TRAPNO=0x000000000000000e

Instructions: (pc=0x00007f424226f40a)
0x00007f424226f3ea:   66 90 66 66 66 90 8b 06 55 41 b9 02 00 00 00 48
0x00007f424226f3fa:   89 e5 ff c8 83 f8 01 77 44 48 8b 05 c6 49 10 00
0x00007f424226f40a:   66 c7 07 49 bb 4c 89 47 0c 66 c7 47 0a 49 ba 48
0x00007f424226f41a:   89 47 02 8b 46 1c 48 89 77 18 48 89 57 20 48 89

Stack: [0x00007f42556a0000,0x00007f42557a1000],  sp=0x00007f425579eae0,  free space=1018k
Native frames: (J=compiled Java code, j=interpreted, Vv=VM code, C=native code)
C  [jna4948368637624641726.tmp+0x1240a]  ffi_prep_closure_loc+0x1a
C  [jna4948368637624641726.tmp+0xd4dd]  Java_com_sun_jna_Native_registerMethod+0x45d
j  com.sun.jna.Native.registerMethod(Ljava/lang/Class;Ljava/lang/String;Ljava/lang/String;[I[J[JIJJLjava/lang/Class;JIZ[Lcom/sun/jna/ToNativeConverter;Lcom/sun/jna/FromNativeConverter;Ljava/lang/String;)J+0
...
j  org.elasticsearch.bootstrap.JNACLibrary.&lt;clinit&gt;()V+45
...
j  org.elasticsearch.bootstrap.JNANatives.definitelyRunningAsRoot()Z+8
</code></pre></div></div>

<p>The header tells us that Elasticsearch received a fatal <code class="language-plaintext highlighter-rouge">SIGSEGV</code> signal while
executing the instruction at <code class="language-plaintext highlighter-rouge">ffi_prep_closure_loc+0x1a</code>, i.e. the one which
starts <code class="language-plaintext highlighter-rouge">0x1a</code> bytes into the function <code class="language-plaintext highlighter-rouge">ffi_prep_closure_loc</code>. This signal
usually means the program attempted to dereference a pointer that doesn’t point
to a valid memory location. The stack trace shows that Elasticsearch was in the
process of executing
<a href="https://github.com/elastic/elasticsearch/blob/36683a4581a8fc2f108701d92af2e4d527e56d02/server/src/main/java/org/elasticsearch/bootstrap/JNANatives.java#L154-L164"><code class="language-plaintext highlighter-rouge">JNANatives#definitelyRunningAsRoot()</code></a>
which looks like this:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>static boolean definitelyRunningAsRoot() {
    if (Constants.WINDOWS) {
        return false; // don't know
    }
    try {
        return JNACLibrary.geteuid() == 0;
    } catch (UnsatisfiedLinkError e) {
        // this will have already been logged by Kernel32Library, no need to repeat it
        return false;
    }
}
</code></pre></div></div>

<p>This method is ultimately trying to call the C library’s <code class="language-plaintext highlighter-rouge">geteuid()</code> function,
and it’s the first time we’ve touched the <code class="language-plaintext highlighter-rouge">JNACLibrary</code> class so we’re running
the static constructor (<code class="language-plaintext highlighter-rouge">&lt;clinit&gt;</code>) which is setting up all the JNA magic.</p>

<p>The dump of instruction memory is useful too: the instruction pointer is at
<code class="language-plaintext highlighter-rouge">0x00007f424226f40a</code> and as mentioned above this is only <code class="language-plaintext highlighter-rouge">0x1a</code> bytes into
executing <code class="language-plaintext highlighter-rouge">ffi_prep_closure_loc</code>, which means the function starts at address
<code class="language-plaintext highlighter-rouge">0x00007f424226f3f0</code> and hence we can disassemble all the instructions in this
function leading up to the one that caused the crash:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>0:  8b 06                   mov    eax,DWORD PTR [rsi]
2:  55                      push   rbp
3:  41 b9 02 00 00 00       mov    r9d,0x2
9:  48 89 e5                mov    rbp,rsp
c:  ff c8                   dec    eax
e:  83 f8 01                cmp    eax,0x1
11: 77 44                   ja     0x57
13: 48 8b 05 c6 49 10 00    mov    rax,QWORD PTR [rip+0x1049c6]        # 0x1049e0
1a: 66 c7 07 49 bb          mov    WORD PTR [rdi],0xbb49
</code></pre></div></div>

<p>The <code class="language-plaintext highlighter-rouge">SIGSEGV</code> happened on the last line which is trying to write something to
the address to which register <code class="language-plaintext highlighter-rouge">RDI</code> points, and the crash dump also tells us
that <code class="language-plaintext highlighter-rouge">RDI</code> is currently <code class="language-plaintext highlighter-rouge">0x0000000000000000</code> which is the null pointer and
definitely not a valid address. This function hasn’t tried to write to <code class="language-plaintext highlighter-rouge">RDI</code>
before it gets to the faulting instruction, which means it must have been
expecting the caller to set <code class="language-plaintext highlighter-rouge">RDI</code> to a valid address.</p>

<p>If a function takes arguments then the caller is responsible for putting them
in appropriate places so that the callee can find them. A <a href="https://en.wikipedia.org/wiki/Calling_convention"><em>calling
convention</em></a> is an agreement
between caller and callee which defines (amongst other things) where the
arguments to a function are when the function is called. Even on a particular
processor architecture there are <a href="https://en.wikipedia.org/wiki/X86_calling_conventions">many possible calling
conventions</a> but in
practice on a 64-bit system running Linux <code class="language-plaintext highlighter-rouge">RDI</code> will contain the first argument
to the function. We can see from the stack dump that the caller is JNA’s
<code class="language-plaintext highlighter-rouge">Java_com_sun_jna_Native_registerMethod</code> function, and <a href="https://github.com/java-native-access/jna/blob/030411b909d5dfd249b1df09a7f24c44babcae64/native/dispatch.c#L3468-L3469">here’s how it calls
<code class="language-plaintext highlighter-rouge">ffi_prep_closure_loc</code></a>:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>closure = ffi_closure_alloc(sizeof(ffi_closure), &amp;code);
status = ffi_prep_closure_loc(closure, closure_cif, dispatch_direct, data, code);
</code></pre></div></div>

<p>The first argument is the <code class="language-plaintext highlighter-rouge">closure</code> which is returned from <code class="language-plaintext highlighter-rouge">ffi_closure_alloc</code>.
The <a href="https://github.com/libffi/libffi/blob/ee1263f7d43bd29b15fc72c4d9520a824e8004df/doc/libffi.texi#L809-L813">docs for this
function</a>
say it allocates and returns a chunk of memory and doesn’t suggest that it
might return <code class="language-plaintext highlighter-rouge">NULL</code>, but if we look at <a href="https://github.com/libffi/libffi/blob/ee1263f7d43bd29b15fc72c4d9520a824e8004df/src/closures.c#L958">its
source</a>
it’s clear that it does return <code class="language-plaintext highlighter-rouge">NULL</code> to indicate various kinds of failure.</p>

<p>Hurrah, we worked it out: the segmentation fault is because JNA isn’t checking
for a failure to allocate this closure, which turns out to be a <a href="https://github.com/java-native-access/jna/issues/1107">known
issue</a> which should be a
<a href="https://github.com/java-native-access/jna/pull/1378">small thing to fix</a>.</p>

<h2 id="but-wait-theres-more">But wait, there’s more</h2>

<p>Although it’s definitely an improvement to throw a Java exception on this
failure instead of a segmentation fault, this doesn’t actually solve anything.
Elasticsearch will still fail to start up even with this fix: it’ll report a
more descriptive message and shut down more gracefully but still users will
need to fiddle around with permissions and ask for help to get Elasticsearch up
and running. Ideally we need to make <code class="language-plaintext highlighter-rouge">ffi_closure_alloc</code> succeed or at least to
understand better why it’s failing.</p>

<blockquote>
  <p>Note that from here on this investigation takes a couple of leaps of faith:
I’ve only read the code, I don’t have a locked-down system on which to run
experiments to verify any of this.</p>
</blockquote>

<p>There are a number of different implementations of <code class="language-plaintext highlighter-rouge">ffi_closure_alloc</code>
depending on operating system and selected by <code class="language-plaintext highlighter-rouge">#ifdef</code> pragmas but they all
ultimately need to allocate some memory into which some machine code can be
written: like JNA, <code class="language-plaintext highlighter-rouge">libffi</code> does some of its magic with dynamically-generated
executable code.</p>

<p>The allocation mechanism is kinda complicated: there’s actually a <a href="https://github.com/libffi/libffi/blob/ee1263f7d43bd29b15fc72c4d9520a824e8004df/src/dlmalloc.c">whole
separate implementation of <code class="language-plaintext highlighter-rouge">malloc()</code> and
friends</a>
which relies on <code class="language-plaintext highlighter-rouge">mmap()</code> to actually acquire memory from the operating system,
but then <code class="language-plaintext highlighter-rouge">mmap()</code> is
<a href="https://github.com/libffi/libffi/blob/ee1263f7d43bd29b15fc72c4d9520a824e8004df/src/closures.c#L541-L542">redefined</a>
to call a custom implementation which does its best to allocate <em>executable</em>
pages using different techniques until it finds one which succeeds. Fortunately
there are some <a href="https://github.com/libffi/libffi/blob/ee1263f7d43bd29b15fc72c4d9520a824e8004df/src/closures.c#L124-L151">helpful comments about how this works on
Linux</a>:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>#if !FFI_MMAP_EXEC_WRIT &amp;&amp; !FFI_EXEC_TRAMPOLINE_TABLE
# if __linux__ &amp;&amp; !defined(__ANDROID__)
/* This macro indicates it may be forbidden to map anonymous memory
   with both write and execute permission.  Code compiled when this
   option is defined will attempt to map such pages once, but if it
   fails, it falls back to creating a temporary file in a writable and
   executable filesystem and mapping pages from it into separate
   locations in the virtual memory space, one location writable and
   another executable.  */
#  define FFI_MMAP_EXEC_WRIT 1
#  define HAVE_MNTENT 1
# endif
...
#if FFI_MMAP_EXEC_WRIT &amp;&amp; !defined FFI_MMAP_EXEC_SELINUX
# if defined(__linux__) &amp;&amp; !defined(__ANDROID__)
/* When defined to 1 check for SELinux and if SELinux is active,
   don't attempt PROT_EXEC|PROT_WRITE mapping at all, as that
   might cause audit messages.  */
#  define FFI_MMAP_EXEC_SELINUX 1
# endif
#endif
</code></pre></div></div>

<p>That tells us that <code class="language-plaintext highlighter-rouge">libffi</code> will sometimes create temporary executable files,
and it will always create them when running under SELinux.  It’s important to
note that this is completely independent of the fact that JNA creates temporary
executable files: <code class="language-plaintext highlighter-rouge">libffi</code> is a language-independent library for calling
foreign functions so it doesn’t know anything about Java and therefore doesn’t
have access to the Java system properties <code class="language-plaintext highlighter-rouge">java.io.tmpdir</code> and <code class="language-plaintext highlighter-rouge">jna.tmpdir</code>
which control where JNA does its work. Instead, on Linux <code class="language-plaintext highlighter-rouge">libffi</code> tries to
create its temporary executable files <a href="https://github.com/libffi/libffi/blob/ee1263f7d43bd29b15fc72c4d9520a824e8004df/src/closures.c#L702-L707">in various places in the following order
of
preference</a>:</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">$LIBFFI_TMPDIR</code></li>
  <li><code class="language-plaintext highlighter-rouge">$TMPDIR</code></li>
  <li><code class="language-plaintext highlighter-rouge">/tmp</code></li>
  <li><code class="language-plaintext highlighter-rouge">/var/tmp</code></li>
  <li><code class="language-plaintext highlighter-rouge">/dev/shm</code></li>
  <li><code class="language-plaintext highlighter-rouge">$HOME</code></li>
</ul>

<p>Elasticsearch doesn’t set any of these environment variables specially, so even
if JNA is creating its temporary files somewhere that permits executables it’s
entirely possible that <code class="language-plaintext highlighter-rouge">libffi</code> does not. This also helpfully resolves the
mystery of why giving the <code class="language-plaintext highlighter-rouge">elasticsearch</code> user a writeable home directory seems
to make the problem go away: when nothing else works, <code class="language-plaintext highlighter-rouge">libffi</code> will try writing
to <code class="language-plaintext highlighter-rouge">$HOME</code> which typically does permit executable code as long as it exists.</p>

<p>Finally, this leads us to a proper fix: we don’t need to give the
<code class="language-plaintext highlighter-rouge">elasticsearch</code> user a whole home directory, instead we should be able to <a href="https://github.com/elastic/elasticsearch/issues/77014">set
<code class="language-plaintext highlighter-rouge">$LIBFFI_TMPDIR</code></a> to
point to the same directory that JNA uses.</p>

<hr />

<p><em>Addendum 2021-08-31</em>: a colleague pointed out that JNA contains a vendored
version of <code class="language-plaintext highlighter-rouge">libffi</code>, and the version of <code class="language-plaintext highlighter-rouge">libffi</code> used by the latest release of
JNA dates back to <a href="https://github.com/java-native-access/jna/blob/030411b909d5dfd249b1df09a7f24c44babcae64/native/libffi/src/closures.c#L691-L695">before support for the <code class="language-plaintext highlighter-rouge">LIBFFI_TMPDIR</code> environment
variable</a>.
This’ll be the right fix eventually, but until then the best we can do is to
set <code class="language-plaintext highlighter-rouge">TMPDIR</code> or <code class="language-plaintext highlighter-rouge">HOME</code>.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[Back in August 2014 a user reported they couldn’t run Elasticsearch: it would immediately crash with a segmentation fault. Elasticsearch is almost entirely written in Java which as a managed language is supposed to protect us from low-level issues like segmentation faults. The “almost” in the previous sentence is the problem: Elasticsearch calls out to native code in a few places, and it was one of these places that was triggering the crash.]]></summary></entry><entry><title type="html">Testing storage for corruption bugs</title><link href="https://davecturner.github.io/2020/12/26/testing-corruption.html" rel="alternate" type="text/html" title="Testing storage for corruption bugs" /><published>2020-12-26T00:00:00+00:00</published><updated>2020-12-26T00:00:00+00:00</updated><id>https://davecturner.github.io/2020/12/26/testing-corruption</id><content type="html" xml:base="https://davecturner.github.io/2020/12/26/testing-corruption.html"><![CDATA[<p>If you are suffering from more than your fair share of <a href="/2020/12/23/lucene-checksums.html">silent corruption</a> then you might have a buggy storage
system. Here’s a couple of tools that exercise your storage with workloads that
are simple and transparent whilst also being quite effective at triggering
corruption bugs.</p>

<p>A successful run of either of these tools does not prove the absence of bugs,
of course, nor does it say anything about corruptions that only occur when the
data has been sitting on disk for an extended period of time. However a failure
definitely indicates that your storage is not working as required. It’s always
worth running tests like these when commissioning a new system.</p>

<h2 id="corruption-on-power-loss">Corruption on power loss</h2>

<p>System calls like Linux’s <code class="language-plaintext highlighter-rouge">fsync()</code> are supposed to guarantee that some data
has genuinely been written to durable storage and will still be there even if
power is lost and subsequently restored. Durable writes can be expensive, so
it’s pretty common to encounter systems which are configured to ignore
<code class="language-plaintext highlighter-rouge">fsync()</code> calls which gives the impression of better performance in benchmarks.
Ignoring <code class="language-plaintext highlighter-rouge">fsync()</code> calls is obviously very dangerous and will lead to data loss
or corruption in a power outage. It’s also a very subtle misconfiguration since
you probably can’t tell whether <code class="language-plaintext highlighter-rouge">fsync()</code> is really working or not without a
genuine power outage. In particular if you’re testing VMs it’s unlikely to be
enough to “power down” the VM and keep the host running: you really need to
abruptly remove power from the host to find bugs of this nature.</p>

<p>The venerable <a href="https://brad.livejournal.com/2116715.html"><code class="language-plaintext highlighter-rouge">diskchecker.pl</code>
script</a> is pretty good at shaking
out cases where <code class="language-plaintext highlighter-rouge">fsync()</code> is not working as it’s supposed to.</p>

<p>This script does quite a bit of I/O and involves pulling power cords out of
things while they’re running so it’s not a great idea to run it on systems
while they’re in production.</p>

<h2 id="write-time-corruption">Write-time corruption</h2>

<p>The excellent <a href="https://kernel.ubuntu.com/~cking/stress-ng/"><code class="language-plaintext highlighter-rouge">stress-ng</code> tool</a>
recently merged a
<a href="https://github.com/ColinIanKing/stress-ng/commit/ecaf4f655d38dfd771a9f6a29bb5f1af64c8aa36">patch</a>
which improves its ability to detect corruption introduced at write time. I
suggest invoking it something like this:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>$ sudo stress-ng --hdd 32 \
                 --hdd-opts wr-seq,rd-rnd \
                 --hdd-write-size 8k \
                 --hdd-bytes 30g \
                 --temp-path $MOUNT_POINT \
                 --verify
</code></pre></div></div>

<p>Some notes on the options:</p>

<p><code class="language-plaintext highlighter-rouge">--hdd 32</code>: This sets the number of concurrent processes writing to disk. I
suggest setting it higher than the number of cores to force some
context-switching which may help surface more bugs.</p>

<p><code class="language-plaintext highlighter-rouge">--hdd-opts wr-seq,rd-rnd</code>: This says to write the file sequentially from start
to finish and then perform random reads during verification. This most closely
matches Lucene’s access pattern, although Lucene may <a href="https://dzone.com/articles/use-lucene’s-mmapdirectory">use <code class="language-plaintext highlighter-rouge">mmap()</code> for reads
instead of <code class="language-plaintext highlighter-rouge">read()</code> and
friends</a>. It’s also
reasonable to use <code class="language-plaintext highlighter-rouge">wr-seq,rd-seq</code>.</p>

<p><code class="language-plaintext highlighter-rouge">--hdd-write-size 8k</code>: Lucene writes in 8kB chunks by default.</p>

<p><code class="language-plaintext highlighter-rouge">--hdd-bytes 30g</code>: This says how large a file each process should write. The
total size of all the files written must of course not exceed the space on the
filesystem, but should be large enough that the reads cannot all be served from
pagecache. In this example there are 32 processes each writing 30GB of data,
which adds up to about 1TB. You can also try dropping pagecache every now and
then during the test with <code class="language-plaintext highlighter-rouge">echo 3 | sudo tee /proc/sys/vm/drop_caches</code> .</p>

<p><code class="language-plaintext highlighter-rouge">--temp-path $MOUNT_POINT</code>: This says where to create the files used for the
test, which must of course be on the filesystem you suspect to be buggy.</p>

<p><code class="language-plaintext highlighter-rouge">--verify</code>: This says to check that the data read is what was written, which is
required to detect cases of silent corruption. Without this option the test
will still go through the motions of reading and writing data but only reports
a failure if any of the I/O actually fails.</p>

<p>This is a pretty aggressive test so it’s not a good idea to run it on systems
while they’re in production. By default it will run for 24 hours which I think
is reasonable.</p>

<p>Make sure that you’re using at least version 0.12.01 of <code class="language-plaintext highlighter-rouge">stress-ng</code> too: older
versions did not test things so thoroughly.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[If you are suffering from more than your fair share of silent corruption then you might have a buggy storage system. Here’s a couple of tools that exercise your storage with workloads that are simple and transparent whilst also being quite effective at triggering corruption bugs.]]></summary></entry><entry><title type="html">Corruption detection in Lucene and Elasticsearch</title><link href="https://davecturner.github.io/2020/12/23/lucene-checksums.html" rel="alternate" type="text/html" title="Corruption detection in Lucene and Elasticsearch" /><published>2020-12-23T00:00:00+00:00</published><updated>2020-12-23T00:00:00+00:00</updated><id>https://davecturner.github.io/2020/12/23/lucene-checksums</id><content type="html" xml:base="https://davecturner.github.io/2020/12/23/lucene-checksums.html"><![CDATA[<p>100% reliable data storage is <a href="http://lamport.azurewebsites.net/pubs/buridan.pdf">fundamentally
impossible</a>, but we can get
pretty close with layers of protection against something going wrong at the
physical level. Databases are typically agnostic to the specific protections
that any given installation is using and mostly just assume that the data they
read from disk is the data they wrote there previously. The protections might
be in the <a href="https://en.wikipedia.org/wiki/ZFS">filesystem itself</a> but it’s more
usual to push them down to the lower layers of RAID controllers and drive
firmware. However it is impossible to truly guarantee protection against
<a href="https://en.wikipedia.org/wiki/Data_corruption#Silent">silent corruption</a> and
this happens often enough that many databases add their own mechanisms for
detecting and correcting those rare cases where the data that they read isn’t
the data that they previously wrote.</p>

<p>Apache Lucene has a simple but effective mechanism for detecting corruption
which the lower layers missed: every relevant file in a Lucene index includes a
<a href="https://en.wikipedia.org/wiki/Cyclic_redundancy_check">CRC32</a> checksum in its
<a href="https://lucene.apache.org/core/8_7_0/core/org/apache/lucene/codecs/CodecUtil.html">footer</a>.
CRC32 is fast to compute and good at detecting the kinds of random corruption
that happen to files on disk. A CRC32 mismatch definitely indicates that
something has gone wrong, although of course a matching checksum doesn’t prove
the absence of corruption. A mismatch is reported with an exception like this:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>org.apache.lucene.index.CorruptIndexException:
    checksum failed (hardware problem?) : expected=335fe5fd actual=399b3f10 ...
</code></pre></div></div>

<p>Elasticsearch, which uses Lucene for most of its interactions with disk, shares
this mechanism for detecting corruption. There are a few places where
Elasticsearch throws its own <code class="language-plaintext highlighter-rouge">CorruptIndexException</code> with a subtly different
message:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>org.apache.lucene.index.CorruptIndexException:
    checksum failed (hardware problem?) : expected=e95zwd actual=fzexao ...
</code></pre></div></div>

<p>This is indicating the same issue except the <code class="language-plaintext highlighter-rouge">expected</code> and <code class="language-plaintext highlighter-rouge">actual</code> values are
reported in base 36 rather than base 16, because that’s how Elasticsearch
tracks these things internally.</p>

<p>Verifying a checksum is expensive since it involves reading every byte of the
file which takes significant effort and might evict more useful data from the
filesystem cache, so systems typically doesn’t verify the checksum on a file
very often. This is why you tend only to encounter a <code class="language-plaintext highlighter-rouge">CorruptIndexException</code>
when something unusual is happening. Users sometimes claim that corruptions are
being <em>caused by</em> merges, or shard movements or snapshots in Elasticsearch, but
this isn’t the case. These activities are examples of the rare times where
reading a whole file is necessary, so they’re a good opportunity to verify the
checksum at the same time, and this is when the corruption is detected and
reported. It doesn’t tell us what actually caused the corruption or when it
happened. It could have been introduced many months earlier.</p>

<p>All the relevant files in a Lucene index are written sequentially from start to
end and then never modified or overwritten. This access pattern is important
because it means the checksum computation is really simple and can happen
on-the-fly as the file is written, and also makes it very unlikely that an
incorrect checksum is due to a userspace bug at the time the file was written.
The <a href="https://github.com/apache/lucene-solr/blob/2dc63e901c60cda27ef3b744bc554f1481b3b067/lucene/core/src/java/org/apache/lucene/store/OutputStreamIndexOutput.java">code that computes the
checksum</a>
is straightforward and very well-tested, so we can have high confidence that a
checksum mismatch really does indicate that the data that Lucene read is not
the data that it previously wrote.</p>

<p>We do get reports of Elasticsearch detecting these silent corruptions in
practice, which is to be expected since our userbase represents a huge quantity
of data running in environments of sometimes-questionable quality. It’s not
always the hardware: often the reports involve a
<a href="http://sbsfaq.com/qnap-fails-to-reveal-data-corruption-bug-that-affects-all-4-bay-and-higher-nas-devices/">nonstandard</a>
setup which hasn’t seen enough real-world use to shake out the bugs. Lucene’s
behaviour during a power outage on properly configured storage is
<a href="http://blog.mikemccandless.com/2014/04/testing-lucenes-index-durability-after.html">well-tested</a>
but many storage systems are not properly configured and may be vulnerable to
<a href="/2020/12/26/testing-corruption.html#corruption-on-power-loss">corrupting data on power loss</a>.
<a href="https://bugzilla.redhat.com/show_bug.cgi?id=1390050">Filesystem</a>
<a href="https://bugzilla.redhat.com/show_bug.cgi?id=1379568">bugs</a>, <a href="https://www.elastic.co/blog/canonical-elastic-and-google-team-up-to-prevent-data-corruption-in-linux">kernel
bugs</a>,
<a href="https://pcper.com/2009/10/intel-halts-downloads-of-new-x25-m-firmware-due-to-corruption/">drive firmware
bugs</a>
and
<a href="https://www.dell.com/community/PowerEdge-HDD-SCSI-RAID/R610-SAS6IR-with-SSD-RAID-1-File-System-Corruption/m-p/4538501#M39288">incompatible</a>
RAID controllers are amongst the many other possibilities. The more unusual or
cutting-edge your storage subsystem is the higher the chances of encountering
this sort of bug. But of course it could genuinely be faulty hardware too,
maybe the <a href="https://community.perforce.com/s/article/2410">RAID controller</a>,
maybe the <a href="https://www.backblaze.com/b2/hard-drive-test-data.html">drive</a>
itself, or maybe even your
<a href="https://discuss.elastic.co/t/253215/4?u=davidturner">RAM</a>. These things do
happen.</p>

<p>The trouble with silent corruption is that it’s <em>silent</em>, it typically doesn’t
result in log entries or other evidence of corruption apart from the checksum
mismatch. In cases where the storage is managed by a separate team from the
applications there’s usually at least one exchange where the storage team says
it must be an application bug because there’s no evidence that anything went
wrong at the storage layer, ignoring of course that a checksum mismatch itself
is pretty convincing evidence that there was a problem with the storage system.
Some <a href="/2020/12/26/testing-corruption.html#write-time-corruption">simple empirical tests</a> are probably a good idea.</p>

<p>When you hear hoofbeats, <a href="https://en.wikipedia.org/wiki/Zebra_%28medicine%29">look for horses not
zebras</a>: Lucene’s mechanism
for writing and checksumming files is simple, sequential and used by everyone,
but there’s a lot of complexity and concurrency and variability on the other
side of the syscalls it makes. <strong>It’s very very likely that the cause of
corruption is infrastructural and out of the application’s control.</strong></p>

<h3 id="technical-details">Technical details</h3>

<p>Here are some technical details of how the checksum works and how to verify it
independently. You can generally rely on Lucene getting this right, but
hopefully the transparency is reassuring.</p>

<p>At time of writing, the last few bytes of a Lucene file will look something
like this:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>...
0000014a: 7265 6446 6965 6c64 7346 6f72 6d61 742e  redFieldsFormat.
0000015a: 6d6f 6465 0a42 4553 545f 5350 4545 4400  mode.BEST_SPEED.
0000016a: c028 93e8 0000 0000 0000 0000 335f e5fd  .(..........3_..
#                                       ^^^^^^^^^ checksum
#                             ^^^^^^^^^ 4 bytes of zeroes (not checksummed)
#                   ^^^^^^^^^ 4 bytes of zeroes (checksummed)
#         ^^^^^^^^^ 4-byte magic number (checksummed)
</code></pre></div></div>

<p>The footer is the last 16 bytes of which the first 12 bytes must be exactly
<code class="language-plaintext highlighter-rouge">c028 93e8 0000 0000 0000 0000</code> and the last 4 bytes are the CRC32 checksum of
the whole file <em>except its last 8 bytes</em>.</p>

<p>There doesn’t seem to be a standard tool to compute and display the CRC32
checksum of a file on Linux, although it’s a very widely-used checksum
algorithm. The most portable method I know is to run the relevant data through
<code class="language-plaintext highlighter-rouge">gzip</code> and then extract the checksum from the footer of the resulting
compressed stream, which can be done without needing to install anything
special if you don’t mind burning a bunch of unnecessary CPU for the
compression:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>$ cat _13.si | head -c-8 | gzip --fast -c | tail -c8 | od -tx4 -N4 -An
 399b3f10
</code></pre></div></div>

<p>There are tools to do this computation too, although they’re not usually
installed by default, and there’s almost certainly a short program in your
language of choice to do the same thing. Reading the expected checksum from the
last 4 bytes is simpler:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>$ tail -c4 _13.si | od -tx4 -N4 -An --endian=big
 335fe5fd
</code></pre></div></div>

<p>These outputs don’t match so the file is corrupt, which was the problem
reported by the example exception shown above. On a file that isn’t corrupt we get matching outputs:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>$ tail -c16 _13.cfs | xxd
00000000: c028 93e8 0000 0000 0000 0000 0064 fc5d  .(...........d.]
$ cat _13.cfs | head -c-8 | gzip --fast -c | tail -c8 | od -tx4 -N4 -An
 0064fc5d
$ tail -c4 _13.cfs | od -tx4 -N4 -An --endian=big
 0064fc5d
</code></pre></div></div>]]></content><author><name></name></author><summary type="html"><![CDATA[100% reliable data storage is fundamentally impossible, but we can get pretty close with layers of protection against something going wrong at the physical level. Databases are typically agnostic to the specific protections that any given installation is using and mostly just assume that the data they read from disk is the data they wrote there previously. The protections might be in the filesystem itself but it’s more usual to push them down to the lower layers of RAID controllers and drive firmware. However it is impossible to truly guarantee protection against silent corruption and this happens often enough that many databases add their own mechanisms for detecting and correcting those rare cases where the data that they read isn’t the data that they previously wrote.]]></summary></entry><entry><title type="html">Rolling packet captures</title><link href="https://davecturner.github.io/2020/12/12/rolling-tcpdump.html" rel="alternate" type="text/html" title="Rolling packet captures" /><published>2020-12-12T00:00:00+00:00</published><updated>2020-12-12T00:00:00+00:00</updated><id>https://davecturner.github.io/2020/12/12/rolling-tcpdump</id><content type="html" xml:base="https://davecturner.github.io/2020/12/12/rolling-tcpdump.html"><![CDATA[<p>I sometimes have a need to collect a packet capture for an extended period of
time. On a busy host this can generate enough data that it needs some special
handling. In particular it’s useful to roll over to a new file every now and
then, and to offload the completed files somewhere else so they don’t fill up
the disk.</p>

<p>This week I discovered that <code class="language-plaintext highlighter-rouge">tcpdump</code> has options <code class="language-plaintext highlighter-rouge">-C</code>, <code class="language-plaintext highlighter-rouge">-G</code> and <code class="language-plaintext highlighter-rouge">-z</code> that let
me do exactly that, rendering obsolete my janky <code class="language-plaintext highlighter-rouge">bash</code> script that tried to do
the same:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>    -C size
        The file size (in multiples of 1e6 bytes) at which to roll over.
    -G age
        The file age (in seconds) at which to roll over.
    -z command
        Runs `command &lt;filename&gt;` on completion of a file.
</code></pre></div></div>

<p>The <a href="https://www.tcpdump.org/manpages/tcpdump.1.html">man page</a> gives more
details including a (somewhat vague) description of how <code class="language-plaintext highlighter-rouge">tcpdump</code> names files
when using these options. I will be using something like this:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>sudo tcpdump -G600 -C1000 -zgzip -Z${USER} -wcapture-%s.pcap -s128 -i eth0 tcp port 12345 ...
#                                adjust these bits according to taste ^^^^^^^^^^^^^^^^^^^^^^^
</code></pre></div></div>

<p>This says to roll over every 10 minutes, and every 1GB, and to run <code class="language-plaintext highlighter-rouge">gzip</code> to
compress the file on rollover. The first file in every 10-minute period is
named something like <code class="language-plaintext highlighter-rouge">capture-1607780545.pcap</code>, and if that file hits 1GB then
subsequent files are named <code class="language-plaintext highlighter-rouge">capture-1607780545.pcap1</code>,
<code class="language-plaintext highlighter-rouge">capture-1607780545.pcap2</code>, etc. The lack of leading zeroes is a bit of a pain:
the eleventh file in each period is called <code class="language-plaintext highlighter-rouge">capture-1607780545.pcap10</code> which
tends to sort between <code class="language-plaintext highlighter-rouge">capture-1607780545.pcap1</code> and
<code class="language-plaintext highlighter-rouge">capture-1607780545.pcap2</code>. At the end of the ten-minute period it starts over
again at <code class="language-plaintext highlighter-rouge">capture-1607781145.pcap</code>. The <code class="language-plaintext highlighter-rouge">-Z${USER}</code> bit means to do the file
writing as the calling user rather than as <code class="language-plaintext highlighter-rouge">root</code>.</p>

<p>Note that <code class="language-plaintext highlighter-rouge">tcpdump</code> will spawn <code class="language-plaintext highlighter-rouge">command</code> each time it rolls over to a new file.
It does not itself limit how many of them are running concurrently, so it’s
your responsibility to make sure that the system can keep up with the traffic.
Make sure to give <code class="language-plaintext highlighter-rouge">tcpdump</code> a suitably restrictive
<a href="https://www.tcpdump.org/manpages/pcap-filter.7.html">filter</a>. If you don’t,
you’ll hit some <a href="https://github.com/lorin/awesome-limits">limit</a> or other
eventually, which probably won’t go well. If the system is under particularly
heavy load then you can alternatively use <code class="language-plaintext highlighter-rouge">tcpdump -w-</code> to send the capture to
<code class="language-plaintext highlighter-rouge">stdout</code> and then pipe it somewhere else that does have the capacity for
further processing. If you do that, make sure to exclude the piped-elsewhere
data from the packets you’re capturing.</p>

<p>The <code class="language-plaintext highlighter-rouge">command</code> is
<a href="https://github.com/the-tcpdump-group/tcpdump/blob/a0e19c0caef95fdcbace674de91e7c181d3bc866/tcpdump.c#L2806">executed</a>
using <code class="language-plaintext highlighter-rouge">execlp(3)</code> so it searches your <code class="language-plaintext highlighter-rouge">$PATH</code> for the executable like a shell,
but you cannot pass any other arguments. Note also that the command above runs
<code class="language-plaintext highlighter-rouge">tcpdump</code> as <code class="language-plaintext highlighter-rouge">root</code> which means that <code class="language-plaintext highlighter-rouge">command</code> runs as <code class="language-plaintext highlighter-rouge">root</code> too.  The man
page suggests writing a script to do some more advanced processing but you may
not want (or be permitted) to run a more complicated script as <code class="language-plaintext highlighter-rouge">root</code>. If you’d
rather do most of the work as a different user then you can have <code class="language-plaintext highlighter-rouge">tcpdump</code> run
<code class="language-plaintext highlighter-rouge">gzip</code> and then use <code class="language-plaintext highlighter-rouge">inotifywait</code> to receive a notification each time a
compressed capture file is ready for further action:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>inotifywait -e delete . -m --format %f \
    | sed -ue '/gz$/d;s/$/.gz/' \
    | xargs -n1 ./done.sh
</code></pre></div></div>

<p>This works because <code class="language-plaintext highlighter-rouge">gzip</code> deletes the original file once the compressed file is
complete which triggers the notification. The <code class="language-plaintext highlighter-rouge">inotifywait</code> command writes out
the names of files that are deleted from the current directory, the <code class="language-plaintext highlighter-rouge">sed</code>
command adds a <code class="language-plaintext highlighter-rouge">.gz</code> to the end of each filename, and then the <code class="language-plaintext highlighter-rouge">xargs -n1</code> runs
the named script on each compressed capture file.</p>

<p>The <code class="language-plaintext highlighter-rouge">done.sh</code> script can do whatever you need, often offloading the data
elsewhere and then removing the original file:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>#!/bin/bash

echo Processing $1
aws s3 cp $1 s3://mybucket/captures/$HOSTNAME/$1
rm -vf $1
</code></pre></div></div>

<p>Invoking <code class="language-plaintext highlighter-rouge">xargs</code> like this will process the files in turn, with no parallelism,
so if the <code class="language-plaintext highlighter-rouge">done.sh</code> script can’t keep up with the traffic then a backlog of
compressed capture files might build up. Consider running these captures in
their own filesystem to bound the disk space they can consume, or else adjust
<code class="language-plaintext highlighter-rouge">done.sh</code> to simply drop the given file with no additional processing if disk
usage is building up.</p>

<p>There’s a couple of feedback gotchas here. Firstly, deleting the compressed
file triggers <code class="language-plaintext highlighter-rouge">inotifywait</code> again, which is why compressed files are skipped by
the <code class="language-plaintext highlighter-rouge">/gz$/d</code> in the <code class="language-plaintext highlighter-rouge">sed</code> script. Secondly, if you’re sending the captured
traffic back out over the network then there’s a risk that <code class="language-plaintext highlighter-rouge">tcpdump</code> will
capture it all over again. Make sure you set up an appropriate filter in the
original capture to avoid that.</p>

<hr />

<p><strong>Addendum 2020-12-15</strong>: An astute colleague pointed out that when we’re doing
these kinds of network trace then we typically only care about the packet
headers, and therefore <code class="language-plaintext highlighter-rouge">-s128</code> is extremely effective at capturing what we need
and dropping all the other junk in the payload. The headers compress pretty
well too since there’s lots of stuff that appears in every one, whereas the
payload is typically encrypted and therefore practically uncompressible. Tools
like <a href="https://www.wireshark.org">Wireshark</a> correctly use the packet length
reported in the header so they handle these truncated packets with no problems.
I added that to the suggested <code class="language-plaintext highlighter-rouge">tcpdump</code> invocation above.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[I sometimes have a need to collect a packet capture for an extended period of time. On a busy host this can generate enough data that it needs some special handling. In particular it’s useful to roll over to a new file every now and then, and to offload the completed files somewhere else so they don’t fill up the disk.]]></summary></entry><entry><title type="html">Ambiguous clocks</title><link href="https://davecturner.github.io/2020/09/02/ambiguous-clocks.html" rel="alternate" type="text/html" title="Ambiguous clocks" /><published>2020-09-02T00:00:00+00:00</published><updated>2020-09-02T00:00:00+00:00</updated><id>https://davecturner.github.io/2020/09/02/ambiguous-clocks</id><content type="html" xml:base="https://davecturner.github.io/2020/09/02/ambiguous-clocks.html"><![CDATA[<p>I recently came across a <a href="https://youtu.be/LT_33XzEX2Q">video from Zach Star</a>
which posed the question of telling the time on a clock whose two hands are
identical lengths. This seems like a fairly well-known puzzle and it’s kinda
entertaining to work out when you can and cannot tell the time on such a clock.
Then I wondered what would happen if you added a second (i.e. third) hand which
is also indistinguishable from the other two hands.</p>

<p><strong>Spoilers follow</strong>. Stop reading if you’d like to try this puzzle yourself
first.</p>

<p>Zach gives a geometric argument for the two-handed case which you can extend to
the three-handed case by adding a dimension, so that you’re looking for
intersections between lines in three-dimensional space which may reasonably
simply not exist.  However actually proving whether they do exist or not was
really rather fiddly given all the different ways that the hands could line up
with each other.</p>

<h2 id="two-handed-case">Two-handed case</h2>

<p>For the two-handed case, the time <code class="language-plaintext highlighter-rouge">t₁</code> is ambiguous if there is a different
time <code class="language-plaintext highlighter-rouge">t₂</code> such that <code class="language-plaintext highlighter-rouge">{h(t₁), m(t₁)} = {h(t₂), m(t₂)}</code>, i.e. <code class="language-plaintext highlighter-rouge">h(t₁) = m(t₂)</code> and
<code class="language-plaintext highlighter-rouge">h(t₂) = m(t₁)</code>, where <code class="language-plaintext highlighter-rouge">h</code> and <code class="language-plaintext highlighter-rouge">m</code> give the respective positions of the hour
and minute hand.  For simplicity’s sake we measure times as a fraction of 12
hours (so <code class="language-plaintext highlighter-rouge">0</code> is midnight and <code class="language-plaintext highlighter-rouge">1</code> is noon) and only consider times in <code class="language-plaintext highlighter-rouge">[0,1)</code>,
and measure the hand position as a fraction of the circle, so that <code class="language-plaintext highlighter-rouge">h(t) = t</code>
and <code class="language-plaintext highlighter-rouge">m(t) = 12t - ⌊12t⌋</code>. Thus we want to solve</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>t₁ = 12t₂ - ⌊12t₂⌋
t₂ = 12t₁ - ⌊12t₁⌋
</code></pre></div></div>

<p>Substituting gives …</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>t₁ = 144t₁ - ⌊144t₁⌋ - ⌊144t₁ - ⌊144t₁⌋⌋
</code></pre></div></div>

<p>… which tells us that the only solutions are where <code class="language-plaintext highlighter-rouge">143t₁</code> must be an
integer, say, <code class="language-plaintext highlighter-rouge">n₁</code>. If <code class="language-plaintext highlighter-rouge">n₁ &lt; 143</code> then <code class="language-plaintext highlighter-rouge">h(t₁) = m(t₂) = n₁/143</code> and <code class="language-plaintext highlighter-rouge">h(t₂) =
m(t₁) = 12n₁ mod 143 / 143</code> is a solution to the equations, but this includes
the cases where the hands coincide and <code class="language-plaintext highlighter-rouge">t₁ = t₂</code> which are not ambiguous.
Excluding those cases gives the answer.</p>

<h2 id="three-handed-case">Three-handed case</h2>

<p>For the three-handed case, the time <code class="language-plaintext highlighter-rouge">t₁</code> is ambiguous if there is a different
time <code class="language-plaintext highlighter-rouge">t₂</code> such that <code class="language-plaintext highlighter-rouge">{h(t₁), m(t₁), s(t₁)} = {h(t₂), m(t₂), s(t₂)}</code> where <code class="language-plaintext highlighter-rouge">h</code>
and <code class="language-plaintext highlighter-rouge">m</code> are as before and <code class="language-plaintext highlighter-rouge">s</code> gives the position of the second hand, i.e. <code class="language-plaintext highlighter-rouge">s(t)
= 720t - ⌊720t⌋</code>. There are a lot of different ways to satisfy the equality
between those sets of hand positions, although at least we know for certain
that <code class="language-plaintext highlighter-rouge">h(t₁) ≠ h(t₂)</code>. Focussing just on the hour hands gives the following four
cases:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>h(t₁) = m(t₂) ∧ h(t₂) = m(t₁)
h(t₁) = s(t₂) ∧ h(t₂) = s(t₁)
h(t₁) = m(t₂) ∧ h(t₂) = s(t₁)
h(t₁) = s(t₂) ∧ h(t₂) = m(t₁)
</code></pre></div></div>

<h3 id="case-1">Case 1</h3>

<p>If <code class="language-plaintext highlighter-rouge">h(t₁) = m(t₂)</code> and <code class="language-plaintext highlighter-rouge">h(t₂) = m(t₁)</code> then this is the same as the two-handed
case, so the solution is of the form <code class="language-plaintext highlighter-rouge">t₁ = n/143</code> and <code class="language-plaintext highlighter-rouge">t₂ = 12n mod 143 / 143</code>.
However,</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>s(t₁) = 720n mod 143 / 143 ∈ {h(t₂), m(t₂), s(t₂)}
                           = {h(t₁), m(t₁), s(t₂)}
                           = {n/143, 12n mod 143 / 143, 8640n mod 143 / 143}
</code></pre></div></div>

<p>Therefore <code class="language-plaintext highlighter-rouge">719n mod 143 = 0</code>, <code class="language-plaintext highlighter-rouge">708n mod 143 = 0</code> or <code class="language-plaintext highlighter-rouge">7920n mod 143 = 0</code>. But
719, 708 and 7920 are all coprime to 143 so in any case <code class="language-plaintext highlighter-rouge">n = 0</code> is the only
possible solution. This implies that <code class="language-plaintext highlighter-rouge">t₁ = t₂</code> (= midnight) which isn’t ambiguous
so this case leads to no ambiguities.</p>

<h3 id="case-2">Case 2</h3>

<p>This case is similar to the previous. If <code class="language-plaintext highlighter-rouge">h(t₁) = s(t₂)</code> and <code class="language-plaintext highlighter-rouge">h(t₂) = s(t₁)</code>
then a similar argument to the two-handed case tells us that the solution is of
the form <code class="language-plaintext highlighter-rouge">t₁ = n/(720×720-1) = n/518399</code> and <code class="language-plaintext highlighter-rouge">t₂ = 720n mod 518399 / 518399</code>.
However,</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>m(t₁) = 12n mod 518399 / 518399 ∈ {h(t₂), m(t₂), s(t₂)}
                                = {h(t₁), m(t₂), s(t₁)}
                                = {n/518399, 8640n mod 518399 / 518399, 720n mod 518399 / 518399}
</code></pre></div></div>

<p>Therefore <code class="language-plaintext highlighter-rouge">11n mod 518399 = 0</code>, <code class="language-plaintext highlighter-rouge">8628n mod 518399 = 0</code> or <code class="language-plaintext highlighter-rouge">708n mod 518399 =
0</code>. But 11, 8628 and 708 are all coprime to 518399 so in any case <code class="language-plaintext highlighter-rouge">n = 0</code> is
the only possible solution. Therefore this case also leads to no ambiguities.</p>

<h3 id="case-3">Case 3</h3>

<p>If <code class="language-plaintext highlighter-rouge">h(t₁) = m(t₂)</code> and <code class="language-plaintext highlighter-rouge">h(t₂) = s(t₁)</code> then</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>t₁ =  12t₂ -  ⌊12t₂⌋
t₂ = 720t₁ - ⌊720t₁⌋
</code></pre></div></div>

<p>Substituting gives</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>t₁ = 8640t₁ - ⌊8640t₁⌋ - ⌊8640t₁ - ⌊8640t₁⌋⌋
</code></pre></div></div>

<p>Therefore the solution is of the form <code class="language-plaintext highlighter-rouge">t₁ = n/8639</code> and <code class="language-plaintext highlighter-rouge">t₂ = 720n mod 8639 / 8639</code>. However,</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>m(t₁) = 12n mod 8639 / 8639 ∈ {h(t₂), m(t₂), s(t₂)}
                            = {h(t₁), s(t₁), s(t₂)}
                            = {n/8639, 720n mod 8639 / 8639, 518400n mod 8639 / 8639}
</code></pre></div></div>

<p>Therefore <code class="language-plaintext highlighter-rouge">11n mod 8639 = 0</code>, <code class="language-plaintext highlighter-rouge">708n mod 8639 = 0</code> or <code class="language-plaintext highlighter-rouge">518388n mod 8639 = 0</code>.
But 11, 708 and 518388 are all coprime to 8639 so in any case <code class="language-plaintext highlighter-rouge">n = 0</code> is the
only possible solution. Therefore this case also leads to no ambiguities.</p>

<h3 id="case-4">Case 4</h3>

<p>This case is the same the previous case, swapping the roles of <code class="language-plaintext highlighter-rouge">t₁</code> and <code class="language-plaintext highlighter-rouge">t₂</code>,
so also leads to no ambiguities.</p>

<h3 id="conclusion">Conclusion</h3>

<p>Since an ambiguity is impossible in all cases, this shows that all positions of
the three hands on the clock uniquely determine the time even if the hands are
indistinguishable.</p>

<p>This was pretty fiddly. There are lots of other ways to split the equality
<code class="language-plaintext highlighter-rouge">{h(t₁), m(t₁), s(t₁)} = {h(t₂), m(t₂), s(t₂)}</code> into cases, many of which look
promising but lead to a lot of duplicated work.  Getting just the right amount
of case splitting took a few goes, and I reckon there’s a more elegant argument
out there somewhere as even this solution really copy-pastes the same argument
four times.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[I recently came across a video from Zach Star which posed the question of telling the time on a clock whose two hands are identical lengths. This seems like a fairly well-known puzzle and it’s kinda entertaining to work out when you can and cannot tell the time on such a clock. Then I wondered what would happen if you added a second (i.e. third) hand which is also indistinguishable from the other two hands.]]></summary></entry></feed>