Skip to content

CoverageData leaks one SQLite file descriptor per terminated thread #2192

Description

@MattLloyd101

Environment

coverage 7.14.1, Python 3.13.7, macOS arm64 (also reproduced with COVERAGE_CORE=sysmon).

Describe the bug

CoverageData keeps one SqliteDb per thread in self._dbs, keyed by threading.get_ident() (sqldata.py: _open_db/_connect). The connection is never closed when the creating thread terminates — entries are only reused if the same ident recurs, and the dict is only fully closed at process end. On workloads with many short-lived threads whose idents are not recycled, _dbs grows without bound and each dead thread permanently leaks one open fd to the data file, eventually hitting OSError: [Errno 24] Too many open files.

Real-world evidence:

A long pytest run showed ~6 live OS threads (ps -M) but 30k open fds to .coverage (lsof) before EMFILE.

To Reproduce

import threading
import coverage

data = coverage.CoverageData(no_disk=True)

# 200 threads that all record coverage at the same time, then all exit.
barrier = threading.Barrier(201)


def work():
    data.add_lines({"foo.py": [1, 2, 3]})  # opens this thread's SqliteDb
    barrier.wait()


threads = [threading.Thread(target=work) for _ in range(200)]
for t in threads:
    t.start()
barrier.wait()          # let every thread open its connection concurrently
for t in threads:
    t.join()            # every worker thread is now dead

print("live threads:", threading.active_count())
print("retained connections in _dbs:", len(data._dbs))
assert threading.active_count() == 1
assert len(data._dbs) == 200, "connections from dead threads were not released"
print("BUG REPRODUCED: 200 dead threads still hold 200 open connections")

Answer the questions below:

  1. What version of Python are you using?

Python 3.13.7

  1. What version of coverage.py shows the problem? The output of coverage debug sys is helpful.

coverage 7.14.1

  1. What versions of what packages do you have installed? The output of pip freeze is helpful.

Our PIP freeze is extensive.

  1. What code shows the problem? Give us a specific commit of a specific repo that we can check out. If you've already worked around the problem, please provide a commit before that fix.

Unfortunately the repo is private, but essentially it is spinning up Django Apps and we are running coverage over our contract tests. The code above is a cut down to the absolute minimum example of the issue.

  1. What commands should we run to reproduce the problem? Be specific. Include everything, even git clone, pip install, and so on. Explain like we're five!
# 1. Make an empty working directory and go into it
mkdir coverage-fd-leak && cd coverage-fd-leak

# 2. Create and activate a fresh virtual environment (Python 3.11+)
python -m venv venv
source venv/bin/activate          # Windows: venv\Scripts\activate

# 3. Install the exact coverage version we saw this on
pip install coverage==7.14.1

# 4. Save the script below as repro.py (see the code block), then run it
python repro.py

repro.py:

import threading
import coverage

data = coverage.CoverageData(no_disk=True)

# 200 threads that all record coverage at the same time, then all exit.
barrier = threading.Barrier(201)


def work():
    data.add_lines({"foo.py": [1, 2, 3]})  # opens this thread's SqliteDb
    barrier.wait()


threads = [threading.Thread(target=work) for _ in range(200)]
for t in threads:
    t.start()
barrier.wait()          # let every thread open its connection concurrently
for t in threads:
    t.join()            # every worker thread is now dead

print("live threads:", threading.active_count())
print("retained connections in _dbs:", len(data._dbs))
assert threading.active_count() == 1
assert len(data._dbs) == 200, "connections from dead threads were not released"
print("BUG REPRODUCED: 200 dead threads still hold 200 open connections")

Running it prints:

live threads: 1
retained connections in _dbs: 200
BUG REPRODUCED: 200 dead threads still hold 200 open connections

Note: the threads are kept briefly concurrent (via the Barrier) so they get distinct thread ids. A naive sequential start()/join() loop can let the OS recycle a single id and hide the bug — so please keep the barrier in.

Expected behavior

When a thread that recorded coverage has terminated, the SqliteDb connection it opened should be closed and removed from CoverageData._dbs, so the number of retained connections (and the file descriptors they hold) stays bounded by the number of live threads.

Concretely, after the 200 worker threads above have all been join()-ed and only the main thread remains alive, len(data._dbs) should fall back to the live-thread count (1), not stay pinned at 200. Today the entry for each dead thread's id is kept forever (it is only reused if that exact id recurs, and the whole dict is only closed at process end), so on a long run with many short-lived threads _dbs grows without bound and leaks one fd per dead thread until the process hits OSError: [Errno 24] Too many open files.

Additional context

  • Real-world symptom: a multi-hour pytest + pytest-cov run failed with OSError: [Errno 24] Too many open files. At the point of failure the process had only ~6 live OS threads (ps -M <pid> | wc -l) but 30,000+ open file descriptors pointing at the single .coverage data file (lsof -p <pid> | grep -c /path/to/.coverage), climbing steadily (11k → 17k → 30k) over the run. The small live-thread count vs. huge fd count is the tell-tale sign that the fds belong to dead threads.

  • Not the tracer: reproduced identically with both the default C tracer and COVERAGE_CORE=sysmon, which points at the data layer (CoverageData) rather than the collector.

  • Source: CoverageData._open_db stores self._dbs[threading.get_ident()] = SqliteDb(...) and _connect only ever adds entries — there is no path that closes a connection when its thread dies (coverage/sqldata.py). Connections are created with sqlite3.connect(..., check_same_thread=False) (coverage/sqlitedb.py), so closing a dead thread's connection from another thread is safe.

Workaround

We are monkey patching coverage to address this:

"""
Workaround for coverage.py leaking one SQLite file descriptor per dead thread.

``coverage.CoverageData`` keeps one ``SqliteDb`` (and thus one open fd to the
``.coverage`` data file) per thread, keyed by ``threading.get_ident()``. The
connection is created lazily in ``_open_db`` and stored in ``self._dbs``, but
it is never closed when the thread that created it terminates. Entries are
only ever reused if the *same* thread id recurs, and the whole dict is only
closed at process end.

Under a real coverage run, thread ids are not recycled (a handful of live
threads, but thousands of short-lived threads over the run), so ``_dbs`` grows
without bound and every terminated thread permanently leaks one fd. This
eventually reaches the per-process file-descriptor limit and fails the run
with ``OSError: [Errno 24] Too many open files``.

Observed on a contracts coverage run: ~6 live threads (``ps -M``) but 11k ->
17k -> 30k open fds to ``src/.coverage`` (``lsof``) before hitting EMFILE.

We patch ``CoverageData._connect`` so that, on the (rare) path where a thread
opens its first connection, we first close and drop any entries belonging to
threads that are no longer alive. This bounds ``_dbs`` to the live-thread
count. Connections are created with ``sqlite3.connect(..,
check_same_thread=False)``, so closing a dead thread's connection from another
thread is safe.

This should be fixed upstream in coverage.py, at which point this plugin can
be removed (the assertions below will fail loudly to remind us).
"""

from __future__ import annotations

import threading

from coverage.sqldata import CoverageData
from coverage.sqlitedb import SqliteDb

# Guard: if a coverage upgrade renames or removes any of the internals we rely
# on, fail loudly at import time rather than silently no-op'ing the workaround.
assert hasattr(CoverageData, "_connect"), (
    "CoverageData._connect no longer exists; revisit this plugin (coverage upgrade?)"
)
assert hasattr(CoverageData, "_open_db"), (
    "CoverageData._open_db no longer exists; revisit this plugin (coverage upgrade?)"
)
assert "_dbs" in CoverageData.__init__.__code__.co_names, (
    "CoverageData no longer initialises _dbs; revisit this plugin (coverage upgrade?)"
)
assert "_lock" in CoverageData.__init__.__code__.co_names, (
    "CoverageData no longer initialises _lock; revisit this plugin (coverage upgrade?)"
)

_original_connect = CoverageData._connect


def _connect_with_reaping(self: CoverageData) -> SqliteDb:
    """
    ``CoverageData._connect`` that reaps connections from dead threads.

    Only the cold path (this thread has no connection yet) does any reaping,
    so the hot path taken on every ``add_lines``/``add_arcs`` is untouched.
    """
    if threading.get_ident() not in self._dbs:
        with self._lock:
            live_idents = {thread.ident for thread in threading.enumerate()}
            dead_idents = [ident for ident in self._dbs if ident not in live_idents]
            for ident in dead_idents:
                db = self._dbs.pop(ident)
                try:
                    db.close(force=True)
                except Exception:  # noqa: BLE001
                    # Closing is best-effort; a failure here must not break
                    # collection. The entry has already been dropped.
                    pass
    return _original_connect(self)


CoverageData._connect = _connect_with_reaping  # type: ignore[method-assign]

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions