Skip to content

SIGSEGV in pcre2_match() from concurrent worker threads sharing rx_rule->md (num_workers>1) #231

Description

@BenkiNew

Version: aide-0.19.2 (el10 build), also plausible on current master — the relevant code path is unchanged.

Config: num_workers=4, local file database (no database_in=http, unrelated to #229/#182).

Crash (coredump backtrace, thread hit inside a worker):

#0  match.constprop.0 (libpcre2-8.so.0)
#1  pcre2_match_8 (libpcre2-8.so.0)
#2  check_list_for_match.isra.0 (aide)
#3  check_node_for_match.isra.0 (aide)
#4  check_seltree (aide)
#5  check_rxtree (aide)
#6  process_path (aide)
#7  process_disk_entries (aide)
#8  worker (aide)
#9  start_thread (libc.so.6)

Memory: 2.9G at crash time, no OOM — this is a straight SIGSEGV, not a resource-limit kill.

Root cause (traced in source):

check_list_for_match() in src/seltree.c calls:

pcre_retval = pcre2_match(rx->crx, (PCRE2_SPTR) file.name, PCRE2_ZERO_TERMINATED, 0, PCRE2_PARTIAL_SOFT, rx->md, NULL);

rx->md (include/rx_rule.h:49) is a single pcre2_match_data* allocated once per rule when the rule is added to the seltree (seltree.c:241, r->md = pcre2_match_data_create_from_pattern(r->crx, NULL)), then reused for the lifetime of the run.

check_rxtree/check_seltree/check_node_for_match/check_list_for_match are called from every worker thread spawned in db_disk.c:worker() (via process_disk_entries → process_path), against the same seltree — so every worker thread checking a file against the same rule calls pcre2_match() with the same shared rx->md concurrently. pcre2_match_data is explicitly documented as not safe for concurrent use by multiple threads (PCRE2 mutates the match_data's internal ovector during matching). With enough concurrent hits on the same rule this corrupts the match_data and segfaults inside PCRE2's internal matcher — consistent with the backtrace above.

Why this isn't a duplicate of existing SIGSEGV reports: #229/#182 are both about database_in=http + libcurl (DNS/remote-fetch related, curl frames in their backtraces). #228 is about get_different_attributes() with percent-encoded filenames. None involve num_workers, PCRE2, or the seltree matching path.

Suggested fix (not yet implemented/tested by me — flagging the shape of it rather than a proposed patch): the correct fix is per-worker-thread match_data, not a mutex around the shared one (a mutex would serialize regex matching across all workers, largely defeating the point of num_workers>1). Concretely: thread the already-existing worker_index (currently plumbed through gen_list.c for hashing, e.g. get_file_attrs(..., int worker_index, ...)) through check_rxtree → check_seltree → check_node_for_match → check_list_for_match as well, and change rx_rule.md from a single pointer to an array sized conf->num_workers + 1 (indices already used elsewhere for the dry-run/single-thread case at index 0), created alongside crx in seltree.c and indexed by worker_index at the pcre2_match() call site.

I didn't submit this as a PR because I can't currently build/test aide's full suite (incl. any thread-sanitizer runs) to verify the fix actually closes the race rather than just moving it — happy to attempt a patch if that's useful, but wanted to report the root cause precisely either way. Can provide the full coredump/journal excerpt or aide.conf if helpful.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions