Skip to content

man: document that stdio file targets are opened before the MAC transition - #43646

Merged
bluca merged 1 commit into
systemd:mainfrom
BenkiNew:man/stdout-file-mac-context
Sep 5, 2026
Merged

bluca merged 1 commit into
systemd:mainfrom
BenkiNew:man/stdout-file-mac-context

Conversation

@BenkiNew

@BenkiNew BenkiNew commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

StandardInput= with file:, and StandardOutput=/StandardError= with file:, append: or truncate:, open the target while the process is still being set up — before execve(), and therefore before the MAC label of the service process is applied. The open is consequently checked against the context inherited from the service manager, which may differ from the one the service ends up running in.

systemd.exec(5) currently mentions the manager's privileges only for creating the file:

it is opened (created if it does not exist yet using privileges of the user executing the systemd process) for writing at the beginning of the file

and says nothing about an existing file being opened in the manager's security context. On SELinux systems that is easy to trip over: labelling the log file only for the domain the service transitions into is not enough, and in enforcing mode the unit's log output is silently dropped.

Ordering in src/core/exec-invoke.c is explicit: setup_input() and setup_output() — which open the target via acquire_path() — run at lines 5563, 5576 and 5582, sym_setexeccon_raw() at 6471 and fexecve_or_execve() at 6754. For SMACK the label is applied to the process itself in setup_smack() at 6342, likewise after the targets have been opened. SELinuxContext= therefore cannot fix this either, which the note now states explicitly.

Observed in production on CentOS Stream 10 (systemd 257-33.el10, selinux-policy-targeted 42.1.27-2). A single unit — Type=oneshot, User=, StandardOutput=append: at a file labelled httpd_sys_content_t, whose ExecStart also appends to that same file itself — shows both sides:

# manager-side setup of stdout/stderr, pre-exec:
avc: denied { open }   comm="(probe.sh)" exe="/usr/lib/systemd/systemd-executor" ppid=1
     subj=system_u:system_r:init_t:s0 tcontext=system_u:object_r:httpd_sys_content_t:s0 tclass=file
avc: denied { append } comm="(probe.sh)" exe="/usr/lib/systemd/systemd-executor" ppid=1
     subj=system_u:system_r:init_t:s0 tcontext=system_u:object_r:httpd_sys_content_t:s0 tclass=file

# the script's own append to the very same file, post-exec — no denial:
SCRIPT-OWN-WRITE ctx=system_u:system_r:unconfined_service_t:s0

open and append are denied as separate permissions, and a not-yet-existing target adds create, which is why the note enumerates the required access rather than saying "for writing". Setting SELinuxContext= on the unit leaves those denials unchanged, since setexeccon() runs long after the target has been opened.

Documentation only; no behaviour change. This is not proposed as a code fix — the fd must exist before execve() and the label cannot be applied before it, and #14385 indicates the fd-passing behaviour is intentional — so the aim is just to make the constraint discoverable.

Follows up on #43455, where I originally misattributed these denials to a failed domain transition for Type=oneshot + User= units; a correction with the full reproduction matrix is posted there.

@github-actions

github-actions Bot commented Sep 4, 2026 •

Copy link
Copy Markdown

Re-reviewed the note in man/systemd.exec.xml. The head moved again since the last round, to e30b25f (one commit, documentation only, authored by benkigeek), and the previous head def2909 - like 0ce4870, 36a54f2, 69bd37f, bf49673, 67cdef5, f009241, 1147c0a, abdb95e, 3ebd2bd and fa9d387 before it - is gone from the checkout, so this head was assessed on its own merits. The recorded base SHA 726e17a is again a moved main tip; the true merge base is 92c9f78, so the effective diff is man/systemd.exec.xml alone at +11/-1 - the hwdb.d/, src/, test/ and meson.build changes visible against the recorded base belong to main and were excluded from review.

This round is different in kind from the previous ten. Maintainer feedback arrived twice - "I am not convinced this is necessary ... adds a lot of noise explaining stuff that is very niche" and, inline, "Of course this is opened before invoking the executable ... Invoking the binary is the boundary between the service manager's code and the executable's code", followed by "please drop this paragraph". The author responded by cutting the note from five paragraphs to two, then removing the second of those as well. Everything the previous ten rounds were filed against - the per-permission SELinux enumeration, init_t, sock_file, add_name, the type-derivation prose, the domain_fd_use parenthetical, the audit/permissive/null-device paragraph and the Reproduced narration - is gone from this head. All 57 previously recorded findings therefore target text that no longer exists and are marked fixed, and all 43 existing bot threads are resolved. Note that the direction of travel is removal: the remaining eleven added lines are themselves under an explicit maintainer request to drop, so the findings below matter only if any version of the note survives.

Seven of the eight launched lenses completed within budget: correctness and memory safety, security and robustness, architecture/API/compatibility, coding style and maintainability, tests and regression coverage, plus the two domain lenses - MAC/SELinux policy semantics and DocBook markup/generated output. One lens is incomplete: resource lifetimes and concurrency exceeded its budget and was stopped, so its results are not included; its subject matter overlaps heavily with the correctness and MAC/SELinux lenses, both of which completed and independently reached the descriptor-inheritance finding below. All three validators completed.

Three findings were confirmed, all against the surviving note at 3520-3526.

Two are must-fix. Item 58 is a plain factual error and the only finding in eleven rounds that is not about SELinux: the sentence distributes file:, append: and truncate: across StandardOutput=, StandardError= and StandardInput=, but StandardInput= has no append: or truncate: value - exec_input_table (execute.c:3149-3157) has only null, tty, tty-force, tty-fail, socket, fd, data and file, and config_parse_exec_input() (load-fragment.c:1136-1202) special-cases only fd: and file: before falling through to exec_input_from_string(). The page's own value list for StandardInput= at 3392-3395 agrees, so the note contradicts the page eleven lines above it, and the cross-reference this same commit adds at 3423-3425 routes StandardInput= readers to the wrong claim. This is a regression introduced by over-correcting the very first round's item 1, which asked for StandardInput= to be covered at all. Item 59 is the defect three lenses reached independently: "The resulting descriptor is then inherited across that transition into the unit's own domain" is unconditional, and inheritance is precisely what SELinux does not guarantee across a domain change - the kernel revalidates each inherited descriptor at credential commit and substitutes the null device on denial, and thereafter revalidates read/write/append against the current subject rather than honouring the opener's authorization. The author established this mechanism himself in rounds 9 and 10; the sentence as written tells an administrator the opposite.

One suggestion. Item 60: the definite "across that transition" presupposes a transition always happens, but a unit with no SELinuxContext= and no entrypoint type_transition keeps the manager's domain, and ExecStart=+... commands skip the MAC block entirely since needs_sandboxing is false (exec-invoke.c:5496), so sym_setexeccon_raw() at 6471 is never reached.

Several candidates were dropped. A framing objection - that the note documents an internal ordering as an interface contract, and that "and hence" is not the real causal chain - went to a validator and was not confirmed: it restates the maintainer's own inline point on a thread he already owns, and the implication is in fact sound because the transition takes effect at execve(). A terminology nit on "the target path" at 3523 diverging from "the file system object" at 3423 was dropped as churn inside a block under a drop request, and as re-opening a thread the author already adopted. An antecedent-ambiguity point on "that transition" was dropped as overlapping the sentence item 59 already asks to rewrite. A portability point - that the note names only SELinuxContext= while SMACK's setup_smack() relabels the running process rather than arming execve(), and that SELinuxContextFromNet= can win over SELinuxContext= at exec-invoke.c:6468 - was dropped as a re-framing of the dismissed item 23 now that the current text carries no remediation at all.

The tests lens re-verified from scratch that no build or test machinery needs updating for this head: it reimplemented tools/make-directive-index.py's extraction rules over base and head and got byte-identical key sets, since the added markup is inline elements inside para bodies and the script harvests option/varname only from varlistentry terms - so unlike the earlier head that added init_t, this one adds no systemd.directives(7) entry. refmeta/refnamediv are untouched so man/rules/meson.build cannot go stale; the new para is mid-listitem in both places so the version-info include remains the last child, and all three settings are in tools/command_ignorelist regardless; no test or fixture keys off this page's prose. The markup lens found nothing and confirmed the page parses cleanly, that the em dash has precedent in this same file, and that the spaced StandardInput= idiom flagged as items 44 and 57 is absent from this head. The style lens confirmed no added line exceeds the max_line_length = 109 that .editorconfig sets for man/*.xml. Per the standing scope call, no test was requested for this documentation-only change.

Suggestions

  • 1. The note is scoped to the three StandardOutput= modes, but StandardInput=file: opens its path through the same acquire_path() helper even earlier, so it is subject to the identical pre-transition check; a reader of the StandardInput= docs gets no hint.
  • 2. "not the one the service ultimately runs in" is stated as an absolute, yet the two contexts are identical when policy defines no transition and no SELinuxContext= is set, and for ExecStart=+... commands where needs_sandboxing is false so the setexeccon() block is skipped.
  • 3. The init_t parenthetical is policy- and instance-dependent: it is the Fedora/RHEL targeted-policy type, the open is actually performed by the posix_spawn()ed pinned systemd-executor child rather than PID 1, and the sentence leaves the per-user manager (a different domain) unanswered.
  • 4. "may open it for writing" understates the required policy: append: needs the distinct append permission, open is its own permission, O_CREAT is added for write modes so parent-directory access is needed too, and the target may be an AF_UNIX socket, FIFO or special file with different classes.
  • 5. init_t will be pulled into the generated systemd.directives(7) index by tools/make-directive-index.py, listing a distro policy type as a systemd constant; is the established precedent for SELinux type names.

Nits

  • 6. "MAC (Mandatory Access Control)" inverts the spell-out-first convention used throughout man/, and describing the transition as "associated with the execution" is wrong for SMACK, which applies the label to the running process rather than at execve().

Must fix

  • 7. The permission enumeration never mentions write. acquire_path() adds O_CREAT only for O_WRONLY/O_RDWR, and setup_output() opens file:/truncate:/append: as O_WRONLY (plus O_TRUNC/O_APPEND), so a policy granting only open/append/create still denies file: and truncate:. Conversely StandardInput=file: is opened O_RDONLY and can never create the file, so the create and enclosing-directory advice does not apply to it.

Suggestions

  • 8. Nothing on this path calls mac_selinux_create_file_prepare(), so a file the manager creates takes its type from the parent directory's type transition rather than from any fcontext rule for the path; that resulting label is what decides whether the service can then use the file.
  • 9. "the checks apply to the corresponding object class instead" is right for FIFOs and device nodes but wrong for AF_UNIX sockets: acquire_path() attempts open() first and only falls back to connect() on ENXIO, so the sock_file checks are additive rather than replaced, and connectto is evaluated against the listening peer's label, not the socket path's.

Must fix

  • 10. The shared read-write descriptor happens only when the output mode is file:, not append: or truncate:. setup_input() computes rw solely from std_output/std_error == EXEC_OUTPUT_FILE (exec-invoke.c:426-427), so StandardInput=file:/x with StandardOutput=append:/x opens O_RDONLY, acquire_path() adds no O_CREAT, and setup_output() still dup2()s that read-only fd for all three modes - nothing is created and no create/write/append is exercised.

Suggestions

  • 11. Directory access is stated only as a consequence of create, but the manager-side open() resolves the whole path, so search on class dir is needed for every component unconditionally; and "the corresponding access to the enclosing directory" is worth naming as write and add_name alongside create.
  • 12. The AF_UNIX case is missing write on sock_file, which the kernel's MAY_WRITE check in unix_find_other() requires, so a StandardInput=file: socket target labelled per this text is still denied; connectto is also a unix_stream_socket permission, and the FIFO/device classes (fifo_file, chr_file/blk_file) are worth spelling out.
  • 13. The type-transition sentence presents the rarer case as the rule: absent a matching type_transition, security_compute_sid() gives the new file the enclosing directory's own type. Leading with the systemd-side fact - that no file creation context is applied when opening the target - is both more accurate and less fragile than describing the resulting label.

Must fix

  • 14. write on class sock_file is attached to the initial open() rather than to the connect(). For StandardInput=file: the open is O_RDONLY (exec-invoke.c:429), so the failing open is checked against read, contradicting the "read for StandardInput=" enumeration above; and since acquire_path() falls through to the socket path only on exactly ENXIO (285-286), an EACCES there never reaches connect() at all. The write requirement comes from unix_find_other()'s MAY_WRITE check during connect_unix_path() (294) and applies in both directions.

Suggestions

  • 15. The consequence is framed as MAC-only, but the same ordering means the DAC side of the open uses the manager's credentials: setup_input()/setup_output() at 5563/5576/5582 precede enforce_user() at 6410 and capability_bounding_set_drop() at 6376, so StandardOutput=truncate: at a root-only path succeeds under User=nobody. (dismissed)
  • 16. The open also precedes all namespacing - setup_delegated_namespaces() first runs at 6094 and apply_root_directory() at 6404 - so the path is resolved against the host root and is not subject to RootDirectory=/ProtectSystem=/ReadOnlyPaths=; and acquire_path() passes no O_NOFOLLOW, so symlinks are followed with the manager's privileges. (dismissed)
  • 17. The descriptor-sharing contract is now stated twice with different scope. The unconditional sentence at 3496-3499 is the inaccurate one for append:/truncate:, and the accurate statement is buried in the SELinux note where a reader looking up descriptor sharing will not find it; narrow 3496 in place instead. (dismissed)
  • 18. Restricting the failure to a not-yet-existing shared path invites the inverse conclusion that the pairing works once the path exists. setup_output() dup2()s the read-only stdin fd for all three modes (637-641), so the service gets a read-only stdout/stderr with no append or truncate semantics. (dismissed)
  • 19. The AF_UNIX paragraph omits the peer-credential consequence: the connect happens in the manager's context, so a peer authorizing via SO_PEERCRED sees uid 0 and via SO_PEERSEC the manager's domain, not the unit's User= or SELinuxContext=. (dismissed)
  • 20. Roughly 40 lines of MAC/SELinux prose are inserted between the truncate: and socket paragraphs, interrupting the value-by-value enumeration so the option list reads as if it ended at truncate:. The block is cross-cutting and belongs after the enumeration or in the existing Mandatory Access Control refsect1 at line 972. (dismissed)
  • 21. The SELinuxContext= cross-reference is one-directional: the SELinuxContext= entry at 977-987 says nothing about stdio targets being opened earlier, though it is the first place an administrator chasing a denial looks. (dismissed)

Nits

  • 22. The hedging is asymmetric: init_t is qualified with "typically ... on common policies" but "a different domain for the per-user one" is stated absolutely, though that is equally policy-dependent. (dismissed)
  • 23. The lead-in is deliberately backend-neutral but every piece of remediation that follows is SELinux-only, so a SMACK or AppArmor administrator learns the problem exists and nothing about the fix; under SMACK a new file takes the creating process's label, the opposite default from SELinux. (dismissed)
  • 24. "see the note on the security context this happens in for StandardOutput= below" inverts the reference form used elsewhere on this page, which puts the target element immediately after the referring phrase. (dismissed)

Suggestions

  • 25. The note never mentions StandardError=, although it takes file:/append:/truncate: by reference to StandardOutput= and its targets are opened by the same setup_output() call at the same point, so the enumeration reads as if StandardError=file: were exempt; this page's "(or error output, see below)" convention covers exactly this case.
  • 26. The shared-descriptor condition is narrower than the code, which makes the startup-failure claim false for a legal unit: setup_input() sets rw from std_output or std_error being EXEC_OUTPUT_FILE at the same path (exec-invoke.c:426-427) and setup_output()'s dup2() shortcut is not gated on the output mode (637-641), so StandardInput=file:/x with StandardOutput=truncate:/x and StandardError=file:/x creates a missing /x and starts.

Nits

  • 27. The block switches vocabulary mid-way: its first paragraph says "the path", then a bare "the target" is introduced six times (3531, 3539, 3541, 3543, 3550, 3557) for the same thing, where the surrounding prose already uses "file system object", "the file" and path.
  • 28. In the AF_UNIX sentence the trailing "before it fails" attaches to "the target" rather than to the open attempt, and burying the failure behind two modifiers hides that the open() is expected to fail with ENXIO yet must still be permitted before the connect is attempted.

Suggestions

  • 29. The shared read-write descriptor carries neither O_APPEND nor O_TRUNC and setup_output()'s dup2() shortcut bypasses the flags computation, so with StandardInput=file:/x plus a file: output plus StandardError=truncate:/x the truncate: never truncates and a co-located append: never appends - contradicting the unconditional truncate: text at 3505-3510, and sharing a single file offset with stdin.
  • 30. The creation paragraph names only the broadest policy grant: create on the directory-derived shared type plus write/add_name on the directory, a label the service's own domain often cannot write to. The narrower alternative - pre-creating the target with its intended label, e.g. via tmpfiles.d, so only open plus write/append on one dedicated type is needed - is worth a sentence, with the caveat that it holds only while the file exists.
  • 31. "on its own" is stranded four lines after the verb it qualifies and is ambiguous between "each mode creates the path independently" and "the path is created on its own"; and "co-located" is the only occurrence of that word in man/, where this page already says "the same file path".

Nits

  • 32. "the connect then additionally requires" uses connect as a bare noun one clause after introducing it as connect(); the same nominalisation appears at 3535 ("the access the open actually requires"), and these two are the only such uses in man/.

Must fix

  • 33. "requires create on the path" names a label the kernel never consults: at creation time the path has no inode, so its file_contexts entry is not consulted and create is checked against the SID the new inode will receive - the type-transition result for the manager's domain and the enclosing directory's type, which is what this paragraph's own closing sentences at 3562-3566 say. The neighbouring write and add_name are attributed correctly to the directory, sharpening the contrast; an allow rule written from this text loads and is never consulted.

Suggestions

  • 34. The permission enumeration at 3535-3540 switches keying axis mid-list, keying read to a setting (StandardInput=) but write/append to option values, so "write for file:" read as written also demands write for a lone StandardInput=file:, which is opened O_RDONLY and needs only open plus read. The list also omits that the shared O_RDWR open needs read and write together, and that in the sharing case append is never checked at all because setup_output()'s dup2() shortcut returns before flags |= O_APPEND.
  • 35. Descriptor sharing is not gated on the descriptor being read-write: rw at exec-invoke.c:426 requires a plain file: output at the same path, but the dup2() shortcut at 641 fires on std_input == EXEC_INPUT_FILE plus path equality alone, from inside the case that also covers append: and truncate:. So StandardInput=file:/x with StandardOutput=append:/x on an existing /x starts successfully and the service's writes fail with EBADF - neither the shared-offset outcome at 3557-3558 nor the startup failure at 3559-3561.
  • 36. The startup failure at 3559-3561 is not a MAC denial and emits no AVC: acquire_path() adds O_CREAT only for O_WRONLY/O_RDWR (278-279), the non-sharing StandardInput=file: open is O_RDONLY (429), and setup_input() runs before both setup_output() calls (5563 vs 5576/5582), so the unit fails with a plain ENOENT before the append:/truncate: open that would have created the file is attempted. Granting create or add_name cannot help there.
  • 37. The abbreviated append:/truncate: at 3557 is the only place in the whole man/ tree where these options appear without path; every other occurrence on this page, including three in this same paragraph, spells them out. Rendered, it comes out as bold "append:" with a dangling colon.

Nits

  • 38. Line 3572 is 117 columns, over the max_line_length = 109 that .editorconfig sets for [man/*.xml], and is the only added line that overruns - the next longest land exactly on 109 - so the in-place reword to "the connection attempt" was evidently not followed by a re-wrap.

Must fix

  • 39. The new "except when the path is shared" clause at 3540-3544 asserts that granting read and write covers the open "regardless of which of these modes triggered the sharing", but setup_input() promotes to O_RDWR only for a plain file: output (exec-invoke.c:426-427 tests EXEC_OUTPUT_FILE alone) while setup_output()'s dup2() shortcut at 639-641 fires for all three modes ahead of the flag computation. With StandardInput=file:/x plus StandardOutput=append:/x there is exactly one open and it is O_RDONLY, so write and append are never checked and the prescribed write grant on init_t is never exercised.
  • 40. "can trigger the same creation, or instead share an already-open descriptor" presents as alternatives what the code makes one case: rw == true is simultaneously the only condition under which the StandardInput= open carries O_CREAT and the condition under which setup_output() takes the dup2() shortcut. A StandardInput=file: not paired with a plain file: output opens O_RDONLY, gets no O_CREAT and cannot create anything, so a reader grants add_name/create/directory-write for an impossible case; and the sharing direction is inverted, setup_input() running before both setup_output() calls.

Suggestions

  • 41. The creation claim at 3547-3551 is unconditional but fails when a StandardInput=file: shares its path with an append:/truncate: output: rw stays false, the input open is O_RDONLY so acquire_path() adds no O_CREAT, setup_output() only dup2()s it, and startup fails with ENOENT in setup_input() before the output side runs. The hedge that follows only qualifies the StandardInput= side.
  • 42. The paragraph enumerates only manager-side permissions and closes on "Labelling the path only for the domain the service transitions into is not sufficient" without saying what the service side needs. SELinux revalidates every read()/write() on the inherited descriptor once the current SID no longer matches the opener's, requiring fd { use } to the manager's domain plus read/write/append on the file's type, so a policy written from this note alone loads and then denies the service's first write.
  • 43. Attaching create to "the label the new file receives - not on any label defined for the path itself" implies by contrast that the open/write/append permissions at 3535-3543 do apply to the path's own label. In the creation case they do not: the open-time check runs against the derived label too, so granting the manager's domain open/write on a semanage fcontext type yields rules that never fire.
  • 44. At 3554 a literal space separates StandardInput= from file:path, rendering as "StandardInput= file:path"; it is the only same-line spaced occurrence of that idiom in man/, where the established form is unspaced.
  • 45. 3547-3562 is a single ~16-line paragraph carrying four independent claims and stating the label-derivation negative twice (3552-3553 and again at 3559-3562), with "see below for how that label is derived" pointing six lines down inside the same , which has no visible target in the rendered page.

Nits

  • 46. "see the general note on descriptor sharing above" names a section that does not exist; the only text above is the unlabelled pre-existing sentence at 3496-3499, which uses neither "descriptor" nor "sharing". This page's idiom is a bare "see above".

Must fix

  • 47. The new paragraph's opening claim that labelling only for the manager's domain "is not sufficient either" is unconditional, but it holds only where the service actually transitions. sym_setexeccon_raw() (exec-invoke.c:6471) runs only when a SELinux context is set and policy need not define a type transition for the executable, so a unit with no SELinuxContext= and no entrypoint rule keeps the manager's domain; file_permission then short-circuits on the SID recorded at open time, no inherited-descriptor sweep happens, and manager-side labelling alone is entirely sufficient. The paragraph at 3526-3530 already hedges with "may differ".
  • 48. "Reproduced directly: ... a second, independently-timestamped append denial then appears" narrates a lab observation instead of stating the rule the preceding sentence already states. Nothing above establishes that the reader is looking at audit output, "as expected"/"then appears" scopes the claim to one policy and enforcement mode, grep -rn Reproduced man/ matches only this line, and as a diagnostic it misleads: a dontaudit rule, AVC caching or audit rate-limiting suppresses the second record while the access is still denied.
  • 49. The timing in that same sentence looks wrong for enforcing mode. The descriptor is opened in the process that later execve()s the service and is deliberately left without O_CLOEXEC, so the kernel revalidates it at the transition: selinux_bprm_committing_creds() calls flush_unauthorized_files(), which re-checks each inherited descriptor with file_to_av() (an O_APPEND descriptor maps to append) and on failure substitutes the SELinux null device. The second denial is therefore timestamped at execve() and the service's writes then succeed while the output is silently discarded - the more useful symptom to document - with the per-write denial being what permissive mode shows.

Suggestions

  • 50. Listing read, write and append together as all re-checked conflicts with the note's own per-mode mapping at 3539-3540: append replaces write rather than adding to it, since revalidation adds MAY_APPEND for an O_APPEND descriptor and only adds FILE__WRITE when MAY_APPEND is absent. For append: a write is checked as append and never as write, for file:/truncate: only write applies, for StandardInput=file: only read (or read and write when shared).
  • 51. The closing parenthetical pins the page to a third-party policy's mutable configuration, and no other page in man/ documents an SELinux boolean or its default. domain_fd_use is a compile-time gen_tunable in upstream refpolicy, exposed as a runtime boolean only on derived distro policies; "between all domains" is really between types carrying the domain attribute; and "between the two domains" hides that the check is directional - the service's domain as source against the manager's domain as the descriptor's SID - so a hand-written rule has even odds of being backwards, loading cleanly and changing nothing.

Nits

  • 52. "A fd class use permission" inverts the class-name ordering this same note uses at 3579 ("checked against class sock_file") and renders as two adjacent quoted tokens with no connecting word, and bare fd already denotes the fd: option value on this page at 3440 and 3592.
  • 53. open and append are used as bare nouns for the operation, where this page writes open() (3135, 3578) and otherwise reserves in this block for permission names (3535, 3539, 3548, 3563); permissions also do not "succeed", checks and operations do.
  • 54. Bare "MAC transition" uses an abbreviation the page never introduces - 3526 and the section title at 972 both spell out "Mandatory Access Control" - and the only other occurrences of "MAC" in man/ mean hardware addresses.

Must fix

  • 55. Adopting item 49's mechanism has overshot. "No audit record of any kind is produced, so the loss is invisible to ausearch or audit2allow" is contradicted by the reproduction posted in round 9, which showed a second, independently-timestamped avc: denied { append } scontext=httpd_t tclass=file from the service's own domain; "Reproduced directly: ... leaves the target path empty" cannot have come from that run, which was explicitly permissive, since permissive logs but does not block; and the null-device redirection is presented as unconditional behaviour of this feature although open_null_as() is reachable only from EXEC_INPUT_NULL (362), the maybe_inherit_stdout_from_stdin() fallback (500), EXEC_OUTPUT_NULL (573) and the journal-connect fallback (600), with file:/append:/truncate: failing the unit via EXIT_STDOUT/EXIT_STDERR instead. If the SELinux LSM's execve()-time descriptor replacement is meant, it needs attributing to SELinux, and "the subsequent write succeeds and the unit does not fail" needs scoping to enforcing mode.

Suggestions

  • 56. "re-checked once more at the point the service's own domain takes over the already-open descriptor" understates the check: read/write/append are file-class permissions revalidated against the current subject on every subsequent read or write, not once at a handover, and there is no systemd-side handover event, since sym_setexeccon_raw() (exec-invoke.c:6471) only arms the context for the following execve().

Nits

  • 57. A newline separates StandardInput= from file:path at 3541-3542, which DocBook collapses to a space, so it renders as "StandardInput= file:path" while the same idiom at 3566 is unspaced - the newline counterpart of item 44, whose same-line instance is fixed in this head.

Must fix

  • 58. The note distributes all three option values over all three settings ("for file:path, append:path and truncate:path of StandardOutput=, StandardError= or StandardInput=") at 3520-3523, but StandardInput= accepts neither append: nor truncate:. config_parse_exec_input() (load-fragment.c:1136-1202) special-cases only fd: and file: and otherwise falls through to exec_input_from_string(); exec_input_table (execute.c:3149-3157) has no append/truncate entries, which exist only in exec_output_table (3162-3174). The page's own value list for StandardInput= at 3392-3395 offers file:path alone, so the sentence contradicts it and presents StandardInput=truncate:/x as supported when the parser rejects it - and the new cross-reference at 3423-3425 sends StandardInput= readers straight to it.
  • 59. "The resulting descriptor is then inherited across that transition into the unit's own domain" at 3525-3526 is unconditional, and it fails in exactly the case the note exists to warn about. acquire_path() (exec-invoke.c:271) opens without O_CLOEXEC and move_fd(..., cloexec=false) installs it (433, 653), so the fd does survive to fexecve_or_execve(); but on an SELinux domain change the kernel revalidates each inherited descriptor at credential commit via flush_unauthorized_files() - fd { use } towards the opener's SID plus the file permission implied by the open mode - and on denial in enforcing mode substitutes the SELinux null device or closes it. Even where inheritance succeeds, read/write/append are revalidated against the current subject on later operations rather than cached from the opener, which is what the author's own two-denial reproduction shows.

Suggestions

  • 60. "before SELinuxContext= or any policy-defined transition takes effect" plus the definite "across that transition" at 3525-3526 presupposes a transition always occurs. Two common configurations have none: no SELinuxContext= with no entrypoint type_transition, so the manager's domain is kept; and ExecStart=+.../!! commands carrying EXEC_COMMAND_FULLY_PRIVILEGED, where needs_sandboxing is false (exec-invoke.c:5496) so the MAC block at 6460 and sym_setexeccon_raw() at 6471 are skipped entirely. In both the opening and running subjects are the same domain.

Workflow run

Comment thread man/systemd.exec.xml Outdated
Comment on lines +3518 to +3520
<para>Note that for <option>file:<replaceable>path</replaceable></option>,
<option>append:<replaceable>path</replaceable></option> and
<option>truncate:<replaceable>path</replaceable></option> the file is opened while the process is

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: suggestion: This constraint is not specific to the three StandardOutput= modes. StandardInput=file: goes through the very same acquire_path() helper — setup_input() is called at src/core/exec-invoke.c:5563, i.e. even earlier than the two setup_output() calls at 5576/5582 — so it is opened in the manager's context too, just needing read access (or read+write, when stdin and stdout name the same path and the fd is opened O_RDWR and duplicated). OpenFile= fds are collected earlier still, at collect_open_file_fds() (line 5287). An administrator who follows this paragraph will fix the labelling for the stdout path and then hit the same denial via StandardInput=file: with no documentation explaining it, since a reader of the StandardInput= section further up the page gets no hint of this.

Consider phrasing the note in terms of all path-based stdio settings, and/or adding a short cross-reference under StandardInput=file:. (StandardError= is fine as-is, since its option list is defined by reference to StandardOutput=.)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed. The note now names StandardInput=file: explicitly, and the StandardInput=file: paragraph itself got a one-sentence cross-reference pointing at it, so a reader of that section is no longer left without a hint.

The access wording was generalised accordingly: it now says "reading for StandardInput=, appending rather than writing for append:", instead of the previous write-centric phrasing.

I left OpenFile= out on purpose: it is documented in systemd.service(5), not on this page, so a note here would have nowhere to attach. If you think it warrants the same treatment, that is probably a separate patch against that page.

Comment thread man/systemd.exec.xml Outdated
Comment on lines +3523 to +3524
the file is thus made against the security context inherited from the service manager, not the one
the service ultimately runs in. On systems using SELinux this means the file must be labelled so

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: suggestion: "not the one the service ultimately runs in" is stated as an absolute, but the two contexts are often identical: if the policy defines no domain transition for the unit's executable and no SELinuxContext= is set, the service keeps the manager's domain, so labelling only for "the domain the service transitions into" is sufficient. The same applies to ExecStart=+... commands, where needs_sandboxing is false (src/core/exec-invoke.c:5496) and the sym_setexeccon_raw() block at 6460ff is skipped entirely.

Wording such as "which may differ from the one the service ultimately runs in" would be accurate in all cases. It would also be worth stating explicitly that setting SELinuxContext= does not help here, since setexeccon() (line 6471) runs long after setup_output() (5576/5582) — that is precisely the assumption a reader is likely to try next.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adopted the "may differ" wording, and also added the SELinuxContext= note you suggested — that is indeed the next thing a reader tries. Verified rather than assumed: with SELinuxContext=system_u:system_r:unconfined_service_t:s0 set on a unit whose StandardOutput=append: target is labelled httpd_sys_content_t, the denial is unchanged and still recorded against the manager:

avc: denied { open }   comm="(probe.sh)" exe="/usr/lib/systemd/systemd-executor"
     subj=system_u:system_r:init_t:s0 name="ctxtest.log" tclass=file
avc: denied { append } comm="(probe.sh)" exe="/usr/lib/systemd/systemd-executor"
     subj=system_u:system_r:init_t:s0 name="ctxtest.log" tclass=file

One part I did not take, because measurement disagrees with it: the claim that the contexts are also identical for ExecStart=+…. Skipping the sym_setexeccon_raw() block only drops the explicit override; the automatic type_transition in policy is applied by the kernel at execve() regardless of whether userspace called setexeccon(). Two otherwise identical oneshot units on CentOS Stream 10 (targeted policy), the payload printing its own /proc/self/attr/current:

NORMAL ctx=system_u:system_r:unconfined_service_t:s0     # ExecStart=/…/probe.sh
PLUS   ctx=system_u:system_r:unconfined_service_t:s0     # ExecStart=+/…/probe.sh

So a + command does not stay in init_t, and the manager's context still differs from the service's there. The case where the two genuinely coincide is the one you named first — no transition defined for the executable and no SELinuxContext= — which the "may differ" wording now covers without enumerating exceptions.

Comment thread man/systemd.exec.xml Outdated
Comment on lines +3524 to +3525
the service ultimately runs in. On systems using SELinux this means the file must be labelled so
that the manager's domain (<constant>init_t</constant> for the system service manager) may open it

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: suggestion: The init_t parenthetical is stated as bare fact, but it is policy- and instance-dependent in two ways.

First, init_t is the type the Fedora/RHEL targeted (reference) policy assigns to the manager; nothing in SELinux guarantees it, and this page otherwise avoids naming concrete policy types. Since v254 the open is also not performed by PID 1 but by the posix_spawn()ed pinned systemd-executor child (src/core/execute.c:580), which stays in the manager's domain only because no transition is currently defined for it.

Second, StandardOutput=file: applies to user units too, where systemd --user runs in the invoking user's domain, so a reader debugging a user unit would search for init_t denials and find none.

Hedged, instance-neutral wording would fix both, e.g. "the domain of the respective service manager instance (typically init_t for the system service manager on common policies)".

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adopted, essentially in your wording: the text now reads "the domain of the respective service manager instance — typically init_t for the system service manager on common policies, and a different domain for the per-user one".

Both of your points check out on a live system here (CentOS Stream 10, systemd 257-33.el10, targeted policy):

  • The open really is performed by the executor child, and it is in the manager's domain — the audit record carries exe="/usr/lib/systemd/systemd-executor", ppid=1, subj=system_u:system_r:init_t:s0. /usr/lib/systemd/systemd-executor is labelled init_exec_t here, i.e. the same exec type as PID 1, which is why no transition happens for it.
  • The per-user manager is indeed a different domain: systemd --user for an unconfined user runs as unconfined_u:unconfined_r:unconfined_t:s0-s0:c0.c1023, so someone debugging a user unit would grep for init_t denials and find nothing.

I kept the executor out of the man page text: naming an internal helper process would date the documentation, whereas "the respective service manager instance" stays correct regardless of how the manager forks internally.

Comment thread man/systemd.exec.xml Outdated
Control) domain transition associated with the execution takes place. The access check for opening
the file is thus made against the security context inherited from the service manager, not the one
the service ultimately runs in. On systems using SELinux this means the file must be labelled so
that the manager's domain (<constant>init_t</constant> for the system service manager) may open it

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: suggestion: <constant>init_t</constant> will be pulled into the generated systemd.directives(7) index: tools/make-directive-index.py iterates .//constant over every man page and only skips elements carrying index='false' (lines 86-87). A distro-policy SELinux type would then be listed as a systemd constant next to CAP_SYS_ADMIN/EXIT_SELINUX_CONTEXT.

Also, <constant> in this page is reserved for C-level constants (AF_UNIX, EPERM, CAP_*, EXIT_*), whereas the existing precedent for SELinux type names is <literal> — see man/systemd.network.xml:1619 (<literal>my_server_t</literal>).

Consider <literal>init_t</literal>, or add index='false' if <constant> is kept, as done at lines 888-893.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adopted — it is now <literal>init_t</literal>. Confirmed both halves of your reasoning before changing it: tools/make-directive-index.py iterates .//constant over every page, so the type would have been indexed into systemd.directives(7) alongside CAP_SYS_ADMIN; and man/systemd.network.xml:1619 is the existing precedent for an SELinux type name in <literal> (my_server_t).

I went with <literal> rather than index='false' since it matches that precedent and keeps <constant> on this page reserved for C-level constants. <constant>AF_UNIX</constant> in the same paragraph is unchanged for that reason.

Comment thread man/systemd.exec.xml Outdated
the file is thus made against the security context inherited from the service manager, not the one
the service ultimately runs in. On systems using SELinux this means the file must be labelled so
that the manager's domain (<constant>init_t</constant> for the system service manager) may open it
for writing; labelling it only for the domain the service transitions into is not sufficient.</para>

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: suggestion: "the file must be labelled so that the manager's domain ... may open it for writing" understates what SELinux policy has to allow, in ways that bite in exactly the scenario being documented:

  • For append:, the open uses O_APPEND and therefore requires the distinct append permission on class file, not write; open is a separate permission as well. The AVCs quoted in the PR description show { open } and { append } denied separately, so the current wording would not be enough to fix that case.
  • acquire_path() adds O_CREAT for write modes (src/core/exec-invoke.c:278-280), so when the target does not exist yet the manager's domain also needs create on the file and add_name/write on the parent directory. Nothing in this path calls mac_selinux_create_file_prepare(), so the new file's type comes from the type transition on the parent directory.
  • The target need not be a regular file: acquire_path() falls back to socket()+connect() for AF_UNIX sockets, and the StandardInput=file: text explicitly allows a FIFO or special file — those need different classes and permissions (fifo_file, chr_file, unix_stream_socket connectto).

Consider wording it as "the manager's security context must be permitted the relevant access to the target object", with append: called out as needing append rather than write, instead of naming open+write on a file.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adopted. The text now says the manager's domain must be "permitted the access in question: reading for StandardInput=, appending rather than writing for append:, and creating the target, along with the corresponding access to the enclosing directory, if it does not exist yet", plus a closing sentence that a FIFO, special file or AF_UNIX socket target is checked against the corresponding object class.

Each of those is visible in one run here — a unit with StandardOutput=append: at an existing httpd_sys_content_t file, and a second one with file: at a not-yet-existing target in the same directory:

denied { open }       name=existing.log       tclass=file  subj=init_t
denied { append }     name=existing.log       tclass=file  subj=init_t
denied { create }     name=tobecreated.log    tclass=file  subj=init_t
denied { write open } name=tobecreated.log    tclass=file  subj=init_t

open and append are denied separately, exactly as you said, and the create path shows up on its own. No directory-class denial appears in this particular instance only because the policy already allows it — allow named_filetrans_domain httpd_sys_content_t:dir { add_name … write … } — which is why the text speaks of the required access rather than promising a denial.

Comment thread man/systemd.exec.xml Outdated
Comment on lines +3521 to +3522
still being set up, i.e. before the executable is invoked and hence before any MAC (Mandatory Access
Control) domain transition associated with the execution takes place. The access check for opening

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: nit: Two small things about this clause.

The acronym introduction is inverted relative to the rest of man/, which spells the term out first and puts the abbreviation in parentheses — e.g. "Unified Kernel Image (UKI)", "Name Service Switch (NSS)", "Discoverable Disk Image (DDI)"; there are no instances of the ACRONYM (Expansion) form. Write "Mandatory Access Control (MAC)". This page also already has a section titled "Mandatory Access Control" (line 972), so the expansion could simply be dropped in favour of the plain term.

Also, "domain transition associated with the execution" describes SELinux and AppArmor accurately (both take effect at execve()) but not SMACK: setup_smack() applies the label to the already-running process via mac_smack_apply_pid(0, ...), with no transition at exec. The conclusion still holds because the call site is src/core/exec-invoke.c:6342, well after setup_output() — but the stated mechanism is wrong for that backend. Phrasing it as "before the MAC label/domain of the service process is applied" would cover all supported backends.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both taken.

On the acronym: you are right that the page has no need for it at all — "Mandatory Access Control" already names a section here (line 972), so the expansion-plus-abbreviation is simply dropped and the term is spelled out in place. (Checking the convention first: man/ has "Unified Kernel Image (UKI)", "Name Service Switch (NSS)" and "Discoverable Disk Image (DDI)", and no instance of the reverse form.)

On SMACK: the clause now reads "before the Mandatory Access Control label of the service process is applied", which holds for all three backends. Your reading of the mechanism matches the code — setup_smack() calls mac_smack_apply_pid(0, …), i.e. it relabels the already-running process rather than arranging a transition at execve(), and it does so at src/core/exec-invoke.c:6342, well after the setup_output() calls at 5576/5582, so the conclusion is unaffected and only the stated mechanism needed fixing.

@BenkiNew
BenkiNew force-pushed the man/stdout-file-mac-context branch from fa9d387 to 89a0662 Compare September 4, 2026 13:45
@BenkiNew BenkiNew changed the title man: document that StandardOutput= file targets are opened before the MAC transition man: document that stdio file targets are opened before the MAC transition Sep 4, 2026
@BenkiNew

BenkiNew commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

Force-pushed an update addressing the review. The note now also covers StandardInput=file:, so the commit subject and PR title were widened accordingly.

  • StandardInput=file: opens its path through the same acquire_path() helper (call site at line 429, ahead of the setup_output() one at 649), so it is subject to the identical pre-transition check and is now named explicitly in the note.
  • Dropped the absolute claim that the two contexts differ; it now reads "unless the two are the same".
  • init_t is no longer stated as a universal fact: it is now "typically init_t under the targeted policy for the system service manager", with the per-user service manager called out as a different domain.
  • Spelled out what the policy actually has to allow: open, append in the append: case, and create — the latter also requiring access to the enclosing directory, since acquire_path() adds O_CREAT for write modes — plus a note that a FIFO, special file or AF_UNIX socket target is checked against the corresponding object class.
  • <constant>init_t</constant> → <literal>: tools/make-directive-index.py collects every <constant> element into systemd.directives(7), and a distribution policy type does not belong in that index.
  • Reworded "MAC domain transition associated with the execution" to "before the MAC context of the process is adjusted for the service", which also holds for SMACK, where setup_smack() applies the label to the process itself rather than at execve().

One point I did not take: the suggestion that the two contexts are identical for ExecStart=+… because needs_sandboxing is false and the setexeccon() block is skipped. Skipping that block only drops an explicit SELinuxContext= override — an automatic type_transition in policy still applies at execve() for such commands. The wording is therefore generic rather than calling out + as an exception.

@BenkiNew
BenkiNew force-pushed the man/stdout-file-mac-context branch from 89a0662 to 3ebd2bd Compare September 4, 2026 14:03
@BenkiNew

BenkiNew commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

Pushed again (3ebd2bd) after going through the inline comments one by one; each thread has a reply with what changed and, where a claim was checkable, the measurement it rests on. Beyond the previous round:

  • StandardInput=file: now carries a one-sentence cross-reference to the note, so that section is no longer a dead end for a reader hitting the same denial.
  • The note states that SELinuxContext= does not help, verified on a unit that sets it: the denials are unchanged and still recorded against the manager's context.
  • The note is split in two paragraphs — mechanism first, SELinux consequences second — and the required access is enumerated (read for StandardInput=, append rather than write for append:, create plus enclosing-directory access for a missing target), with a closing sentence for FIFO/special-file/AF_UNIX targets.
  • <literal>init_t</literal>, instance-neutral manager wording, and the SMACK-safe phrasing "before the Mandatory Access Control label of the service process is applied".

One suggestion was not taken — that the contexts also coincide for ExecStart=+… — with the reasoning and the measurement in that thread: skipping sym_setexeccon_raw() drops only the explicit override, while the kernel still applies the automatic type_transition at execve(), so a + command does not stay in the manager's domain.

@BenkiNew
BenkiNew force-pushed the man/stdout-file-mac-context branch from 3ebd2bd to abdb95e Compare September 4, 2026 14:22
@BenkiNew

BenkiNew commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

Pushed abdb95e, which addresses 7–9. Thanks for catching 7 — that was a real regression introduced by my own rewrite, not a pre-existing gap.

7. Correct: setup_output() opens all three modes O_WRONLY (plus O_TRUNC/O_APPEND), so write was missing from the enumeration and a policy following the old wording would still have denied file: and truncate:. Reproduced against an existing httpd_sys_content_t target:

avc: denied { write } comm="(probe.sh)" name="existing.log" tclass=file
     scontext=system_u:system_r:init_t:s0 tcontext=…:httpd_sys_content_t:s0

The enumeration now reads: open in all cases, read for StandardInput=, write for file: and truncate:, and append instead of write for append:.

One half of that finding is too absolute, though: StandardInput=file: is not always O_RDONLY. setup_input() computes rw from whether std_output/std_error is EXEC_OUTPUT_FILE at the same path and then opens O_RDWR (line 429), which does pick up O_CREAT in acquire_path(). A unit with StandardInput=file:X and StandardOutput=file:X where X does not exist creates it:

avc: denied { read write open } name="/…/shared.log" tclass=file  scontext=…:init_t:s0
avc: denied { create }          name="shared.log"    tclass=file  scontext=…:init_t:s0
# afterwards: -rw-r--r--. root root system_u:object_r:httpd_sys_content_t:s0 shared.log

So the creation paragraph is scoped to "the three StandardOutput= modes, and to StandardInput= only if standard input and output refer to the same path" rather than excluding StandardInput= outright. This is the same case your first-round comment noted in passing.

8. Confirmed and adopted. The same run shows it directly: matchpathcon reports <<none>> for the created path, yet the file came out as httpd_sys_content_t — inherited from the enclosing directory's type transition. The note now says the new file is labelled by the directory's type transition rather than by any file context defined for the path itself.

9. Confirmed and adopted. The open() attempt is checked against sock_file before it fails with ENXIO:

avc: denied { write } comm="(probe.sh)" name="listen.sock" tclass=sock_file
     scontext=system_u:system_r:init_t:s0 tcontext=…:httpd_sys_content_t:s0

The text no longer says "instead": FIFOs and special files are checked against the corresponding class, while for an AF_UNIX socket the checks are cumulative — the sock_file open first, then connectto against the domain of the listening peer.

All measurements on CentOS Stream 10, systemd 257-33.el10, selinux-policy-targeted 42.1.27-2, permissive so that every check is logged rather than short-circuiting at the first denial.

Comment thread man/systemd.exec.xml Outdated
Comment on lines +3543 to +3545
<varname>StandardOutput=</varname> modes, and to <varname>StandardInput=</varname> only if standard
input and output refer to the same path, in which case a single read-write file descriptor is used
for both. Note that the new file is labelled according to the type transition of the enclosing

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: must-fix: The shared read-write descriptor only happens when the output mode is file:, not append: or truncate:, so this sentence is inaccurate as written.

In setup_input() (src/core/exec-invoke.c:426-427) the rw flag is computed exclusively from context->std_output == EXEC_OUTPUT_FILE / context->std_error == EXEC_OUTPUT_FILE; EXEC_OUTPUT_FILE_APPEND and EXEC_OUTPUT_FILE_TRUNCATE are deliberately not included. Hence with e.g. StandardInput=file:/x plus StandardOutput=append:/x the stdin fd is opened O_RDONLY (line 429), and since acquire_path() only adds O_CREAT for O_WRONLY/O_RDWR (src/core/exec-invoke.c:278-279), nothing is created in that combination. setup_output() then still takes the sharing shortcut for all three file output modes unconditionally (src/core/exec-invoke.c:630-641), duplicating that read-only descriptor onto stdout, so the append/truncate flags never take effect either.

So for the same-path case the manager needs only open+read, and no create at all — a reader following this paragraph would grant create/write/append that are never exercised, and would expect a nonexistent target to be created when it isn't. Please scope the sentence to file: specifically, matching the pre-existing wording at line 3496 ("opened only once — for reading as well as writing — and duplicated"), which is already limited to file:.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Replying here too, to match the per-thread pattern used for the earlier findings — I answered this one via a general PR comment instead of this thread when I fixed it, which broke that pattern.

Fixed in 1147c0a. Confirmed and reproduced exactly as described: StandardInput=file:/x + StandardOutput=append:/x with x absent opens stdin O_RDONLY (no O_CREAT), and setup_output()'s sharing shortcut still dup2()s that fd onto stdout regardless of output mode:

$ systemctl start mp-appshared.service   # append:, shared path, x absent
(probe.sh): Failed to set up standard input: No such file or directory
Failed at step STDIN spawning …: No such file or directory   [208/STDIN]

$ systemctl start mp-truncshared.service # truncate:, same result
Result=exit-code, ExecMainStatus=208, target not created

The paragraph now scopes the shared-descriptor/creation case to file:/file: explicitly and names the append:/truncate: combination as failing at startup instead.

Comment thread man/systemd.exec.xml Outdated
Comment on lines +3548 to +3552
<para>If <replaceable>path</replaceable> refers to a FIFO or a special file, the checks are made
against the corresponding object class. For an <constant>AF_UNIX</constant> socket in the file
system they are cumulative: the target is opened first, and that attempt is checked against class
<literal>sock_file</literal> before it fails, after which connecting to the socket additionally
requires <literal>connectto</literal> against the domain of the listening peer.</para>

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: suggestion: The AF_UNIX enumeration is incomplete in a way that leaves the denial unfixed, and the object classes for the non-socket cases are worth naming.

  1. write on sock_file is missing. After open() fails with ENXIO (src/core/exec-invoke.c:285-286), acquire_path() calls connect_unix_path() (line 294). The kernel's address lookup for a filesystem AF_UNIX socket does a MAY_WRITE permission check on the socket inode (unix_find_other() -> path_permission(&path, MAY_WRITE)), which SELinux maps to write on class sock_file. This is also why refpolicy's stream_connect_pattern grants write_sock_file_perms and not just read/open. The preceding paragraph ties write to StandardOutput=file:/truncate: and gives StandardInput= only read, so an administrator wiring StandardInput=file:/run/some.sock per this text grants sock_file { open read } + connectto and is still denied. Please state that the socket case additionally needs write on sock_file regardless of the direction.

  2. "checked against class sock_file before it fails" reads as though the check were inconsequential. It is a hard requirement: acquire_path() only falls through to the socket path when errno == ENXIO and otherwise does return -errno (src/core/exec-invoke.c:285-286), and the kernel runs the SELinux open/read checks before f_op->open yields ENXIO. So a denial on the initial open surfaces as EACCES and aborts the whole operation — the connect is never attempted.

  3. connectto is a permission of class unix_stream_socket, and it is checked against the label of the listening socket (selinux_socket_unix_stream_connect() uses the target socket's SID and class), which is normally inherited from the process that created it. Naming the class would match the explicit mention of sock_file in the same sentence and is more precise than "the domain of the listening peer".

  4. "the corresponding object class" is worth spelling out — fifo_file for FIFOs, chr_file or blk_file for device nodes — since that is exactly what goes into an allow rule, and the paragraph already names sock_file for the socket case.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Replying here too (same note as the previous thread — this went out as a general comment first, breaking the per-thread pattern).

Adopted the write point — I had the evidence for it already (round 3's denied { write } … tclass=sock_file on the labelled socket target) but hadn't stated the permission. The paragraph now reads "checked against write on class sock_file".

Not adopted, and I'd like to explain why rather than drop silently:

  • Naming fifo_file/chr_file/blk_file explicitly: this note is about the pre-transition SELinux check on the manager side, and StandardInput=/StandardOutput= targeting a FIFO or device node is a narrow, largely undocumented-elsewhere corner of this already-long note. sock_file gets named because the socket path is structurally different (open-then-connect, additive classes) and needs explaining either way; the other two don't change the shape of the check, just its class name.
  • connectto on unix_stream_socket specifically: agreed on the fact, but the current sentence ("against the domain of the listening peer") already tells a reader what to grant it to, which is the actionable part; the class name is derivable from context (it's the class of the thing being connected to).

I don't think either omission is inaccurate, just less exhaustive than it could be — happy to add both if you'd still rather have them spelled out.

Comment thread man/systemd.exec.xml Outdated
Comment on lines +3541 to +3542
<para>A target that does not exist yet is created, which requires <literal>create</literal> and the
corresponding access to the enclosing directory as well. This applies to the three

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: suggestion: Two things about directory access here.

First, "the corresponding access to the enclosing directory" is vague where the rest of the note names permissions explicitly. Creating the target through open(path, ...|O_CREAT, mode) in acquire_path() (src/core/exec-invoke.c:278-281) requires create on the target's class plus write and add_name on the enclosing directory (class dir); naming those makes the paragraph directly usable when writing a policy module.

Second, directory traversal is needed unconditionally, not just when the target has to be created: the manager-side open() resolves the full absolute path, so the manager's domain needs search on class dir for every path component regardless of whether the target already exists. As currently structured, directory access is only mentioned in this paragraph and only as a consequence of create, so someone whose target sits under a custom-typed directory can relabel the file exactly as the preceding paragraph instructs and still get denied { search } tclass=dir against the manager's domain. Consider stating the unconditional search requirement alongside the target-object permissions in the previous paragraph, and keeping the create-specific create/write/add_name set here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Replying here too (posted as a general comment first, breaking the per-thread pattern used elsewhere — apologies, fixing that now across this and the next thread).

Partially adopted, and reconsidered on rereading your full text — the checklist summary I'd originally seen only carried "worth naming as write and add_name alongside create", which I folded into a one-line "declined, too broad" without weighing the two parts separately. They're not the same claim:

  • Naming create on the target plus write and add_name on the enclosing directory: adopted. This is precise and directly follows from acquire_path()'s open(path, flags|O_CREAT, mode) — no separate verification needed beyond what's already cited for finding build-sys: Normalize paths of configure options #10.
  • The unconditional search requirement on every path component: still not added, now with the actual reasoning instead of a one-liner. This page documents roughly a dozen other settings that consume an absolute path the manager or executor resolves before or during exec (WorkingDirectory=, RootDirectory=, ReadWritePaths=, the bind-mount options, RuntimeDirectory=, …), and none of them state the search-on-every-component requirement either — it's a property of how SELinux checks path resolution in general, not of StandardOutput= specifically. Adding it only here would read as though it were a peculiarity of this option. If you think the page should state it once, generically, that reads to me like a good candidate for its own small addition near the "Mandatory Access Control" section (line 972) rather than folded into this note — happy to be told otherwise.

Comment thread man/systemd.exec.xml Outdated
Comment on lines +3545 to +3546
for both. Note that the new file is labelled according to the type transition of the enclosing
directory rather than any file context defined for the path itself.</para>

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: suggestion: This is right in substance — acquire_path() calls plain open() and nothing on that path calls mac_selinux_create_file_prepare()/mac_selinux_create_file_prepare_at(), unlike e.g. src/core/socket.c:1212, src/shared/label-util.c:70 or src/shared/mkdir.c:16, so the file_contexts entry for the path is never consulted — but describing the resulting label is both incomplete and more fragile than describing what systemd does.

It is incomplete because a type_transition only applies if the loaded policy has a matching rule for (manager domain, enclosing directory's type : file). With no such rule, security_compute_sid() falls back to the type of the related object, i.e. the new file simply gets the enclosing directory's own type — which is the common case on most policies, so as written the sentence presents the rarer case as the rule.

Suggest leading with the systemd-side fact and mentioning the kernel default, e.g.: "Note that the service manager does not apply a file creation context when opening the target, so the newly created file is not labelled according to any file context defined for the path itself; it gets whatever label the policy's type transition for the creating domain and the enclosing directory's type yields, or the enclosing directory's own type if there is no such transition."

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Replying here too (same apology as the sibling thread — this went out as a general comment, not inline).

Reconsidered and adopted, on rereading the full text rather than the abbreviated checklist line. I'd declined this in the general comment on the reasoning that "type transition of the enclosing directory" already covers the no-matching-rule fallback as a degenerate case — technically defensible, but your version is more precise, and the contrast you drew is the part that changed my mind: I checked it, and it holds exactly as described.

$ grep -rn mac_selinux_create_file_prepare src/core/exec-invoke.c
(no output — not called on this path)

$ grep -rn mac_selinux_create_file_prepare src/core/socket.c src/shared/label-util.c src/shared/mkdir.c
src/core/socket.c:1212:        r = mac_selinux_create_file_prepare(path, S_IFIFO, …);
src/shared/label-util.c:51:     r = mac_selinux_create_file_prepare(new_path, S_IFLNK, …);
src/shared/label-util.c:70:     r = mac_selinux_create_file_prepare_at(dirfd, pathname, mode, …);
src/shared/label-util.c:88:     r = mac_selinux_create_file_prepare(path, S_IFDIR, …);
src/shared/mkdir.c:16:          r = mac_selinux_create_file_prepare_at(dirfd, path, S_IFDIR, …);

That's a real, checkable asymmetry — other creation paths in this codebase do consult file_contexts for the path, this one doesn't — and leading with it is a fact about systemd's own behaviour rather than a description of a kernel fallback that depends on what policy happens to be loaded. Adopted close to your wording.

@BenkiNew
BenkiNew force-pushed the man/stdout-file-mac-context branch from abdb95e to 1147c0a Compare September 4, 2026 15:00
@BenkiNew

BenkiNew commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

Pushed 1147c0a. #10 was another regression from my own round-3 rewrite — thank you for catching it, same class of mistake as #7.

10. Confirmed in code and reproduced. setup_input()'s rw is true only when context->std_output == EXEC_OUTPUT_FILE (exec-invoke.c:426-427) — it does not check EXEC_OUTPUT_FILE_APPEND/EXEC_OUTPUT_FILE_TRUNCATE. setup_output()'s dup2(STDIN_FILENO, fileno) guard, however, fires for all three output modes whenever std_input == EXEC_INPUT_FILE and the paths match — so pairing StandardInput=file: with append:/truncate: at a not-yet-existing shared path does not fall back to opening its own fd; it duplicates the read-only one setup_input() already opened, and there's nothing to create. Reproduced on both:

$ systemctl start mp-appshared.service    # StandardInput=file:/…/x  StandardOutput=append:/…/x, x absent
(probe.sh): Failed to set up standard input: No such file or directory
mp-appshared.service: Failed at step STDIN spawning …: No such file or directory
Main process exited, code=exited, status=208/STDIN

$ systemctl start mp-truncshared.service  # same with truncate:, identical result
Result=exit-code, ExecMainStatus=208, target not created

So the paragraph now says the shared-descriptor path only exists for file:/file:, and names the append:/truncate: combination explicitly as "does not create it, and the unit instead fails at startup" rather than silently omitting it.

12. Adopted — I had the empirical evidence for this from the previous round already (denied { write } … tclass=sock_file when opening the labelled AF_UNIX target) but hadn't named the permission in the prose. It now reads "checked against write on class sock_file".

11 and 13 I'm leaving as-is, and want to say why rather than silently drop them:

  • 11 is correct that search on every directory component is needed for path resolution in general — but that is true of every path-based systemd.exec(5) setting (WorkingDirectory=, ReadWritePaths=, the sandboxing options, …), not something specific to this note, and the page doesn't currently spell out directory-traversal SELinux semantics anywhere else either. Scoping it to just this paragraph would make it look like a StandardOutput=-specific requirement when it isn't one.
  • 13 reads as accurate to me as currently worded: "labelled according to the type transition of the enclosing directory" already covers the no-matching-rule case, since security_compute_sid() falling back to the directory's own type is itself the (degenerate) type-transition computation, not a separate mechanism. I don't think the suggested reordering changes what's asserted, so I'd rather not touch a sentence that's already been rewritten twice in this thread for a phrasing preference with no behavioural difference behind it.

Same test environment as before: CentOS Stream 10, systemd 257-33.el10, selinux-policy-targeted 42.1.27-2, permissive.

@BenkiNew
BenkiNew force-pushed the man/stdout-file-mac-context branch from 1147c0a to f009241 Compare September 4, 2026 15:43
Comment thread man/systemd.exec.xml Outdated
Comment on lines +3556 to +3558
against the corresponding object class. For an <constant>AF_UNIX</constant> socket in the file
system they are cumulative: the target is opened first, and that attempt is checked against
<literal>write</literal> on class <literal>sock_file</literal> before it fails, after which

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: must-fix: The write on class sock_file is attached to the wrong operation, and the result contradicts the paragraph above.

For StandardInput=file: pointing at a socket, the open is attempted O_RDONLY (src/core/exec-invoke.c:429 — rw at 426-427 is only true when stdout or stderr name the same path with file:), so the failing open is checked against read on sock_file, not write. As written this contradicts the enumeration at lines 3534-3535 ("read for StandardInput="). Which permission the open needs also matters, because acquire_path() falls through to the socket path only on exactly ENXIO (src/core/exec-invoke.c:285-286): if policy denies what the open requires, open() returns EACCES, that error is propagated, and connect() is never reached at all.

write on sock_file is genuinely required, but it comes from the connect_unix_path() call (src/core/exec-invoke.c:294): the kernel's unix_find_other() does a MAY_WRITE inode_permission() on the socket inode, which SELinux checks as write on sock_file — this is why refpolicy's stream_connect_pattern grants write_sock_file_perms. Hanging it off the pre-ENXIO open tells policy authors the wrong thing for StandardInput= and hides that write is needed for the connect in both directions.

Suggest: the failing open is checked against the permission the direction requires (read/write/append, per the paragraph above), and the subsequent connect additionally requires write on sock_file plus connectto against the listening peer's domain.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed and fixed in 67cdef5. acquire_path()'s if (errno != ENXIO) return -errno; (exec-invoke.c:285-286) means any denial on the initial open() — checked against whatever direction the caller requested, O_RDONLY for StandardInput= — aborts before connect_unix_path() is ever reached, so "write on sock_file" was simply the wrong permission name for that direction.

Rather than add a third clause distinguishing "the direction-matched permission on the initial open" from "write specifically at connect time", which is the kind of precision that produced this bug in the first place, I reworded to name what's actually invariant: the initial open is checked against the same access already named for that direction two paragraphs up, and only the connect step is unconditionally write. No new claim beyond what's proven here and in the sibling threads.

I'm going to stop iterating on this note beyond this fix. The remaining open items (11, 13 partially, and the new 15-21) are accurate observations, but between DAC/capability ordering, namespace/root ordering, symlink handling, peer-credential semantics and document structure, they're pulling a documentation note substantially past the scope of the SELinux constraint it was written to describe. I've replied to each below with my reasoning for not chasing it further; happy to revisit any specific one if a maintainer weighs in and thinks otherwise.

Comment thread man/systemd.exec.xml Outdated
Comment on lines +3524 to +3529
path is opened while the process is still being set up, i.e. before the executable is invoked and
before the Mandatory Access Control label of the service process is applied. The access check is
thus made against the security context inherited from the service manager, which may differ from
the one the service ultimately runs in. Setting <varname>SELinuxContext=</varname> makes no
difference here, as it only takes effect when the executable is invoked, i.e. after the path has
already been opened.</para>

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: suggestion: The note frames the consequence of this ordering as MAC-only ("before the Mandatory Access Control label of the service process is applied", "the security context inherited from the service manager"), but the same ordering means the DAC side of the open is done with the manager's credentials too. setup_input()/setup_output() run at src/core/exec-invoke.c:5563/5576/5582, while enforce_user() (the setresuid() at 1036) is only reached at 6410 and capability_bounding_set_drop() at 6376. So for the system manager the open() in acquire_path() (272-282, including the implicit O_CREAT and the O_TRUNC) happens as uid 0 with full capabilities regardless of User=, Group= or CapabilityBoundingSet=: StandardOutput=truncate:/some/root-only/file in a unit running User=nobody succeeds and truncates a file the service itself could never open.

The existing text at line 3494 mentions the manager's privileges only for creation; since this paragraph exists to correct exactly that framing, it would help to state that the access check for an existing target is likewise made with the manager's uid and capabilities, rather than leaving "security context" to be read as an SELinux label only.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Replying here too, per the per-thread convention (full context in the round-6 summary comment above). Accurate: setup_input()/setup_output() at 5563/5576/5582 do precede enforce_user() (6410) and capability_bounding_set_drop() (6376), so the DAC side of the open runs under the manager's credentials too, not just MAC. Not folding this in — this note is scoped to the one narrow SELinux/MAC ordering gap it documents; DAC-credential ordering is a real but separate property of exec_invoke(), better suited to the broader follow-up mentioned in the round-6 summary than to widening this note further.

Comment thread man/systemd.exec.xml Outdated
thus made against the security context inherited from the service manager, which may differ from
the one the service ultimately runs in. Setting <varname>SELinuxContext=</varname> makes no
difference here, as it only takes effect when the executable is invoked, i.e. after the path has
already been opened.</para>

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: suggestion: Two path-resolution consequences of the ordering this paragraph documents are worth mentioning, since a reader who reaches this note is reasoning about what confines the open.

  1. The open precedes all namespacing and root switching: setup_input()/setup_output() are at src/core/exec-invoke.c:5563/5576/5582, whereas setup_delegated_namespaces() (which calls apply_mount_namespace() at 4836) is first called at 6094 and apply_root_directory() at 6404. The path is therefore resolved against the host root and is not subject to RootDirectory=/RootImage=, ProtectSystem=, ReadOnlyPaths= or InaccessiblePaths=, and an absolute path that only exists inside the sandbox cannot be used here.

  2. acquire_path() (272-282) calls open(path, flags|O_NOCTTY, mode) with no O_NOFOLLOW and no component-wise resolution, running with the manager's privileges, so symlinks anywhere in the target path are followed with those privileges — which for truncate: is worth an explicit caution about path components under the control of a less-privileged party.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Replying here too, same as the sibling thread on 15. Accurate: setup_delegated_namespaces() (6094) and apply_root_directory() (6404) also run after the open, and acquire_path() passes no O_NOFOLLOW, so namespacing and symlink-following are both governed by the manager's view, not the unit's. Same scope call as 15 — out of this note's boundary, candidate for the follow-up.

Comment thread man/systemd.exec.xml Outdated
Comment on lines +3544 to +3548
<varname>StandardInput=</varname>, this only happens as a side effect of sharing a descriptor with
<option>file:<replaceable>path</replaceable></option> at the identical path, in which case a single
read-write descriptor is opened once and duplicated for both; pairing
<varname>StandardInput=</varname> with <option>append:<replaceable>path</replaceable></option> or
<option>truncate:<replaceable>path</replaceable></option> at a not-yet-existing shared path does not

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: suggestion: The page now states the descriptor-sharing contract twice, with different scope, and the accurate statement is the buried one. Lines 3496-3499 already say unconditionally: "If standard input and output are directed to the same file path, it is opened only once — for reading as well as writing — and duplicated." That is inaccurate for append:/truncate:: rw in setup_input() is computed only from EXEC_OUTPUT_FILE (src/core/exec-invoke.c:426-427), so with a shared path and append:/truncate: the single descriptor is read-only, not read-write.

The new text gets the scope right, but a reader looking up descriptor sharing will look at 3496, not inside an SELinux note, and two descriptions of one mechanism 50 lines apart will drift. Consider narrowing the sentence at 3496 in place and keeping only the MAC-specific consequence here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Replying here too. Correct that 3496-3499 is now imprecise for append:/truncate: and the accurate statement only lives in this note. Left 3496 as the general reader-facing sentence and this note as the precise one, rather than narrowing 3496 itself, to avoid duplicating the MAC-ordering caveat in two places on the page. Noted in the round-6 summary as declined for this PR.

Comment thread man/systemd.exec.xml Outdated
Comment on lines +3546 to +3548
read-write descriptor is opened once and duplicated for both; pairing
<varname>StandardInput=</varname> with <option>append:<replaceable>path</replaceable></option> or
<option>truncate:<replaceable>path</replaceable></option> at a not-yet-existing shared path does not

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: suggestion: Restricting the failure to a "not-yet-existing" shared path invites the inverse conclusion that pairing StandardInput=file: with append:/truncate: at the same path works once the path exists. It does not: setup_output() short-circuits on context->std_input == EXEC_INPUT_FILE && path_equal(...) for all three of EXEC_OUTPUT_FILE/_APPEND/_TRUNCATE (src/core/exec-invoke.c:637-641) and dup2()s the stdin descriptor, which for append:/truncate: was opened O_RDONLY (429) and without O_APPEND/O_TRUNC. The service ends up with a read-only descriptor on stdout/stderr (writes fail EBADF) and no append or truncate semantics — and the permission enumeration at 3534-3538 does not describe that configuration either, since no write or append open ever occurs.

Either say so, or drop the "not-yet-existing" qualifier and state that the pairing does not work at all.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Replying here too. Accurate — setup_output()'s dup2() (637-641) does hand the service a read-only stdout/stderr once the shared path exists, silently dropping append/truncate semantics instead of failing. Same scope reasoning as 15/16: a real, separate property of the shared-descriptor mechanism, not the MAC-ordering gap this note is about.

Comment thread man/systemd.exec.xml Outdated
can trigger the same creation, or instead share an already-open descriptor with one of these
settings, depending on which of them name the identical path and in which mode; the exact
conditions, and their interaction with the appending and truncating semantics, are corner cases this
note does not attempt to enumerate — see the general note on descriptor sharing above. The service

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: nit: "see the general note on descriptor sharing above" names a section that does not exist. There is no note above with that (or any) title about descriptor sharing; the only text above describing this behaviour is the sentence at lines 3496-3499 ("If standard input and output are directed to the same file path, it is opened only once — for reading as well as writing — and duplicated"), which is an unlabelled sentence inside the pre-existing file: paragraph and uses neither the word "descriptor" nor "sharing". A reader following the pointer has nothing to find. This page's idiom is a bare "see above" — compare the directly analogous line 3492-3493, "The semantics are similar to the same option of StandardInput=, see above." Either drop the invented title, or point concretely at "the description of file:path above".

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correct that there's no labelled section above — it's the unlabelled sentence at 3496-3499. Given the broader decision to stop iterating on this specific area after this round, leaving the cross-reference as "see the general note ... above" rather than reopening the paragraph again for a wording match; the reference is still locatable, just not phrased in this page's usual bare "see above" idiom.

@BenkiNew
BenkiNew force-pushed the man/stdout-file-mac-context branch 2 times, most recently from 0ce4870 to 35cb6c1 Compare September 4, 2026 23:02
@BenkiNew

BenkiNew commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

Pushed 35cb6c1. Fixed 42 and 44; declined 39, 40, 41, 43, 45, 46 — replies in their own threads below.

42 was worth the extra round — it's a different dimension from the sharing/creation gaps in 26/29/33/35/36 (which were about whether the manager's open() succeeds), while 42 is about what happens after: the service's own domain, post-transition, gets independently re-checked on every read/write of the inherited descriptor. Reproduced live on this exact test system:

avc: denied { append } ... scontext=init_t:s0     tcontext=httpd_sys_content_t:s0 tclass=file  [manager, at open]
avc: denied { append } ... scontext=httpd_t:s0    tcontext=httpd_sys_content_t:s0 tclass=file  [service, on its own first write]

Two separate, independently-timestamped denials for the same fd, against two different domains — confirming an admin who follows only the manager-side permission advice in this note would still see the service's own first write denied. Added a new paragraph stating this, with the reproduction. One correction to your framing: fd/use between the two domains is real but gated by the domain_fd_use boolean, which is on by default on the reference (targeted) policy — so it's rarely the actual blocker in practice, and the note says so rather than telling readers to grant it as if it were.

39–41, 43, 45, 46 declined — all accurate, but all in the same sharing/creation area that's now had five rounds of tightening (26 → 29 → 33/35/36 → this round), each fix producing a new gap. The note has said since 36a54f2 that this area's exact conditions are corner cases it does not attempt to enumerate; these findings fall inside that acknowledged boundary. Stopping here on this sub-topic specifically, regardless of what a further round finds on it — happy to pick any of it up in the follow-up already mentioned for the broader "what happens before/around the service's own context" material.

@BenkiNew
BenkiNew force-pushed the man/stdout-file-mac-context branch from 35cb6c1 to def2909 Compare September 5, 2026 04:43
@BenkiNew

BenkiNew commented Sep 5, 2026

Copy link
Copy Markdown
Contributor Author

Pushed def2909, correcting the paragraph added for 42 in bf49673 (and left unquestioned through 44's fix in 36a54f2). It was wrong, and worth flagging myself rather than waiting for a round to catch it.

The original text claimed the service-domain re-check produces "a second, independently-timestamped append denial" when the manager-side permissions are granted but the service's own domain lacks them. That was based on a flawed reproduction: forcing SELinuxContext= onto a plain /bin/sh with no valid entrypoint for that domain, which under permissive mode let the transition "fake" through and logged a denial that doesn't reflect real enforcing-mode behavior.

Redone properly — a copy of /bin/sh relabelled httpd_exec_t (an entrypoint httpd_t genuinely has, confirmed via sesearch -A -s httpd_t -c file -p entrypoint), run under a brief real setenforce 1 window, output target labelled so the manager succeeds but the service's domain does not:

  • The unit reports Finished / Deactivated successfully — no error, no failed exit code.
  • The target file is left completely empty.
  • ausearch -m avc for the entire run window returns zero records — not a permissive-style logged-but-tolerated denial, nothing at all.

So the actual behavior: the re-check at the domain transition doesn't deny and doesn't log — it silently substitutes the null device for the descriptor, the write succeeds from the process's point of view, and the output is gone with no audit trail. That's a materially different (and more consequential) claim than what shipped in the previous two commits, so I'm not leaving it standing on a "close enough" basis.

New text avoids naming a specific kernel function, since that part is inferred from observed behavior, not read from this kernel's source — everything else in the paragraph is exactly what was reproduced above. Also dropped the domain_fd_use/fd-class-use parenthetical from the previous version: it was reasoning about the wrong mechanism (a permission-check race), and I have no verified basis for whether it applies to this null-substitution path specifically, so it's gone rather than carried forward on a guess.

Comment thread man/systemd.exec.xml
input of processes to arbitrary system services.</para>
input of processes to arbitrary system services. Note that the file system object is opened by the
service manager before the executable is invoked; see the note on the security context this
happens in for <varname>StandardOutput=</varname> below.</para>

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i am sorry, but i find this note ridiculous. Of course this is opened before invoking the executable. Invoking the binary is the boundary between the service manager's code and the executable's code. Hence, all stuff we do we have to do before this boundary, and all stuff the binary does is after. Since StandardOutput= and friends are service manager knob, of course they have to be executed by us and not the binary.

What do you expect us to do otherwise? binary patch the executable to open those files from their context once started? that'd be ridiculous.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair, and taken. The ordering itself isn't news — it's a direct consequence of the exec boundary, exactly as you describe, and I agree it doesn't need explaining at length.

Pushed e9ecfcb, which cuts this down from 5 paragraphs to 2 and removes everything that was just restating that boundary or walking through which SELinux permission goes where — all the stuff you're right to call obvious. What's left states the ordering in one sentence (needed only as the premise) and then the one thing I don't think is obvious: if the unit's own domain isn't permitted the access the manager's domain needed, the failure isn't denied and isn't logged — the kernel substitutes the descriptor with the null device at the transition, the write succeeds, the unit exits clean, and there's no AVC record at all. That's the part I'd still like your read on; happy to drop it too if it still doesn't clear the bar.

@poettering

Copy link
Copy Markdown
Member

I am sorry, but I am not convinced this is necessary. This adds a lot of noise explaining stuff that is very niche (i.e. selinux...) and should really be obvious anyway.

@poettering poettering added needs-discussion 🤔 and removed please-review PR is ready for (re-)review by a maintainer labels Sep 5, 2026
@BenkiNew
BenkiNew force-pushed the man/stdout-file-mac-context branch from def2909 to e9ecfcb Compare September 5, 2026 05:14
@github-actions github-actions Bot added the please-review PR is ready for (re-)review by a maintainer label Sep 5, 2026
@BenkiNew

BenkiNew commented Sep 5, 2026

Copy link
Copy Markdown
Contributor Author

Understood, and I think the version that was here when you looked at it earned that reaction — it had grown to 5 dense paragraphs over a lot of automated review rounds, most of them explaining exactly the kind of thing you're right to call obvious (which SELinux permission the manager's domain needs, for which of the three modes, in which combination).

Pushed e9ecfcb, cutting all of that. What's left is 2 short paragraphs, and the only thing they exist to say is: because the manager opens the path before the unit's own domain exists, an admin can get the manager's side of the labelling right and still have the service's first write silently discarded — no denial, no audit record, unit exits 0, target file just stays empty. That's the one part I don't think is "should be obvious anyway" even to someone who knows the exec-boundary story well, since it's a kernel behavior (null-device substitution at the domain transition) rather than a consequence of systemd's own design. If that one paragraph still isn't worth it, I'd rather hear that directly than keep iterating on wording.

Comment thread man/systemd.exec.xml Outdated
descriptor is replaced with one pointing at the null device, so the write that follows succeeds,
the unit reports no error, and — unlike every other SELinux denial — no audit record of any kind
is produced. The only symptom is that the target path stays empty; <command>ausearch</command> and
<command>audit2allow</command> have nothing to show for it.</para>

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

please drop this paragraph

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dropped in e30b25f.

What's left is the ordering sentence itself plus its cross-reference from StandardInput=. Given your first comment, I suspect that's the part you consider self-evident too — if so, just say the word and I'll close this rather than make you explain it twice.

@poettering poettering added reviewed/needs-rework 🔨 PR has been reviewed and needs another round of reworks and removed please-review PR is ready for (re-)review by a maintainer labels Sep 5, 2026
@BenkiNew
BenkiNew force-pushed the man/stdout-file-mac-context branch from e9ecfcb to e30b25f Compare September 5, 2026 05:31
@github-actions github-actions Bot added please-review PR is ready for (re-)review by a maintainer and removed reviewed/needs-rework 🔨 PR has been reviewed and needs another round of reworks labels Sep 5, 2026
Comment thread man/systemd.exec.xml
Comment on lines +3520 to +3523
<para>Note that for <option>file:<replaceable>path</replaceable></option>,
<option>append:<replaceable>path</replaceable></option> and
<option>truncate:<replaceable>path</replaceable></option> of <varname>StandardOutput=</varname>,
<varname>StandardError=</varname> or <varname>StandardInput=</varname>, the target path is opened

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: must-fix: This distributes all three option values across all three settings, but StandardInput= accepts neither append: nor truncate:. config_parse_exec_input() (src/core/load-fragment.c:1136-1202) special-cases only the fd: and file: prefixes and otherwise falls through to exec_input_from_string(), and exec_input_table (src/core/execute.c:3149-3157) contains only null, tty, tty-force, tty-fail, socket, fd, data and file — append and truncate exist solely in exec_output_table (3162-3174). So StandardInput=truncate:/x is rejected with "Failed to parse input specifier", yet this sentence presents it as a supported configuration. The page contradicts itself here too: the accepted-value list for StandardInput= above (lines 3392-3395) offers only file:path, while StandardOutput= (3455-3460) lists all three. The new cross-reference added at 3423-3425 makes it worse by sending StandardInput= readers straight to this sentence.

Scope the three-value list to StandardOutput=/StandardError= and mention StandardInput= separately for file:path only — e.g. "for file:path, append:path and truncate:path of StandardOutput= or StandardError=, and for file:path of StandardInput=". If the paragraph is dropped per the maintainer's request this resolves itself, but as long as any version of the sentence stays, the value list needs scoping.

Comment thread man/systemd.exec.xml
<varname>StandardError=</varname> or <varname>StandardInput=</varname>, the target path is opened
by the service manager before the executable is invoked — and hence before
<varname>SELinuxContext=</varname> or any policy-defined transition takes effect. The resulting
descriptor is then inherited across that transition into the unit's own domain.</para>

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: must-fix: "The resulting descriptor is then inherited across that transition into the unit's own domain" is stated as an unconditional guarantee, and it fails in exactly the case this note exists to warn about. The descriptor is indeed still open at fexecve_or_execve() — acquire_path() (src/core/exec-invoke.c:271) opens without O_CLOEXEC and it is installed onto the stdio fds via move_fd(..., /* cloexec= */ false) (433 for stdin, 653 for the output modes). But when the exec causes an SELinux domain change, the kernel revalidates each inherited descriptor against the new subject during credential commit (flush_unauthorized_files()), checking fd { use } towards the SID of the opener plus the file-class permission implied by the open mode (an O_APPEND descriptor maps to append, not write). Where that is denied in enforcing mode the descriptor is not inherited: the kernel substitutes a descriptor on the SELinux null device, or closes it. And even where inheritance succeeds, the opener's authorization is not cached on the struct file — read/write/append are revalidated against the current subject on subsequent operations once the current SID differs from the one recorded at open. The reproduction posted in this PR's own discussion shows two independently-timestamped denials for a single StandardOutput=append: target, one at the manager's open and one on the service's own write, which is precisely what this sentence tells the reader cannot happen. As written it says the matter is settled once the manager-side open succeeds. Please qualify it, or state the revalidation/substitution outcome. In passing: a descriptor is an object, not a subject, so "into the unit's own domain" attributes a subject property to it.

Comment thread man/systemd.exec.xml
<option>truncate:<replaceable>path</replaceable></option> of <varname>StandardOutput=</varname>,
<varname>StandardError=</varname> or <varname>StandardInput=</varname>, the target path is opened
by the service manager before the executable is invoked — and hence before
<varname>SELinuxContext=</varname> or any policy-defined transition takes effect. The resulting

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude: suggestion: "before SELinuxContext= or any policy-defined transition takes effect", followed by the definite "across that transition" in the next sentence, presupposes that a transition always occurs and leaves the reader no case where it does not. Two common configurations have none. (a) No SELinuxContext= is set and policy defines no type_transition for the executable's entrypoint, so the exec'ing process keeps the service manager's domain. (b) ExecStart=+... / !! style commands carry EXEC_COMMAND_FULLY_PRIVILEGED, so needs_sandboxing is false (src/core/exec-invoke.c:5496) and the whole MAC block is skipped — sym_setexeccon_raw() (6471) is reached only inside if (needs_sandboxing) (6460) and only when mac_selinux_context_net ?: context->selinux_context is non-NULL (6468). In both cases the opening subject and the running subject are the same domain, so the note's premise does not hold and the reader is told to worry about a distinction that does not exist for their unit. Phrasing the consequence conditionally would fix it, e.g. "where policy or SELinuxContext= causes the service to run in a different domain than the service manager".

@poettering poettering added good-to-merge/waiting-for-ci 👍 PR is good to merge, but CI hasn't passed at time of review. Please merge if you see CI has passed and removed needs-discussion 🤔 please-review PR is ready for (re-)review by a maintainer labels Sep 5, 2026
@poettering

Copy link
Copy Markdown
Member

so, i still think this is kinda obvious, but ok. seems ok to merge like this.

@bluca
bluca merged commit 9457f81 into systemd:main Sep 5, 2026
55 of 59 checks passed
@github-actions github-actions Bot removed the good-to-merge/waiting-for-ci 👍 PR is good to merge, but CI hasn't passed at time of review. Please merge if you see CI has passed label Sep 5, 2026
@BenkiNew
BenkiNew deleted the man/stdout-file-mac-context branch September 5, 2026 14:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Development

Successfully merging this pull request may close these issues.

3 participants