|

bugs: close 004 (fixed in 0.2.7) — archive

Author: Chris Tusa <chris.tusa@leafscale.com>
Date: Jul 21, 2026 14:52
Changeset: deaa4c646836d3cf1ec83956066acb17ff183369
Branch: default
Tags: tip

Changed files:

D bugs/004-nonpid1-invocation-clobbers-live-control-socket.md

Diff

diff -r dbf0a653d282 -r deaa4c646836 bugs/004-nonpid1-invocation-clobbers-live-control-socket.md
--- a/bugs/004-nonpid1-invocation-clobbers-live-control-socket.md	Tue Jul 21 14:52:10 2026 -0500
+++ /dev/null	Thu Jan 01 00:00:00 1970 +0000
@@ -1,105 +0,0 @@
-# Bug 004: running the zyginit binary as a non-PID-1 process (e.g. `zyginit -v`) re-runs init and clobbers the live control socket
-
-## Summary
-
-On a system where zyginit is already running as PID 1, invoking the zyginit
-binary directly from a shell — e.g. an admin running **`zyginit -v`** to check
-the version — does **not** short-circuit and exit. Instead it falls through
-into the full supervisor sequence ("non-init mode"), which creates the control
-socket. `ctlsocket.create_socket()` **unlinks the existing
-`/var/run/zyginit.sock` and binds a new one**, clobbering the socket owned by
-the live PID 1. From then on `zygctl` cannot reach the real init:
-
-```
-zygctl: cannot connect to zyginit at /var/run/zyginit.sock
-zygctl: is zyginit running?
-```
-
-The system otherwise keeps running (PID 1 is untouched in memory), but it is no
-longer controllable via `zygctl`, and per `main.reef:534` recovery "requires
-reboot."
-
-Two distinct defects combine to make this an easy foot-gun:
-
-1. **`-v` is not recognized as a version flag.** Only `--version` and `-V`
-   (capital) short-circuit (`main.reef:943`). Lowercase `-v` — the conventional
-   "show version" flag for most CLI tools — is unmatched, so `zyginit -v` runs
-   the init/supervisor path. An admin's muscle-memory `zyginit -v` is exactly
-   the wrong thing.
-
-2. **Non-PID-1 invocation runs the full supervisor with no live-instance
-   guard.** The `getpid() != 1` branch is labeled "supervisor mode for testing"
-   (`main.reef` ~979/1015), but on a production host it still loads services,
-   enters the event loop, and calls the un-guarded socket setup (~`main.reef:1124`,
-   not wrapped in `is_pid_1()`), so `create_socket()` clobbers the production
-   socket. There is no check for "a live PID-1 zyginit already owns this socket."
-
-## Severity
-
-**Medium.** Not a boot-path defect, but a sharp operational foot-gun: a
-completely reasonable admin action (`zyginit -v` to read the version) silently
-breaks `zygctl` control of the running system until the next reboot. It is easy
-to trigger and non-obvious to diagnose (the socket just "disappears").
-
-## Affected Version
-
-zyginit 0.2.6 (confirmed by source inspection) and 0.2.5 (observed on
-Hammerhead hh-alpha9 — an admin ran `zyginit -v` and `zygctl` then reported the
-socket unreachable; system was fine but uncontrollable until reboot).
-
-## Suggested Fix (for the zyginit team)
-
-1. **Make `zyginit -v` safe.** Treat `-v` as a version alias (or, more
-   conservatively, reject unknown flags with a usage message) so `zyginit -v`
-   prints the version and exits **before** any PID-1/supervisor/socket setup —
-   the same short-circuit that already exists for `--version`/`-V`.
-2. **Guard the socket against a live instance.** Before `create_socket()`
-   unlinks/binds `/var/run/zyginit.sock`, detect whether a live PID-1 zyginit
-   already owns it — e.g. try to `connect()` first (a successful connect means
-   a live server is present → refuse and exit), or check `/proc/1` is zyginit.
-   Never unlink a socket that a live server is actively listening on.
-3. **Gate "supervisor test mode" behind an explicit opt-in.** Non-PID-1
-   supervisor runs (for integration testing) should require an explicit flag
-   (e.g. `--supervisor-test`) and/or use a distinct socket path (e.g.
-   `$ZYGINIT_SOCK` / a temp path) so a bare `zyginit ...` on a production host
-   can never touch the production socket.
-
-## Reported by
-
-Hammerhead team (Chris Tusa), 2026-07-21. Triggered on hh-alpha9 (zyginit 0.2.5,
-PID 1) by running `zyginit -v` from a root shell; `zygctl` subsequently could
-not connect to `/var/run/zyginit.sock`. Cross-ref Hammerhead memory
-`zyginit-no-exec-init-binary` (the known "don't exec the init binary on a
-zyginit host — clobbers the socket, recover via `virsh destroy && virsh start`"
-note).
-
-## Resolution (2026-07-21, zyginit 0.2.7-dev)
-
-All three suggested fixes implemented (defense in depth):
-
-1. **`-v` / unknown-flag safety** (`src/main.reef`). `-v`, `-V`, and
-   `--version` now print the version and exit — **but only when
-   `not is_pid_1()`**. This is a deliberately narrower gate than the report
-   proposed: the Hammerhead kernel passes `-v` (verbose) as a boot-arg to PID 1
-   (`boot-args=-v -m verbose`), so treating `-v` as "print version and exit" in
-   the PID-1 case would make init return from `main()` and the kernel would
-   re-exec it in a tight loop. Shell invocations (getpid() != 1) short-circuit;
-   PID 1 never does.
-2. **Live-instance socket guard** (`src/ctlsocket.reef::create_socket`). Before
-   unlinking/binding, it `connect()`s the path; if a live server answers it
-   logs `REFUSING` and returns -1 instead of clobbering. A stale socket file
-   (no listener) fails to connect and is cleaned up as before. `main.reef`
-   additionally exits cleanly when a non-PID-1 instance cannot own the socket
-   (PID 1 keeps running without a control socket — losing zygctl beats halting
-   init).
-3. **Opt-in for non-PID-1 supervisor mode** (`src/main.reef`). A non-PID-1 run
-   now requires `--supervisor-test`; a bare `zyginit ...` on a host where a real
-   zyginit owns PID 1 refuses and exits before any setup. The integration
-   harnesses (`run_tests.sh`, `test_shutdown.sh`) and the CLAUDE.md dev
-   invocation pass the flag.
-
-Verified by integration suite (`run_tests.sh`, Suite "BUG-004: non-PID-1 socket
-safety", 104/104 passing): `zyginit -v` prints the version and creates no
-socket; a bare `zyginit` refuses without the flag; and a second instance
-launched against a live daemon's socket refuses to clobber it, leaving the
-original reachable.
diff -r dbf0a653d282 -r deaa4c646836 bugs/archive/004-nonpid1-invocation-clobbers-live-control-socket.md
--- /dev/null	Thu Jan 01 00:00:00 1970 +0000
+++ b/bugs/archive/004-nonpid1-invocation-clobbers-live-control-socket.md	Tue Jul 21 14:52:38 2026 -0500
@@ -0,0 +1,109 @@
+> **CLOSED (2026-07-21):** Fixed and released in **zyginit 0.2.7** (tag
+> v0.2.7, rev 227). All three suggested fixes landed — see the Resolution
+> section below. Non-PID-1 runs now require `--supervisor-test`.
+
+# Bug 004: running the zyginit binary as a non-PID-1 process (e.g. `zyginit -v`) re-runs init and clobbers the live control socket
+
+## Summary
+
+On a system where zyginit is already running as PID 1, invoking the zyginit
+binary directly from a shell — e.g. an admin running **`zyginit -v`** to check
+the version — does **not** short-circuit and exit. Instead it falls through
+into the full supervisor sequence ("non-init mode"), which creates the control
+socket. `ctlsocket.create_socket()` **unlinks the existing
+`/var/run/zyginit.sock` and binds a new one**, clobbering the socket owned by
+the live PID 1. From then on `zygctl` cannot reach the real init:
+
+```
+zygctl: cannot connect to zyginit at /var/run/zyginit.sock
+zygctl: is zyginit running?
+```
+
+The system otherwise keeps running (PID 1 is untouched in memory), but it is no
+longer controllable via `zygctl`, and per `main.reef:534` recovery "requires
+reboot."
+
+Two distinct defects combine to make this an easy foot-gun:
+
+1. **`-v` is not recognized as a version flag.** Only `--version` and `-V`
+   (capital) short-circuit (`main.reef:943`). Lowercase `-v` — the conventional
+   "show version" flag for most CLI tools — is unmatched, so `zyginit -v` runs
+   the init/supervisor path. An admin's muscle-memory `zyginit -v` is exactly
+   the wrong thing.
+
+2. **Non-PID-1 invocation runs the full supervisor with no live-instance
+   guard.** The `getpid() != 1` branch is labeled "supervisor mode for testing"
+   (`main.reef` ~979/1015), but on a production host it still loads services,
+   enters the event loop, and calls the un-guarded socket setup (~`main.reef:1124`,
+   not wrapped in `is_pid_1()`), so `create_socket()` clobbers the production
+   socket. There is no check for "a live PID-1 zyginit already owns this socket."
+
+## Severity
+
+**Medium.** Not a boot-path defect, but a sharp operational foot-gun: a
+completely reasonable admin action (`zyginit -v` to read the version) silently
+breaks `zygctl` control of the running system until the next reboot. It is easy
+to trigger and non-obvious to diagnose (the socket just "disappears").
+
+## Affected Version
+
+zyginit 0.2.6 (confirmed by source inspection) and 0.2.5 (observed on
+Hammerhead hh-alpha9 — an admin ran `zyginit -v` and `zygctl` then reported the
+socket unreachable; system was fine but uncontrollable until reboot).
+
+## Suggested Fix (for the zyginit team)
+
+1. **Make `zyginit -v` safe.** Treat `-v` as a version alias (or, more
+   conservatively, reject unknown flags with a usage message) so `zyginit -v`
+   prints the version and exits **before** any PID-1/supervisor/socket setup —
+   the same short-circuit that already exists for `--version`/`-V`.
+2. **Guard the socket against a live instance.** Before `create_socket()`
+   unlinks/binds `/var/run/zyginit.sock`, detect whether a live PID-1 zyginit
+   already owns it — e.g. try to `connect()` first (a successful connect means
+   a live server is present → refuse and exit), or check `/proc/1` is zyginit.
+   Never unlink a socket that a live server is actively listening on.
+3. **Gate "supervisor test mode" behind an explicit opt-in.** Non-PID-1
+   supervisor runs (for integration testing) should require an explicit flag
+   (e.g. `--supervisor-test`) and/or use a distinct socket path (e.g.
+   `$ZYGINIT_SOCK` / a temp path) so a bare `zyginit ...` on a production host
+   can never touch the production socket.
+
+## Reported by
+
+Hammerhead team (Chris Tusa), 2026-07-21. Triggered on hh-alpha9 (zyginit 0.2.5,
+PID 1) by running `zyginit -v` from a root shell; `zygctl` subsequently could
+not connect to `/var/run/zyginit.sock`. Cross-ref Hammerhead memory
+`zyginit-no-exec-init-binary` (the known "don't exec the init binary on a
+zyginit host — clobbers the socket, recover via `virsh destroy && virsh start`"
+note).
+
+## Resolution (2026-07-21, zyginit 0.2.7)
+
+All three suggested fixes implemented (defense in depth):
+
+1. **`-v` / unknown-flag safety** (`src/main.reef`). `-v`, `-V`, and
+   `--version` now print the version and exit — **but only when
+   `not is_pid_1()`**. This is a deliberately narrower gate than the report
+   proposed: the Hammerhead kernel passes `-v` (verbose) as a boot-arg to PID 1
+   (`boot-args=-v -m verbose`), so treating `-v` as "print version and exit" in
+   the PID-1 case would make init return from `main()` and the kernel would
+   re-exec it in a tight loop. Shell invocations (getpid() != 1) short-circuit;
+   PID 1 never does.
+2. **Live-instance socket guard** (`src/ctlsocket.reef::create_socket`). Before
+   unlinking/binding, it `connect()`s the path; if a live server answers it
+   logs `REFUSING` and returns -1 instead of clobbering. A stale socket file
+   (no listener) fails to connect and is cleaned up as before. `main.reef`
+   additionally exits cleanly when a non-PID-1 instance cannot own the socket
+   (PID 1 keeps running without a control socket — losing zygctl beats halting
+   init).
+3. **Opt-in for non-PID-1 supervisor mode** (`src/main.reef`). A non-PID-1 run
+   now requires `--supervisor-test`; a bare `zyginit ...` on a host where a real
+   zyginit owns PID 1 refuses and exits before any setup. The integration
+   harnesses (`run_tests.sh`, `test_shutdown.sh`) and the CLAUDE.md dev
+   invocation pass the flag.
+
+Verified by integration suite (`run_tests.sh`, Suite "BUG-004: non-PID-1 socket
+safety", 104/104 passing): `zyginit -v` prints the version and creates no
+socket; a bare `zyginit` refuses without the flag; and a second instance
+launched against a live daemon's socket refuses to clobber it, leaving the
+original reachable.