Skip to content

diagnostics

diagnostics

On-host supervisor + sidecar artifact paths for a container.

Single source of truth for the human-facing debug layout. When an operator needs the supervisor log, the wrapper, the PID file, or the sidecar bundle for a container, the paths come from here rather than being re-derived by each frontend — terok task status -v is the first consumer, pointing a human at the file to send back instead of making them hand-assemble it from podman inspect annotations.

The artifacts live under three different keys, which is exactly why a shared resolver earns its keep:

  • log + PID key on the immutable container ID — podman assigns it at create time, and the supervisor names both after it.
  • sidecar keys on the container name — known before the ID exists, so the launch path can write it pre-podman run.
  • the wrapper is install-global (one per state root), not per container.

Paths are computed, never probed: a file may be absent (the sidecar is removed at teardown, the supervisor may never have logged). Callers that care about existence check it themselves. The two exceptions are supervisor_liveness, which probes the root-cause question "is this container's supervisor actually running?", and respawn_supervisor, its remediation — both act on the supervisor the path bundle only points at.

__all__ = ['ContainerDiagnostics', 'SupervisorLiveness', 'container_diagnostics', 'respawn_supervisor', 'supervisor_liveness'] module-attribute

ContainerDiagnostics(container_id, log, pid, wrapper, sidecar, hook_log) dataclass

Resolved on-host artifact paths for one container.

Every field is an absolute Path; none is guaranteed to exist on disk (see the module docstring).

container_id instance-attribute

log instance-attribute

Persistent supervisor log: <state>/logs/<id>.log.

pid instance-attribute

Supervisor PID file: <state>/pids/supervisor-<id>.pid.

wrapper instance-attribute

Install-global supervisor wrapper: <state>/supervisor_wrapper.py.

sidecar instance-attribute

Per-container sidecar bundle: <state>/sidecar/<name>.json.

hook_log instance-attribute

Install-global OCI-hook diary: <state>/logs/hook.log.

Shared across every container (each line is container-tagged) — an absent or empty file means the supervisor hook never fired, which is the first thing to check when a container comes up unsupervised.

SupervisorLiveness(alive, pid, detail, services=()) dataclass

Whether a container's per-container supervisor is currently running.

The result of probing the recorded PID file and /proc — the same signal the OCI hook's idempotent-respawn guard uses. pid is the PID the file records (kept even when the process turns out dead, so a caller can name it); detail is a short human phrase for a status line.

alive instance-attribute

pid instance-attribute

detail instance-attribute

services = () class-attribute instance-attribute

container_diagnostics(container_id, container_name, *, state_dir=None)

Resolve the artifact-path bundle for one container.

state_dir defaults to state_root — the operator's resolved paths.root — so the bundle moves with a relocated state tree just like every other terok artifact. Pass an explicit state_dir (e.g. a caller's SandboxConfig.state_dir) to pin resolution to a specific config rather than the layered default.

Source code in src/terok_sandbox/diagnostics.py
def container_diagnostics(
    container_id: str,
    container_name: str,
    *,
    state_dir: Path | None = None,
) -> ContainerDiagnostics:
    """Resolve the artifact-path bundle for one container.

    *state_dir* defaults to [`state_root`][terok_sandbox.paths.state_root]
    — the operator's resolved ``paths.root`` — so the bundle moves with
    a relocated state tree just like every other terok artifact.  Pass
    an explicit *state_dir* (e.g. a caller's
    [`SandboxConfig.state_dir`][terok_sandbox.config.SandboxConfig]) to
    pin resolution to a specific config rather than the layered default.
    """
    root = state_dir or state_root()
    return ContainerDiagnostics(
        container_id=container_id,
        log=root / _LOGS_DIR_NAME / f"{container_id}.log",
        pid=root / _PIDS_DIR_NAME / f"supervisor-{container_id}.pid",
        wrapper=root / _WRAPPER_NAME,
        sidecar=root / _SIDECAR_DIR_NAME / f"{container_name}.json",
        hook_log=root / _LOGS_DIR_NAME / _HOOK_LOG_NAME,
    )

supervisor_liveness(container_id, *, state_dir=None)

Probe whether the supervisor wrapper for container_id is running.

Reads <state>/pids/supervisor-<id>.pid and confirms the recorded process is alive and is our wrapper for this container — the recorded wrapper path and the container id must both appear in its argv, the double-mark that guards against a recycled PID (the same check the OCI hook makes before deciding whether to respawn).

Also reports which service children are running, because the parent outliving them is a real state and a common one: a child that cannot resolve the vault passphrase exits at once and the parent stays up, so a container with no vault and no SSH agent answers alive=True. The two facts are separate — alive is what a respawn can fix, and a bundle missing children is not that.

Never raises: a missing PID file (the hook never spawned a supervisor) and a stale one (the wrapper died, or the PID was recycled) both come back alive=False with a distinguishing detail. Pass state_dir to pin resolution to a specific config (mirrors container_diagnostics).

Source code in src/terok_sandbox/diagnostics.py
def supervisor_liveness(
    container_id: str,
    *,
    state_dir: Path | None = None,
) -> SupervisorLiveness:
    """Probe whether the supervisor wrapper for *container_id* is running.

    Reads ``<state>/pids/supervisor-<id>.pid`` and confirms the recorded
    process is alive **and** is our wrapper for this container — the
    recorded wrapper path and the container id must both appear in its
    argv, the double-mark that guards against a recycled PID (the same
    check the OCI hook makes before deciding whether to respawn).

    Also reports which service children are running, because the parent
    outliving them is a real state and a common one: a child that cannot
    resolve the vault passphrase exits at once and the parent stays up,
    so a container with no vault and no SSH agent answers ``alive=True``.
    The two facts are separate — ``alive`` is what a respawn can fix, and
    a bundle missing children is not that.

    Never raises: a missing PID file (the hook never spawned a supervisor)
    and a stale one (the wrapper died, or the PID was recycled) both come
    back ``alive=False`` with a distinguishing *detail*.  Pass *state_dir*
    to pin resolution to a specific config (mirrors
    [`container_diagnostics`][terok_sandbox.diagnostics.container_diagnostics]).
    """
    root = state_dir or state_root()
    pid_file = root / _PIDS_DIR_NAME / f"supervisor-{container_id}.pid"
    wrapper = root / _WRAPPER_NAME
    try:
        pid = int(pid_file.read_text().strip())
    except (OSError, ValueError):
        return SupervisorLiveness(
            alive=False,
            pid=None,
            detail="no PID file — the supervisor hook never spawned a supervisor",
        )
    if _pid_alive(pid) and _wrapper_argv_matches(pid, wrapper, container_id):
        services = _proc.service_children(container_id)
        running = ", ".join(services) if services else "none"
        unit = _supervisor_state.unit_name(container_id)
        placement = (
            f"user unit {unit}" if _supervisor_state.unit_active(unit) else "namespace daemon"
        )
        return SupervisorLiveness(
            alive=True,
            pid=pid,
            detail=f"supervisor pid {pid} alive as {placement}; children running: {running}",
            services=services,
        )
    return SupervisorLiveness(
        alive=False, pid=pid, detail=f"stale PID file — pid {pid} is dead or recycled"
    )

respawn_supervisor(container_id, container_name, *, state_dir=None, container_pid=None)

Re-fire the OCI hook's supervisor spawn for a running container.

Idempotent remediation for a container that came up unsupervised: it re-invokes the installed createRuntime hook with a synthesized OCI state — exactly what crun feeds it at container start. The hook reconstructs the same env and paths, no-ops when a supervisor is already alive, and writes the PID file, so re-running it is the safe, canonical respawn (the same mechanism the poststop/stray machinery already trusts) rather than a second copy of the spawn logic.

container_pid (the container init's host PID, if the caller has it) lets the respawned supervisor watch the container directly; omitted, the supervisor falls back to its podman wait watch.

Returns the liveness after the attempt. A missing hook or sidecar (setup never ran, or the container was torn down) comes back alive=False without spawning anything. Never raises.

Source code in src/terok_sandbox/diagnostics.py
def respawn_supervisor(
    container_id: str,
    container_name: str,
    *,
    state_dir: Path | None = None,
    container_pid: int | None = None,
) -> SupervisorLiveness:
    """Re-fire the OCI hook's supervisor spawn for a running container.

    Idempotent remediation for a container that came up unsupervised: it
    re-invokes the installed ``createRuntime`` hook with a synthesized OCI
    state — exactly what crun feeds it at container start.  The hook
    reconstructs the same env and paths, no-ops when a supervisor is
    already alive, and writes the PID file, so re-running it is the safe,
    canonical respawn (the same mechanism the poststop/stray machinery
    already trusts) rather than a second copy of the spawn logic.

    *container_pid* (the container init's host PID, if the caller has it)
    lets the respawned supervisor watch the container directly; omitted,
    the supervisor falls back to its ``podman wait`` watch.

    Returns the liveness *after* the attempt.  A missing hook or sidecar
    (setup never ran, or the container was torn down) comes back
    ``alive=False`` without spawning anything.  Never raises.
    """
    root = state_dir or state_root()
    hook = root / _HOOKS_DIR_NAME / _HOOK_SCRIPT_NAME
    sidecar = root / _SIDECAR_DIR_NAME / f"{container_name}.json"
    if not (hook.is_file() and sidecar.is_file()):
        return supervisor_liveness(container_id, state_dir=root)
    try:
        descriptor = json.loads((hook.parent / _descriptor_name("createRuntime")).read_text())
        installed_hook = descriptor["hook"]
        argv = [installed_hook["path"], *installed_hook["args"][1:]]
        if not all(isinstance(arg, str) for arg in argv):
            raise ValueError("Invalid installed hook arguments")
    except (OSError, ValueError, KeyError, TypeError):
        return supervisor_liveness(container_id, state_dir=root)

    state: dict[str, object] = {
        "id": container_id,
        "annotations": {_TRIGGER_ANNOTATION: str(sidecar)},
    }
    if container_pid is not None:
        state["pid"] = container_pid
    with contextlib.suppress(OSError, subprocess.SubprocessError):
        subprocess.run(  # nosec B603 — fixed argv, our own installed hook script
            argv,
            input=json.dumps(state),
            text=True,
            timeout=_RESPAWN_TIMEOUT_S,
            check=False,
        )

    # The hook writes the PID file before returning, but the detached
    # wrapper may still be mid-``exec`` — poll briefly so a healthy respawn
    # isn't reported as a failure on that race.
    deadline = time.monotonic() + _RESPAWN_SETTLE_S
    while True:
        live = supervisor_liveness(container_id, state_dir=root)
        if live.alive or time.monotonic() >= deadline:
            return live
        time.sleep(_RESPAWN_POLL_INTERVAL_S)