Skip to content

linux: process lifecycle and signals as Linux has them - #585

Open
eKisNonos wants to merge 27 commits into
linux/go-guestsfrom
linux/guest-lifecycle
Open

eKisNonos wants to merge 27 commits into
linux/go-guestsfrom
linux/guest-lifecycle

Conversation

@eKisNonos

@eKisNonos eKisNonos commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #582. This PR targets linux/go-guests so it shows only its own 27 commits: 26 new ones and one merge of linux/go-guests at 24cafe103. When #582 merges into main, retarget this PR to main. It cannot stand on main alone: the signal frame, sigreturn and delivery code it changes only exist on linux/go-guests.

Before this, Go could not start another program, musl's posix_spawn never returned, a leader's plain exit ended the whole process, wait4 accepted options it does not know, waitid was not served, SIGCHLD and SIGPIPE were never raised, and a signal caught while a thread waited in sleep or pause landed only when the wait ran out, if it ever did. The signal frame put sigcontext at the wrong offset, so a handler that read its ucontext read garbage.

What Linux guests can do now

  • A leader's plain exit ends only the leader. The process lives on in its other threads and reports the status of the last one to end.
  • Go's syscall.ForkExec and os/exec run as written: clone(CLONE_VFORK) makes a child and holds the parent until the child execs or exits, and /dev/null, /dev/zero, /dev/full, /dev/random and /dev/urandom exist.
  • musl's posix_spawn, system and popen work, through a kernel fork that can start the child on a stack the caller names.
  • wait4 and waitid answer with Linux's status words and siginfo, including WNOHANG, WNOWAIT, P_ALL, P_PID, P_PGID, EINVAL for an unknown option and ECHILD. A child's end raises SIGCHLD at its parent with the child's pid, code and status.
  • Signals follow per-thread masks, sa_mask, SA_NODEFER, SA_RESETHAND, SA_ONSTACK with alternate stacks, and SA_RESTART. The frame is Linux's rt_sigframe byte for byte.
  • A caught signal ends the wait a thread is parked in, with EINTR or a restart as Linux decides for that call.
  • kill reaches every process of the family, not only the caller.
  • A write, writev, sendto or sendmsg that answers EPIPE raises SIGPIPE, unless MSG_NOSIGNAL is set.
  • alarm, setitimer, getitimer, pause, rt_sigsuspend, rt_sigtimedwait, rt_sigpending, rt_sigqueueinfo, signalfd4, signalfd and the POSIX timer_* calls are served, with overruns counted as Linux counts them.
  • The Linux-guest test store can pack only the guests LINUX_GUEST_SET names.

What changed

The signal frame and the signal state

  • linux dff31857f: the signal frame is Linux's rt_sigframe byte for byte: ucontext at +8, uc_mcontext at +40 inside it, uc_sigmask at +296, siginfo at +312, 440 bytes in all. It had sigcontext at 48.
  • linux 1b2ea5796: any handler can turn a kernel pid into the guest's number, through one pid space the whole family shares.
  • linux 3bd641ce6: signals honour per-thread masks, alternate stacks and each signal's default action.
  • linux 3af1d3cc3: kill reaches every process of the family. A signal for another process goes out through the family's outbox instead of being queued in the sender.

Processes start and end as on Linux

  • linux 294461d13: a leader's plain exit ends only the leader, and the process's status is Linux's.
  • linux 88bd9a132: fork copies a span larger than 1 MiB in pieces. A single peer copy is capped at 1 MiB, so forking a Go process answered ENOMEM.
  • linux 4c5f99d3e: clone can make a process, and CLONE_VFORK holds its parent until the child execs or exits.
  • linux aaed19326: wait4 and waitid answer as Linux's do, and SIGCHLD is raised.
  • foreign 478e293f9: MkForeignFork takes an optional stack for the child, checked against the user half of the address space. Zero keeps the parent's stack pointer.
  • libc 5f35d3b5b: mk_foreign_fork_at forks a guest onto a stack it names.
  • linux ac08871b1: a process clone that names a stack starts the child on it. musl's posix_spawn does this, and the child used to run on its parent's stack pointer.

Signals that end waits, timers and signal waits

  • linux 0c9b8fdc4: the sigpipe guest checks SIGPIPE for a write to a pipe with no reader.
  • linux b27092ce9: a caught signal ends the wait a thread is parked in.
  • linux f2ce1fcdc: alarm and setitimer arm ITIMER_REAL, which raises SIGALRM.
  • linux 6e6f70515: pause, sigsuspend, sigtimedwait and sigpending are served.
  • linux 949b45224: POSIX timers are served, with overruns counted as Linux does.
  • linux 3f26b6d11: five guests check process lifecycle and signals against Linux: leaderexit, goexec, sigchld, sigpipe and alarm.

On top of #582

  • linux 2ed909690: merges linux/go-guests at 24cafe103, so the waits, eventfd and pipe ends of linux: run Go and musl threads in a guest, let them wait as on Linux, and end the guest whole #582 meet the signal code here.
  • linux 75d14d10a: the family answers lifecycle and signal calls before its general dispatch.
  • linux 240ee685e: a new thread starts with its creator's signal mask.
  • linux 244566980: the serve loop's wait ends at the next timer or signal wait, not only at the next parked call's deadline.
  • linux 2c6af3098: a thread's death by a signal keeps Linux's wait status.
  • linux 1b1e6aee9: SIGPIPE is raised at a thread whose write answers EPIPE, in both answer paths.
  • linux 3117b517f: /dev/null, /dev/zero, /dev/full and the random devices are served for open, read, write, seek, stat, statx, poll and epoll.
  • linux 9b551ebf6: signalfd4 and signalfd are served, readable through read, readv, poll and epoll.
  • mk 408e336f7: the store packs only the Linux guests LINUX_GUEST_SET names. The whole set is 19.8 MB and the vfs loads at most 16 MiB.
  • linux 15067eb12: goexec checks the devices and os/exec, and alarm checks signalfd.

Evidence

x86_64, q35 under TCG, one vCPU, built from the committed tree. Each boot packed the guest and what it starts through LINUX_GUEST_SET.

guest proves line calls served unserved triple faults
leaderexit the leader's SYS_exit leaves the worker running, and the status is the worker's [C] leaderexit PASS: worker ran 1 s after the leader's exit, [LINUX] process 2 exited, status 42 15 0 0
goexec Go's syscall.ForkExec and os/exec (cmd.Output, and Run with nil streams through /dev/null): statuses 0 and 42, ENOENT for a missing program; /dev/null, /dev/zero, /dev/full [GO] goexec PASS: 10 of 10 parts held 1515 1 (pidfd_open, which Go probes and falls back from) 0
sigchld SIGCHLD caught with pid, CLD_EXITED and status 7; waitid P_ALL, WNOHANG, WNOWAIT; wait4 WNOHANG, WUNTRACED, EINVAL, death by SIGTERM; posix_spawn, system through busybox sh (status 42) and popen through a pipe; ECHILD [C] sigchld PASS: 13 of 13 parts held 365 0 0
sigpipe caught: the handler runs and the write answers EPIPE; ignored: EPIPE; default: the writer dies of signal 13 [C] sigpipe PASS: 3 of 3 parts held 33 0 0
alarm alarm(1) ends sleep(10) at about 1010 ms; a 200 ms setitimer fires three times through pause in 606 ms; sigsuspend and sigpending; the handler's ucontext at Linux's offsets; sigqueue into sigtimedwait; a sigtimedwait timeout; a POSIX timer three times; signalfd: EAGAIN, poll, the record, the sigqueue value [C] alarm PASS: 15 of 15 parts held 89 0 0
gohello regression [GO] hello PASS 270 0 0
goconc regression [GO] conc PASS: 8 goroutines summed 3199960000 251 0 0
cthreads regression [C] cthreads PASS: 8 pthreads joined, summed 3199960000 87 0 0
threadfault regression, with its death line [LINUX] process 2 ended by signal 11 7 0 0
gopoll regression [GO] poll PASS: 3 parts 357 0 0
cwait regression [C] cwait PASS: 16 parts 345 0 0

gopreempt still times out, as it does on #582: it needs a signal to reach a thread that is running user code, which this PR does not add.

Each fix was also run with its change removed:

change removed guest result without it result with it
all of this PR (linux/go-guests alone) leaderexit the process ends at the leader's exit, no PASS line, 13 calls PASS, status 42
all of this PR (linux/go-guests alone) sigchld the SIGCHLD part fails, waitid is unserved nr=247 three times, an unknown option is accepted, and it hangs in wait4 on a child killed by SIGTERM (that guest had the first 10 parts) 13 of 13
the route for a leader's plain exit leaderexit process 2 exited, status 0, no PASS line PASS, status 42
the process clone and vfork routes goexec ForkExec cthreads: ... function not implemented, 0 of 3 3 of 3
the fork copy in 1 MiB pieces goexec ForkExec cthreads: ... cannot allocate memory, 0 of 3 3 of 3
the SIGCHLD raise and the waitid route sigchld 4 of 10, unserved nr=247 four times (before the three spawn parts existed) 10 of 10 at the time
the SIGPIPE raise sigpipe 1 of 3: no handler runs and the writer exits 0 3 of 3
a caught signal ending a parked wait alarm sleep(10) returns after 10008 ms, then pause never ends (the boot timed out at 300 s) 1008 ms, 11 of 11
the frame layout (sigcontext at 48 again) alarm 10 of 11: the ucontext part fails 11 of 11
the frame layout (sigcontext at 48 again) host proof the_frame_is_linux_rt_sigframe_byte_for_byte ... FAILED 111 passed
the child's own stack (the fork always given 0) sigchld the posix_spawn part never returns: the child runs on its parent's stack pointer and makes no call (the boot timed out at 400 s after 8 parts) 13 of 13
the timer and signal-wait routes (alarm, setitimer, getitimer, pause, sigsuspend, sigtimedwait, sigpending, sigqueueinfo, timer_*) alarm 2 of 11, 38 calls served and 69 unserved 11 of 11, 0 unserved

The alarm and sigchld counts in the second table are from before their later parts were added, which is why they read 11 and 10 there and 15 and 13 above.

Checks

  • x86_64 kernel crate: 43 warnings, the same set before and after. The personality has 0 warnings.
  • check_stubs, check_allows, check_dark_features and check_syscall_abi print the same as on linux/go-guests, with 0 new sites. check_unreachable prints the same as there too; its one new site against main is has_children, which linux: run Go and musl threads in a guest, let them wait as on Linux, and end the guest whole #582 already reports.
  • Host proofs (capsule_linux_proofs): 116 passed, 0 failed.
  • Every commit builds on its own: cargo check gave 0 errors at each of the 27. The kernel commit was also built and booted in the full image, and the libc commit was checked at that commit.
  • Every file this PR adds is at most 75 lines. Files that were already longer and gain a line or two keep their length: call/mod.rs, file/mod.rs, file/epoll.rs, serve/family_waits.rs, call/spawn/clone.rs, libc src/lib.rs, the kernel's dispatch/process.rs and Guests.mk. The new guests' rules live in the new LifeGuests.mk. Comments added are /* */, and no #[allow] was added.
  • About 60 boots in all. One stalled in the kernel at Microkernel init, before any capsule ran, and passed when run again (alarm).

Not done here

  • A signal sent to a thread that is running user code waits until the thread next makes a call. Delivering it sooner needs a kernel entry that stops a running thread for its supervisor. gopreempt shows the gap: SIGURG is raised at its spinning thread and never lands.
  • No guest is ever stopped. SIGSTOP and the default action of SIGTSTP, SIGTTIN and SIGTTOU are refused by name, so WUNTRACED and WCONTINUED never report anything. Stopping needs the same kernel entry.
  • ITIMER_VIRTUAL, ITIMER_PROF and timers on a CPU-time clock are refused by name, and wait4 and waitid write rusage as zeros with a log line saying so. The kernel counts ticks in 256 slots by pid, so it cannot say how much time one guest used.
  • The whole Linux-guest store is 19.8 MB, over the 16 MiB the vfs loads. LINUX_GUEST_SET packs a subset until that limit is decided.
  • A signal unblocked by rt_sigreturn waits for the thread's next call instead of being delivered on the way out.
  • Timers and sigtimedwait deadlines are kept in milliseconds, the finest step the family's clock keeps.
  • A signalfd read through readv, or polled, looks at the process and its first thread. read looks at the reading thread, as Linux does.
  • clone with CLONE_VM but without CLONE_VFORK or CLONE_THREAD is refused by name with ENOSYS.

The ucontext in the frame a handler is entered through put the sigcontext
at offset 48 where Linux has it at 40, and the mask right after the 18
register words where Linux has uc_sigmask at 296. A handler that only
returned through rt_sigreturn never noticed, since the frame was read back
with the same wrong offsets. A handler that reads its ucontext did: Go's
preemption handler reads the interrupted rip and rsp from it and saw every
register one word off.

The frame is now struct rt_sigframe as Linux lays it out: pretcode, then
the ucontext with uc_stack at +16, the sigcontext at +40 with oldmask as
its 22nd word, uc_sigmask at +296, then the 128-byte siginfo at +312. The
handler is entered with the interrupted registers left as they were, rax
zero and the direction, resume and trap flags cleared. sigframe holds the
layout and the Entry a handler is entered with, which can also place the
frame at the top of an alternate stack and carries a whole siginfo;
sigframe_build writes the frame and sigframe_read reads it back. The host
proofs mount all three and check the offsets against Linux's.
The numbers a guest sees for its processes and threads were kept inside
the family's PidNs, which only the serve loop holds. A handler writing a
pid into memory rather than returning it, as a siginfo or a waitid answer
does, had no way to translate it, and would have handed the guest a
kernel pid it cannot name.

A personality hosts one family, so the numbering now lives in one place,
pid_space, that PidNs fronts, and serve exports guest_pid and kernel_pid
for any handler to use. The numbers given are the same as before. The two
exports have no caller yet; the siginfo and waitid work that follows uses
them.
rt_sigprocmask ignored the mask it was given and read back an empty one,
and sigaltstack ignored the stack. A handler was entered on the thread's
own stack whatever SA_ONSTACK said, which crashes Go as soon as a signal
lands on a small goroutine stack. sa_mask, SA_NODEFER and SA_RESETHAND
were ignored, so a handler could be re-entered by its own signal. The
siginfo carried only the signal's number, and rt_sigreturn restored the
registers but not the mask. A signal that could not be delivered was
raised again and again, and one the process did not catch waited forever.

Each thread now has its own mask and alternate stack. A new thread's mask
is its creator's, a forked child's is the forking thread's, and exec
resets caught handlers to their default while keeping the mask and what is
pending. A handler is entered on the alternate stack when it asked for
one, with the thread's mask plus sa_mask plus the signal itself unless
SA_NODEFER, and SA_RESETHAND puts the default back. The frame carries a
full siginfo with the sender's pid in the guest's numbering. rt_sigreturn
restores the mask and the alternate stack from the ucontext, and a frame
that cannot be written or read back ends the process with SIGSEGV, as on
Linux. A signal with no handler does its Linux default: ignore, or end the
whole process. Standard signals coalesce, realtime ones queue, and the
lowest number is taken first.

Signals keeps the state the later calls need, so some fields and types
are written but not yet read. born has its caller in clone, which is kept
as it is here.
kill, tkill and tgkill looked up the target's disposition in the calling
process and queued the signal in the caller's own queue under the
target's pid. A signal sent to a child was therefore judged by the
parent's handlers and never seen by the child. A signal whose default is
fatal killed only the one thread it named, and left the rest of its
process running with the process never marked as ended.

A signal for the caller's own process is now queued there, for a thread
or for the whole process, and taken by whichever thread may take it. A
signal for any other process leaves through an outbox with the caller
parked; the family raises it in every process the target names, a pid, a
thread, a process group or all of them, and answers the caller 0 or
ESRCH, counting an ended child not yet waited for as Linux counts a
zombie. The target's own disposition decides what happens, and a fatal
default ends the whole process. rt_sigqueueinfo and rt_tgsigqueueinfo
carry the caller's siginfo and value, with Linux's rule that only a
process signalling itself may claim a kernel si_code. pid_map translates
the pids these calls name.

The new routes in route_life are not asked yet, so kill keeps a form for
callers that cannot park, and tgkill_from and the sigqueue calls have no
caller until they are.
A plain exit from the leader was treated as exit_group: the whole process
ended with the leader, taking every other thread with it, where Linux
ends only the leader and lets the process run until its last thread
exits. A program whose main thread leaves with a plain exit while other
threads still run was cut short. The status a process ended with was
kept as a bare exit code, or as 128 plus the signal, so a waiting parent
could not tell an exit from a signal.

exit_one ends only the calling thread. The leader cannot be killed while
other threads run, since every peer call names its pid as the address
space, so it clears its tid word, wakes its joiner and stays parked as a
zombie until the last thread exits; that thread's code is the process's
status. What a process ended with is now kept as Linux's wait status
word: the exit code in the second byte, or the signal's number in the
first. The reap logs which it was, and the first guest's code is taken
from it, as 128 plus the signal for a signal. A fault reported by the
kernel is still kept as 128 plus the signal, which reads as that signal.

exit_one is routed from route_life, which dispatch does not ask yet.
fork mapped and copied each span of the parent into the child with one
peer call. The kernel maps and copies at most MAX_SPAN, 1 MiB, in one
call and refuses more, so a parent with any larger span could not fork:
a Go process, whose heap arenas are far larger, got ENOMEM from fork and
from the clone os/exec uses.

A span is now mapped and copied in pieces of at most MAX_SPAN each, with
the same protection, so a span of any size crosses whole.
clone without CLONE_THREAD answered ENOSYS, and vfork was served as a
plain fork that let the parent run on at once. Go's os/exec starts every
child with clone(CLONE_VFORK|CLONE_VM|SIGCHLD) on no new stack, so no Go
program could run another program.

clone without CLONE_THREAD now makes a process copied from its parent, as
fork does, and honours CSIGNAL, CLONE_VFORK, CLONE_PARENT_SETTID,
CLONE_CHILD_SETTID and CLONE_CHILD_CLEARTID, writing the tids in the
guest's numbering. A copy stands in for CLONE_VM with CLONE_VFORK, since
such a child may only exec or exit and the parent is held until it does,
so neither can see the difference. Anything else is refused by name.
vfork, and clone with CLONE_VFORK, park the calling thread until the child
execs, answered from the successful exec, or ends, answered from the
reap, in both cases with the child's pid. A new process takes its signal
state from the thread that made it and keeps the signal it raises at its
parent when it ends. These routes are in route_life, which dispatch does
not ask yet, so they have no caller until it does.
wait4 held one parked waiter per process, and a second waiting thread
replaced the first. It wrote every status as an exit code, so a child
ended by a signal read as a normal exit, and it treated a wait by process
group as a wait for any child. It never wrote the rusage, ignored its
options, and waitid was not served. Nothing was raised at a parent when a
child ended, so a SIGCHLD handler never ran.

wait4 and waitid now park the calling thread, any number of them, and
the family answers each once it can: at once with a child that has ended
and fits, 0 under WNOHANG, ECHILD when no child could ever fit, and
otherwise when one ends. A wait can name any child, one pid, the caller's
own group or another group, with __WALL and __WCLONE choosing by the
signal a child raises. The status word is the one the child ended with,
waitid reports CLD_EXITED or CLD_KILLED in a siginfo and can look without
reaping under WNOWAIT. The rusage is written as zeros, since no CPU time
is reported for a guest, and the log says so once per family run in a
"[LINUX] rusage: ..." line, so the zeros are not read as a measurement. A
child that ends raises its exit signal at its parent, SIGCHLD unless clone
named another, with its pid and status, and a parent that ignores SIGCHLD
or set SA_NOCLDWAIT has its children reaped as they end. Reaping goes
round again while answering ends anything more.

wait4 keeps its caller in dispatch. waitid is routed from route_life,
which dispatch does not ask yet. The guest's old parked wait4 field is no
longer read and is left for a separate change to remove.
A write to a pipe no process can read answered EPIPE and nothing more.
Linux raises SIGPIPE at the writing thread first, so a program that
relies on the default to end it, as a shell pipeline does, kept writing
into nothing, and a SIGPIPE handler never ran.

sigpipe raises SIGPIPE at the writing thread, or at the process when
given 0, with SI_USER and the writer's own pid, as Linux's pipe_write
does. Caught, the handler runs over the EPIPE the write answers; ignored,
only EPIPE is seen; at its default, the process ends. It is exported for
the pipe write to call before it answers EPIPE; that write belongs to a
change not made here, so sigpipe and SIGPIPE have no caller yet.
A signal was only acted on when its thread next returned from a call. A
thread parked in nanosleep, a futex wait, a pipe read or wait4 kept the
signal until the wait ended by itself, so a handler for a signal sent to
a thread in sleep(10) ran up to ten seconds late, and a signal whose
default is fatal did not end a process whose threads were all parked.

After every answer the family now looks at what each process has pending.
An uncaught signal does its default at once, as Linux does when it is
sent. A caught one is taken by a parked thread that does not block it:
the thread leaves its wait, which is then never answered, and enters the
handler over the parked call. The call answers EINTR, or under SA_RESTART
runs again once the handler returns for the calls Linux restarts: reads
and writes, the socket calls, wait4, waitid and an untimed futex wait. An
interrupted relative nanosleep writes the time it had left.
alarm, setitimer and getitimer were not served, so a program that bounds
a wait with SIGALRM, or keeps time with a periodic ITIMER_REAL, never saw
the signal.

ITIMER_REAL is now kept per process on the family's monotonic clock, in
milliseconds, the finest step it keeps. alarm arms it once and answers
the whole seconds the last one had left, rounded as Linux rounds them;
setitimer arms it once or with a period and hands back the old value;
getitimer reads what is left. When it runs out the family raises SIGALRM
at the process and moves a periodic timer past every period that ended.
Exec keeps it and fork does not, as on Linux. ITIMER_VIRTUAL and
ITIMER_PROF count CPU time, which the kernel does not report to a
supervisor: arming one is refused by name and reading one reports it
disarmed. The host proofs check how a periodic timer moves on and that a
one-shot one stops. Signals::next_due gives the nearest deadline for the
serve loop's wait to end by; that wait is computed outside this change,
so it has no caller yet and a timer fires when the loop next wakes.
pause, rt_sigsuspend, rt_sigtimedwait and rt_sigpending were not served,
so a program could not wait for a signal, take one without a handler,
or see what it held blocked.

pause parks the thread until a handler runs and answers EINTR through
it. rt_sigsuspend does the same under the mask it is given, and the
handler's return puts the old mask back, since the frame carries it.
rt_sigtimedwait takes a pending signal of its set at once, or parks until
one arrives, answered with its number and siginfo, or until its timeout
runs out, answered EAGAIN. rt_sigpending names what is pending and
blocked for the calling thread. The sigtimedwait deadlines join the
timers in next_due. The three waits are routed from route_life, which
dispatch does not ask yet, so they have no caller until it does.
timer_create, timer_settime, timer_gettime, timer_getoverrun and
timer_delete were not served, so a program using POSIX timers failed at
the first call.

A process can now make timers on the realtime, monotonic, boottime,
alarm and TAI clocks, numbered from 0 as Linux numbers them. A timer
raises its signal at the process, or at the one thread SIGEV_THREAD_ID
names, with SI_TIMER, its id and sigev_value; a NULL sigevent means
SIGALRM with the id as its value, and SIGEV_NONE raises nothing.
timer_settime arms it relative or absolute, once or with a period, and
hands back the old setting. An expiry while the timer's signal still
waits counts as an overrun, carried by the signal when it is taken and
reported by timer_getoverrun. Deleting a timer drops its queued signal,
and a timer whose signal was discarded may queue again. Exec drops every
timer and fork gives the child none. Timers on a CPU-time clock are
refused by name, since the kernel does not report a guest's CPU time to
its supervisor. The host proofs check the overruns a late timer counts.
MkForeignFork always started the child on its parent's stack pointer.
A supervisor serving a clone that makes a process on a stack of its own,
as musl's posix_spawn does, had no way to ask for that stack, so it could
only refuse the call.

MkForeignFork now takes the child's stack pointer as its second argument.
Zero keeps the parent's, as every caller passes today, and anything else
is where the child starts; a stack outside the user half is refused with
EINVAL. The parked frame is read where it was before, only without the
one-line wrapper around it.
nonos_libc could only ask MkForeignFork for a child on its parent's stack,
so a supervisor had no way to use the stack argument the kernel now takes.

mk_foreign_fork_at passes a stack pointer for the child, and
mk_foreign_fork is that with zero, as before. Both live in foreign_fork,
next to the rest of the foreign calls, and are exported as the other
foreign calls are.
musl's posix_spawn, system and popen clone a vfork child onto a stack of
its own, not the parent's. That clone was refused with ENOSYS, so each of
them failed and no C program could start another through them.

A clone that makes a process now hands its stack to the fork, and the
child starts on it with the rest of its registers copied from the parent,
as Linux starts it. With no stack named, the child runs on its parent's,
as before.
Nothing in the guest set exercised a leader's plain exit, a program
starting another, a parent told about its children, SIGPIPE, or a signal
arriving while a thread waits, so each of the changes before this one
could only be judged by reading it.

Five guests now check them against what Linux does, each printing a line
per part. leaderexit ends its leader with a plain exit and expects its
worker to run on and set the status to 42. goexec starts programs from
Go through syscall.ForkExec and os/exec, which use clone(CLONE_VFORK),
and reads each child's status. sigchld catches SIGCHLD and checks its
siginfo, then waitid and wait4 with their options and ECHILD, then
starts programs through posix_spawn, system and popen. sigpipe checks a
write to a widowed pipe caught, ignored and at its default. alarm checks
that SIGALRM ends a sleep early, then setitimer, getitimer, pause,
sigsuspend with sigpending, sigqueue into sigtimedwait, a timed out
sigtimedwait and a POSIX timer, and reads its own ucontext at Linux's
offsets. LifeGuests.mk, which Guests.mk includes, builds the C ones with
musl and adds all five to the guest set.
The base gained descriptor waits, eventfd, pipe ends with EPIPE, futex
timeouts and the scheduler calls while this branch taught signals and
process lifecycle to Linux. Both changed the same lists of parked threads.

A caught signal now ends every wait the base can park a thread in: the
sleepers, the futex waits and their timeouts, and each blocked descriptor
wait, through forget_waits, so a pipe read, an eventfd read, epoll_wait,
poll and select end with EINTR or restart as Linux decides. The base's own
tgkill row in pid_map also maps rt_tgsigqueueinfo.
The calls that can leave a thread parked in a new way, a leader's plain
exit, a process clone, vfork, waitid, kill across the family, pause,
sigsuspend, sigtimedwait and sigqueueinfo, were served in route_life,
but nothing asked it: dispatch answered them the old way, or named them
unserved. The family now asks route_life for each trap first, which
counts the call, and dispatch answers everything else. wait4 now parks
in the signal state's list, so the guest's single waiting slot is gone.
A thread clone made started with every signal unblocked, whatever its
creator had blocked. Go and musl block every signal around clone so a
new thread cannot take one before its runtime is ready; here it could,
on its first call. The new thread now takes its creator's mask, as
Linux's clone gives it.
The serve loop waited up to 250 ms unless a sleeper was due sooner, so
SIGALRM, a POSIX timer's signal or a sigtimedwait timeout could come up
to 250 ms late. The nearest of those deadlines now ends the wait too.
A process whose thread died on a signal was given 128 plus the signal
as its status, a shell's convention, which a parent reading it with
WIFSIGNALED saw as an exit with a code of 139 or 137. It is now the
signal's number, as Linux reports it.
A write to a pipe no process can read answered EPIPE and nothing more,
so a program relying on SIGPIPE's default to end it, as a shell
pipeline's writer does, kept going, and a SIGPIPE handler never ran. A
write, writev, or a send without MSG_NOSIGNAL that answers EPIPE now
raises SIGPIPE at the writing thread before its answer, whether it is
answered at once or after waiting, so the thread takes it on its way
out as on Linux: caught, the handler runs over the EPIPE; ignored, only
EPIPE is seen; at its default, the process ends.
No path under /dev existed, so Go's os/exec, which opens /dev/null for
every stream a command leaves nil, failed with ENOENT before it forked,
and so did any program writing to /dev/null. The five character devices
every Linux program may assume are now descriptors this capsule answers
itself: null reads end of file and takes every write, zero reads zeros,
full reads zeros and refuses writes with ENOSPC, random and urandom
read the kernel's random bytes. fstat, stat and statx report them as
character devices with Linux's major and minor numbers, lseek answers
0, and epoll refuses null, zero and full with EPERM as Linux's does.
signalfd4 was not served, so a program that reads its signals from a
descriptor instead of running handlers, as event loops built on epoll
do, could not start. A signalfd now takes a mask; a read by a thread
takes the pending signals of that mask for the thread or its process as
signalfd_siginfo records, laid out as Linux's, answers EAGAIN on a
non-blocking descriptor when none is pending, and otherwise parks until
one comes. poll and epoll see it readable once a signal of its mask
waits. A second signalfd4 on it changes the mask, and a fork gives the
child a copy.
The Linux-guest test store packs every guest, and vfs loads a store of
16 MiB at most. With the waiting guests and the lifecycle guests
together it holds 19.8 MB, and nonos-store-pack refuses it, so no guest
boots. LINUX_GUEST_SET, when given, names the guests to pack; left
empty, every guest is packed as before. Each guest is still built,
signed and enrolled.
goexec now leaves each os/exec command's streams nil, as most Go
programs do, so cmd.Output and Run open /dev/null, and checks the
devices directly: /dev/null stats as a character device and reads end
of file, /dev/zero reads zeros, /dev/full refuses a write with ENOSPC.
alarm now checks a signalfd: EAGAIN with nothing pending, readable in
poll once SIGUSR1 waits, the record's signal and pid, and a sigqueue
value through a blocking read.
@eKisNonos
eKisNonos changed the base branch from main to linux/go-guests September 29, 2026 09:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant