Skip to content

linux: run Go and musl threads in a guest, let them wait as on Linux, and end the guest whole - #582

Closed
eKisNonos wants to merge 254 commits into
mainfrom
linux/go-guests
Closed

eKisNonos wants to merge 254 commits into
mainfrom
linux/go-guests

Conversation

@eKisNonos

@eKisNonos eKisNonos commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Targets main. It also carries the 209 commits of #567 (linux/store-install), which come before it and hold the Linux personality these commits change, so it merges after #567 and shows those commits until #567 lands. What it adds is thirty-six commits.

Before this, a Linux-guest test image started no guest at all. Once it did, no guest could run a second thread: Go's threads called address zero, musl's pthread_create failed outright, and a thread's exit returned to it. A guest that did get threads reset the machine when it exited, left its threads running, and hung if one of them faulted. And nothing could wait: eventfd2 was not served, so every Go program with a timer died at its first one, while epoll_wait and a futex with a timeout returned at once or never.

What Linux guests can do now

  • Go programs start their runtime threads and run to a clean exit.
  • musl programs create pthreads, lock mutexes and join every thread.
  • A thread that faults ends the whole process, as on Linux; nothing is left running or waiting.
  • A fork from any thread gives the child that thread's registers and thread pointer.
  • Go's timers and poller work: time.Sleep, a timer woken early through Go's eventfd, and a pipe read through the poller to end of file.
  • musl programs wait as on Linux: timed condition variables, blocking eventfd reads, epoll_wait, poll and select timeouts, non-blocking pipes, several readers or a writer waiting on one pipe, edge-triggered epoll, and timerfd one-shot, periodic and absolute.
  • The descriptor ioctls, the scheduler calls, epoll_create, epoll_pwait2 and tgkill are served as Linux serves them to an unprivileged process on one CPU.

What changed

The boot guest starts

  • mk bf9d9dbf3: the Linux-guest test image (NONOS_LINUX_GUESTS=1) builds microkernel-desktop-gui with nonos-stark-attest, without first-boot setup. Under the setup profile every app, the personality included, waits for setup to exit, and setup waits for keys, so the unattended image never started its guest: 0 [APP-LINUX] or [LINUX] lines in 180 s. Every other build of nonos-mk-desktop-gui-prod is unchanged.

Threads start, run and end

  • foreign 3fc1ac19d: a clone child starts on a copy of the calling thread's registers, with rax 0, the new stack, and the parent's thread pointer unless a TLS is named. MkForeignThread takes the calling thread as a fifth argument; it must be supervised by the caller and in the same thread group as the target, so no guest receives another's registers. Zero keeps the fresh start. Go's child calls r12 and musl's calls r9, so both called address 0 before. A TLS is taken only when CLONE_SETTLS is set. abi/syscalls.toml records the new argument.
  • linux addec49af: clone writes the new tid where CLONE_PARENT_SETTID and CLONE_CHILD_SETTID ask. musl's thread-list lock stores its owner's tid, and a thread that never learned its tid locked the list as 0, which the lock reads as free.
  • linux 1e6b8a673: mprotect commits a PROT_NONE reservation it is asked to open. musl reserves every pthread stack with PROT_NONE and opens the usable part with mprotect; MkPeerProtect reprotects only pages that exist, so pthread_create failed. A backed region is reprotected as before; a reservation is backed with the asked protection, keeping any page the guest already touched, and recorded as backed so fork copies it. A span outside every region is refused with ENOMEM, as on Linux.
  • linux 10947aa96: a thread's exit ends the thread and never returns to it. The word it named through CLONE_CHILD_CLEARTID or set_tid_address is zeroed and one waiter woken, as Linux does. musl names its thread-list lock there and pthread_join waits for that lock, so every join hung before. Go's exitThread falls into INT3 when exit returns, which the fault report below would have turned into the end of the process.

A process ends whole

  • exit 5ab0a0912: a thread group's page tables stay until its last member leaves the process table. Release freed them when the leader was finalized, under a Go thread still running on them, and the machine reset: CR3 was the guest's, FS held the Go thread's TLS, CR2 was the IDT entry for vector 8. The tables now pass to a process still running on them, whose capability token is rebound to the ASID it now owns.
  • kill 874f3ecfc: MkKill admits the caller the foreign registry names as the target's supervisor. A thread's parent is its leader, not the supervisor, so every thread kill at a guest's exit returned EPERM and the threads ran on. A refused kill now prints [LINUX] kill refused: pid <n> outlives its process, errno <e>.
  • foreign 2311afc39: when a guest thread ends on a signal, and only when it ends itself so a supervisor's own kill cannot loop back, the kernel leaves a one-way notice for its supervisor, collected on its next MkForeignWait as a frame numbered FOREIGN_NR_DIED. The kernel reports; what a dead thread means is the personality's policy.
  • linux 959b57f98: the personality ends the process on that notice, as Linux ends a thread group on an unhandled fault, and reports the signal as 128+signo.

Fork

  • foreign ffc1b542b: a fork copies the calling thread's parked frame, and the kernel carries that thread's own thread pointer to the child. Before, every fork took the leader's registers and the personality's single fs_base, which is only the last thread to set one.

Waiting, as Linux waits

  • linux 69f3b48d4: a pipe has Linux's ends. The family notes, each time it lends its pipes to the guest being answered, which ends are still open anywhere in the family. From that note: a read end is readable while it holds bytes and hung up once no write end is left; a write end is writable while there is room and in error once no read end is left. An empty pipe with no write end reads as end of file, and a write no reader can take is refused with EPIPE. poll, epoll and select report hang-up and error unasked. Before, every pipe was readable and writable to all three.
  • linux 3e48edbaf: a descriptor keeps O_NONBLOCK. F_SETFL sets it, F_GETFL reports it with the access mode, pipe2 takes it, dup and fork carry it, and a non-blocking read of an empty pipe answers EAGAIN. Before, F_SETFL was dropped, which is how Go's runtime prepares every pipe for its poller.
  • linux bfb07892f: eventfd2 and eventfd are served. The counter is kept with the pipes as the family's, so dup, fork and every thread reach one. EFD_SEMAPHORE, EFD_NONBLOCK and EFD_CLOEXEC as Linux; all ones and a short buffer EINVAL. Go's runtime makes one at its first timer and threw when it could not.
  • linux 06a266ac5: epoll_wait and epoll_pwait wait for their timeout (an int of ms; negative waits until ready). A blocking eventfd read or write and a write to a full pipe wait too. A waiting call is left parked in its trap and tried again after every answer and at its deadline, so a write from another thread ends the wait at once. A wait on a socket or timer is looked at every 10 ms. A pipe write of up to PIPE_BUF goes in whole or waits. maxevents of zero or less is EINVAL.
  • linux 32d2a1a06: EPOLLET entries are reported as readiness rises and re-armed by an EAGAIN; EPOLLONESHOT reports once. epoll_ctl refuses as Linux: EBADF, EPERM for a regular file or directory, EINVAL, EEXIST, ENOENT. Go registers every pipe and socket with EPOLLET and EPOLLOUT, so a level-triggered report of an always-writable end would answer its poller at once, every time.
  • linux aa0087d75: a futex wait ends with ETIMEDOUT at its timeout. FUTEX_WAIT_BITSET (absolute, MONOTONIC or REALTIME), FUTEX_WAKE_BITSET, FUTEX_REQUEUE and FUTEX_CMP_REQUEUE are served. The value is compared as 32 bits, so musl's sign-extended -1 matches. musl's timed waits and condition variables, and Go's sysmon, use these.
  • linux b1ead15c5: two proof guests, gopoll and cwait (below).

More waiting, timers, descriptors and threads

  • linux b716a34af: any number of threads can wait to read one pipe. A read parked in one slot per process, so a second reader of an empty pipe took the slot and the first was never answered. Pipe reads now wait with the other parked calls; the family settles them after it reaps, so a writer that left with its process reads as end of file at once. A zero-byte read answers zero.
  • linux 439ed66c6: poll, ppoll, select and pselect6 wait for their timeout (poll's int of ms, ppoll's and pselect6's timespec, select's timeval; null or negative waits until ready; bad fields EINVAL). poll(NULL, 0, ms) sleeps. select writes its sets back only when something is ready, copies only the longs nfds covers, always empties the exception set, and empties all three when its time runs out. A negative poll descriptor is never ready.
  • linux 85b3d7724: a closed descriptor leaves every epoll interest list, as on Linux, so its number can be added again; before, that add was EEXIST and the old entry reported the new file under the old token. dup2 closes an open target first (a file's buffered bytes were lost, a socket's handle kept) and refuses a number past the table.
  • linux b9efc7f96: timerfd has Linux's semantics and is kept with the family like the eventfd counter, so dup and fork share it. A periodic timer keeps firing and a read returns how many times it fired; the clock is Linux's set (else EINVAL); TFD_NONBLOCK, TFD_CLOEXEC and TFD_TIMER_ABSTIME are honoured; the old setting is written when asked; timerfd_gettime is served; a blocking read waits for the firing; a wait watching a timer is woken when it fires.
  • linux 262c06280: FIONREAD (a pipe's bytes, a file's remainder), FIONBIO, FIOCLEX and FIONCLEX. Any other request is ENOTTY, as for a descriptor with no terminal or device behind it.
  • linux 8f778075b: a dup'd or inherited epoll descriptor carries the interest list it had; before, it started empty.
  • linux ff778c704: sched_getscheduler, sched_setscheduler, sched_getparam, sched_setparam, sched_get_priority_max and _min, sched_setaffinity, epoll_create, epoll_pwait2. Every thread is SCHED_OTHER at zero; SCHED_FIFO and SCHED_RR are EINVAL outside 1 to 99 and EPERM inside it, the range checked first as Linux checks it; a mask without the one CPU is EINVAL; pid arguments are mapped like kill's.
  • linux 9039b893d: tgkill's thread group and thread are mapped into the family's pids. gettid and getpid answer with the family's numbers, so tgkill(getpid(), gettid(), sig), which Go's runtime and glibc's pthread_kill use, was ESRCH.
  • linux 84e10f538, 7c3967f18, 468f5aa97, 2f0db3b25, a045bec06, 66b55faf8, afd96ef58, 24cafe103: the guest gopreempt, and seven more cwait parts, one per change above; cwait now runs every part and names each one that fails.

Cleanups

  • linux 289272374: delete call/clone.rs, a copy left behind when clone moved to call/spawn/; no module included it.
  • mm ce395f7ff: the [PF] demand fill line is printed only after a fill, not before the handler refuses the null page or the kernel half.

Evidence

q35 under TCG, one vCPU. Each guest is the boot guest of its own store.

Threads, faults and exit (the first thirteen commits)

guest proves line calls served unserved triple faults
gohello Go runtime threads, GC, clean exit [GO] hello PASS 271 0 0
goconc 8 goroutines on runtime threads [GO] conc PASS: 8 goroutines summed 3199960000 236 0 0
cthreads 8 musl pthreads, a mutex, 8 joins [C] cthreads PASS: 8 pthreads joined, summed 3199960000 87 0 0
threadfault a worker faults while main waits [LINUX] guest thread 51 ended on a signal; ending the process, then [LINUX] guest exited 7 0 0

threadfault's worker runs on a stack mapped read-write outright through clone(), so it does not depend on the pthread path. Its SURVIVED line, printed only if main outlives the faulting thread, never appears.

Waiting (the next seven commits)

Both guests pass on host Linux, which is what they are measured against. On one build carrying all twenty commits:

guest proves line calls served unserved triple faults
gopoll Go's timer and poller: a 50 ms sleep, a 10 ms timer set while the poller waits on a 3 s one, a pipe read through the poller [GO] poll PASS: 3 parts, 3454ms in all: slept 69 ms, the early timer ended in 18 ms, the pipe read to end of file in 339 ms 399 0 0
cwait musl waiting in 9 parts [C] cwait PASS: 9 parts in 951 ms: cond_timedwait ETIMEDOUT after 102 ms, a blocking eventfd read woken after 147 ms, epoll_wait timed out after 100 ms and woken after 153 ms, POLLHUP on a closed pipe, a full-pipe write waited 148 ms, edge-triggered counts as Linux 196 0 0

The same build ran the first four guests again: gohello 283 calls, goconc 253, cthreads 87, threadfault 7 with its death line; 0 unserved in each.

More waiting, timers, descriptors and threads (the next sixteen commits)

On one build carrying all thirty-six commits:

guest proves line calls served unserved triple faults
cwait 16 parts, the 9 above and 7 more [C] cwait PASS: 16 parts in 2313 ms; SCHED_FIFO at priority 1 errno 1 and a CPU-1-only mask errno 22; a 50 ms periodic timer fired 3 times in 180 ms, a one-shot read waited 103 ms, an absolute timer woke epoll after 105 ms 345 0 0
gopoll as above [GO] poll PASS: 3 parts, 3437ms in all 386 0 0
gopreempt a goroutine spinning with no call beside main, on one P no line in 360 s: this set does not reach a running thread, the gap named below 0

The same build ran the regression set: gohello 282 calls, goconc 248, cthreads 87, threadfault 7 with its death line; 0 unserved in each.

Each fix was also run with its change removed, on the same guest:

change removed guest result without it result with it
the supervisor clause in MkKill goconc 3 kill refused ... errno 1 lines; guest threads fault after guest exited 0 and 0
the kernel's death notice threadfault SURVIVED printed, no death line death line, no SURVIVED
the clear-tid write and wake cthreads no PASS, no FAIL, no exit in 241 s after the guest started PASS
the mprotect commit (the pthread build of threadfault, before the fix) threadfault [C] threadfault FAIL: no worker pthread_create succeeds (cthreads)
a parked wait tried again when something changes (only at its deadline instead) gopoll a 10ms timer set under a 3s one took 2823ms, FAIL 18 ms, PASS
the same cwait hangs in its blocking eventfd read; nothing after part 3 in 420 s PASS
eventfd2 served gopoll [LINUX] unserved nr=290, then fatal error: runtime: eventfd failed PASS
the futex timeout cwait hangs in cond_timedwait; no line in 420 s ETIMEDOUT after 102 ms
EPOLLET (reported level-triggered instead) cwait FAIL: edge-triggered counts (222, 1): the second look counts 2, not 1 PASS
the same gopoll PASS, with 1282 calls served: the poller answered at once while the pipe was open 399 calls
a pipe's hang-up cwait FAIL: end of file after the write end closed (0, 0): poll sees nothing POLLHUP
O_NONBLOCK kept by fcntl and pipe2 cwait hangs in the non-blocking pipe read; nothing after part 6 in 420 s EAGAIN at once
the personality before the next sixteen commits cwait eight parts pass, then it hangs in pipe-readers: the first of two blocked readers is never answered PASS
poll, ppoll and select waiting (answering at once instead), a closed descriptor leaving epoll, periodic timers, FIONREAD, sched_getscheduler, the tgkill mapping, all six removed on one build cwait FAIL: 6 parts failed, 10 passed: poll and ppoll returned after 2 ms, the re-add was EEXIST, the periodic timer counted 1, FIONREAD failed, [LINUX] unserved nr=145, tgkill ESRCH 16 parts pass
the epoll list carried through fork cwait FAIL: epoll list through fork (13, 256): the child finds nothing ready and exits 1 PASS

Each of these ran on its own build with only that change removed, except the six that fail different cwait parts, which were removed together. gopoll passes with the hang-up removed, as its writer closes before its reader waits, and with O_NONBLOCK dropped, as Go's pipe reads then park instead of polling.

The store settled in 12.8 to 84.2 s across the 22 guest boots for the first thirteen commits, in 38.0 to 104.8 s across the 16 for the next seven, and in 17.2 to 117.7 s across the 10 for the sixteen after; the personality waits up to 300 s.

Checks

  • x86_64 kernel crate: 0 errors; 43 warnings, the same set before and after.
  • check_stubs, check_allows, check_dark_features: 0 new sites. check_syscall_abi: 111 published syscalls reach a handler.
  • check_unreachable: 1 new site, has_children, which linux/store-install reports too; this branch has 1257 sites against its 1258.
  • Each of the first 13 commits builds on its own: the kernel (microkernel-core) and the personality were checked at every one, with 0 errors. The next twenty-three change only the personality and the guests; the personality was checked at each with 0 errors and 0 warnings, and each guest passes on host Linux.

Not done here

  • A PROT_NONE reservation is not enforced: the kernel demand-fills any guest page on first touch, so a guard page does not stop a stack overflow.
  • mprotect on a backed region does not update the recorded protection, so a fork after it gives the child the protection the region was mapped with.
  • A fork does not copy a reserved region the guest has touched, since only backed regions are copied.
  • The leader's own exit (not exit_group) still ends the whole process; on Linux the others would run on.
  • A caught signal reaches a guest thread only when it next makes a call. A thread parked in a wait sees it when the wait ends, and a thread running user code never does, so Go cannot preempt a goroutine that spins without a call (gopreempt, above). The kernel change that stops a running thread for its supervisor follows this set.
  • A caught signal sent with kill to a child process is queued in the sender, not the child, so the child's handler never runs; an uncaught one ends the child as it should.
  • A write to a pipe with no reader is EPIPE, but SIGPIPE is not raised.
  • A wait on a socket is looked at again every 10 ms; nothing tells the family sooner that a socket became ready.
  • A dup'd or inherited epoll descriptor gets a copy of the interest list; on Linux both share one, so a change made on one side after the dup or fork is not seen on the other.
  • TFD_TIMER_CANCEL_ON_SET is accepted and has no effect.

senseix21 and others added 30 commits September 24, 2026 22:00
IconId::Store points table.rs at assets/icons/store.a8, which only
existed on the app-store branch, so every build of this branch failed
to read it. Add the mask and its SVG source here so the table stands
on its own.
text::line returns the drawn width, so the bare match evaluated to i32
where the function body expects (), failing the capsule build.
nonos-data/marketplace/index.bin has no make rule, so naming it as a
hard prerequisite failed every build on a checkout without it (CI:
No rule to make target). Wrapping it in $(wildcard) keeps the rebuild
on a newer catalogue where it exists and drops the prerequisite where
it does not.
The branch's manifest predated the switch to in-process Ed25519 and
dropped the dependency while verify/crypto.rs imports it, so the
capsule failed with an unresolved import. Restore main's manifest and
add only the app_skeleton dependency boot_index.rs needs.
Twenty-five syscalls were published as caps = ["valid_token"] while the
cap table demands a hardware, dev-root or time capability. Twenty-two take
any one of Admin or a hardware capability, now published with caps_any; the
dev-root and time calls need one capability, published with caps.
The cap table gates MDRO with MDRQ and MDRC on can_enrol_dev_root. The old
syscall caps check cannot resolve that predicate and wants valid_token, so
this fails it until abi/caps-check-fail-closed lands; the fixed check passes.
Every lane was pinned to -accel hvf -cpu host, which only macOS has, so
no Linux host could boot an image. KVM when /dev/kvm opens read-write,
hvf on macOS, TCG otherwise, the rule the boot matrix already uses; the
display and audio backends follow the host too.
The trailer's magic picked the verifier for every root, so a Pedersen
trailer was checked against the vendor root too. A local root's leaf is a
commitment to a secret this kernel holds; the vendor root's is not, so there
the trailer may no longer choose the weaker proof.
MkLocalSign let a LocalSign holder prove any capability it held, including
LocalSign itself, so one signer could hand out the right to sign. A local
proof now names nothing beyond AMBIENT_CAPS. crypto_proofs checks a minted
trailer and its refusals against the kernel's own verifier files.
An Alpine package index is signed that way, and the capsule offered SHA-1
only without the prefix, so the index could not be checked at all. The
rsa crate rebuilds the whole padded block and compares it. Nothing here
signs, so offering SHA-1 verification mints nothing new with it.
resolve.rs names super::root, which the crate never mounted, so it did not
build and nothing noticed because no workflow ran it. It mounts root.rs,
follows the rename of absolute to visible, adds the /linux confinement and
Phdr::file_range tests, and joins the proof-crate matrix.
A signature over a package covers the compressed bytes of one member, so a
verifier needs to know where each one starts and ends. members() refuses a
file with any byte that belongs to no member, where gunzip ends the stream
and ignores the rest.
The index's .SIGN.RSA entry is checked against Alpine's x86_64 keys by the
crypto service, the package's control member against the index's C: SHA-1,
and its data against the control member's datahash. Only a Verified value
reaches the store, so unauthenticated bytes are refused, not kept unvouched.
Without the operator-key rotation: NOX_OPERATOR_V1 keeps main's key, now
read from .keys/marketplace_operator_ed25519.pub. The capsule embedded an
index nothing built, so tools/nonos-market-index writes one, signed and
verified when the operator seed is present and empty otherwise. nonos-mk
moves to 78dae45 for the zk_trailer_hash field; the icon table is 49 long.
A release naming x86_64-linux counted as having its attestation, so any
release could claim the exemption for itself. It now needs the linux.
namespace too, which is where the store routes it: to the installer that
authenticates the bytes before the machine mints their proof.
Conflicts were both sides adding: the install and app-store capsules sit
together in Cargo.toml, mk and userspace, and init/mod.rs names the
install queue once. The init loop and install queue are taken as the PR
wrote them; the commit after this replaces its wake path.
The scheduler takes every ready process's priority lock from the timer
interrupt. wake.rs and boost_init_for_drain took init's with interrupts on
from syscall context, so a tick inside either spun forever on one CPU. The
install queue now raises through the guarded setter the window queue uses.
MkAppInstall took a package name and the store's own readiness flag. It now
takes a listing and release; init asks the market for readiness and the
release's package hash, and the installer refuses bytes of any other BLAKE3.
The index keeps each record's D: and p: lines, so dependencies come too.
…each

CryptoMachineKey derives the machine key for any label a Crypto holder
names, so a key the kernel keeps for itself needs a label no syscall can
ask for. Kernel labels start with a zero byte, and the syscall refuses any
label that does.
The local signing identity was random each boot, so consent was too. It is
now derived from the machine key, and first-boot setup, which alone holds
EnrolDevRoot, grants the local root as a named step and keeps a token only
this machine can make; later boots restore it. The desktop profile now
includes setup and the market, and builds every capsule it embeds.
MkAppLaunch queues a run for init, which spawns the personality to start
the program the installer recorded outside /linux. Whether it may start is
the exec gate's answer. The store drops its console-code enrolment, which
it never held the capability for, and gains an o key to open.
Every boot resolved validator.nymtech.net through net.dns, so the first
query of a session went out in the clear naming the service in use. The
bootstrap set is pinned by address in this capsule's attested image; the
name is kept for the certificate check and never resolved.
The installer connected to dl-cdn.alpinelinux.org, which the socket service
resolved in the clear. It now takes a mirror by address, Alpine's CDN by
default or NONOS_ALPINE_MIRROR at build, sends Alpine's name as the Host
line, and refuses a host that is not a literal. The signatures still decide.
The generator hashed packages from v3.21 while the installer downloads
from v3.20, so every Linux listing pinned bytes the installer never sees
and the kernel's package-hash check would refuse every install.
The spawn installs an empty token, and nothing on the console said so.
The line reads the bits back from the process table rather than printing
the value it meant to install, so it is evidence and not a restatement.
ioctl answered ENOTTY to every request, so a program asking how many bytes
a pipe or file held, setting a descriptor non-blocking the BSD way, or
marking it close-on-exec without fcntl was refused.

FIONREAD answers what a pipe's read end holds, or what is left of a file
past its offset; FIONBIO sets or clears O_NONBLOCK from the int it is
given; FIOCLEX and FIONCLEX set and clear close-on-exec. A socket's
FIONREAD, and every device request, is still ENOTTY: there is no terminal
or device behind a descriptor here, which is also how isatty says no.
A duplicated or inherited epoll descriptor started with an empty interest
list, so a child that waited on the epoll its parent set up before fork
waited on nothing. The list is now copied with the descriptor. Linux shares
one list between them; the copy holds what was registered at the dup or
fork, and a change made afterwards on one side is not seen on the other.
A fourteenth part: FIONREAD counts three bytes in a pipe, FIONBIO makes its
read end answer EAGAIN once drained, FIOCLEX sets close-on-exec, and a
forked child finds a ready entry in the epoll list its parent built.
Passes on host Linux.
sched_getscheduler, sched_setscheduler, sched_getparam, sched_setparam,
sched_get_priority_max and _min and sched_setaffinity were not served, nor
epoll_create and epoll_pwait2, so a program that sets a thread's policy or
pins it, or a libc that reaches for the older or newer epoll form, died
there.

They are answered as Linux answers an unprivileged process on the one CPU
the guest is shown: every thread is SCHED_OTHER at priority zero; asking
for SCHED_OTHER, BATCH or IDLE at zero is accepted and changes nothing,
since scheduling is the kernel's; asking for FIFO or RR is EINVAL outside
the priority range 1 to 99 and EPERM inside it, the range checked first as
Linux checks it, since no guest holds the privilege Linux asks for; the
priority ranges are Linux's. An affinity mask that includes the one CPU is
accepted and one that leaves it out is EINVAL. A pid argument is mapped
from the guest's namespace like kill's, and one outside the family is
ESRCH. epoll_create checks its size is positive and makes a list;
epoll_pwait2 waits like epoll_pwait with a timespec timeout.
A fifteenth part: the policy is SCHED_OTHER, SCHED_FIFO at priority zero is
EINVAL, SCHED_OTHER is accepted, the real-time range is 1 to 99, CPU 0 can
be pinned, epoll_create refuses a size of zero, and epoll_pwait2 waits its
50 ms timespec. SCHED_FIFO at priority one and a CPU-1-only mask depend on
privilege and CPU count, so they are printed, not checked. Passes on host
Linux.
gettid and getpid answer with the family's own numbers, and kill and tkill
map theirs back, but tgkill's were passed through as they came. A thread
signalling itself with tgkill(getpid(), gettid(), sig), which is how Go's
runtime preempts a goroutine and how glibc's pthread_kill and raise reach a
thread, named numbers no kernel thread has and was answered ESRCH.

Both of tgkill's pid arguments are now mapped, and one outside the family
is ESRCH, as for kill.
A sixteenth part: a caught SIGUSR1 sent with tgkill(getpid(), gettid())
runs its handler, and a thread the family does not have is ESRCH. Passes
on host Linux.
cwait stopped at its first failing part, so a run with more than one
change removed showed only the first. Every part now runs, each failure is
printed, and the last line counts them. A hang still stops it at the part
it hangs in.
A signal reached a guest thread only as the answer to a call it had made,
so a thread running its own code never received one. Go preempts a
goroutine that spins without a call by sending SIGURG to its thread, and a
guest is shown one CPU, so such a goroutine held the only P for good: no
other goroutine ran and a garbage collection never stopped the world.
SIGALRM, SIGINT and any other signal for a busy thread waited the same way.

MkForeignInterrupt lets a supervisor mark one of its guest threads. The
timer trampoline, after a tick that interrupted user mode, parks a marked
thread with its whole register file as a frame numbered NR_INTERRUPTED and
wakes the supervisor; the thread sleeps as a parked call does. A signal
answer rewrites the trampoline's frame to enter the handler, keeping the
thread's FPU state for its return as a delivered call does; any other
answer lets the thread run on exactly where it was. A thread already
parked in a call is not marked, since that call's answer can carry the
signal. Only the supervisor recorded for the thread may mark it; the mark
is dropped with the thread; the handler's context is checked as for any
signal answer. The trampoline's frame is its 160 bytes and no more: it is
read and written as those 20 words, never as the whole SavedUser, whose
TLS words would lie past the top of the kernel stack. A tick with nothing
marked costs one load. The libc gains mk_foreign_interrupt and
FOREIGN_NR_INTERRUPTED.
A caught signal raised for another thread waited in the queue until that
thread next made a call, so a thread running its own code never saw it.

Raising one now also asks the kernel to stop the target at its next tick.
The stopped thread arrives as a FOREIGN_NR_INTERRUPTED frame and is
delivered what is pending, with its handler returning to the thread's own
rax since no call is being answered; with nothing pending it runs on where
it was. A thread parked in a call, the caller included, still gets the
signal with that call's answer.
cpreempt spins a thread on a flag only its SIGUSR1 handler sets, sends it
SIGUSR1 with pthread_kill, and joins it. The thread makes no call while it
spins, so the join returns only if the handler runs inside the spin. It
passes on host Linux, pinned to one CPU as well.
@eKisNonos eKisNonos changed the title linux: run Go and musl threads in a guest, and end the guest whole linux: run Go and musl threads in a guest, let them wait as on Linux, and end the guest whole Sep 28, 2026
The frame a handler is entered on wrote the 18 saved registers at
ucontext + 48 and the blocked mask right after them. Linux, musl, glibc
and Go all read uc_mcontext at ucontext + 40 and uc_sigmask at + 296
(measured with offsetof on musl and glibc). Every register a handler
read from its context was therefore the one before it: Go's preemption
handler reads rip and rewrites rip and rsp, so it would have read the
saved rsp as the pc and resumed into garbage. The C handlers so far
never read their context, and the frame tests only read it back the way
it was written, so nothing showed it.

The registers now start at + 40 and the mask sits at + 296. rt_sigreturn
reads the same constant, so a frame still round-trips, and a new test
pins r8, rsp, rip and the mask to the offsets musl and glibc use. With
the old offset that test fails and the four round-trip tests pass.
… handlers on it

sigaltstack read back an empty stack and ignored the one a program set,
so every handler was entered on whatever stack the thread was using.
Go installs every handler with SA_ONSTACK and gives each thread its own
signal stack; its handler checks that it runs on that stack or on g0's,
and a frame on a goroutine's stack sends it into a path that waits for a
spare M a program without cgo never has. That is where gopreempt stopped:
one SIGURG delivered to tid 50, then no further line in 420 s.

Each thread's alternate stack is now kept, and sigaltstack answers as
Linux's do_sigaltstack: the old setting is what was there before the
call, SS_ONSTACK while the thread runs on it and SS_DISABLE when none is
set; changing it from on it is EPERM, a stack under MINSIGSTKSZ (2048) is
ENOMEM, and a mode other than 0, SS_ONSTACK or SS_DISABLE is EINVAL.
SS_AUTODISARM is named unserved and answered EINVAL, as a Linux before 4.7
answers it. A handler with SA_ONSTACK is entered at the top of the
alternate stack when the thread is not already on it, and below the
interrupted rsp when it is; a frame that would run off the bottom is not
written. uc_stack carries the stack and its flags for the interrupted rsp.
A fork gives the child the forking thread's stack, a thread's exit drops
its own, and execve clears them all.
A seventeenth part sets a 16 KiB alternate stack through the raw call,
so the answers are the kernel's and not musl's own checks, and raises a
SIGUSR2 caught with SA_ONSTACK. It passes only if: none set reads back
SS_DISABLE; 1024 bytes is ENOMEM and flags 4 is EINVAL; the stack set
reads back with flags 0; the handler's own local is on that stack while
the saved rsp is not; inside the handler the stack reads SS_ONSTACK and
changing it is EPERM; uc_stack names it with flags 0; and SS_DISABLE
turns it off again. On the container's Linux it prints
"cwait altstack ok: ... flags inside 1" and "cwait PASS: 17 parts".
MkForeignInterrupt marks a running thread so that the next tick in user
mode stops it and hands it to its supervisor. When the thread makes a
call before that tick, the call's answer is where the supervisor
delivers the signal, but the mark stayed, and a later tick stopped the
thread once more with nothing to deliver. The instrumented gopreempt run
showed it: after "signal 23 to tid 50" on a call's return, the next tick
took tid 50 again and the supervisor answered it with a plain resume.

A thread's mark is now dropped as it traps into a call. The check costs
one atomic load while nothing is marked.
clone returned the new thread's tid in the guest's pid namespace, but
wrote the kernel's pid for it to the CLONE_PARENT_SETTID and
CLONE_CHILD_SETTID words. musl keeps the parent's copy as the thread's
own tid and hands it to tkill, so pthread_kill asked for a tid the guest
does not have and got ESRCH. An instrumented cpreempt showed it,
"[C] dbg kill rc 3", then a join that waited for a signal never sent.
gettid and Go's tgkill were right, since both use the translated number.

The family now writes those words once the reply has been put in the
guest's terms, with the same number clone returns. clone itself keeps
recording CLONE_CHILD_CLEARTID, which is keyed by the kernel's pid.
An eighteenth part starts a thread that yields until a SIGUSR1 handler
sets its flag, then sends pthread_kill to it and joins. It passes only if
pthread_kill answers 0 and the handler ran. The thread gives up after
3 s, and sooner when the kill failed, so a lost signal fails the part
and does not hang it. On the container's Linux it prints
"cwait pthread-kill ok ... handled 1" and "cwait PASS: 18 parts".
eKisNonos added a commit that referenced this pull request Sep 30, 2026
#582 gained nine commits on 30 September. Four are already here as the
same commits. The other five are earlier forms of 07d9c92, d7b2019,
f5fe356, 2b5c726 and 5b40bbc, which this branch carries in the split
form linux/next-waits gave them. The tree here already holds all of it, so
the merge keeps it as it is, and #582 and this branch now merge into main
in either order.
@senseix21

Copy link
Copy Markdown
Collaborator

Reviewed at b8985e07f (not a draft, 254 commits, 1114 files, +35246/−2978, 127 commits behind main, last pushed 2026-09-30). 209 of those commits and most of that diff are #567 underneath; this PR's own delta over 1910df85d is 45 commits, +472/−26 in src/.

Verdict: Request changes. One defect, in the mechanism this PR adds: the death notice a guest is not supposed to be able to send, it can send, with three instructions. Everything else is either inherited red from the trunk or a gap worth naming.

Before that, the part that deserves to be said first.

The hardest bug here was found and fixed properly

5ab0a0912 is the one I'd have expected to go wrong. Freeing a thread group's page tables when the leader is finalized, while a Go thread is still running on them, is the kind of fault that reports as noise — and the commit message records exactly the right evidence for it: CR3 the guest's, FS holding the Go thread's TLS, CR2 the IDT entry for vector 8. A machine reset diagnosed down to "we freed the tables under a live thread and then triple-faulted on the IDT" is not a guess. The fix passes the tables to a process still running on them and rebinds its capability token to the ASID it now owns, rather than deferring the free or refcounting it, which is the version that stays correct when the leader is not the last out.

And the machine still comes up. From nonos-benchmarks-36712450455-1/boot-log.json: zk_attest_ok: 24, zk_attest_fail: 0, fatal: 0, panic: 0, 24 of 24 capsules through attest, preflight and runqueue, userspace_entry_ms 2565. On a change that rewrites page-table release, thread teardown and the timer trampoline's user-mode tail, nothing regressed on the ordinary boot.

Critical

1. A guest forges its own death notice with three instructions, and the supervisor believes it.

2311afc39's own description is the claim I checked: "the kernel leaves a one-way notice for its supervisor, collected on its next MkForeignWait as a frame numbered FOREIGN_NR_DIED". The notice is one-way, but the channel it arrives on is not.

NR_DIED is u64::MAX and NR_INTERRUPTED is u64::MAX - 1 (src/process/foreign/frame.rs:28,32). A death notice and a parked guest syscall are delivered over the same MkForeignWait channel, and wait.rs:63 says so: "A claimed guest call if one is waiting, else a one-shot death notice built into a frame the supervisor recognises by its nr." The supervisor distinguishes them by frame.nr alone (userland/capsule_linux/src/linux/serve/family.rs:55,59).

But frame.nr for a parked call is the guest's raw rax, unfiltered:

  • src/arch/x86_64/syscall/manager/entry.rs:38 — SyscallNumber::from_u64(number) is a linear scan of registered ids (src/syscall/abi/mod.rs:31). Every NONOS id is a four-byte ASCII magic, so u64::MAX is not one and never will be.
  • :49 — the unknown number goes to redirect(number, args, saved).
  • src/process/foreign/trap.rs:36 — park(ForeignFrame::new(pid, nr, args, saved.rip)), with nr passed straight through.

So a guest executing mov rax, -1; mov rdi, <anything>; syscall lands in Family::answer on the FOREIGN_NR_DIED branch, and thread_died(frame.pid, frame.arg0 as i32) runs with arg0 fully guest-chosen. The consequences, in family.rs:88:

  • g.exited = Some(128 + signo_of(rdi)) — the guest picks its own reported wait status, which its parent reads back through wait.
  • the thread is dropped from g.threads, and per the doc comment "reap then kills its other threads and answers any waiter", so the whole thread group ends.
  • thread_died returns without mk_foreign_reply, so the thread — which is not dead, it is parked in wait_for_answer — is never answered. The supervisor's view and the kernel's disagree permanently.

Scope: self-inflicted. frame.pid is kernel-stamped from current_pid(), and owns(pid) scopes the lookup to the guest's own Guest, so this is not a cross-guest read. What it breaks is the supervisor's authority over guest lifecycle — the one property the foreign-process design exists to hold. u64::MAX - 1 reaches interrupted() the same way, which is benign by comparison (it replies and resumes).

Smallest corrective action: reserve the top of the number space in redirect, since that is the one place every guest-chosen nr passes through. if nr >= NR_INTERRUPTED { return None; } at trap.rs:28 returns ENOSYS via entry.rs:52 — which is what Linux returns for rax = -1 anyway, so no guest loses a behaviour it had. One guard, and it keeps the sentinels out of reach as more of them are added; a second sentinel already arrived in this PR, so there will be a third.

2. Nothing in CI boots a Linux guest, and the four gates that would judge this change did not run.

bf9d9dbf3 builds the Linux-guest test image under NONOS_LINUX_GUESTS=1, and the body's evidence is all from it. That variable appears only in mk/20-build.mk:623,1197 and mk/40-run.mk:38,48 — nowhere in .github/, nonos-ci/ or benchmarks/. The benchmark lane's boot is the desktop image: grep -c LINUX over boot/qemu-serial.log is 0. So on a 1114-file PR whose subject is Go and musl threads inside a guest, no lane runs a guest.

The gates that exist for exactly this also did not run. abi-contracts is sequential steps in one job with no if: always() (.github/workflows/verify.yml:97-129, unchanged by this PR), and step 9 — tools/nonos-assumptions — fails, so steps 10–14 skip: "No new control that exists and does not run", "Linux syscall coverage does not fall", "Wayland coverage does not fall", "Every served Linux call says what it discloses", "Every mutant still removes the control it names". The job log contains no output from any of them.

The assumption gate's failure is inherited from #567 and is the same shape — 122 found, 10 stated, 53 unlisted, 0 stale, every unlisted id an Aeneas opaque model of a core/alloc primitive (core.sync.atomic.compiler_fence, core.sync.atomic.private.Align1..8, memory.addr.phys.PhysAddr...ge). Worth adding here, because it is new information since #567's run two days earlier: that run reported 33 unlisted, this one 53, with no change to the register between them. CI checks out refs/pull/N/merge, so the count tracks main's Lean corpus, which went 5 → 591 extraction files in the 127 commits this branch is behind. The number grows on its own.

Taken together: this PR's subject matter is currently judged by author-run evidence and nothing else. The two fixes are independent and both small — if: always() on the steps after the register, and a lane that runs the guest image the author already built.

Important

3. Ring 0 grows another +316 on top of the trunk's +1237, and the baseline is still not moved.

From ci-reports-build/build/tcb-budget.txt: 134217 against baseline 132664, delta +1553. #567 at 1910df85d reported +1237, so this PR's own share is +316, nearly all of it process 8338 → 8565. That is src/process/foreign/interrupt.rs (114), interrupt_frame.rs (70), notice.rs (55) and thread.rs (+60) — the stop-at-a-tick mechanism and the clone register copy, both of which have to be in ring 0, so the resolution is to move the baseline.

The gate asks for the justification in the PR. #567's body does not mention ring 0 growth and neither does this one, and since both lanes fail on the same check, whichever lands first has to carry the argument for both. Worth deciding here which PR moves nonos-ci/baselines/tcb-x86_64-capsules.txt rather than having both move it and conflict.

4. The confinement check on MkForeignThread's new argument is the right check and nothing tests it.

parked_parent (src/process/foreign/thread.rs:99) is the security-relevant part of 3fc1ac19d, and it is correct: supervisor_of(from) == caller or EPERM, same thread_group_id as pid or EPERM, parked_frame(from) present or ENOENT. Three refusals, and the comment states the reason the middle one matters — "a copy across groups would hand one guest another's register contents".

That is the one refusal in this PR a reader would most want to see held down by a test, and there is none. This PR adds 74 lines to userland/capsule_linux_proofs/src/tests/sigframe_tests.rs and zero lines to verification/ or userland/kernel_proofs/. Not a regression — git grep foreign userland/kernel_proofs/src is empty at both heads, so the whole foreign subsystem is outside the proof crate — but the house pattern in recent main is a kernel_proofs: an X is refused commit beside the kernel change (175d2e542 is the last one), and a register copy across thread groups is precisely that shape.

5. Six inherited causes account for the nine real failures, and five are one line each.

I checked each against the trunk rather than assuming:

  • hygiene: one violation, userland/capsule_wallet_nonos/src/wallet/etna/parts/fact.rs:18 — a doc comment flagged placeholder-comment for containing the word while stating the rule the scanner enforces. Identical to linux: personality, store install, desktop and kernel fixes #567. Reword to "stand-in value".
  • crypto-proofs: the same five dead-code errors in userland/crypto_proofs/src/security/capsule_attest/ — POLICY_EPOCH, POLICY_TREE_DEPTH, TRAILER_MAGIC, parse, verify. The mirror #[path]-includes against_pedersen.rs but not against_root.rs, its only non-test caller. pub use against_pedersen::verify; in the mirror's mod.rs.
  • proof-crates (capsule_linux_proofs): one lint, "digits of hex, binary or octal literal not in groups of equal size" — mutation_tests.rs:35, 0xDEB5_EEDu64. 0x0DEB_5EED_u64 is the same value and keeps the joke.
  • runnable-proofs: the same two tests at usage_tests.rs:98 and :111, byte-identical. The include_str! parser over capsule_terminal's help.rs did not survive that file's restructure into GROUPS/DEEPER, so the 80-column bound and the alignment rule are both unchecked. Better fixed as a capsule_terminal unit test calling help::run against a fake Output.
  • evidence-manifest: stale EVIDENCE.json; the error states the fix.
  • build / production-build / benchmark: all three are nonos-verify build's tcb-budget, item 3.

None of these are this PR's doing, and none need to be fixed here. They matter because hygiene has the same skip cascade as abi-contracts — steps 6–10 skip, including "No new exported function that nothing calls", which is the ratchet that would have caught the crypto-proofs failure one lane earlier.

Minor

  1. The body says "What it adds is thirty-six commits"; git rev-list --count b8985e07f ^1910df85d is 45.
  2. The Cleanups section names one deletion, and three files are gone: call/clone.rs plus call/pipe_wait.rs and serve/family_pipes.rs. The latter two are the pipe rework described under b716a34af and 69f3b48d4 rather than removals — read_or_park was replaced by waiting with the other parked calls — but a reader checking the deletions against the body will come up two short.
  3. The 5–7 second failures on extraction, lean, hygiene, seven dark-features-compile variants and a dozen proof-crates entries are all from runs 36712442920/36712442940, which were superseded and report conclusion: cancelled. gh pr checks renders them as "fail". Not findings; the real set is the nine from 36712450024/36712450455.
  4. This PR is not a draft and targets main, but its body correctly says "it merges after linux: personality, store install, desktop and kernel fixes #567", which is a draft. Nothing downstream can land before linux: personality, store install, desktop and kernel fixes #567 either way — every one of linux: a guest's sockets are the family's own, on 127.0.0.0/8 #584, linux: process lifecycle and signals as Linux has them #585, linux: a guest holds only the memory it asked for #586, linux: Go's own standard-library tests as the oracle for Linux guests #588 and linux: land the Linux lane on main; terminal linux command, signals to running threads, pipe waits, madvise, VFS store, Process Manager #593 has its merge base at 1910df85d.

Questions

  1. on_user_tick (src/process/foreign/interrupt.rs:85) parks and then blocks in wait_raw on the timer trampoline's user-mode tail (src/interrupts/isr/timer_trampoline/handler.rs:125). The call site drops _ctx_guard first and is gated on from_user, and MARKED is released in take() before the wait, so I could not find a lock cycle — but what answers the guest if its supervisor dies while it is parked there? park returning false is handled; the supervisor vanishing mid-wait is the case I could not trace.
  2. MFIN is a new syscall number in abi/syscalls.toml with caps = ["ForeignExec"] and one argument. abi-contracts steps 3–8 pass, so the table and the kernel agree — but steps 10–14 never ran, so nothing checked that the new call declares what it discloses. It hands a supervisor the complete register file of a guest thread at an arbitrary instruction. Is that covered by ForeignExec alone, or does it want its own row in the disclosure ratchet?

Verified correct

  • The page-table handoff is the right fix, and the boot proves it. 24 of 24 capsules attest, fatal: 0, panic: 0, after a change to address_space/lifecycle/release.rs and exit/teardown.rs. This is the lane that would have gone red if the handoff were wrong.
  • No allocation on the timer path. MARKED is a Vec<u32> reachable from the ISR tail, which in this codebase is how you get a heap deadlock. sys_foreign_interrupt is the only path that pushes, and it is a syscall; on_user_tick reaches take(), which only reads and swap_removes. The ANY: AtomicBool fast path means a tick with nothing marked costs one load, as the comment claims.
  • parked_parent checks the pid argument before it trusts it. from goes through pid_arg first, then the three refusals, so a from that is not a live supervised thread of this caller never reaches parked_frame. The ordering matters: the thread-group comparison reads both pids out of the process table through with_process and refuses when either is absent, rather than defaulting to equal.
  • The sentinels agree across the boundary. NR_DIED/NR_INTERRUPTED in src/process/foreign/frame.rs:28,32 and FOREIGN_NR_DIED/FOREIGN_NR_INTERRUPTED in userland/libc/src/foreign_frame.rs:27,32 are the same two values. Hand-synced constants across the kernel/libc boundary have drifted in this repo before; these have not. Critical 1 is not a drift, it is a reachability problem with values that are correctly mirrored.
  • The sigaltstack tests assert placement, not source text. The six new tests in sigframe_tests.rs check that an SA_ONSTACK handler enters at the top of the alternate stack, that one without it stays on the thread stack, that a signal already on the alternate stack nests below itself with SS_ONSTACK, that no alternate stack reads back SS_DISABLE, and that a frame which would run off the alternate stack is refused. That is the kind of test that survives a restructure — the opposite of item 5's help harness.
  • The proof floor held. proof-coverage passes against scripts/baselines/proof-coverage.txt, so none of this was paid for out of extracted-code theorem lines.
  • 289272374's claim that no module included call/clone.rs is true. call/mod.rs at the trunk has no mod clone; — only pub use spawn::{clone, ...}, which resolves to the live call/spawn/clone.rs. The deleted file was genuinely orphaned, not a module someone stopped calling.
  • boot-smoke, adversarial, attestation, attestation-attack, boot-proofs, reproducible, supply-chain, trust-chain, trust-ledger, evidence tooling, symbol-scan, section-size and build-x86_64-capsules all pass, along with 58 of 59 proof-crates matrix entries.

CI at this head

Nine real failures from runs 36712450024 and 36712450455; everything else either passes or is cancelled noise from the two superseded runs. Six causes: tcb-budget covers three lanes, and hygiene, evidence-manifest, proof-crates (capsule_linux_proofs), crypto-proofs, runnable-proofs and abi-contracts are one each. All six are inherited from #567 and all are small. The thing to fix in this PR is Critical 1, which no lane reports.

@eKisNonos

Copy link
Copy Markdown
Contributor Author

Superseded. This work is integrated into the 0.9.2 release and ships in the current tree. Closing as part of the 0.9.2 consolidation.

@eKisNonos eKisNonos closed this Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants