Skip to content

linux: a guest's sockets are the family's own, on 127.0.0.0/8 - #584

Open
eKisNonos wants to merge 14 commits into
linux/go-guestsfrom
linux/guest-sockets
Open

eKisNonos wants to merge 14 commits into
linux/go-guestsfrom
linux/guest-sockets

Conversation

@eKisNonos

@eKisNonos eKisNonos commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

This lets a Linux program running on NONOS use sockets the way it would on Linux, as long as it talks to itself or to the other processes it started. A server can bind to 127.0.0.1, listen, accept and answer; a client in the same family of processes can connect to it over TCP, UDP or a named Unix socket. None of that traffic leaves the capsule.

This PR sits on top of #582 (linux/go-guests) and only makes sense after it. Once #582 is in main, the base here moves to main.

What a guest can now do:

  • socket, bind, listen, accept and accept4, connect, getsockname and getpeername, shutdown with a real half close, and socketpair.
  • sendto, recvfrom, sendmsg, recvmsg, sendmmsg and recvmmsg, with MSG_PEEK, MSG_TRUNC, MSG_WAITALL and MSG_DONTWAIT.
  • setsockopt and getsockopt for the options Linux keeps, read back the way Linux reports them. Unknown options are refused with ENOPROTOOPT and a log line naming them.
  • Unix sockets bound to a path or an abstract name, including autobind. A path stays behind after close until it is unlinked, as on Linux.
  • Blocking calls wait without polling. SO_RCVTIMEO and SO_SNDTIMEO put a limit on the wait.
  • A non-blocking connect to a full listener gets EINPROGRESS and completes once accept makes room.
  • Several listeners can share one port with SO_REUSEPORT.
  • Sockets follow fork, and they are released when the process exits.

What a guest still cannot do:

  • It cannot listen on anything but 127.0.0.0/8. Binding to 0.0.0.0 or to another address gets EACCES.
  • It cannot send UDP outside the family. Streams to the outside still go through net.sockets over the Nym mixnet, as before.
  • It cannot open raw sockets.

Each refusal logs a [LINUX] refused line saying why.

The idle cost also went down. Before this change, every socket wait woke the serve loop every 10 ms. Now only streams that go through net.sockets do. With one guest blocked in accept for 10 seconds, the loop ran 548 times before and 67 times after.

How it was tested:

Five small C programs in userland/linux_guests (csock, cudp, cunix, cpolicy, cidle) were first built static and run on plain Linux, then booted on NONOS. The Linux run is the reference.

  • csock passes all 18 of its parts, on Linux and on NONOS.
  • cudp passes 6 of 6, and cunix passes 6 of 6.
  • cidle blocks 10 seconds in accept and wakes on the connect.
  • cpolicy checks that every outside address, raw sockets and forged descriptors are refused. It passes on NONOS. On Linux it reports "confined 0", which is what an unconfined root process should see there.
  • Every NONOS run reported 0 unserved calls.

For each feature I also removed the code that provides it, rebuilt and booted again, and checked that the matching part failed and named itself. That covered socketpair, half close, EINPROGRESS, readiness on a listener, MSG_TRUNC, the loopback check, the wait path, fork, missing Unix paths, the full-listener case, the kept options and SO_REUSEPORT. One removal broke two parts rather than one: taking away readiness on a listener also broke the SO_REUSEPORT part, because that part polls its listener.

The guests from #582 still behave as they did: the window demo still draws, gohello, goconc, cthreads and cwait still pass, and threadfault still ends the whole process when its thread faults. Every commit builds on its own with no warnings, and the kernel is not touched.

Known problems:

  • The Go guest gohttp, and gopoll from linux: run Go and musl threads in a guest, let them wait as on Linux, and end the guest whole #582, can hang at the first SIGURG. The personality does not implement sigaltstack yet, so Go's preemption signal lands on the goroutine stack. With preemption turned off, gohttp runs 20 GETs over one connection and passes. gopoll hangs without these commits too, but in my boots it hung more often with them: 3 of 5 boots against 1 of 6. I have not found the reason.
  • A send to a closed peer returns EPIPE but does not raise SIGPIPE yet.
  • Passing descriptors with SCM_RIGHTS is not in this PR.

bind, listen, getsockname, getpeername, socketpair, setsockopt,
getsockopt, recvmmsg and sendmmsg had no numbers here, and the errnos a
socket call answers with (ENOPROTOOPT, EADDRINUSE, EISCONN, EDOM and the
rest) had no names, so a call that needed them could not be written.

They are now in abi/nr_sock.rs and abi/errno_sock.rs, taken from
arch/x86/entry/syscalls/syscall_64.tbl and include/uapi/asm-generic; nr
and errno re-export them with the calls that first use them.
A run started in the libc default heap of 16 MiB and read the program
it runs whole from the store into it. A 6 MB Go program (net/http)
outgrew it while it was read: the personality died with "memory
allocation of 8388608 bytes failed" before the guest started.

A run now asks for 64 MiB, the largest program the personality accepts
(source::MAX_IMAGE), and falls back to the default when the kernel will
not give it. An install keeps its 320 MiB.
Every guest went into the store the vfs loads, and the vfs loads at most
16 MiB. The set on this branch already used 16161792 bytes, so a new
guest, or a Go guest of any size, made the store step fail with
"nonos-store-pack: ... vfs loads at most 16777216".

LINUX_GUEST_STORE_ONLY, when set, names the guests the store carries;
every guest is still built, signed and proven. Unset, the store carries
them all, as before.
A guest could not use a socket as Linux has one. Every AF_INET stream
asked net.sockets for a mixnet socket, which the Linux capsule may not
call (it holds no Network capability), so socket answered EIO; a
datagram socket was always the in-capsule resolver; bind, listen,
accept, accept4, getsockname, getpeername, socketpair, sendmmsg and
recvmmsg were unserved; sendmsg and recvmsg served only the display
socket; and shutdown closed the descriptor, so SHUT_WR ended both
directions and a dup'd descriptor lost its socket.

Now every socket a guest opens is an entry in one table the personality
keeps for the family (net/sock). On 127.0.0.0/8 nothing leaves the
capsule:
- bind and listen, only on loopback: any other address is EACCES with a
  [LINUX] refused line naming it, and raw sockets are EPERM;
- connect to a listener links the two ends in the caller's own call; a
  non-blocking one answers EINPROGRESS then POLLOUT; a port with no
  listener is ECONNREFUSED, or EINPROGRESS then POLLERR with the error
  kept for SO_ERROR, as Linux's loopback answers;
- accept and accept4 hand out the oldest connection, with SOCK_NONBLOCK
  and SOCK_CLOEXEC, and a listener with one pending reads POLLIN;
- bytes wait in the peer's queue: end of file after the peer closes or
  shuts its side, ECONNRESET when it closed with bytes unread, and after
  the peer is gone the first send is taken and the next is EPIPE;
- shutdown shuts one direction or both of the socket and leaves the
  descriptor open; socketpair gives two connected Unix sockets;
- datagrams go to whichever socket holds the port, cut to the buffer
  with MSG_TRUNC; a connected socket keeps only its peer's, and learns
  of a closed port as ECONNREFUSED; sendmsg, recvmsg, sendmmsg and
  recvmmsg carry iovecs and addresses;
- a datagram to port 53 that no family socket holds still becomes the
  resolver; one to anywhere else outside is ENETUNREACH, said by name.
A stream to anywhere outside the family still goes over the mixnet
through net.sockets, as before.

A socket belongs to the processes that hold a descriptor naming it: a
forked child is added (fork_state), close lets go when the process has
no other descriptor on it (call/io.rs now hands close the guest and the
descriptor), and a process that ends lets go of everything it held (the
Holder field at the end of Guest, dropped with the process).

abi/nr.rs and abi/errno.rs re-export the socket numbers and errnos
added in abi/nr_sock.rs and abi/errno_sock.rs.
Both were unserved, so every Go listener and dialer, which sets
SO_REUSEADDR, TCP_NODELAY and the keepalive options, and every program
that reads SO_ERROR after a non-blocking connect, got ENOSYS.

Now a socket keeps SO_REUSEADDR, SO_REUSEPORT, SO_KEEPALIVE,
SO_BROADCAST, SO_RCVBUF and SO_SNDBUF (doubled and bounded as Linux
does), SO_LINGER, SO_RCVTIMEO and SO_SNDTIMEO (EDOM for microseconds out
of range), TCP_NODELAY and TCP_KEEPIDLE, TCP_KEEPINTVL and TCP_KEEPCNT
(EINVAL outside Linux's range), each with Linux's default, and reads
them back; SO_TYPE, SO_DOMAIN, SO_PROTOCOL, SO_ACCEPTCONN and SO_ERROR,
which reports a refused connect once, are read. A TCP option on a
datagram socket and an IPv6 option on an AF_INET one answer as Linux
does. Any other option is ENOPROTOOPT with a [LINUX] refused line
naming its level and number. SO_REUSEADDR decides whether a bind may
share a port, SO_RCVBUF bounds the socket's queue, and SO_LINGER of
zero seconds makes close reset the peer, as on Linux.
A blocking accept on an empty listener, a blocking receive with nothing
queued, and a blocking send to a full peer answered EAGAIN at once,
since only read and write on a pipe or an eventfd could park; and any
wait that watched a socket was looked at every 10 ms, because
net.sockets does not say when its streams change.

Now accept, accept4, connect, the sends and receives and read, write,
readv and writev on a socket are routed to waits_sock (one arm in
serve/dispatch.rs, one in waits::attempt). Each is tried at once; a
blocking one that cannot finish is parked and tried again after every
answer, and a non-blocking one or MSG_DONTWAIT answers EAGAIN. A
blocking stream send returns once every byte is queued, a receive with
MSG_WAITALL once its buffer is full, and recvmmsg once vlen messages
have come unless MSG_WAITFORONE; the count so far is kept per thread
between tries. SO_RCVTIMEO and SO_SNDTIMEO bound the wait, which then
answers EAGAIN, or the count moved, as Linux does.

A family socket changes only in an answer, after which every parked
call is tried, so the 10 ms tick (family_waits::next_wait_ms) now runs
only while a wait watches a stream net.sockets holds.
Nothing here proved a guest's sockets, so none of what the family's
socket table does could be shown or held to Linux.

Five guests, ids 5000 to 5009, each built static and run on a Linux
host first as the oracle:
- csock (C), fifteen parts, each named when it fails: socketpair both
  ways, a listener's name and accept4's flags, a non-blocking connect,
  a refused port, half-close, epoll on a listener, end of file, EAGAIN,
  MSG_PEEK, EPIPE, an accept and a receive that wait for another
  thread, the options a server sets, and a socket a forked child
  shares. Linux: "[C] csock PASS: 15 parts".
- cudp (C), six parts: an echo, a connected socket that drops a
  stranger's datagram, MSG_TRUNC, a refused port, sendmmsg and
  recvmmsg, and no peer. Linux: "[C] cudp PASS: 6 parts".
- cpolicy (C): the answers the capsule's policy gives, which an
  unconfined Linux does not (its line records what Linux allows), and
  what descriptor numbers the guest never opened get.
- cidle (C): accept blocked while nothing happens for the seconds given,
  to measure what the serve loop does meanwhile.
- gohttp (Go): net/http, a server on 127.0.0.1:0 and its client in one
  guest, twenty GETs over one kept-alive connection. Linux: "[GO]
  gohttp PASS: 20 GETs answered 200 over 1 connection".
A guest's AF_UNIX socket could only reach the display: socket made a
display descriptor, a datagram one was ENOSYS, bind, listen and accept
had nothing to do on it, and a connect to any other path was
ECONNREFUSED. A Go net.Listen("unix", ...), a local daemon, or two
processes meeting at a path could not run.

Now an AF_UNIX socket is an entry in the family's table like any other:
- bind to a path makes the path a file, as Linux does, EADDRINUSE if
  anything is there; the file stays after the socket closes, so a
  rebind is EADDRINUSE and a connect ECONNREFUSED until the guest
  unlinks it; an abstract name has no file and goes with its socket;
  bind with the family alone chooses a NUL and five hex digits;
- listen needs a name (EINVAL), and a connect to a listener completes
  in the caller's call, blocking or not; a path with nothing there is
  ENOENT, one with no listener ECONNREFUSED, a socket of the other type
  EPROTOTYPE;
- getsockname, getpeername, accept and recvfrom report sockaddr_un at
  Linux's lengths: a path with its NUL, an abstract name without, an
  unnamed peer as the family alone, an unnamed sender as nothing;
- a datagram goes to the socket bound to the name it is sent to; a
  connected datagram socket sends to its peer, answers ECONNREFUSED once
  the peer is gone and ENOTCONN if it never had one, and refuses a
  stranger's datagram with EPERM;
- SOCK_RAW is a datagram socket, as on Linux; SOCK_SEQPACKET is refused
  by name.
A connect to the display's path still turns the descriptor into the
display connection this capsule serves (unix/), so a Wayland client is
unchanged.

A socket that goes now clears every socket that pointed at it, not only
its own peer: a connected datagram socket points one way, and it kept
an index that a later socket could reuse.
Nothing proved a named Unix socket.

cunix (C, ids 5010 and 5011), six parts, each named when it fails: a
listener on a path with its name, its client's and a connect that
completes at once; ENOENT, ECONNREFUSED and EADDRINUSE for what a path
leaves behind, and a rebind after unlink; abstract datagrams and the
sender each reports; a connected datagram socket that refuses a
stranger with EPERM; autobind; and a connection across fork. Linux:
"[C] cunix PASS: 6 parts".
A non-blocking connect to a listener whose queue was full answered
EAGAIN. Linux answers EINPROGRESS: the connect stays in SYN_SENT, not
writable, a second connect on it is EALREADY, and it completes once an
accept makes room.

Now the listener keeps such connects in order (net/sock/syn.rs) and
completes them at the accept that makes room; Linux completes them at
its next SYN retransmit, so only the timing differs. Closing the
listener, or shutting its reading side, refuses them with ECONNREFUSED.
A blocking connect to a full listener still waits, as before.

csock gains a sixteenth part, backlog, which shows it. Linux: "[C]
csock PASS: 16 parts".
IP_TOS, IP_TTL, SO_PRIORITY, TCP_USER_TIMEOUT, TCP_QUICKACK and
TCP_FASTOPEN were ENOPROTOOPT with a refused line, though Linux keeps
each of them and nothing on a loopback connection changes with them; a
TCP option on a Unix socket was kept, where Linux answers EOPNOTSUPP.

Now each is kept with Linux's starting value and checked as Linux
checks it (a TCP socket drops the two ECN bits of IP_TOS; IP_TTL takes
1 to 255, and -1 for the default of 64; a negative user timeout or Fast
Open queue is EINVAL), and read back. A Unix socket answers EOPNOTSUPP
for any level but SOL_SOCKET, and a datagram socket keeps answering
ENOPROTOOPT for a TCP option.

csock gains a seventeenth part, quiet_options. Linux: "[C] csock PASS:
17 parts".
A second listener on an address answered EADDRINUSE with SO_REUSEPORT
set on both, which Linux allows: a server that runs one listener per
worker could not start its second.

Now sockets that each set SO_REUSEPORT may bind and listen on one
address, and a connect goes to one of them by its own port, as Linux
spreads connections by a hash of each; one without the option still
answers EADDRINUSE, and when one listener closes the others take what
comes next.

csock gains an eighteenth part, reuseport. Linux: "[C] csock PASS: 18
parts".
Every file under net/ and serve/waits_sock*.rs now holds at most 75
lines. The Unix name calls move to net/named/, the options to net/opt/,
and the longer table methods split by what they do (deliver, put, take,
unlisten, gram_target, kinds). mod.rs files hold module declarations
and re-exports only; net/api.rs carries the re-exports the rest of the
personality calls. Plain comments are /* */; doc comments stay /// so
rustdoc still reads them.

The lines this branch adds to guest/handle.rs, serve/dispatch.rs and
serve/family_waits.rs are shortened; those three files were over 75
lines before this branch and are not split here.

No call answers differently: csock, cudp, cunix, cpolicy and cidle pass
on NONOS as they did before the split.
csock, cudp, cunix, cpolicy and cidle keep their parts and their
output; each is now <guest>.c with its parts in <guest>_N.h, none over
72 lines, and comments in /* */, as are gohttp's. Guests.mk picks the
parts up with a wildcard, so a changed part rebuilds its guest.

Every write and pipe result is now checked, so a guest built with
glibc's -Wall -Wextra builds without a warning; a write that fails is
named as its part's failure instead of passing unseen.

backlog's second accept now waits at most 3 s for the held connect;
before, a listener that never held it blocked csock there for good
instead of failing the part by name.

On Linux: csock PASS 18 parts, cudp PASS 6, cunix PASS
6, cidle PASS; cpolicy reports "confined 0" there, as root on an
unconfined host must.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant