Conversation
bind, listen, getsockname, getpeername, socketpair, setsockopt, getsockopt, recvmmsg and sendmmsg had no numbers here, and the errnos a socket call answers with (ENOPROTOOPT, EADDRINUSE, EISCONN, EDOM and the rest) had no names, so a call that needed them could not be written. They are now in abi/nr_sock.rs and abi/errno_sock.rs, taken from arch/x86/entry/syscalls/syscall_64.tbl and include/uapi/asm-generic; nr and errno re-export them with the calls that first use them.
A run started in the libc default heap of 16 MiB and read the program it runs whole from the store into it. A 6 MB Go program (net/http) outgrew it while it was read: the personality died with "memory allocation of 8388608 bytes failed" before the guest started. A run now asks for 64 MiB, the largest program the personality accepts (source::MAX_IMAGE), and falls back to the default when the kernel will not give it. An install keeps its 320 MiB.
Every guest went into the store the vfs loads, and the vfs loads at most 16 MiB. The set on this branch already used 16161792 bytes, so a new guest, or a Go guest of any size, made the store step fail with "nonos-store-pack: ... vfs loads at most 16777216". LINUX_GUEST_STORE_ONLY, when set, names the guests the store carries; every guest is still built, signed and proven. Unset, the store carries them all, as before.
A guest could not use a socket as Linux has one. Every AF_INET stream asked net.sockets for a mixnet socket, which the Linux capsule may not call (it holds no Network capability), so socket answered EIO; a datagram socket was always the in-capsule resolver; bind, listen, accept, accept4, getsockname, getpeername, socketpair, sendmmsg and recvmmsg were unserved; sendmsg and recvmsg served only the display socket; and shutdown closed the descriptor, so SHUT_WR ended both directions and a dup'd descriptor lost its socket. Now every socket a guest opens is an entry in one table the personality keeps for the family (net/sock). On 127.0.0.0/8 nothing leaves the capsule: - bind and listen, only on loopback: any other address is EACCES with a [LINUX] refused line naming it, and raw sockets are EPERM; - connect to a listener links the two ends in the caller's own call; a non-blocking one answers EINPROGRESS then POLLOUT; a port with no listener is ECONNREFUSED, or EINPROGRESS then POLLERR with the error kept for SO_ERROR, as Linux's loopback answers; - accept and accept4 hand out the oldest connection, with SOCK_NONBLOCK and SOCK_CLOEXEC, and a listener with one pending reads POLLIN; - bytes wait in the peer's queue: end of file after the peer closes or shuts its side, ECONNRESET when it closed with bytes unread, and after the peer is gone the first send is taken and the next is EPIPE; - shutdown shuts one direction or both of the socket and leaves the descriptor open; socketpair gives two connected Unix sockets; - datagrams go to whichever socket holds the port, cut to the buffer with MSG_TRUNC; a connected socket keeps only its peer's, and learns of a closed port as ECONNREFUSED; sendmsg, recvmsg, sendmmsg and recvmmsg carry iovecs and addresses; - a datagram to port 53 that no family socket holds still becomes the resolver; one to anywhere else outside is ENETUNREACH, said by name. A stream to anywhere outside the family still goes over the mixnet through net.sockets, as before. A socket belongs to the processes that hold a descriptor naming it: a forked child is added (fork_state), close lets go when the process has no other descriptor on it (call/io.rs now hands close the guest and the descriptor), and a process that ends lets go of everything it held (the Holder field at the end of Guest, dropped with the process). abi/nr.rs and abi/errno.rs re-export the socket numbers and errnos added in abi/nr_sock.rs and abi/errno_sock.rs.
Both were unserved, so every Go listener and dialer, which sets SO_REUSEADDR, TCP_NODELAY and the keepalive options, and every program that reads SO_ERROR after a non-blocking connect, got ENOSYS. Now a socket keeps SO_REUSEADDR, SO_REUSEPORT, SO_KEEPALIVE, SO_BROADCAST, SO_RCVBUF and SO_SNDBUF (doubled and bounded as Linux does), SO_LINGER, SO_RCVTIMEO and SO_SNDTIMEO (EDOM for microseconds out of range), TCP_NODELAY and TCP_KEEPIDLE, TCP_KEEPINTVL and TCP_KEEPCNT (EINVAL outside Linux's range), each with Linux's default, and reads them back; SO_TYPE, SO_DOMAIN, SO_PROTOCOL, SO_ACCEPTCONN and SO_ERROR, which reports a refused connect once, are read. A TCP option on a datagram socket and an IPv6 option on an AF_INET one answer as Linux does. Any other option is ENOPROTOOPT with a [LINUX] refused line naming its level and number. SO_REUSEADDR decides whether a bind may share a port, SO_RCVBUF bounds the socket's queue, and SO_LINGER of zero seconds makes close reset the peer, as on Linux.
A blocking accept on an empty listener, a blocking receive with nothing queued, and a blocking send to a full peer answered EAGAIN at once, since only read and write on a pipe or an eventfd could park; and any wait that watched a socket was looked at every 10 ms, because net.sockets does not say when its streams change. Now accept, accept4, connect, the sends and receives and read, write, readv and writev on a socket are routed to waits_sock (one arm in serve/dispatch.rs, one in waits::attempt). Each is tried at once; a blocking one that cannot finish is parked and tried again after every answer, and a non-blocking one or MSG_DONTWAIT answers EAGAIN. A blocking stream send returns once every byte is queued, a receive with MSG_WAITALL once its buffer is full, and recvmmsg once vlen messages have come unless MSG_WAITFORONE; the count so far is kept per thread between tries. SO_RCVTIMEO and SO_SNDTIMEO bound the wait, which then answers EAGAIN, or the count moved, as Linux does. A family socket changes only in an answer, after which every parked call is tried, so the 10 ms tick (family_waits::next_wait_ms) now runs only while a wait watches a stream net.sockets holds.
Nothing here proved a guest's sockets, so none of what the family's socket table does could be shown or held to Linux. Five guests, ids 5000 to 5009, each built static and run on a Linux host first as the oracle: - csock (C), fifteen parts, each named when it fails: socketpair both ways, a listener's name and accept4's flags, a non-blocking connect, a refused port, half-close, epoll on a listener, end of file, EAGAIN, MSG_PEEK, EPIPE, an accept and a receive that wait for another thread, the options a server sets, and a socket a forked child shares. Linux: "[C] csock PASS: 15 parts". - cudp (C), six parts: an echo, a connected socket that drops a stranger's datagram, MSG_TRUNC, a refused port, sendmmsg and recvmmsg, and no peer. Linux: "[C] cudp PASS: 6 parts". - cpolicy (C): the answers the capsule's policy gives, which an unconfined Linux does not (its line records what Linux allows), and what descriptor numbers the guest never opened get. - cidle (C): accept blocked while nothing happens for the seconds given, to measure what the serve loop does meanwhile. - gohttp (Go): net/http, a server on 127.0.0.1:0 and its client in one guest, twenty GETs over one kept-alive connection. Linux: "[GO] gohttp PASS: 20 GETs answered 200 over 1 connection".
A guest's AF_UNIX socket could only reach the display: socket made a
display descriptor, a datagram one was ENOSYS, bind, listen and accept
had nothing to do on it, and a connect to any other path was
ECONNREFUSED. A Go net.Listen("unix", ...), a local daemon, or two
processes meeting at a path could not run.
Now an AF_UNIX socket is an entry in the family's table like any other:
- bind to a path makes the path a file, as Linux does, EADDRINUSE if
anything is there; the file stays after the socket closes, so a
rebind is EADDRINUSE and a connect ECONNREFUSED until the guest
unlinks it; an abstract name has no file and goes with its socket;
bind with the family alone chooses a NUL and five hex digits;
- listen needs a name (EINVAL), and a connect to a listener completes
in the caller's call, blocking or not; a path with nothing there is
ENOENT, one with no listener ECONNREFUSED, a socket of the other type
EPROTOTYPE;
- getsockname, getpeername, accept and recvfrom report sockaddr_un at
Linux's lengths: a path with its NUL, an abstract name without, an
unnamed peer as the family alone, an unnamed sender as nothing;
- a datagram goes to the socket bound to the name it is sent to; a
connected datagram socket sends to its peer, answers ECONNREFUSED once
the peer is gone and ENOTCONN if it never had one, and refuses a
stranger's datagram with EPERM;
- SOCK_RAW is a datagram socket, as on Linux; SOCK_SEQPACKET is refused
by name.
A connect to the display's path still turns the descriptor into the
display connection this capsule serves (unix/), so a Wayland client is
unchanged.
A socket that goes now clears every socket that pointed at it, not only
its own peer: a connected datagram socket points one way, and it kept
an index that a later socket could reuse.
Nothing proved a named Unix socket. cunix (C, ids 5010 and 5011), six parts, each named when it fails: a listener on a path with its name, its client's and a connect that completes at once; ENOENT, ECONNREFUSED and EADDRINUSE for what a path leaves behind, and a rebind after unlink; abstract datagrams and the sender each reports; a connected datagram socket that refuses a stranger with EPERM; autobind; and a connection across fork. Linux: "[C] cunix PASS: 6 parts".
A non-blocking connect to a listener whose queue was full answered EAGAIN. Linux answers EINPROGRESS: the connect stays in SYN_SENT, not writable, a second connect on it is EALREADY, and it completes once an accept makes room. Now the listener keeps such connects in order (net/sock/syn.rs) and completes them at the accept that makes room; Linux completes them at its next SYN retransmit, so only the timing differs. Closing the listener, or shutting its reading side, refuses them with ECONNREFUSED. A blocking connect to a full listener still waits, as before. csock gains a sixteenth part, backlog, which shows it. Linux: "[C] csock PASS: 16 parts".
IP_TOS, IP_TTL, SO_PRIORITY, TCP_USER_TIMEOUT, TCP_QUICKACK and TCP_FASTOPEN were ENOPROTOOPT with a refused line, though Linux keeps each of them and nothing on a loopback connection changes with them; a TCP option on a Unix socket was kept, where Linux answers EOPNOTSUPP. Now each is kept with Linux's starting value and checked as Linux checks it (a TCP socket drops the two ECN bits of IP_TOS; IP_TTL takes 1 to 255, and -1 for the default of 64; a negative user timeout or Fast Open queue is EINVAL), and read back. A Unix socket answers EOPNOTSUPP for any level but SOL_SOCKET, and a datagram socket keeps answering ENOPROTOOPT for a TCP option. csock gains a seventeenth part, quiet_options. Linux: "[C] csock PASS: 17 parts".
A second listener on an address answered EADDRINUSE with SO_REUSEPORT set on both, which Linux allows: a server that runs one listener per worker could not start its second. Now sockets that each set SO_REUSEPORT may bind and listen on one address, and a connect goes to one of them by its own port, as Linux spreads connections by a hash of each; one without the option still answers EADDRINUSE, and when one listener closes the others take what comes next. csock gains an eighteenth part, reuseport. Linux: "[C] csock PASS: 18 parts".
Every file under net/ and serve/waits_sock*.rs now holds at most 75 lines. The Unix name calls move to net/named/, the options to net/opt/, and the longer table methods split by what they do (deliver, put, take, unlisten, gram_target, kinds). mod.rs files hold module declarations and re-exports only; net/api.rs carries the re-exports the rest of the personality calls. Plain comments are /* */; doc comments stay /// so rustdoc still reads them. The lines this branch adds to guest/handle.rs, serve/dispatch.rs and serve/family_waits.rs are shortened; those three files were over 75 lines before this branch and are not split here. No call answers differently: csock, cudp, cunix, cpolicy and cidle pass on NONOS as they did before the split.
csock, cudp, cunix, cpolicy and cidle keep their parts and their output; each is now <guest>.c with its parts in <guest>_N.h, none over 72 lines, and comments in /* */, as are gohttp's. Guests.mk picks the parts up with a wildcard, so a changed part rebuilds its guest. Every write and pipe result is now checked, so a guest built with glibc's -Wall -Wextra builds without a warning; a write that fails is named as its part's failure instead of passing unseen. backlog's second accept now waits at most 3 s for the held connect; before, a listener that never held it blocked csock there for good instead of failing the part by name. On Linux: csock PASS 18 parts, cudp PASS 6, cunix PASS 6, cidle PASS; cpolicy reports "confined 0" there, as root on an unconfined host must.
eKisNonos
force-pushed
the
linux/guest-sockets
branch
from
September 29, 2026 09:08
ec9daec to
de69c36
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This lets a Linux program running on NONOS use sockets the way it would on Linux, as long as it talks to itself or to the other processes it started. A server can bind to 127.0.0.1, listen, accept and answer; a client in the same family of processes can connect to it over TCP, UDP or a named Unix socket. None of that traffic leaves the capsule.
This PR sits on top of #582 (linux/go-guests) and only makes sense after it. Once #582 is in main, the base here moves to main.
What a guest can now do:
What a guest still cannot do:
Each refusal logs a [LINUX] refused line saying why.
The idle cost also went down. Before this change, every socket wait woke the serve loop every 10 ms. Now only streams that go through net.sockets do. With one guest blocked in accept for 10 seconds, the loop ran 548 times before and 67 times after.
How it was tested:
Five small C programs in userland/linux_guests (csock, cudp, cunix, cpolicy, cidle) were first built static and run on plain Linux, then booted on NONOS. The Linux run is the reference.
For each feature I also removed the code that provides it, rebuilt and booted again, and checked that the matching part failed and named itself. That covered socketpair, half close, EINPROGRESS, readiness on a listener, MSG_TRUNC, the loopback check, the wait path, fork, missing Unix paths, the full-listener case, the kept options and SO_REUSEPORT. One removal broke two parts rather than one: taking away readiness on a listener also broke the SO_REUSEPORT part, because that part polls its listener.
The guests from #582 still behave as they did: the window demo still draws, gohello, goconc, cthreads and cwait still pass, and threadfault still ends the whole process when its thread faults. Every commit builds on its own with no warnings, and the kernel is not touched.
Known problems: