NOTE: Initially, I built this using Kiro, using a spec based approach. Now using Claude to build further.
Initial thoughts and ideas: https://blog.tovganesh.in/2012/01/kosh-building-mobile-user-experience.html
At this point this is my personal toy project. I currently have no idea if this will actually work. If you are interested in writing a rust based OS with or without AI tools you can poke me.
An x86-64 operating system kernel written in Rust. It boots on QEMU from a Multiboot2 ISO, runs programs in ring 3 with per-process address spaces, reads a FAT32 disk, and gives you a shell.
ksh: the Kosh shell, in ring 3. Type 'help'.
ksh:/$ ls
d - DOCS
- 199 README.TXT
- 22000 BIG.TXT
- 26 A Long File Name.txt
6 entries, 22280 bytes
ksh:/$ date
2026-07-30 11:42:25 UTC
ksh:/$ hello
hello from a loaded ELF binary
my pid is 3
mmap gave me 8192 usable bytes at 0x0000000010000000
exiting cleanly
ksh:/$ ksh
ksh: the Kosh shell, in ring 3. Type 'help'.
ksh:/$ getpid
pid 4
ksh:/$ exit
ksh:/$ getpid
pid 3
This project has spent most of its life describing things it could not do. The
kernel had a "virtual memory manager" that read the bootloader's page tables and
described them; a shell whose read_line replayed a hardcoded array of six
commands; a sys_open that returned the literal 3 for every path. For 27
commits the kernel never executed a single instruction, because Multiboot2 hands
off in 32-bit mode and nothing bridged the gap.
Unwinding that is most of the work that has happened since, so this file tries to
be accurate about what runs and blunt about what doesn't. If something is listed
under Works, there is a marker in scripts/run.sh that fails if it stops
working.
docs/BOOT.md is the long version: how the boot chain, paging, scheduling,
ring 3, the filesystem, the address-space work and the driver interface actually
function, including the bugs found along the way and how each was caught.
One command builds the kernel, the seven userspace programs, a GRUB ISO and a FAT32 test disk, then boots the lot:
./scripts/run.sh # boot it, serial on stdio (ctrl-a x to quit)
./scripts/run.sh --check # boot headless, assert 64 serial markers (CI)
./scripts/run.sh --check-cli # drive the shell through QEMU's monitor, assert 46
./scripts/run.sh --debug # same as plain run, plus a gdb stub on :1234The kernel and userspace are built for different targets: userspace uses
x86_64-kosh.json, the kernel uses the built-in x86_64-unknown-none, whose
spec disables SSE. A kernel that uses xmm registers corrupts them for whichever
ring-3 program made the system call — see docs/BOOT.md.
You need cargo (nightly), grub-mkrescue, xorriso, qemu-system-x86_64,
mkfs.vfat and mcopy. scripts/run.sh checks for all of them and tells you
what to install.
The other scripts in scripts/ predate the current build and are stale —
build.sh and build-iso.sh in particular do not build the two userspace
binaries that actually ship, and do not handle the -Z build-std flags the
kernel now needs. run.sh is the only supported entry point.
.github/workflows/boot.yml runs --check on every push, so "it boots" is a
merge gate rather than a claim.
Each of these is asserted by a serial marker in --check or a scripted console
interaction in --check-cli.
Boot
- Multiboot2 hand-off, a 32-bit trampoline into long mode, and a higher-half
kernel linked at
0xFFFFFFFF80100000and loaded at 1 MiB - Early COM1 output from assembly, so a failure before Rust is visible
- GRUB's own output on the serial port too, because a kernel GRUB refuses to load otherwise looks identical to one that hung
Memory
- Page tables the kernel builds itself, with W^X enforced —
.textis read-execute,.rodataand everything else is NX, andCR0.WPplusEFER.NXEare set so those bits actually bind in ring 0 - A physmap of all RAM at
0xFFFF800000000000, a heap window at0xFFFF900000000000, and page 0 left unmapped as a null guard - A bitmap frame allocator that excludes the running kernel and the boot tables
- A first-fit kernel heap with coalescing, alignment handling and stats
Processes and scheduling
- Preemptive round-robin over kernel threads, driven by the PIT at 100 Hz
- Per-process address spaces: a PML4 each, kernel's upper half shared. Two
programs can be — and
hello,hello2andkshall are — linked at the same address forkandexec: the child returns from a syscall it never executed;execreplaces the whole image- Copy-on-write: a
forkcopies no page data at all. Both sides go read-only, the frame reference count goes up, and the first write from either side faults into a private copy - Demand paging:
.bss, user stacks and anonymousmmapare reserved rather than allocated — a page table entry with no frame behind it until the program touches it. A boot that runs five programs reserves 322 pages and touches 38 - Per-thread kernel stacks, so more than one thread can be inside a syscall
- Real blocking:
waitparks a thread inState::Blockedrather than spinning
Ring 3
SYSCALL/SYSRETwith aswapgs-free per-CPU block reached throughgs:- User-pointer validation that checks range and page permissions, tested from the user side by a payload that deliberately passes a kernel address
- A ring-3 fault of any kind kills the process and the kernel carries on — not just a page fault, which is all it used to be
- A static ELF64 loader:
PT_LOADsegments mapped at theirp_vaddrwith per-segment permissions and a zeroed.bsstail
Devices and storage
- 8259 PIC remap, PIT timer, CMOS RTC
- The disk driver and the filesystem both run in ring 3.
ata-driverreads the disk from an unprivileged process usinginandoutdirectly — no syscall per port — andfs-servicemounts FAT32 on top of it and answersopen/read/stat/getdentsfor the shell. The kernel has no filesystem and no disk driver at all:fs/,block/, the descriptor table and nine system calls were deleted, 2,071 lines against 164 added Its ports come from the TSS I/O permission bitmap, 9 bits granted to that thread and denied to every other; a process without the grant that touches 0x1F7 is killed and the system carries on - The keyboard driver runs in ring 3.
kbd-driverclaimskbd0(port0x60), sleeps on IRQ 1 viawait_irq, translates Set 1 scancodes into ASCII/ANSI sequences, and answers read requests over IPC as the"input"service. The kernel does not read port0x60whilekbd0is claimed, preserving the byte for userspace, and reclaims it for the fallback console only after userspace exits - Devices are named, not addressed:
request_device("ata0")andrequest_device("kbd0")are checked against aDeviceAccesscapability, and the kernel decides which ports the name means — so a keyboard driver cannot touch the disk or pulse CPU reset via 0x64 - One driver per device at a time: while a ring-3 driver holds
ata0orkbd0, the kernel refuses other claims, and the claim is released even if the driver crashes - Read-only FAT32 — BPB validation, cluster-chain walking, long filenames — in ring 3, reading sectors over IPC
- CMOS RTC, so
dateprints the real date
Userspace
init— process 1, and the only program the kernel starts. Brings up the block driver and the filesystem in dependency order, hands the console to the shell, and shuts them down afterwardsfs-service— read-only FAT32 in ring 3, reading sectors fromata-driverover IPC and serving the shell over IPC- Services find each other by name.
register_serviceclaims a name;lookup_servicereturns the pid and grants a capability for it, which is what lets two processes that are not parent and child exchange messages at all hello— a static ELF that proves the loader, and exercisesmmap,munmap,clock_gettimeanddebug_printksh— a shell in ring 3 with line editing, history, arrow keys, a recursive-descent parser, andls/cat/cd/stat/date/getpidkshcan launch programs: unknown commands becomespawn+wait, andcmd &starts one in the backgroundata-driver— the ATA driver as an ordinary ring-3 process, driving a real disk withinandoutand answering block reads over IPCkbd-driver— the PS/2 keyboard driver, in ring 3, waking on IRQ 1 and delivering scancodes and decoded characters tokshover IPC as"input"- The SSE register file survives both a system call and a context switch, which is what a userspace driver and a userspace filesystem running concurrently needs and what neither used to get
Processes and IPC
- A process is a ring-3 thread with an address space, registered in the process table at the same id — one namespace, not two
send_message/receive_messagecarry real bytes, with a blocking receive that parks the thread rather than spinning- Capabilities that can refuse:
forkandspawnopen a channel between a process and its parent, and nothing else. A message to any other process isPermissionDenied, and the test checks it
In-kernel console
- A fallback shell on the same keyboard, which takes over when
kshexits — so there is a prompt on a system whose userspace has died
39 numbers are defined; 21 do the work and the other 18 return an error saying so. Nothing returns success for work it did not do.
Nine went away this phase rather than staying and refusing: open, close,
lseek, stat, fstat, mkdir, rmdir, unlink, getdents. A file is a
message to the fs service now, and a syscall number that exists and always
fails is a worse answer than one that does not exist.
Working (all 21 exercised from ring 3, not just by the kernel checking itself):
exit fork exec wait getpid getppid yield spawn · mmap
munmap · read write · time clock_gettime · send_message
receive_message · request_device release_device · register_service
lookup_service · debug_print debug_dump
read is stdin (with ksh reading from the "input" service over IPC and falling
back to the syscall) and write is the console — the display and timer remain
inside the kernel.
Refuses with NotSupported, honestly:
kill · mprotect brk sbrk · fstat mkdir rmdir unlink ·
reply_message create_channel destroy_channel · all 4 driver calls · all 4
capability calls · uname sysinfo
Being explicit about this, because the earlier version of this file claimed most of it.
- Demand paging is anonymous-only. A reserved page is always filled with
zeros. File-backed mappings would need the fault handler to read through the
VFS, so
mmapstill refuses anything butMAP_ANONYMOUS. There is also no reclaim: nothing ever takes a page back, so the only pressure valve is a process exiting. exectakes noargv.spawndoes —kshpasses command-line arguments andinitstarts the driver asata-driver ata0— butexecstill replaces an image with the command line it already had.- The filesystem is read-only.
kshrefuses redirection rather than pretending; there is nomkdir,unlinkor write path. - IPC is send/receive only.
reply_message,create_channelanddestroy_channelstill refuse, and there is no timeout on a blocking receive: a process waiting for a message that never comes waits forever. - Capabilities are not exposed to userspace. They are enforced on the IPC
path — a process may message its parent and its children, and nothing else —
but the four capability syscalls still return
NotSupported, so a program cannot inspect or delegate what it holds. - The fallback console has no file commands. It runs only because
userspace has stopped, so it cannot ask the
fsservice either — a fallback that needs the thing it is a fallback for is not one.lsfrom the kernel prompt says so, and the disk commands live inksh. - IPC has no reply port, only a sender filter.
receive_messagecan wait for a named process, which is enough for request/reply and is what the services use. It is not enough to tell two outstanding requests to the same server apart — that needs a per-request token, and nothing here has two in flight yet. - Drivers are trusted by boot-module name.
DRIVER_IMAGESinusermode.rsmapsata-drivertoata0, so the trust root is "GRUB loaded it from the ISO". A real capability system delegates frominit; this is a two-entry table standing in for one. - Interrupt delivery is a wake, not a message. A driver can sleep on a line
it owns —
ata-driverdoes, and 104 of 104 sector waits are woken by IRQ 14 — but an interrupt cannot carry anything, and a line cannot be handed to a process that does not own the whole device. userspace/driver-managerdoes not run. Not built, not on the ISO, no linker script.initandfs-servicewere in the same state until this phase; both are real now.drivers/*are not drivers.storageandnetworkare trait skeletons whoseinitsets a bool and returnsOk(());keyboarddeclares the PS/2 ports and then returns0fromread_datawith a comment saying it would use them in a real implementation;graphicsdoes touch0xB8000but is built for the host and never loaded. The kernel's own ATA and keyboard drivers are the real ones.shared/kosh-ipcandshared/kosh-driverare used by nothing that ships.- No ARM64.
aarch64-kosh.jsonis a valid target spec and nothing more: the build does not compile for it (the panic handler,serial,vga_bufferandmemoryare all x86-only and not gated), there is no boot path, andkernel/src/platform/aarch64/says "stub implementations" in its own module docs. - No network, no graphics beyond VGA text, no UEFI, no SMP.
- Swap and power management are disabled in
boot.rs, with the reason in-line: swap tried to allocate 8 MiB from a 1 MiB heap, and power management was entirely simulated.
kernel/src/
boot32.rs 32-bit Multiboot2 trampoline; builds the bootstrap map
boot.rs init_kernel: brings each subsystem up, in order
gdt.rs GDT/TSS, with the descriptor order sysret forces
percpu.rs per-CPU block reached through gs:, for the syscall stub
interrupts/ IDT, exception handlers, PIC, PIT, keyboard
memory/
physical.rs bitmap frame allocator
paging.rs the kernel's own page tables, W^X, physmap
address_space.rs per-process PML4s
heap.rs first-fit allocator with coalescing
task/ preemptive kernel threads and the context switch
syscall/ SYSCALL entry, dispatcher, user-pointer checks, files
elf.rs static ELF64 loader
ipc/services.rs name registry; a lookup is also a capability grant
console/ the in-kernel fallback shell
platform/rtc.rs CMOS real-time clock
platform/devports.rs which ports a named device is, and who holds it
userspace/
hello/ static ELF that proves the loader and the newer syscalls
hello2/ what hello execs into — a different image at the same address
ata-driver/ the ATA driver, in ring 3, talking to the disk over IPC
kbd-driver/ the PS/2 keyboard driver, in ring 3, delivering keys over IPC
fs-service/ read-only FAT32, in ring 3, over IPC in both directions
init/ process 1: starts the services, then the shell
shell/ ksh
docs/BOOT.md how all of the above actually works
scripts/run.sh build, ISO, disk image, boot, and the two CI gates
Everything else in the tree — drivers/, userspace/init, fs-service,
driver-manager, shared/, and the other scripts — is either vestigial or not
yet wired up. See What does not work.
Roughly in order, each unblocking the next:
- Boot to long mode, higher-half kernel
- Page tables, W^X, physmap, heap
- Preemptive scheduling
- Ring 3,
SYSCALL/SYSRET, ELF loading - ATA + read-only FAT32
- A shell in userspace that can launch programs
- Per-process address spaces
-
forkandexec - Copy-on-write, so a
forkcosts a page table rather than a program - Demand paging, so an untouched
.bsscosts nothing at all - One process-id namespace, so IPC and capabilities became reachable
- The disk driver out of the kernel, with port permissions to make it possible
- The filesystem out of the kernel, and an
initthat starts both -
fs/andblock/deleted — the kernel cannot read a disk at all - Selective receive, so a service can be asked something while it waits
-
argvforspawn, so a driver is told what to serve rather than knowing - Interrupts delivered to ring 3, so a driver sleeps instead of polling
- The keyboard out of the kernel — the last driver inside it
- FAT32 writes
-
userspace/initdoing its job
(The first boot)
Memory safety without a runtime, and unsafe as a marker rather than a mode:
the places where this kernel talks to hardware are visible in the source because
they have to be spelled out. It has not prevented a single one of the bugs in
docs/BOOT.md — those were all logic, not memory safety — but it did keep them
to logic.
MIT. See LICENSE.