Run AmigaOS 4 PowerPC binaries inside a sandbox process so that buggy guests crash the guest world but leave the host AmigaOS 4 system running. Targeted use case: iterating on drivers, libraries and Intuition apps that would normally cost a reboot per crash.
Non-goals:
- Running a non-OS4 OS as a guest. The guest ABI is OS4 PPC.
- Full instruction-level isolation. We are not QEMU.
- Running guest code that requires supervisor mode (drivers that touch real
hardware, custom interrupt handlers,
Supervisor()-mode work). Those will be refused at the shim boundary or trapped.
Three OS4 features make a soft sandbox practical:
-
Interface-based ABI. Every function call into the system goes through a vtable (
IExec,IDOS,IIntuition, …). We can hand the guest a replacementIExecand replacement library bases. The guest cannot tell the difference because the cross-call ABI is identical. -
Extended Memory (ExtMem) above the 2 GB barrier. The
ASOT_EXTMEMsystem object plusExtMemIFace::Map/Unmaplets us allocate physical pages above 2 GB and map a window of them into virtual address space on demand. We back the entire guest world from one extmem object so the guest never consumes scarce sub-2 GB virtual address space — and the host's working set is unaffected. -
AllocSysObjectparametrisation. Tasks, mutexes, message ports, lists are all created viaAllocSysObject(type, tags). We can intercept the allocator and own the lifetime of every kernel object the guest creates.
Combine these: a guest's code runs natively, calls a vtable we control, allocates from a pool we own, creates kernel objects we own, and lives in pages backed by extmem we manage. When the guest blows up, we tear down its pool, its kernel objects, its mapped window — host stays up.
A wild store from guest code can still hit a host address. We do not control the MMU from userspace. So the isolation we offer is at the API boundary, not the memory boundary:
- Guest doing
IExec->FreeVec(<garbage>)— caught (we own the pool). - Guest opening intuition.library and
CloseScreen()-ing a real host screen — caught (we hand it our fake intuition base; CloseScreen routes to our shim, which only closes screens our intuition opened). - Guest leaking tasks/semaphores/ports on crash — caught (we tracked every
AllocSysObjectand tear them down on guest abort). - Guest doing
*(uint32_t *)0xDEADBEEF = 0— not caught. Host may corrupt or alert. Same blast radius as today. This sandbox does not pretend to solve memory protection; it solves resource and API protection, which empirically is where most "had to reboot" failures actually come from.
The guest world lives in extended memory (above 2 GB physical)
except for executable code pages, which AOS4 maps non-executable
in the extmem region (verified live on real X5000: jumping to
extmem-backed code traps with ISI "Instruction fetch in non-execute
page"). Code therefore has to come from host RAM via
IExec->AllocVecTags(AVT_Type, MEMF_EXECUTABLE, …). Everything else
stays in extmem.
What lives where:
| Allocation | Pool |
|---|---|
Guest .text (PT_LOAD with PF_X) |
host RAM, MEMF_EXECUTABLE |
Guest .rodata / .data / .bss |
extmem |
Guest heap (every guest-side IExec->AllocVecTags/AllocSysObject) |
extmem (g->vmem) |
| Guest stack | extmem (when per-task stacks land); currently inherits host stack |
Private IExec clone |
extmem (vmem_alloc in private_iexec_create) |
vm.host.library synth Library + IVMHost interface |
extmem (vmem_alloc in vmhost_open) |
g2h / h2g rings |
extmem (vmring_create) |
| ELF phdrs / symtab / rela buffers (loader scratch) | extmem, freed before elf_load returns |
Resource ledger (struct GuestResource) |
extmem |
MEMF_EXECUTABLE is its own memory type per exec.doc (line 1146,
"defines a new memory 'type' (like MEMF_CHIP, MEMF_FAST)") — it must
be passed alone, not OR'd with MEMF_PRIVATE / MEMF_SHARED, or
the alloc fails with no diagnostic.
Concretely:
Host AS (4 GB virtual)
┌──────────────────────────────────────────────────────────────┐
│ 0x0000_0000 .. 0x7FFF_FFFF normal OS4 layout (host code, │
│ host heap, host stacks, libs) │
│ + guest CODE pages (MEMF_EXEC) │
├──────────────────────────────────────────────────────────────┤
│ extmem window (kernel-picked address, e.g. 0xF500_0000) │
│ ↕ 256 MB sliding window into the extmem object │
│ ↕ guest sees virtual addresses INSIDE this window │
│ ↕ holds: guest data, heap, IExec clone, IVMHost, rings │
├──────────────────────────────────────────────────────────────┤
│ 0xC000_0000+ host libraries, MMIO, etc │
└──────────────────────────────────────────────────────────────┘
ExtMem object (up to N GB physical, backing store)
┌────┬────┬────┬────┬────┬────┬────┬────┬────┬────┬────┬───────┐
│ pg0│ pg1│ pg2│ pg3│ pg4│ … guest heap … │ pgK│ … guest data … │
└────┴────┴────┴────┴────┴────┴────┴────┴────┴────┴────┴───────┘
For typical OS4 apps .text is a small fraction of total memory (a few
MB at the high end vs. tens-to-hundreds of MB of heap), so sub-2GB
pressure stays modest even running many guests in parallel — the bulk
of each guest's working set still lives above 2 GB.
The window is mapped and unmapped at page granularity through
IExtMem->Map(baseAddress, length, offset, flags). Pass NULL for
baseAddress to let the kernel pick — EXTMEMF_FAIL_UNAVAIL with a
preferred VA is rejected outright on this kernel rather than slid.
A naive loader could allocate the entire vaddr_lo..vaddr_hi span
(text + rodata + data + bss + zero-fill gaps) in MEMF_EXECUTABLE
host RAM. That would work but waste sub-2GB pages — for an app with
1 MB .text and 50 MB .bss, 51 MB of sub-2GB pool gets consumed
unnecessarily.
elf_loader.c therefore splits per PT_LOAD:
- PT_LOADs with
p_flags & PF_X→ host RAM,MEMF_EXECUTABLE - PT_LOADs without
PF_X→ extmem (vmem_alloc)
Wrinkle: per-segment loading means each section sits at a different
slide from its linked address. Relocations therefore can't share a
single delta across the whole image; each symbol's slide depends on
which PT_LOAD it lives in. Implementation:
- Walk PT_LOADs once, place each in the right pool, record per-segment
(linked_lo, linked_hi, host_addr). - For each
Elf32_Sym, look upst_shndx→ section header → which PT_LOAD covers that section's[sh_addr, sh_addr+sh_size). Computesym_delta = pt_load.host_addr - pt_load.linked_lo. - For each rela, target =
sym.st_value + r_addend + sym_delta, patch as before (ADDR32/16/HA/LO/RELATIVE). - R_PPC_REL24/REL14 stay slide-invariant only WITHIN the same PT_LOAD; if a branch crosses PT_LOADs that ended up at different slides, it needs an explicit relocation (the linker emits ones that span sections; ones that don't span are slide-invariant).
For the bundled bin/hello (~1.5 KB code, single REL24 within text)
the per-PT_LOAD split saves ~64 KB of sub-2GB per launch.
EXTMEMPOLICY_DELAYEDfor the backing object — physical pages are pulled from the >2 GB pool only when we touch them.- Window size at startup: defaults to 256 MB virtual (configurable). Size of
the extmem object: defaults to 1 GB physical (configurable up to
SysBase->MaxExtMem). - Inside the mapped window we run a simple slab + free-list allocator
(
vmem.c). Slabs come in 4 KB, 16 KB, 64 KB, 1 MB sizes. - Pages in the window that haven't been touched in a while can be unmapped
and remapped at a different offset, giving a paging-style overlay if the
guest world grows beyond
WINDOW_SIZE. (Current configurations size window == backing store and skip paging; see §10.2 for the cooperative paging design.)
MEMF_PRIVATE is task-scoped but still in the same physical pool as
everyone else. ExtMem above 2 GB is in a separate physical pool; an OOM in
the guest world cannot starve the host's <2 GB allocations, and vice-versa.
That's a real isolation win the OS already gives us for free.
The forced-into-host-RAM .text allocation only weakens the OOM-isolation
story for code pages — a very small fraction of typical guest working sets,
and one that does not grow at runtime. Heap exhaustion (which is the
realistic OOM scenario) still hits extmem, never the host pool.
SysBase->MaxLocMemandSysBase->MaxExtMemare legacy 68k-era fields and not reliably populated on PPC OS4 —MaxExtMemreads NULL on a 4 GB X5000 even though ExtMem fully works. Don't gate on them.- The modern surface is
IExpansion->GetMachineInfoTags()from<libraries/expansion.h>:GMIT_TotalPhysicalMemory→uint64storage; the only API that surfaces total RAM including the >2 GB pool.GMIT_Machine→uint32storage;MACHINETYPE_*enum (X5000/20=9, X5000/40=10, A1222=11, PEGASOS2=5, …). Use it to refuse Pegasos2 up-front (per AmigaOS wiki, ExtMem unsupported).GMIT_MachineString→STRPTR; human-readable name.GMIT_MemoryAvailableis not populated on the kernel currently in use (returns the int64-min sentinel0x80000000_00000000).
AvailMem(MEMF_TOTAL)is bounded by 32-bitULONGand reflects only the AOS-virtual pool; it cannot represent ≥ 4 GiB and never reflects the >2 GB pool.ASOEXTMEM_Size'sti_Datais a pointer to auint64, not the value inline (struct TagItem.ti_Dataisuint32). Passing inline reads as 0 on PPC big-endian; the kernel silently accepts that with bogus behaviour (8 GiB DELAYED reservations on a 4 GB box "succeed"). Always pass&uint64.
ExecIFace is a vtable. We AllocSysObject(ASOT_INTERFACE, …) for a fresh
interface, then memcpy the host IExec's function pointers into it as a
default. We then overwrite specific function pointers with our own
implementations:
static struct ExecIFace *priv_iexec_make(void) {
struct ExecIFace *p = AllocVecTagsExt(sizeof(*p), AVT_Type, MEMF_SHARED, ...);
memcpy(p, IExec, sizeof(*p)); /* default = pass through */
p->AllocVecTags = priv_AllocVecTags;
p->AllocVecTagList = priv_AllocVecTagList;
p->FreeVec = priv_FreeVec;
p->AllocPooled = priv_AllocPooled;
p->FreePooled = priv_FreePooled;
p->OpenLibrary = priv_OpenLibrary;
p->CloseLibrary = priv_CloseLibrary;
p->GetInterface = priv_GetInterface;
p->DropInterface = priv_DropInterface;
p->CreateTask = priv_CreateTask;
p->CreateTaskTags = priv_CreateTaskTags;
p->DeleteTask = priv_DeleteTask;
p->AllocSysObject = priv_AllocSysObject;
p->FreeSysObject = priv_FreeSysObject;
p->Forbid = priv_Forbid; /* downgraded — see §5.3 */
p->Permit = priv_Permit;
p->ColdReboot = priv_RefuseReboot;
p->IceColdReboot = priv_RefuseReboot;
p->Supervisor = priv_RefuseSupervisor;
p->SuperState = priv_RefuseSupervisor;
p->CacheControl = priv_NopCacheControl;
p->SetIntVector = priv_RefuseIntVector;
p->AddIntServer = priv_RefuseIntServer;
return p;
}Roughly 40 of the ~250 ExecIFace methods need overrides. The rest are either pure (list ops, AVL, RawDoFmt, CopyMem) or already-safe queries.
A normal OS4 ELF gets IExec resolved at link time via the standard startup
glue, which reads SysBase->MainInterface (or equivalent). Two options:
- Option A — modify startup glue. Build guests against a custom CRT0 that
takes
IExecfrom a known address (slot in TaskStorage, register r2 pointer passed via a calling convention we control). Cleanest, but requires building guests with our toolchain. - Option B — patch on load. When loading the guest ELF, scan for the
_main_interfaceimport /IExecglobal and replace it with our pointer during relocation. Works for already-built binaries.
The current loader ships option A with a thin custom crt0.o
(crt0_guest.c). Option B is future work.
Forbid() halts the host scheduler — a guest that wedges inside Forbid()
locks the host. We replace these with a per-guest mutex + counter:
void priv_Forbid(struct ExecIFace *self) {
Guest *g = guest_from_iexec(self);
MutexObtain(g->forbid_mutex);
g->forbid_depth++;
}This loses cross-task atomicity for guest tasks against host tasks (which is fine — guest cannot expect to lock the host scheduler) but preserves it within the guest world.
For every OS4 library a guest opens, we return a synthesized Library *
whose lib_NegSize jump table consists of ba <trampoline> instructions.
Each trampoline pushes the function index and tail-calls shim_dispatch,
which routes to one of three things:
- Pass-through — call the real host library function.
- Pass-through with bookkeeping — call host, record the returned handle in the guest's resource ledger so we can free it on guest abort.
- Replacement — service entirely from inside the sandbox.
Same trick for the GetInterface() result: we memcpy the real interface,
overwrite specific slots.
| library | strategy |
|---|---|
dos.library |
Pass-through with bookkeeping for Open/Close, Lock/UnLock, AllocDosObject*/FreeDosObject, ObtainDirContext*/ReleaseDirContext. Read/Write/IoErr pass through unchanged. See src/shims/dos_shim.c. |
intuition.library |
Pass-through with bookkeeping for OpenScreen*, OpenWindow*, LockPubScreen. CloseScreen / CloseWindow / UnlockPubScreen validate that the handle is in the guest's ledger via guest_find() and refuse if not — a guest cannot close a screen/window the host owns. See src/shims/intuition_shim.c. |
expansion.library |
IPCI clone with FindDevice / FindDeviceTags overridden so synthetic PCI devices can be returned in place of real hardware. See src/shims/expansion_shim.c and src/shims/pci_synth.c. |
Per exec.doc/CreateTaskTags: "An Exec task may not call dos.library
functions or any function which might cause the loading of a disk-
resident library, device, or file." Without pr_CurrentDir /
pr_FileSystemTask / the rest of the Process state, dos.library
returns [DOS] ERROR - <op> called by task <name> and the call fails.
SandboxVM therefore spawns the guest via IDOS->CreateNewProcTags()
(NOT IExec->CreateTaskTags), passing args via
NP_UserData → pr_Task.tc_UserData:
struct Process *p = IDOS->CreateNewProcTags(
NP_Entry, guest_runner_proc,
NP_Name, g->name,
NP_StackSize, 64 * 1024,
NP_UserData, &args,
NP_Child, TRUE,
TAG_END);struct Process starts with struct Task, so the returned Process *
is downcastable to Task * for DeleteTask / FindTask /
SetTaskTrap calls. tc_UserData doubles as the Guest pointer
slot once the runner has consumed args.
graphics.library, utility.library, asl.library, application.library
follow the same pattern. Locale, iffparse, datatypes are mostly
pass-through.
| library | strategy | rationale |
|---|---|---|
exec.library |
private IExec clone with overrides | Memory, libraries, devices, IO, tasks, interrupts, DMA all sandboxed. Legacy AllocVec / AllocMem / Allocate route through vmem; SetTaskPri clamped to <=0 (gap 4/7). |
dos.library |
shim with overrides | Handle tracking for Open/Lock/AllocDosObject/ObtainDirContext. Open refuses MODE_NEWFILE / MODE_READWRITE on system roots (gap 2/7). SystemTags / System / CreateNewProc* / RunCommand refused outright (gap 3/7). |
intuition.library |
shim with overrides | Tracks OpenScreen / OpenWindow / LockPubScreen. CloseScreen / CloseWindow validates ownership against the ledger. |
expansion.library |
shim with synthetic PCI | FindDevice* returns synthesized devices in place of real PCI. |
graphics.library |
pass-through clone with trailer guard | API surface is rendering-only; no host-side handles to track beyond what intuition already tracks. |
utility.library |
pass-through clone with trailer guard | Tag-list / string helpers; no kernel-state side effects. |
iffparse.library |
shim with handle tracking | Tracks AllocIFF and OpenClipboard (gap 5/7). OpenIFF wraps an already-tracked dos BPTR. |
ogles2.library |
pass-through clone with aglCreateContext2 tracking |
GL context is host-owned; ledger pairs with aglDestroyContext. |
keymap.library |
pure pass-through clone | Read-only key-map lookups; no resources. |
textclip.library |
shim with ReadIFFClip tracking |
Host allocates the returned buffer; ledger frees it. |
icon.library |
shim with DiskObject tracking | GetDiskObject returns host-allocated icon state; ledger pairs with FreeDiskObject. |
workbench.library |
shim with AppIcon/AppWindow/AppMenu tracking | Each Add*/Remove* pair is ledger-tracked. |
AmigaInput.library |
shim with handle tracking | Tracks ObtainDevices / ReadEvent registrations. |
math*.library (mathieeedoubbas, mathieeesingbas, mathffp, mathtrans, etc.) |
explicit pass-through, no shim | These are pure-function libraries -- no kernel objects, no host state, no resources to leak. The guest opens the host's real Library*; method calls execute on the host's vtable. Sandbox value of shimming would be zero. |
locale.library |
explicit pass-through, no shim | Locale catalogs are managed by OpenCatalog/CloseCatalog, but the host catalogs are read-only shared resources opened from LOCALE:Catalogs/. A guest leaking a catalog handle bumps a refcount; nothing the host can't recover at process exit. Not worth a shim's overhead. |
realtime.library |
explicit pass-through, no shim | Time-source ticks only; no per-guest state beyond what tasks already own (and guest tasks are ledger-tracked via priv_CreateTask). |
datatypes.library |
pass-through (handle tracking deferred) | NewDTObject returns a host-allocated object; eventual ledger tracking would parallel iffparse. No current guest exercises this surface. |
"Explicit pass-through, no shim" rule. A library lands in that column if and only if all three hold: (a) the library exposes no kernel handles, no host buffer ownership, no spawn surface; (b) any state it owns is process-global, refcounted, or otherwise reclaimed by the host at process exit; (c) the overhead of cloning + trailer guard would dominate the API cost. Adding a shim for a library that fits all three is pure noise -- the shim has nothing to do.
Every shimmed call that returns a kernel handle adds a row to the guest's ledger:
struct GuestResource {
struct MinNode node;
enum { RES_WIN, RES_SCREEN, RES_FILE, RES_LOCK, RES_PORT, RES_SEM,
RES_TASK, RES_LIB, RES_IFACE, RES_SIGNAL, RES_DOSOBJ } kind;
APTR handle;
APTR release_fn; /* function used to free this thing */
APTR release_arg; /* extra arg if the release fn needs it */
};
Teardown walks the list in reverse order and calls each release_fn. This
is the single most important property the sandbox provides: deterministic
resource recovery on guest crash.
host main()
├─ extmem_init() (alloc ASOT_EXTMEM, map window)
├─ vmem_init() (heap inside the window)
├─ guest = guest_create() (private IExec, ledger, forbid mutex)
├─ elf = elf_load("PROGDIR:hello", guest)
├─ task = IExec->CreateTaskTags(
│ "SandboxVM guest #1",
│ 0, elf->entry, 64*1024,
│ TASKTAG_USERDATA, guest, /* lookup hook */
│ CHILD_Launcher, guest_trampoline,
│ TAG_END);
├─ Wait(SIGBREAKF_CTRL_C | guest->done_sig)
└─ guest_destroy(guest) (walks ledger, unmaps window, frees extmem)
guest_trampoline is what the host task lands in. It:
- Stashes
guestintc_UserData. - Replaces the implicit IExec lookup so further calls from this task hit
the private vtable. (Mechanism: the guest CRT0 reads IExec from
tc_UserData->iexecinstead of fromSysBase.) - Calls into the loaded ELF entry point.
- On return / on a
LONGJMPfrom a fatal-handler hook, signals the host.
Loads a stripped, statically-linked OS4 PPC ELF.
- Parse
Elf32_Ehdr, walk PT_LOAD program headers. - For each PT_LOAD: allocate from the appropriate pool (host RAM
for
PF_Xsegments viaIExec->AllocVecTags(MEMF_EXECUTABLE), extmem viavmem_allocotherwise — see §4.1). Readp_fileszbytes from disk; zero the BSS tail. - Apply
R_PPC_ADDR32,R_PPC_ADDR16_*,R_PPC_REL24,R_PPC_REL14,R_PPC_REL32, andR_PPC_RELATIVErelocations from every.rela.*section. Each relocation uses a per-segment slide (host_addr − linked_lo) so PT_LOADs that landed in different pools resolve correctly. - Resolve
_SDA_BASE_from the symbol table for use in synthesised driver interfaces. - Cache-flush every executable segment with
IExec->CacheClearE.
Dynamic linking against shimmed library bases is out of scope:
guests call IExec->OpenLibrary explicitly through the private
IExec.
Two layers, modelled on virtio's split-ring + config-space pattern.
The guest reaches the host via a synthesised library:
struct Library *lb = IExec->OpenLibrary("vm.host.library", 0);
struct VMHostIFace *IVMHost =
(struct VMHostIFace *)IExec->GetInterface(lb, "main", 1, NULL);Both calls are intercepted by the private IExec. The Library and
Interface are crafted by hand in extmem (not via MakeInterface,
which allocates from the system pool) so the comms structures live
above 2 GB along with everything else guest-visible. Per
exec.library/MakeInterface (autodocs), the interface layout is
fixed by struct InterfaceData followed by function pointers; we
just zero-init Data, set LibBase and Version, and fill the
function-pointer slots manually. The guest cannot tell the difference.
Methods:
| method | purpose |
|---|---|
Hello(guest_abi) → host_abi |
version negotiation; host returns 0 if guest below VMHOST_ABI_MIN |
Log(text) |
cheap pre-ring log path; DebugPrintFs through host |
GetHostTask() → Task * |
guest stashes this for direct Signal() calls |
RingNegotiate(guest_bit, &host_bit, &g2h, &h2g) |
exchange signal bits, hand back the ring pair |
Doorbell() |
Signal(host_worker, 1<<host_bit) |
RequestShutdown(rc) |
polite teardown |
HostCall(verb, data, size) |
catch-all RPC for forward-compat |
Latency is one indirect call per method.
A struct VMRing is a single contiguous block in extmem laid out as
+-----------------+
| header (magic, |
| capacity, etc) |
+-----------------+ <-- producer cache line
| volatile head |
+-----------------+ <-- consumer cache line
| volatile tail |
+-----------------+ <-- doorbell wiring
| target_task |
| target_sigmask |
+-----------------+
| slots[capacity] |
| ... |
+-----------------+
head and tail are free-running uint32; the in-buffer index is
head & (capacity - 1). capacity is a power of two so the mask is
free.
Producer/consumer are SPSC, so the only synchronisation needed is a
release/acquire pair around the index bumps. PowerPC has weakly-
ordered memory; we use lwsync for release and isync for acquire
(and rely on the natural address-dependency for the slot read on the
consumer side).
A pair of rings is created per guest:
g2h— produced by guest task, consumed by host workerh2g— produced by host worker, consumed by guest task
Per the exec.library/AllocSignal autodoc, "allocated signals are
only valid for use with the task that allocated them." That rules
out a single shared signal bit. Each side must AllocSignal() in its
own context and publish the bit number to the other through
RingNegotiate().
Per the exec.library/Signal autodoc, Signal is "safe to call
from interrupts" and may target any task in any state. So the
producer can poke the consumer freely from anywhere — including from
inside the guest's library shim trampolines, where Wait() is not
permitted.
vmring_push and vmring_pop deliberately do not signal — the
caller batches enqueues and rings the doorbell once at the end. This
is the same pattern virtio uses to avoid one MSI per descriptor.
A producer that wedges inside vmring_push between the slot writes
and the head bump will never increment head. A consumer waiting
forever would deadlock the host worker too. Mitigations:
- The host worker uses
Waitwith a timeout. (Waititself has no timeout argument; we set up atimer.devicerequest alongside the signal mask, exactly the pattern from the autodoc comment underWaitIO().) - On timeout, the worker compares the current
g2h->headagainst the last-seen value. If unchanged across N consecutive timeouts AND the guest task is in a wait state with no signals pending, the guest is declared dead andguest_destroy()runs. - A
tc_TrapCodeinstalled on the guest task aborts the guest and signals the host worker via the doorbell on PPC machine-check / DSI / ISI / alignment / illegal-instruction / privilege traps.
- The shim already has
IExec->Signalavailable — no need for a custom doorbell mechanism. AllocSysObject(ASOT_PORT)is intercepted, so the guest can't accidentally publish a public message port that lets host code reach into the sandbox. All guest-visible communication channels flow throughIVMHost.- The ring lives in extmem, which means the host cannot accidentally read or write it from a sub-2 GB virtual address — it has to map the window first. That's an extra cheap lock against host code scribbling on the comms area.
Cross-compile with the existing walkero/amigagccondocker:os4-gcc11 image.
Top-level Makefile shells out to a thin docker-cc wrapper that mounts
the project + the SDK read-only into the container and invokes
ppc-amigaos-gcc.
make # builds bin/sandboxvm
make hello # builds the bundled guest test (test/hello.c)
make run # SCP to a real OS4 box (target IP from SANDBOXVM_HOST env)
make clean
The host binary uses -mcrt=newlib and -lauto. The guest test binary uses
our custom crt0_guest.o.
sandboxvm [opts] guest1 guest2 ... runs each ELF in turn. Per-guest
state (ASOT_EXTMEM pool, private IExec clone, IVMHost worker task,
resource ledger, DOS Process) is freshly created and torn down for
each path. Names are <base>.<idx> so debug traces stay
disambiguated. The host process exits with the last non-zero guest
rc (or 0 if every guest succeeded). A trap-killed guest does not
poison the next guest.
Concurrent multi-guest is not yet implemented. The architecture
already supports it (per-guest extmem and worker task), but no
shared global is wired up — parallelising is a matter of converting
main.c's loop into N parent threads and choosing a global
window-allocator policy.
Tier 1 — IVMHost->RemapWindow / ReleaseWindow. The simplest
useful subset: cooperative auxiliary mappings beyond the primary
heap window. The heap allocator (vmem.c) is unchanged; two
additional vtable slots let a guest reach ASOT_EXTMEM backing past
window_size by asking the host to IExtMem->Map an additional
range and hand back a host VA. Each aux mapping is ledger-tracked
(RES_MEM with a release_aux_mapping callback) so a forgetful
guest doesn't leak — guest_destroy reaps.
This tier explicitly has no eviction: every RemapWindow takes
a fresh ExtMem range. Useful for streaming I/O buffers and one-shot
large allocations; not enough for a heap that exceeds the window.
Tier 2 — slot-aware vmem.c over Window.
The allocator is built on top of a Window module
(src/window.c, include/window.h) that divides the host VA window
into WINDOW_SLOT_SIZE-byte slots (default 64 KiB, equal to the slab
arena size). Each slot:
- records the backing range it currently maps,
- keeps a
live_countrefcount of allocations that live in it, - is eligible for eviction (Unmap + Map to fresh backing) only when
live_count == 0— the cooperative-paging invariant is "a live pointer keeps its slot mapped."
vmem.c calls window_acquire_arena instead of bumping a single
heap cursor whenever a slab class needs a new arena or a one-off
large allocation (sized > SLAB_MAX). Each slab/large alloc does
window_inc on its slot; vmem_free does window_dec. A "phantom"
refcount of 1 is held by the slab class on its current arena, then
dropped when the class moves on; the slot's mapped state is from
that point governed entirely by its remaining live chunks.
Slot acquisition policy is two-pass:
- Reuse a slot that's already mapped AND has refcount 0 (free arena recycling, no remap).
- Otherwise pick any refcount-0 slot, Unmap it (if mapped), Map a
fresh range from
next_backing_off, advance. - If every slot has refcount > 0 OR fresh backing is exhausted,
vmem_allocreturns NULL.
Honest limits of the slot-aware tier:
- Pass-1 reuse means the heap recycles slot mappings rather than
driving new mappings up to
backing_size. In a stable workload the working set still has to fit inwindow_size. To actually exceed the window the guest has to either let allocations fully drain (so all slots hit refcount 0) and force the next acquire to land in a remapped slot, OR a future stage needs to add cooperative pin/unpin so the policy can prefer "stale" slots. - "Freed but still-valid backing" is not retained beyond eviction — a slot that gets remapped loses access to its old backing range forever (future work: backing-range recycling).
- Multi-slot allocations are not supported.
large_alloccaps atWINDOW_SLOT_SIZE - sizeof(Chunk). Guests that want jumbo buffers should useIVMHost->RemapWindowinstead.
What the slot-aware tier does deliver: the load-bearing structural piece. The heap is no longer a single contiguous bump arena — it's a set of slots, each independently mappable. Future work can add smarter eviction policies (LRU on refcount-0 slots that prefer recently-touched), pin/unpin verbs, range recycling, or multi-arena large allocs without breaking the chunk/free-list contracts that current callers rely on.
When window_size >= extmem_size, both tiers degenerate to no-ops
and the guest sees a single permanently-mapped contiguous heap —
which is the configuration current guests run on.
- Hypervisor backstop. On X5000 (P5020) we could eventually back this with the e500mc hypervisor for true MMU isolation. Out of scope for the current implementation, but the architecture above doesn't preclude it — the IExec shim becomes a hypercall front end.
- PLT /
bctrlretargeting for late-bound libraries. Sidestepped today by requiring the guest to callOpenLibraryexplicitly through our shim. tc_TrapCodereach. OS4 lets you set a trap handler per task that fires on PPC machine-check / DSI / ISI. We will set one per guest task that does an immediatelongjmpback intoguest_trampoline. This is how the real "host stays up when guest crashes" guarantee gets delivered for the wild-store case the API shim cannot catch.