A standalone image build engine: boot a machine, run ordered provisioning steps against it over SSH, capture the result as an image, tear the machine down.
Proof of concept for frappe/atlas#74.
Runs on its own — no Frappe, no bench, no Atlas. See SPEC.md for the
design and the plan to fold the engine back into Atlas later.
Status: verified end to end. uv run pytest → 40 passed, and a real
kiln run has actually booted a QEMU VM, provisioned it over SSH, and produced
a genuine bootable disk image — not simulated, not mocked. See
Testing guide for exactly what that run did and how to
reproduce it, and for what is not yet verified (Linux, the ssh-target
builder against a real host, an x86_64 guest).
kiln run <template.toml> build an image
kiln retry <build-id> re-run a finished build from a clean machine
kiln status [<build-id>] list builds, or show one
kiln logs <build-id> [--step S] print step output
kiln gc [--older-than 30m] destroy machines left by crashed/kept builds
kiln validate <template.toml> check a template without running it
Two builders:
--builder |
What it does | Needs |
|---|---|---|
qemu (default) |
Kiln boots a real VM from a base cloud image, provisions it, and qemu-img converts the result to a bootable rootfs.img. |
qemu, an ISO tool, a cloud image |
ssh-target |
Kiln boots nothing — you give it a reachable SSH host; it runs the steps and streams back a tar of the paths the template names. |
an SSH host |
cd kiln
uv sync
uv run kiln --helpNeeds Python 3.12+ and uv. No runtime dependencies.
brew install qemu xorriso # macOS
# Debian/Ubuntu: apt install qemu-system xorrisobrew install qemu also ships the edk2-aarch64-code.fd UEFI firmware an
arm64 guest needs; Kiln finds it automatically under share/qemu/.
Download an Ubuntu cloud image that matches your template's [base].platform:
mkdir -p ~/kiln-images
# Apple Silicon → arm64:
curl -Lo ~/kiln-images/noble-arm64.img \
https://cloud-images.ubuntu.com/noble/current/noble-server-cloudimg-arm64.img
# Intel → amd64:
curl -Lo ~/kiln-images/noble-amd64.img \
https://cloud-images.ubuntu.com/noble/current/noble-server-cloudimg-amd64.imgKiln makes a copy-on-write overlay per build, so the base image stays untouched.
uv run kiln validate examples/hello.toml
uv run kiln run examples/hello.toml \
--var base_image=$HOME/kiln-images/noble-arm64.imgOn success:
.kiln/output/hello-<build-id>/
rootfs.img a raw, bootable disk
manifest.json sha256, sizes, platform, os — the Atlas Virtual Machine Image field set
plus .kiln/builds/<build-id>/build.json recording every phase and step, and
.kiln/builds/<build-id>/machine/console.log with the guest's boot output.
uv run kiln run examples/hello.toml --builder ssh-target \
--var ssh_host=203.0.113.10 --var ssh_user=ubuntuThe machine must let you ssh in without a password prompt, and sudo without
one if your steps need root. A local VM (UTM, Lima, multipass) or a cloud VM both
work. The artifact is a rootfs.tar.gz, not a bootable disk.
uv run kiln status # all builds, newest first
uv run kiln status <build-id> # one build, phase + per-step detail
uv run kiln logs <build-id> # every step's output
uv run kiln logs <build-id> --step verify-health
uv run kiln gc --dry-run # show machines a crashed/kept build left behind
uv run kiln gc --older-than 30m # destroy themTeardown runs automatically at the end of every build. gc is the backstop for a
kiln process that was killed before its finally ran.
Three tiers: unit tests (hermetic, run every time), a real end-to-end QEMU run (verified once, reproducible by anyone with QEMU installed), and things that are designed to work but not yet actually run by anyone.
uv run pytest -q # 40 passedEverything here runs against FakeBuilder/FakePublisher, or a plain cat/
sleep subprocess standing in for SSH — no real machine, fully hermetic.
| File | Covers |
|---|---|
test_engine.py |
Every phase transition, persisted before the next step begins. A failing step stops the run and marks the rest skipped, naming the step and its exit code. Teardown runs on success and failure, from a finally, and a teardown that itself raises doesn't mask the real result. A build that never got a machine handle never calls destroy. A builder that fails to construct (before create() even runs) still reaches a terminal phase instead of leaving build.json stuck pending. A transport-level OSError from a step still fails the build cleanly instead of crashing the CLI. --keep-on-failure leaves the machine. The concurrency cap refuses a second run. gc reaps a stale build, --dry-run changes nothing, and one build's builder failing to construct doesn't stop gc from reaping the rest. A step's run =/script body gets {{var|file|env.*}} resolved, not just its env values. |
test_template.py |
Timeout unit parsing; a step needs exactly one of script/run; a missing script file; duplicate step names; empty snapshot_paths rejected outright rather than silently defaulting; a template or step name containing /, \, or .. rejected (it would otherwise reach a filesystem path). |
test_environment.py |
{{var.*}}, {{file.*}}, {{env.*}} resolve; an unknown placeholder fails loudly; a non-string value passes through untouched. |
test_ssh.py |
_pump round-trips a small payload and a 200KB one (deliberately larger than a pipe buffer, to prove the concurrent stdin/stdout loop doesn't deadlock the way write-everything-then-read would); the timeout path actually kills the process; the incremental UTF-8 decoder doesn't corrupt a multi-byte character split across a read boundary. |
test_ssh_target.py |
A snapshot_paths entry containing a space or shell metacharacters is shlex.quoted before reaching the remote command, not spliced in raw. |
This was run for real while building Kiln, on an Apple Silicon Mac, with only
hdiutil available as an ISO tool (no xorriso/genisoimage installed) —
deliberately the harder path, not the one the setup guide recommends:
curl -Lo /tmp/noble-arm64.img \
https://cloud-images.ubuntu.com/noble/current/noble-server-cloudimg-arm64.img
uv run kiln run examples/hello.toml --var base_image=/tmp/noble-arm64.img --boot-timeout 8mResult: success in under two minutes total (download not included). Verified
by hand afterward, not just by the exit code:
.kiln/builds/<id>/steps/02-read-it-back.logcontains real guest output —Linux kiln 6.8.0-139-generic ... aarch64 GNU/Linux— proving the VM actually booted this exact kernel and ran the step, not a stub.file .kiln/output/hello-<id>/rootfs.imgreports a real DOS/MBR boot sector with a GPT protective partition table — a genuinely bootable disk, not a corrupted or truncated file.manifest.json'simage_sha256/image_size_mibmatched the actual file.- The QEMU process (
builder_handle.data.pid) was gone frompsafterward,disk.qcow2/seed.isowere deleted from the machine directory, andkiln gc --dry-runreported nothing to reap — teardown left nothing behind.
Reproduce this yourself with any Ubuntu cloud image matching your host's architecture (arm64 on Apple Silicon, amd64 on Intel or most Linux hosts) — the setup guide above has the exact download commands. A first run takes as long as VM boot + cloud-init (roughly 1–2 minutes); repeat runs against the same base image are much faster once the base image itself is cached by your OS's disk cache.
To verify a failure path, point a template's last step at a command that
exits non-zero (e.g. run = "exit 1") and confirm: kiln status shows
failed with that step named, kiln logs <id> shows its output, and
kiln gc --dry-run still reports nothing leaked — teardown ran anyway.
Be honest with yourself about this list before trusting these paths blind:
- Linux, on any distro. Nothing in the codebase is macOS-specific —
selectors.DefaultSelectoruses epoll automatically,QemuBuilderalready branches onplatform.system() == "Linux"forkvmacceleration and includes the standard Debian/Ubuntu (qemu-efi-aarch64, path/usr/share/AAVMF/AAVMF_CODE.fd) and Fedora/RHEL (edk2-aarch64, path/usr/share/edk2/aarch64/QEMU_EFI.fd) UEFI firmware locations — but none of it has actually been run on a Linux box. Two things worth checking on first try: your user needs to be in thekvmgroup (or root) for/dev/kvmto actually be usable, not just present — a permissions-only failure surfaces as an opaque QEMU launch error, not a clean pre-check; and the exact UEFI firmware package/path can vary by distro version, so--var uefi=<path>is there as an escape hatch if the built-in candidates don't match yours. - The
ssh-targetbuilder against a real host. Its command-quoting and handle plumbing are unit-tested with a mockedsubprocess.run, but nobody has pointed it at an actual reachable machine and confirmed the tarball it produces is what you'd expect. - An x86_64 QEMU guest. The one real run so far was arm64-on-arm64 (HVF-accelerated, native speed). An amd64 guest on an Intel host, or any cross-arch combination (which falls back to full emulation, no accelerator, and is expected to be much slower), hasn't been tried.
_launch_with_port_retry's actual retry path. The free-port TOCTOU it guards against has never been forced to happen in practice.- Two builds running concurrently (the default concurrency cap is 2) — never actually run at the same time to confirm they don't collide on anything besides the port issue above.
An automated pytest -m integration that drives SshTargetBuilder against
KILN_TEST_SSH_HOST, and a Linux CI job that does the QEMU run above for
real, are the obvious next additions once this is worth wiring into CI.
uv run ruff check kiln tests
uv run ruff format --check kiln testsRuff settings mirror the Atlas app so code moves cleanly when the engine ports.
The engine and the Builder / Publisher interfaces are the stable part;
everything Atlas-specific is meant to live behind them. SPEC.md has the full
map — in short: template.toml → Image Build Template doctype, build.json →
Image Build doctype, kiln/engine.py → atlas.vm.core.image_build almost
unchanged, QemuBuilder → AtlasBuilder calling virtual_machine.create /
create_machine_image / terminate, LocalPublisher dropped because
create_machine_image already is the publisher.