Skip to content

Cleanup targets found on disk, Podman host-disk reclaim, and readable container images - #37

Merged
rducom merged 5 commits into
mainfrom
feat/discovered-cleanup-targets
Aug 16, 2026
Merged

rducom merged 5 commits into
mainfrom
feat/discovered-cleanup-targets

Conversation

@rducom

@rducom rducom commented Aug 16, 2026 •

Copy link
Copy Markdown
Collaborator

Low-Hanging Fruits only ever offered paths written down by hand. This adds a second kind of target, whose paths are found on the disk, plus a host-side reclaim for Podman and a readable container image list.

Figures below come from one developer Mac and are there to show scale, not as general claims.

New: discovered targets

CleanupDiscovery finds targets whose location depends on what is installed. Each one pairs a candidate directory with a fact on disk that explains it, and offers nothing it cannot pair:

Target Offered when Found here
bin/, obj/ a .csproj/.sln sits beside it 8,360 dirs, 20.8 GiB
target/ a Cargo.toml sits beside it —
node_modules a package.json beside it and no install for 6 months 5 projects, 2.0 GiB
workspace state its workspace.json names a folder that no longer exists 6 of 193
Chrome/Firefox caches inside the browser's own cache tree 5 profiles

The marker rule is what makes this safe. A find -name bin sweep over a home directory deletes every Python virtualenv on it; the rule rejected all of them across 8,360 real matches.

Two behaviours are new and worth knowing about:

  • removal: .directory — discovered targets delete the directory itself, where caches only ever had their contents emptied. Build tools recreate it.
  • regenerable: false — targets whose bytes never come back (Trash, orphaned editor state). Excluded from Select all, badged not a cache, named individually in the confirmation. Replaces the hardcoded id == "trash".

Podman: reclaim host disk

Pruning inside the VM frees guest blocks but leaves the host image allocated until the guest discards them, which CoreOS only does weekly. A new action runs fstrim on demand and reports the change in the image file's st_blocks.

fstrim's own output is not usable here — it printed 39.5 GiB for 2.06 GB actually returned — so the figure is measured before/after, never reported.

Container images

227 rows, 200 of them <none>, most the same size. Two changes:

  • Identity. A displaced image keeps its former tag in History; 198 of the 200 resolve. Full ladder: live tag → repository by digest → former tag → OCI labels → <none>. Everything below the first rung is inference, so the row carries a badge saying so. Resolved names group, collapsing the list to ~12.
  • Size. An image's size includes layers its neighbours also hold; the rows summed to 100.8 GiB against 17.3 GiB of real images. podman system df -v gives the bytes an image holds alone, shown when it diverges (906 MiB listed, 11.6 KiB alone).

No reclaim figure is offered per group. For 38 rebuilds of one tag, size sums to 42.6 GiB and per-image unique size to 0 B; neither is what removing the set frees, so the group says the layers are shared and defers to system df.

Fixes

  • Container prune buttons never opened their confirmation. Three .alert modifiers on one view; SwiftUI presents one. Verified against 8aa8fcd, so this predates the branch — three destructive actions have been unreachable in shipped builds. All dialogs now route through one enum and one modifier.
  • The installed skill was never refreshed. plan() treated "directory exists" as "up to date", so corrections shipped in the app never reached the assistant reading it (86 lines installed vs 157 bundled on this machine). Now compared by content, with an Update action in the panel.
  • Catalog paths. uv keeps its cache under XDG on macOS, so ~/Library/Caches/uv alone found nothing while ~/.cache/uv held 4.19 GiB. Same for pre-commit/prek, gh, opengrep, opencode. ~/.npm/_npx joins the npm target.
  • ToolActivity gap: go mapped to go-build but not go-mod.
  • The journal would have written one 8,360-path line per clean; capped at 32 with the true count kept.

Review focus

  • CleanupDiscovery.swift — the marker rules, the prune list, and isOldEnough. This decides what gets deleted, and it is the only place in the app that does so from disk contents rather than from a constant.
  • CleanupEngine.clean — the .directory branch. It deletes a whole directory, fenced by passesFence, which discovery reuses unchanged.
  • ContainerQueries.identify — every rung below the first is a guess presented as an image's name.
  • ContainerQueries.parseUniqueSizes — parses a table because podman refuses --format json with --verbose. Read from both ends; CREATED is a human duration of unpredictable width.

Testing

232 tests (+22). The refusals are pinned before the successes: virtualenv bin/ in four spellings, unmarked bin, node_modules not descended into, the depth bound, the fence on discovered paths, live vs orphaned workspaces, unresolvable workspace.json, the buildah hex-stem rule, a three-token CREATED column, and the group that must not invent a total. explain is covered through a real tools/call.

Every UI change was opened against the running app and read on screen; the four container dialogs were opened and cancelled rather than confirmed.

Known gaps

  • The "installed skill is outdated" banner has test coverage but its rendering was never seen — syncing the copy on this machine makes the state unreachable.
  • denseBurstCoalescesToItsClusterAncestor fails about 1 full-suite run in 5, alternating between two assertions, never in isolation. Pre-existing timing race in the test, untouched here.
  • Discovery adds ~10 s to the mode's load on a large home. It runs off the main actor, alongside catalog sizing rather than before it, so rows still stream in.

rducom added 5 commits August 16, 2026 21:47
…ever runs

The three largest reclaimable findings on a developer's Mac cannot be written
down as constants: per-project build output, per-workspace editor state, and
per-profile browser caches. Their paths depend on what is on the disk, so the
hand-picked catalog — safe because a human argued for each path in prose —
has nothing to say about them.

CleanupDiscovery substitutes a structural argument for the written one. Each
candidate is paired with a fact on disk that explains it, and nothing is
offered that cannot be paired: a bin/ only beside a project file, a workspace
folder only once its workspace is gone, a browser cache only inside the
browser's own cache tree. The refusals are the point — a bare `find -name bin`
over a home directory destroys every Python virtualenv on it. Measured on a
real disk: 8,360 folders, 20.8 GiB, zero virtualenvs.

Not offered as `dotnet clean`. That command cleans one configuration of one
project, leaves Release/ and project.assets.json, needs a restore and the SDK
a global.json pins; measured, it frees 28% of bin+obj. This is the opposite of
the go-mod case, where the vendor tool is the only thing that works at all.

workspaceStorage is the finding this corrects rather than adds. The usual
advice — empty it, it is only editor layout — is wrong twice: 85% of the bytes
are AI chat transcripts that nothing regenerates, and most folders belong to
repositories still on disk (6 of 193 here are actually orphaned). Each folder
names its project in workspace.json, so orphan status is decidable. Only the
orphans are offered, and `regenerable` now carries the Trash's protection to
them instead of a second hardcoded id.

Podman: pruning inside a VM frees guest blocks and leaves the host image
allocated, which is where "podman disks only grow, recreate the machine" comes
from. It is wrong — applehv passes discard through and CoreOS trims weekly, so
the blocks return up to seven days late. A Reclaim host disk action collapses
that to now. The freed figure is measured from the image's st_blocks, never
taken from fstrim, which reported 39.5 GiB for 2.06 GB actually returned.

Also: ownerBundleID generalises the hardcoded Notion gate to every app cache
(Chrome and Firefox are usually several GB and usually open); _npx joins the
npm target, routinely the larger of the two; explain now answers "NOT a
cleanup target" with the reason, because the reference doc already said the
right thing about workspaceStorage and that did not stop the wrong advice.

Two defects surfaced by the new scale: the journal would have written an
8,360-path line per clean (capped at 32, count kept exact), and ToolActivity
did not warn about builds writing the output they were about to lose — which
exposed a pre-existing gap where `go` mapped to go-build but not go-mod.
… what it frees

The image list was 227 rows on this machine, 200 of them called `<none>` and
most of them the same size. Nothing in it could be read, and the two numbers it
did show were both wrong in the same direction.

The engine knows more than the list was asking. When a build or a pull moves a
tag, the displaced image keeps the tag it used to carry in its `History`, so
198 of those 200 resolve to a real name. The rest fall through a ladder — live
tag, repository by digest, former tag, OCI labels — with `<none>` reserved for
the images nothing identifies. Each rung below the first is inference, and a
badge says so on the row: a name presented bare would claim a tag that now
points somewhere else. Buildah scratch names (`<64 hex>-tmp:latest`) become
"build leftover" rather than a repository that never existed, and digest
references lose 71 characters of hash to a badge.

Resolved names then group: forty rebuilds of one tag arrive as one line saying
so, collapsed, largest first, and open to the individual images — which now
carry an age, the only thing telling two otherwise identical rebuilds apart.
227 rows become about a dozen.

The sizes were the other half. An image's size counts every layer it holds,
including those its neighbours hold too, so the rows summed to 100.8 GiB
against 17.3 GiB of actual images. `podman system df -v` reports the bytes an
image holds alone — it refuses `--format json`, hence the table parse, read
from both ends because CREATED is a human duration of unpredictable width — and
a row whose size mostly is not its own now says what deleting it really frees:
906 MiB listed, 11.6 KiB alone.

No such figure is offered per group, deliberately. Measured here, 38 rebuilds
of harness/runner-local:latest sum to 42.6 GiB of size and 0 B of per-image
unique size, because every layer is shared with a sibling: one number is far
too large, the other far too small, and neither is what removing the set frees.
The group says the layers are shared and to remove them together; the section
header points at the engine's own reclaimable for the total. Three wrong
numbers would not have been better than two.
…open

Clicking "Reclaim" on the Images, Containers or Volumes card did nothing. Not
slowly, not with an error — nothing. Verified against the pre-existing build at
8aa8fcd, so this predates the reclaim work rather than arriving with it: those
prune buttons have never opened their confirmation, which means the container
mode has shipped three destructive actions that were unreachable.

SwiftUI presents one alert per view. The view had three `.alert` modifiers
stacked on the same VStack — two of the deprecated `item:`-returning-`Alert`
form, one of the modern `isPresented:` form — and only one of them survives.
Adding the host-disk confirmation made it four and moved which one won, which
is how the symptom surfaced.

So the fix is not to reorder them. Every dialog this view can raise —
three prunes, a single-image removal, the host-disk trim, and the action-failure
report — now goes through one `Dialog` enum and one alert modifier, with the
controller's error channel merged into the same binding. A failure supersedes a
pending confirmation because it can only arrive after an action started, by
which point the confirmation that launched it is gone.

Each of the four was then opened by hand against the running app to confirm it
presents, and cancelled rather than confirmed.
…d a skill frozen at install

Three misses, all found by pointing an assistant at the same disk again.

**uv keeps its cache under XDG on macOS too.** The catalog only listed
`~/Library/Caches/uv`, which does not exist here, so the largest single
reclaimable location on the machine — 4.19 GiB in `~/.cache/uv` — was invisible.
The Apple-shaped path stays for older uv builds; `detect` drops whichever is
absent. `UV_CACHE_DIR` is deliberately not read: a GUI app inherits no shell
environment, so it would be empty far more often than right, the same reasoning
already written down for the Go module cache. Two more tools hiding in the same
place join it: pre-commit/prek (477 MiB) and the GitHub CLI (82 MiB).

**Cold node_modules.** The marker rule generalises as predicted, but not for
free: build output is recreated in seconds from the source beside it, while a
dependency tree only comes back over the network, and only while the registry
still serves the versions the lockfile pins. So the rule gains an age gate —
six months of no install, measured on the tree's own mtime, which package
managers rewrite and nothing else does. That yields the same five projects and
the same ~2 GiB a human reviewing this disk picked out. It is a separate row
from the build-output ones, because it is a materially different bargain.

**The skill was frozen at first install.** `plan()` treated "the directory
exists" as "the skill is current", so a correction shipped in the app never
reached the assistant reading it. That is not hypothetical: this machine's
installed copy was 86 lines against the bundle's 157, and it is why a fresh
analysis still reported that a podman disk image "never shrinks on its own"
hours after that exact sentence was corrected — and why it recommended
`podman machine reset` for a file that had already given back 17.7 GiB. The
plan now compares the two trees by content and the panel says so, with an
Update button, rather than reporting a green check over stale guidance.
A content comparison also catches a hand-edited copy, which is a reason to ask
before overwriting rather than a reason to overwrite.
…its sessions

Both were seen earlier and left out for want of knowing what was inside them.
Opened, and they are different shapes:

`~/.cache/opengrep` is one extracted tree per engine version ever run — a
bundled Python runtime and the analyser, no rules and no results — so old
versions are pure waste and the current one is re-downloaded on the next scan.

`~/.cache/opencode` needed the check more than it looked. It holds downloaded
helper binaries (a C# language server, ripgrep) and the model list, all
re-fetched. But the tool keeps a second directory, `~/.local/share/opencode`,
whose name is nearly the same and whose contents are the opposite: `auth.json`
and `opencode.db` — the login token and every conversation. Only the cache half
is in the catalog, and a test asserts no entry anywhere reaches the other one,
because the failure mode is a plausible generalisation from a matching name
rather than a typo.

That split is worth more than the 368 MiB the two rows add, so it also goes in
the skill's reference: for XDG-shaped tools `.cache` is disposable while
`.local/share` and `.config` are not, and a verdict must never carry from one
to the other on the strength of the name matching.

Both tools run out of the very tree being emptied, so both join the active-tool
warnings — which is why the coverage test accepted them without an exemption.
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Aug 16, 2026
@rducom rducom added enhancement New feature or request and removed documentation Improvements or additions to documentation labels Aug 16, 2026
@rducom rducom changed the title Cleanup targets the catalog cannot name, honest container sizes, and the trim podman never runs Cleanup targets found on disk, Podman host-disk reclaim, and readable container images Aug 16, 2026
@github-actions github-actions Bot added documentation Improvements or additions to documentation and removed documentation Improvements or additions to documentation labels Aug 16, 2026
@rducom
rducom merged commit aa49370 into main Aug 16, 2026
12 checks passed
@rducom
rducom deleted the feat/discovered-cleanup-targets branch August 16, 2026 21:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant