Skip to content

discovery: unique host-side UDP port per QEMU target, and find them again - #7

Merged
derfsss merged 1 commit into
mainfrom
develop
Sep 11, 2026
Merged

derfsss merged 1 commit into
mainfrom
develop

Conversation

@derfsss

@derfsss derfsss commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner

Finishes in-flight work on the discovery port forward. The half that was done was right; the half that was missing made QEMU guests undiscoverable.

The bug being fixed

Every QEMU target forwarded the discovery port as hostfwd=udp::4323-:4323. The guest side is correct — the daemon always binds 4323 in there — but the host side was the same number for all seven configured targets, and QEMU refuses to start on a duplicate rule. Two guests could not run at once.

The host port now follows the target's endpoint port, which is already unique per target, with [targets.<name>.channels.mcpd] discovery_port to override when that number is taken.

The half that was missing

discover() probes 127.0.0.1 on the default port only. Once the forward moved, nothing answered. Verified against a live guest before fixing it:

1) fleet.discover as shipped:  NOTHING FOUND
2) direct probe to 127.0.0.1:4432:  reply from ('127.0.0.1', 4432)
   {'server': 'mcpd/1.3', 'tcp_port': 4322, 'methods': 59, ...}

Fixing the collision had quietly traded it for no discovery at all. fleet.discover now probes the forwarded port of every configured QEMU target as well as broadcasting — only the configuration knows those numbers. Real hardware still answers the broadcast and is unaffected.

A third problem the same test exposed

The daemon announces "tcp_port":4322 — the port it binds inside the guest. Reporting that back hands the caller 127.0.0.1:4322 for a guest reachable on 127.0.0.1:4432: an endpoint that fails to connect or, on a workstation running several guests, silently connects to a different machine's daemon.

Replies arriving from a known forward are now reported with the endpoint the host actually has, plus the configured target's name. Responders that match no forward are still listed with target: null — the useful case for "what else is on this network?".

Tests

The UDP forward had no coverage at all, which is why the partial change passed cleanly. Eight tests now cover the rule itself (default, pinned, reported ports), which targets get probed (QEMU only, skipping disabled channels), and the endpoint a caller is handed for both a forwarded and an unknown responder. Reverting the cmdline fix fails two of them — checked.

Verified live

On qemu-pegasos2-sm501: endpoint=127.0.0.1:4432 target=qemu-pegasos2-sm501 server=mcpd/1.3 methods=59, and none of the 15 generated hostfwd rules across all seven QEMU targets collide.

Deliberately not changed

The daemon hardcodes its announced TCP port rather than reporting the port it was started with, so MCPd --port N misreports itself to a host with no forward configured for it. Host-side mapping covers every configured target, so this only affects an unknown daemon on a non-default port — and changing it means a new binary, which the just-shipped v1.3 release does not need.

🤖 Generated with Claude Code

https://claude.ai/code/session_01VF4wwBXaKNK2Y2N5sYtt71

…gain

Finishes the in-flight `discovery_port` work. The half that was done
was right; the half that was missing made QEMU guests undiscoverable.

**The bug being fixed.** Every QEMU target forwarded the discovery
port as `hostfwd=udp::4323-:4323`. The guest side is correct -- the
daemon always binds 4323 in there -- but the host side was the same
number for all seven configured targets, and QEMU refuses to start on
a duplicate rule. Two guests could not run at once. The host port now
follows the target's `endpoint` port, which is already unique per
target, and `discovery_port` on the mcpd channel overrides it when
that number is taken.

**The half that was missing.** `discover()` probes 127.0.0.1 on the
default port only, so once the forward moved, nothing answered.
Verified against a live guest: `fleet.discover` returned nothing at
all, while a hand-rolled probe to the forwarded port got a normal
reply. Fixing the collision had quietly traded it for no discovery.

`fleet.discover` now also probes the forwarded port of every
configured QEMU target. Only the configuration knows those numbers,
so the knowledge has to come from there; real hardware still answers
the broadcast and is unaffected.

**A third problem the same test exposed.** The daemon announces
`"tcp_port":4322` -- the port it binds *inside* the guest. Reporting
that back hands the caller `127.0.0.1:4322` for a guest reachable on
`127.0.0.1:4432`: an endpoint that fails to connect, or worse, on a
workstation running several guests, silently connects to a different
machine. Replies arriving from a known forward are now reported with
the endpoint the host actually has, and carry the configured target's
name. Daemons that answer without matching a forward are still listed
with `target: null`, which is the useful case for "what else is out
there?".

Tests: the UDP forward had no coverage at all, which is why the
partial change passed cleanly. Eight tests now cover the rule itself
(default, pinned, reported ports), which targets get probed (QEMU
only, skipping disabled channels), and the endpoint a caller is
handed for both a forwarded and an unknown responder. Reverting the
cmdline fix fails two of them.

Still open, and deliberately not changed here: the daemon hardcodes
its announced TCP port rather than reporting the port it was started
with, so `MCPd --port N` would misreport itself to a host that has no
forward configured for it. Host-side mapping covers every configured
target, so this only affects an unknown daemon on a non-default port.
Changing it means a new binary, which the just-shipped v1.3 release
does not need.

Verified live on qemu-pegasos2-sm501: discovery returns
`endpoint=127.0.0.1:4432 target=qemu-pegasos2-sm501 server=mcpd/1.3`,
and no two of the 15 generated hostfwd rules collide.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VF4wwBXaKNK2Y2N5sYtt71
@derfsss
derfsss merged commit 08ed84c into main Sep 11, 2026
8 checks passed
derfsss added a commit that referenced this pull request Sep 11, 2026
Comments under mcpd/src carried section signs and an arrow, which the
project's ASCII-only rule for on-Amiga source exists to keep out, and
pointed at "19.3 P1 #7" -- a section of a design document that is not
in this repository, so the reference resolved to nothing for anyone
reading the shipped source.

The substance of each comment already stood on its own; only the
citation is gone. No code change -- the rebuilt binary is byte-identical
in size and the build is warning-clean apart from the pre-existing
b64.c -Woverride-init noise.

Also re-checks-out three files (fs.c, methods.h, rpc.c) that a prior
session had left CRLF in the working tree. They were stored as LF
either way thanks to .gitattributes; this just makes the tree agree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VF4wwBXaKNK2Y2N5sYtt71
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant