Skip to content

MXL shared-memory flows on the FFmpeg 8.x series - #37

Open
jpietek wants to merge 19 commits into
developfrom
mxl-ffmpeg8
Open

jpietek wants to merge 19 commits into
developfrom
mxl-ffmpeg8

Conversation

@jpietek

@jpietek jpietek commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator

Continuation of #36 by @matiaspl, which GitHub auto-closed when its base branch ffmpeg8 was deleted after #33 merged. Same eight commits rebased onto develop (after the FFmpeg 8 series #34 and the HDR mixer #33), plus one commit resolving the rebase: the MXL patch is now number 0010 of the series (develop carries nine), bases.env re-pinned (both n8.0 and n8.1 verified with verify.sh), the demo imports pyplumber.mixer, and the mixer Dockerfile keeps nv-codec-headers 13. Supersedes #32 (older MXL work on the pre-FFmpeg-8 stack).


Adds MXL (Media eXchange Layer) shared-memory transport support on top of the FFmpeg 8.x patch series, plus a round-trip demo that exercises it end-to-end.

What's here

  • FFmpeg patch 8/0010-avformat-libmxl-demuxer-muxer.patch — libmxl demuxer, muxer, URI parser, JSON/diagnostic helpers, FATE coverage and --enable-libmxl glue, squashed from cbcrc/FFmpeg branch dmf-mxl/8.1 (pinned at 9eddb90). patch_count in 8/bases.env bumped to 10.
  • Patch regeneration tooling — deps/ffmpeg/Dockerfile.mkpatch, mkpatch-0010-mxl.sh and mkpatch-finish.sh replay the fork's history in a container (no compiler, works on Linux and macOS hosts) so the squashed patch can be re-cut against a newer fork branch or FFmpeg base.
  • MXL SDK in the shared demo image — demos/mixer/Dockerfile builds dmf-mxl/mxl @ v1.1.0 with gcc-13 (the SDK needs C++20; FFmpeg and avplumber keep gcc-11) via vcpkg, and configures FFmpeg with --enable-libmxl --enable-demuxer=mxl --enable-muxer=mxl --enable-protocol=mxl. The build asserts the mxl demuxer and muxer are registered.
  • demos/mxl — writer (lavfi testsrc or a looped file → v210 → MXL) and reader (MXL → v210 unpack → mp4) graphs in one process, using avplumber's generic input / output nodes with format="mxl". No new node types. The reader unpacks straight onto the GPU with v210_to_cuda + NVENC when a CUDA device is present, and falls back to the CPU v210 decoder otherwise. Grains are taken zero-copy out of the shared-memory ring by default.
  • Framework: input and input_rec now log a warning for input options libavformat did not consume. Silently dropped options cost real debugging time on this demo (blocking=1 was being swallowed); this is the only change under src/.

Validation

Verified end-to-end on x86_64 Fedora with an RTX 4000 Ada (driver 615.71, Docker + NVIDIA Container Toolkit) against FFmpeg n8.1-12-g93aafbb from this stack: GPU zero-copy, GPU with --no-zero-copy, --gpu-unpack off, from both lavfi testsrc and a looped file, at 320x240p25 and 640x480p25. All runs started at PTS 0, reported 25.0 fps (25.06 on NVENC) and decoded back to the source pattern.

Known gaps (Ctrl-C hanging in teardown, the unpaced reader, the write path still going through swscale and the CPU v210 encoder) are documented in demos/mxl/README.md.

🤖 Generated with Claude Code

matiaspl and others added 18 commits September 21, 2026 15:18
Adds 8/0011-avformat-libmxl-demuxer-muxer.patch, squashed from the two
MXL commits on cbcrc/FFmpeg branch dmf-mxl/8.1 (pinned at 9eddb90).
That branch forks from the same n8.1 commit deps/ffmpeg/8/bases.env
pins, so both picks apply without conflicts — unlike the earlier
7.1.5-based version of this patch, which force-fitted an 8.x codebase
onto 7.1.5 with -X theirs and dragged fork test scaffolding along.

The configure hunks are re-anchored so one copy serves both bases: the
fork puts its `require_pkg_config libmxl` next to libmpeghdec, which
only exists in 8.1, and carries a whitespace-only reindent of the mmal
check. The generator moves the require next to the mxl_* deps lines and
drops the reindent; verify.sh now reproduces the pinned tree for both
n8.0 and n8.1 with patch_count=11.

The generator lives at deps/ffmpeg/Dockerfile.mkpatch and needs no
compiler — it replays git history only, so it runs on macOS hosts too.
MXL_REMOTE_REF/MXL_PIN/FFMPEG_TAG move it to a future base (the fork
already carries a dmf-mxl/9.0 branch).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds demos/mxl/, a demo that publishes into and reads back from an MXL
shared-memory domain through avplumber's generic input/output nodes
(format="mxl"), with no MXL-specific node types.

The Dockerfile mirrors demos/mixer/Dockerfile: bases on
nvidia/cuda:11.7.1-devel-ubuntu22.04, installs gcc-13 (required by the
MXL SDK) alongside the distro gcc-11, brings up vcpkg + Rust 1.88.0,
builds libmxl v1.1.0-beta-1, applies the FFmpeg patch series, and
configures FFmpeg with the CUDA composition suite plus
--enable-libmxl --enable-demuxer=mxl --enable-muxer=mxl
--enable-protocol=mxl. avplumber's python_module is built with
HAVE_CUDA=1 HAVE_NVCC=1 so the same image also runs mixer-class
workloads on real NVIDIA hosts. The demo runs without a GPU too — CUDA
init fails silently, MXL round-trip works.

mxl_demo.py builds two graphs in one process:
  writer: lavfi testsrc -> decode -> rescale(yuv422p10le) ->
          AssumeVideoFormat -> enc(v210) -> mux -> Output(format="mxl")
  reader: Input(format="mxl") -> demux -> dec(v210) -> rescale(yuv420p)
          -> AssumeVideoFormat -> enc(mpeg4) -> mux -> fragmented mp4

Verified end-to-end on colima (aarch64, no GPU): the reader produces
28k+ mpeg4 packets of the testsrc pattern in a 10-second run. The
fragmented-mp4 output stays playable on SIGINT.

Known gaps (documented in demos/mxl/README.md):
  * options={"blocking":"1"} does not currently reach the MXL demuxer's
    private AVOptions through avplumber's input node; worked around
    with auto_restart:"group" on the reader input.
  * Reader runs at wall-clock max — a RealtimeVideoFrame node between
    decode and encode would cap it to the source frame rate.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Deletes demos/mxl/Dockerfile and merges its libmxl SDK build, vcpkg /
Rust / gcc-13 setup, cmake upgrade, MXL FFmpeg configure flags, and MXL
demuxer/muxer assertions into demos/mixer/Dockerfile. The mixer image is
now a superset of what the mixer, playlist, replay, and mxl demos need
— one image, one build.

FFmpeg comes from deps/ffmpeg/apply.sh like the other demo images. The
7.1.5-era tests/Makefile and fate/mxl.mak cleanups are gone: the 8.1
patch carries only MXL files plus configure/Makefile/allformats glue, so
there is no fork scaffolding left to strip.

Also COPYs demos/mxl into the image so the MXL demo can be run through
the shared tag by overriding the entrypoint:

  docker run --rm --ipc=host --entrypoint python3 \
      -v /dev/shm/mxl:/dev/shm/mxl \
      -v "$PWD/demos/mxl/test-media:/media" \
      avplumber-mixer:local /build/demos/mxl/mxl_demo.py

demos/mxl/README.md is retargeted at the FFmpeg 8.x series (patch
8/0011, deps/ffmpeg/verify.sh, the dmf-mxl/8.1 fork branch) and notes
that the end-to-end run predates the 8.1 migration. It also records
where this is heading: the demuxer emits one packed v210 frame per
grain, which is exactly what the v210_to_cuda node consumes, so the read
path can drop swscale and the CPU v210 decoder entirely.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
v1.1.0 was cut on 2026-09-09, 125 commits ahead of the beta pin we
shipped. Notable changes in that window: audio RMA samples are packed
in channel-major order, fabrics/OFI got batch grain drain and
configurable CQ depth, the audio testsrc no longer busy-spins at
100% CPU, and Rust deps were bumped to fix RUSTSEC-2026-0204.

Verified end-to-end: MXL demuxer/muxer still register in FFmpeg, the
CUDA composition suite still passes assertions, and the mxl_demo
round-trip publishes into and reads back from the v1.1.0 SDK,
producing 16,753 mpeg4 packets in a short run.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
`input` and `input_rec` forward the JSON `options` dictionary to
`avformat_open_input`, which strips out any AVOptions it recognizes
and leaves unrecognized entries behind. Nothing was inspecting those
leftovers, so misnamed options or ones targeting the wrong scope
were silently ignored. Log each remaining entry as a warning.

Verified against the MXL round-trip demo that `blocking=0/1` and
`grain_index_init=head` are consumed by the MXL demuxer as advertised
(`blocking=0` produces the expected EAGAIN + auto_restart storm,
`blocking=1` matches the auto default for video-only). Drop the
"known gap" bullet in `demos/mxl/README.md` and the matching inline
comment in `mxl_demo.py`.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The MXL demuxer hands out one packed v210 frame per grain, which is
exactly what v210_to_cuda consumes, so the reader no longer needs the
CPU v210 decoder or swscale: the grain is staged once through pinned
memory, unpacked to CUDA p210 by the PTX kernel, and finished by
scale_cuda + h264_nvenc. NVENC also drops the mpeg4 fallback the image
needed for lack of libx264.

Selected by --gpu-unpack, defaulting to auto: on when /dev/nvidiactl is
mapped in, CPU path otherwise, so the GPU-less runs the demo was
verified with keep working.

Also enable the demuxer's zero_copy by default, which points each
AVPacket at the MXL ring in /dev/shm instead of copying the grain out.
Nothing holds a reference on the grain, so the packet queue is planned
down to one frame and the documented history_duration goes from 100 ms
(~2 grains) to 1 s; --no-zero-copy trades the memcpy back for immunity.

The geometry is now a CLI contract (--width/--height/--fps) that the
writer rescales to, since packed v210 carries no dimensions and the
reader derives the row stride from the width.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three fixes found while running the MXL round-trip on an RTX 4000 Ada
with FFmpeg 8.1:

The MXL muxer ignores PTS and increments its own grain index per
packet, so an unpaced lavfi source published ~2300 grains/s, ran the
flow into the future and left every reader "too late". Insert
realtime(set_pts) + force_fps ahead of the v210 encoder.

The demuxer picks its grain index during avformat_open_input's probe,
i.e. seconds before the group starts, with CUDA and NVENC init in
between; by then it is behind the ring tail and the EAGAIN is fatal to
input.cpp. on_too_late=reset re-derives it, and reset_on_drop rebases
the timestamps so the output does not open with a multi-second hole.

The reader's mp4 was frequently 28 bytes: fragmented mp4 only helps if
the data leaves the muxer's 256 KiB avio buffer, which a low-bitrate
run killed after 25 s never fills. Add flush_packets=1, plus a
one-second NVENC GOP so fragments are cut at a useful rate.

README records what was actually verified on the hardware and the
SIGINT teardown hang, which reproduces on the CPU path too and so is
not from this wiring.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Trim the MXL demo README to the facts a reader needs: fold the
duplicated build recipe into what is MXL-specific about the shared
image, merge the two teardown-hang gaps into one, and drop the
in-image verification snippet the Dockerfile already asserts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…overs

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
--bench-seconds samples enqueued_total on the writer and reader encoder
edges once a second and prints a summary; polling that counter keeps the
measurement off the media path, unlike a wiretap callback. It exits with
os._exit because graceful teardown hangs.

--writer-pace off drops realtime + force_fps to find the writer's
ceiling. Running it against the CPU reader exposed an abort: with
reset_on_drop=1 a mid-stream "too late" reset sets mxl_start_grain_index
to the current grain, so the derived PTS drops back to 0 and the
monotonic-PTS av_assert0 in mxldec.c:1393 kills the process. Documented
for now, not fixed.

Numbers on 1080p59.94 smptehdbars in the README.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The fork's demuxer rebases the pts origin onto the grain the reader
resumes at. At startup that is correct; mid-stream it drives the derived
pts back to 0 and trips the monotonic-pts av_assert0 a few lines below,
aborting the process. Rebase so the first packet after a reset continues
one past the last one delivered instead, on both the video and audio
paths.

Reproducible in seconds with demos/mxl and
`--writer-pace off --gpu-unpack off`, where the reader cannot keep up
with the flood and on_too_late=reset fires repeatedly.

The change is carried in mkpatch-fixups/ and git am-ed by
mkpatch-0010-mxl.sh after the cherry-picks, so it lands inside the
squashed patch 0010 with its commit message listed as provenance, and
retiring it once the fork carries a fix is one file deletion plus a
regeneration. Both base trees in 8/bases.env refreshed accordingly
(verify.sh passes for n8.0 and n8.1).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The earlier table was taken while an unrelated container was using ~2 of
the 4 vCPU, so every ceiling in it was understated: the writer alone now
reaches 567 fps rather than 351, and with the GPU reader attached both
sides sustain 507 fps. Realtime figures barely move (the target was met
either way) but drop to 0.27 cores.

The unpaced CPU-unpack case no longer aborts now that patch 0010 keeps
the pts monotonic across a reset, so it is a row in the table instead of
a warning, and it shows what the CPU reader actually costs: 192 fps
against the writer's 419, with all four cores busy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The node parsed "flags" into sws_flags_ and then built av::VideoRescaler
without it, so every graph in the tree scaled with avcpp's default pick
(bicubic when upscaling, area otherwise) no matter what it asked for.
Graphs that request SWS_LANCZOS -- examples/complicated_transcoder,
examples/from_dmabuf_hwdownload, pyplumber/examples/simple-node -- now
get it, and pay for it where they actually resize.

While there: converting a frame to the geometry and format it already
has still costs swscale a full-frame copy, 1.0 ms and 8 MB per frame at
1080p 10-bit 4:2:2. Hand such frames straight to the sink instead. Edges
carry reference-counted frames and no software node in src/ writes
pixels in place, so nothing downstream can tell.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…coder

Measuring "the MXL reader" with the full tail mostly measured mpeg4,
which is 3.58 of the reader's 5.9 ms of CPU per frame. Add --reader-tail
{demux,unpack,encode} to drop grains at either end of the unpack, and a
per-node CPU table taken from /proc/self/task: every node runs in a
thread named after it, so the profiler is free and stays off the media
path. Add --sws-flags to pick the swscale algorithm for the two rescale
nodes, now that rescale_video honours it.

Fix the bench edge counters, which is what made --reader-tail unpack fail
on the GPU path: startNodes() only queues the group's state change, so
getEdge() ran while the nodes were still being created and *created*
r_vcuda as a Packet edge, after which the real node failed forever. Wait
for the edge in the stats JSON, which never creates one.

What the numbers say, in README: MXL transport is not the cost centre
(publish is one memcpy at 0.27 ms/frame, zero-copy read is 0.07 against
0.38), v210_to_cuda costs as much host CPU as the libavcodec v210
decoder it replaces and wins only by skipping swscale and the encoder,
and unpaced the writer is the limit until the output encoder enters and
pins one core.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The two rescale_video nodes are the largest per-frame CPU cost after the
output encoder, so add --gpu-scale {writer,reader,both} to hand them to
scale_cuda instead (hwupload,scale_cuda,hwdownload, with --cuda-interp as
its --sws-flags). What that actually saves depends on how many times the
frame crosses PCIe, so add --reader-encoder {mpeg4,nvenc} to decouple the
output encoder from --gpu-unpack: the same conversion can then be timed
with an upload, a download, both, or neither. --reader-tail gains "scale",
which drops the frame right after the conversion.

Reorganise the reader around that: _build_convert and _build_encoder
replace the two whole-tail builders, each tail exits through one _drop,
and hwaccel.init now runs before both graphs are built, gated on whether
any chosen node actually wants the device. v210_to_cuda stamps the frame
period rather than 90 kHz when mpeg4 is the output, because mpeg4 rejects
any timebase denominator above 65535.

What the measurements say, in README: the host CPU cost of a GPU
conversion is 0.04 ms plus every transferred byte at ~19 GB/s -- unpinned
frames are staged through the driver's own pinned buffer, and the
conversion itself adds nothing measurable, which nvidia-smi dmon
corroborates byte for byte. So the GPU wins outright where it is handed a
frame it does not have to move (0.04 and 0.22 against swscale's 1.6-1.7
ms/frame) and wins less than it looks like where it pays both transfers
(0.61, a third of which mpeg4 takes back on frames arriving cold from
DMA). A resize is the clear case: lanczos to 720p costs 0.36 ms/frame on
the GPU against 3.18 in swscale, and a third of what swscale's *area*
resize costs. Both paths agree to SSIM 0.9996.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A GPU graph publishing uncompressed video had to come back to the host
twice: hwdownload after the conversion, then the libavcodec v210 encoder,
together 0.73 of the MXL writer's 1.15 ms of CPU per frame. Add
cuda_to_v210, the mirror of v210_to_cuda: a kernel packs the CUDA frame
into v210 and the bytes leave the device by DMA into pinned memory that
the muxer reads directly, so nothing on the host touches a pixel.

The node is the encoder as far as the muxer below it is concerned. It
implements IEncoder with a hand-built AVCodecParameters rather than a
libavcodec context (as nvjpeg_enc does), which means geometry, frame rate
and time base have to be answerable at node creation -- output opens its
whole format context then -- so they come from the nodes above via
IVideoFormatSource / IFrameRateSource / ITimeBaseSource, with parameters
to override or stand in. The MXL muxer takes the flow's grain rate from
the stream's frame rate, hence a hard error when nothing above reports
one. auto_eof_ is off so the EOF marker is forwarded as a packet.

Packets are handed out of an AVBufferPool of cuMemHostAlloc'd buffers
that cuMemcpyDtoHAsync writes into, and the pool holds its own reference
on the CUDA device context so buffers may outlive the node. The pack runs
on the device context's stream, which is also whatever produced the frame
used, so no extra synchronization is needed beyond waiting for the copy.

Everything both v210 nodes share now lives in v210_cuda.hpp -- the error
check, the context guard, the geometry/stride layout, the format
predicates -- and v210_to_cuda is refactored onto it instead of keeping a
second copy.

The output is bit-exact: grains from the kernel are byte-for-byte what
the libavcodec v210 encoder produces at 1920x1080, 1280x720 and
1918x1080, the last two exercising rows that end in a partial 6-pixel
block and all three the padding to a 128-byte row pitch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Add --writer-pack {auto,cpu,gpu}: gpu drops the writer's hwdownload and
its v210 encoder for a single cuda_to_v210, which needs the conversion on
the GPU too, so it requires --gpu-scale writer|both and auto turns it on
exactly when that holds.

What the measurements say, in README: the conversion node loses precisely
its download (0.62 -> 0.21 ms/frame, the rest being the 3.1 MB upload at
~19 GB/s) and the pack costs the fixed 0.04 alone, because the 5.5 MB
leaves the device by DMA and the CPU pays nothing per byte. The writer
falls from 1.15 to 0.59 ms of CPU per frame, the paced writer process
from 0.09 to 0.05 cores, and an all-GPU round trip from 0.17 to 0.09.
Unpaced the writer's peak nearly doubles, 855 -> 1583 fps on the same
1.2-1.3 cores at 82% SM.

That moves the bottom of the write path somewhere new. w_output -- the
muxer's memcpy into the opened grain -- gets more expensive (0.23 -> 0.34
ms/frame) reading bytes no CPU has touched, and at 1583 fps it is 8.8
GB/s into /dev/shm and the writer's actual cap: the first time in this
demo that MXL itself limits anything. Removing that copy means having the
muxer hand out the grain's address before the packet is built, which is a
change to patch 0010, so it goes to Known gaps next to the read side's
cudaHostRegister note.

Also correct the SM figure on the unpaced --gpu-scale writer row, which
had the paced --gpu-scale both dmon reading (7%) rather than its own; the
case re-measures at 51% on an idle GPU, same 855 fps.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
band_blur_cuda took 0010 on develop, so the MXL patch is 0011 here -- the
file, the mkpatch scripts and the series list move with it in the commit
that adds them. Regenerate the patch against 0001-0010 instead of leaving
its hunks anchored a line off what they now apply to: nothing in the
FFmpeg change moves, only the provenance line, the squash's local commit
ids and the configure offsets band blur shifts by one.

Refresh 8/bases.env for the combined series -- patch_count 11 and the
trees it actually produces, which verify.sh reproduces for both bases:

  FFmpeg n8.1 patch series verified: 93dd33926bd56b56cab82a747ae91c39f31e84ba
  FFmpeg n8.0 patch series verified: 3f3cacd2a72b5a43a06bbc4cb212307bd7e1a9d6

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every number in demos/mxl/README.md came from an ad-hoc wrapper around
mxl_demo.py --bench-seconds that lived only on the test host, so the
tables were reproducible in principle and not in practice. Collect them
under demos/mxl/bench/ as one harness: bench.sh runs a single case
(fresh MXL domain, nvidia-smi and cgroup samplers, log triage, ffprobe
of the output, one digest instead of a log), cases.sh holds the case
lists as named groups, and bench/README.md maps each group to the table
it produced.

Two checks are not an fps number and stay separate. pack_bitexact.sh
publishes the same converted frames with the CPU and the GPU v210 pack,
dumps the grains with ffmpeg -c copy and compares them byte for byte at
1920x1080, 1280x720 and 1918x1080, exiting nonzero on a difference so it
can gate. pixel_compare.sh covers the rest of "Conversion on the GPU":
PSNR/SSIM between the two writer conversions, nvidia-smi dmon for the SM
and PCIe columns, and which formats this build's scale_cuda accepts.

The scripts reach the demo two ways, because measuring and developing
want different things. --image starts a container per case, which is the
only mode whose cgroup holds nothing but the run and the only one that
can drop the GPU (--no-gpu); --exec uses a container that is already up,
for a /build tree you keep rebuilding, where per-node CPU is still exact
(the demo reads its own /proc/self/task) but whole-container CPU is not.
Anything the harness does not recognise goes to the demo, so its flags
work unchanged and beat the group defaults.

Verified on the NVIDIA host in both modes: all 54 cases of pack,
swsflags, srcfmt, tails, throughput and conversion exit 0 with only the
benign "setting edge more than once" triage entry; smoke 11/11; nodecpu,
zerocopy, pack_bitexact.sh and all three pixel_compare.sh sections
reproduce the documented per-node ms/frame, IDENTICAL grains, and the
PSNR/SSIM and dmon figures.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants