Conversation
Adds 8/0011-avformat-libmxl-demuxer-muxer.patch, squashed from the two MXL commits on cbcrc/FFmpeg branch dmf-mxl/8.1 (pinned at 9eddb90). That branch forks from the same n8.1 commit deps/ffmpeg/8/bases.env pins, so both picks apply without conflicts — unlike the earlier 7.1.5-based version of this patch, which force-fitted an 8.x codebase onto 7.1.5 with -X theirs and dragged fork test scaffolding along. The configure hunks are re-anchored so one copy serves both bases: the fork puts its `require_pkg_config libmxl` next to libmpeghdec, which only exists in 8.1, and carries a whitespace-only reindent of the mmal check. The generator moves the require next to the mxl_* deps lines and drops the reindent; verify.sh now reproduces the pinned tree for both n8.0 and n8.1 with patch_count=11. The generator lives at deps/ffmpeg/Dockerfile.mkpatch and needs no compiler — it replays git history only, so it runs on macOS hosts too. MXL_REMOTE_REF/MXL_PIN/FFMPEG_TAG move it to a future base (the fork already carries a dmf-mxl/9.0 branch). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds demos/mxl/, a demo that publishes into and reads back from an MXL
shared-memory domain through avplumber's generic input/output nodes
(format="mxl"), with no MXL-specific node types.
The Dockerfile mirrors demos/mixer/Dockerfile: bases on
nvidia/cuda:11.7.1-devel-ubuntu22.04, installs gcc-13 (required by the
MXL SDK) alongside the distro gcc-11, brings up vcpkg + Rust 1.88.0,
builds libmxl v1.1.0-beta-1, applies the FFmpeg patch series, and
configures FFmpeg with the CUDA composition suite plus
--enable-libmxl --enable-demuxer=mxl --enable-muxer=mxl
--enable-protocol=mxl. avplumber's python_module is built with
HAVE_CUDA=1 HAVE_NVCC=1 so the same image also runs mixer-class
workloads on real NVIDIA hosts. The demo runs without a GPU too — CUDA
init fails silently, MXL round-trip works.
mxl_demo.py builds two graphs in one process:
writer: lavfi testsrc -> decode -> rescale(yuv422p10le) ->
AssumeVideoFormat -> enc(v210) -> mux -> Output(format="mxl")
reader: Input(format="mxl") -> demux -> dec(v210) -> rescale(yuv420p)
-> AssumeVideoFormat -> enc(mpeg4) -> mux -> fragmented mp4
Verified end-to-end on colima (aarch64, no GPU): the reader produces
28k+ mpeg4 packets of the testsrc pattern in a 10-second run. The
fragmented-mp4 output stays playable on SIGINT.
Known gaps (documented in demos/mxl/README.md):
* options={"blocking":"1"} does not currently reach the MXL demuxer's
private AVOptions through avplumber's input node; worked around
with auto_restart:"group" on the reader input.
* Reader runs at wall-clock max — a RealtimeVideoFrame node between
decode and encode would cap it to the source frame rate.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Deletes demos/mxl/Dockerfile and merges its libmxl SDK build, vcpkg /
Rust / gcc-13 setup, cmake upgrade, MXL FFmpeg configure flags, and MXL
demuxer/muxer assertions into demos/mixer/Dockerfile. The mixer image is
now a superset of what the mixer, playlist, replay, and mxl demos need
— one image, one build.
FFmpeg comes from deps/ffmpeg/apply.sh like the other demo images. The
7.1.5-era tests/Makefile and fate/mxl.mak cleanups are gone: the 8.1
patch carries only MXL files plus configure/Makefile/allformats glue, so
there is no fork scaffolding left to strip.
Also COPYs demos/mxl into the image so the MXL demo can be run through
the shared tag by overriding the entrypoint:
docker run --rm --ipc=host --entrypoint python3 \
-v /dev/shm/mxl:/dev/shm/mxl \
-v "$PWD/demos/mxl/test-media:/media" \
avplumber-mixer:local /build/demos/mxl/mxl_demo.py
demos/mxl/README.md is retargeted at the FFmpeg 8.x series (patch
8/0011, deps/ffmpeg/verify.sh, the dmf-mxl/8.1 fork branch) and notes
that the end-to-end run predates the 8.1 migration. It also records
where this is heading: the demuxer emits one packed v210 frame per
grain, which is exactly what the v210_to_cuda node consumes, so the read
path can drop swscale and the CPU v210 decoder entirely.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
v1.1.0 was cut on 2026-09-09, 125 commits ahead of the beta pin we shipped. Notable changes in that window: audio RMA samples are packed in channel-major order, fabrics/OFI got batch grain drain and configurable CQ depth, the audio testsrc no longer busy-spins at 100% CPU, and Rust deps were bumped to fix RUSTSEC-2026-0204. Verified end-to-end: MXL demuxer/muxer still register in FFmpeg, the CUDA composition suite still passes assertions, and the mxl_demo round-trip publishes into and reads back from the v1.1.0 SDK, producing 16,753 mpeg4 packets in a short run. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
`input` and `input_rec` forward the JSON `options` dictionary to `avformat_open_input`, which strips out any AVOptions it recognizes and leaves unrecognized entries behind. Nothing was inspecting those leftovers, so misnamed options or ones targeting the wrong scope were silently ignored. Log each remaining entry as a warning. Verified against the MXL round-trip demo that `blocking=0/1` and `grain_index_init=head` are consumed by the MXL demuxer as advertised (`blocking=0` produces the expected EAGAIN + auto_restart storm, `blocking=1` matches the auto default for video-only). Drop the "known gap" bullet in `demos/mxl/README.md` and the matching inline comment in `mxl_demo.py`. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The MXL demuxer hands out one packed v210 frame per grain, which is exactly what v210_to_cuda consumes, so the reader no longer needs the CPU v210 decoder or swscale: the grain is staged once through pinned memory, unpacked to CUDA p210 by the PTX kernel, and finished by scale_cuda + h264_nvenc. NVENC also drops the mpeg4 fallback the image needed for lack of libx264. Selected by --gpu-unpack, defaulting to auto: on when /dev/nvidiactl is mapped in, CPU path otherwise, so the GPU-less runs the demo was verified with keep working. Also enable the demuxer's zero_copy by default, which points each AVPacket at the MXL ring in /dev/shm instead of copying the grain out. Nothing holds a reference on the grain, so the packet queue is planned down to one frame and the documented history_duration goes from 100 ms (~2 grains) to 1 s; --no-zero-copy trades the memcpy back for immunity. The geometry is now a CLI contract (--width/--height/--fps) that the writer rescales to, since packed v210 carries no dimensions and the reader derives the row stride from the width. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three fixes found while running the MXL round-trip on an RTX 4000 Ada with FFmpeg 8.1: The MXL muxer ignores PTS and increments its own grain index per packet, so an unpaced lavfi source published ~2300 grains/s, ran the flow into the future and left every reader "too late". Insert realtime(set_pts) + force_fps ahead of the v210 encoder. The demuxer picks its grain index during avformat_open_input's probe, i.e. seconds before the group starts, with CUDA and NVENC init in between; by then it is behind the ring tail and the EAGAIN is fatal to input.cpp. on_too_late=reset re-derives it, and reset_on_drop rebases the timestamps so the output does not open with a multi-second hole. The reader's mp4 was frequently 28 bytes: fragmented mp4 only helps if the data leaves the muxer's 256 KiB avio buffer, which a low-bitrate run killed after 25 s never fills. Add flush_packets=1, plus a one-second NVENC GOP so fragments are cut at a useful rate. README records what was actually verified on the hardware and the SIGINT teardown hang, which reproduces on the CPU path too and so is not from this wiring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Trim the MXL demo README to the facts a reader needs: fold the duplicated build recipe into what is MXL-specific about the shared image, merge the two teardown-hang gaps into one, and drop the in-image verification snippet the Dockerfile already asserts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…overs Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
--bench-seconds samples enqueued_total on the writer and reader encoder edges once a second and prints a summary; polling that counter keeps the measurement off the media path, unlike a wiretap callback. It exits with os._exit because graceful teardown hangs. --writer-pace off drops realtime + force_fps to find the writer's ceiling. Running it against the CPU reader exposed an abort: with reset_on_drop=1 a mid-stream "too late" reset sets mxl_start_grain_index to the current grain, so the derived PTS drops back to 0 and the monotonic-PTS av_assert0 in mxldec.c:1393 kills the process. Documented for now, not fixed. Numbers on 1080p59.94 smptehdbars in the README. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The fork's demuxer rebases the pts origin onto the grain the reader resumes at. At startup that is correct; mid-stream it drives the derived pts back to 0 and trips the monotonic-pts av_assert0 a few lines below, aborting the process. Rebase so the first packet after a reset continues one past the last one delivered instead, on both the video and audio paths. Reproducible in seconds with demos/mxl and `--writer-pace off --gpu-unpack off`, where the reader cannot keep up with the flood and on_too_late=reset fires repeatedly. The change is carried in mkpatch-fixups/ and git am-ed by mkpatch-0010-mxl.sh after the cherry-picks, so it lands inside the squashed patch 0010 with its commit message listed as provenance, and retiring it once the fork carries a fix is one file deletion plus a regeneration. Both base trees in 8/bases.env refreshed accordingly (verify.sh passes for n8.0 and n8.1). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The earlier table was taken while an unrelated container was using ~2 of the 4 vCPU, so every ceiling in it was understated: the writer alone now reaches 567 fps rather than 351, and with the GPU reader attached both sides sustain 507 fps. Realtime figures barely move (the target was met either way) but drop to 0.27 cores. The unpaced CPU-unpack case no longer aborts now that patch 0010 keeps the pts monotonic across a reset, so it is a row in the table instead of a warning, and it shows what the CPU reader actually costs: 192 fps against the writer's 419, with all four cores busy. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The node parsed "flags" into sws_flags_ and then built av::VideoRescaler without it, so every graph in the tree scaled with avcpp's default pick (bicubic when upscaling, area otherwise) no matter what it asked for. Graphs that request SWS_LANCZOS -- examples/complicated_transcoder, examples/from_dmabuf_hwdownload, pyplumber/examples/simple-node -- now get it, and pay for it where they actually resize. While there: converting a frame to the geometry and format it already has still costs swscale a full-frame copy, 1.0 ms and 8 MB per frame at 1080p 10-bit 4:2:2. Hand such frames straight to the sink instead. Edges carry reference-counted frames and no software node in src/ writes pixels in place, so nothing downstream can tell. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…coder
Measuring "the MXL reader" with the full tail mostly measured mpeg4,
which is 3.58 of the reader's 5.9 ms of CPU per frame. Add --reader-tail
{demux,unpack,encode} to drop grains at either end of the unpack, and a
per-node CPU table taken from /proc/self/task: every node runs in a
thread named after it, so the profiler is free and stays off the media
path. Add --sws-flags to pick the swscale algorithm for the two rescale
nodes, now that rescale_video honours it.
Fix the bench edge counters, which is what made --reader-tail unpack fail
on the GPU path: startNodes() only queues the group's state change, so
getEdge() ran while the nodes were still being created and *created*
r_vcuda as a Packet edge, after which the real node failed forever. Wait
for the edge in the stats JSON, which never creates one.
What the numbers say, in README: MXL transport is not the cost centre
(publish is one memcpy at 0.27 ms/frame, zero-copy read is 0.07 against
0.38), v210_to_cuda costs as much host CPU as the libavcodec v210
decoder it replaces and wins only by skipping swscale and the encoder,
and unpaced the writer is the limit until the output encoder enters and
pins one core.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The two rescale_video nodes are the largest per-frame CPU cost after the
output encoder, so add --gpu-scale {writer,reader,both} to hand them to
scale_cuda instead (hwupload,scale_cuda,hwdownload, with --cuda-interp as
its --sws-flags). What that actually saves depends on how many times the
frame crosses PCIe, so add --reader-encoder {mpeg4,nvenc} to decouple the
output encoder from --gpu-unpack: the same conversion can then be timed
with an upload, a download, both, or neither. --reader-tail gains "scale",
which drops the frame right after the conversion.
Reorganise the reader around that: _build_convert and _build_encoder
replace the two whole-tail builders, each tail exits through one _drop,
and hwaccel.init now runs before both graphs are built, gated on whether
any chosen node actually wants the device. v210_to_cuda stamps the frame
period rather than 90 kHz when mpeg4 is the output, because mpeg4 rejects
any timebase denominator above 65535.
What the measurements say, in README: the host CPU cost of a GPU
conversion is 0.04 ms plus every transferred byte at ~19 GB/s -- unpinned
frames are staged through the driver's own pinned buffer, and the
conversion itself adds nothing measurable, which nvidia-smi dmon
corroborates byte for byte. So the GPU wins outright where it is handed a
frame it does not have to move (0.04 and 0.22 against swscale's 1.6-1.7
ms/frame) and wins less than it looks like where it pays both transfers
(0.61, a third of which mpeg4 takes back on frames arriving cold from
DMA). A resize is the clear case: lanczos to 720p costs 0.36 ms/frame on
the GPU against 3.18 in swscale, and a third of what swscale's *area*
resize costs. Both paths agree to SSIM 0.9996.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A GPU graph publishing uncompressed video had to come back to the host twice: hwdownload after the conversion, then the libavcodec v210 encoder, together 0.73 of the MXL writer's 1.15 ms of CPU per frame. Add cuda_to_v210, the mirror of v210_to_cuda: a kernel packs the CUDA frame into v210 and the bytes leave the device by DMA into pinned memory that the muxer reads directly, so nothing on the host touches a pixel. The node is the encoder as far as the muxer below it is concerned. It implements IEncoder with a hand-built AVCodecParameters rather than a libavcodec context (as nvjpeg_enc does), which means geometry, frame rate and time base have to be answerable at node creation -- output opens its whole format context then -- so they come from the nodes above via IVideoFormatSource / IFrameRateSource / ITimeBaseSource, with parameters to override or stand in. The MXL muxer takes the flow's grain rate from the stream's frame rate, hence a hard error when nothing above reports one. auto_eof_ is off so the EOF marker is forwarded as a packet. Packets are handed out of an AVBufferPool of cuMemHostAlloc'd buffers that cuMemcpyDtoHAsync writes into, and the pool holds its own reference on the CUDA device context so buffers may outlive the node. The pack runs on the device context's stream, which is also whatever produced the frame used, so no extra synchronization is needed beyond waiting for the copy. Everything both v210 nodes share now lives in v210_cuda.hpp -- the error check, the context guard, the geometry/stride layout, the format predicates -- and v210_to_cuda is refactored onto it instead of keeping a second copy. The output is bit-exact: grains from the kernel are byte-for-byte what the libavcodec v210 encoder produces at 1920x1080, 1280x720 and 1918x1080, the last two exercising rows that end in a partial 6-pixel block and all three the padding to a 128-byte row pitch. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Add --writer-pack {auto,cpu,gpu}: gpu drops the writer's hwdownload and
its v210 encoder for a single cuda_to_v210, which needs the conversion on
the GPU too, so it requires --gpu-scale writer|both and auto turns it on
exactly when that holds.
What the measurements say, in README: the conversion node loses precisely
its download (0.62 -> 0.21 ms/frame, the rest being the 3.1 MB upload at
~19 GB/s) and the pack costs the fixed 0.04 alone, because the 5.5 MB
leaves the device by DMA and the CPU pays nothing per byte. The writer
falls from 1.15 to 0.59 ms of CPU per frame, the paced writer process
from 0.09 to 0.05 cores, and an all-GPU round trip from 0.17 to 0.09.
Unpaced the writer's peak nearly doubles, 855 -> 1583 fps on the same
1.2-1.3 cores at 82% SM.
That moves the bottom of the write path somewhere new. w_output -- the
muxer's memcpy into the opened grain -- gets more expensive (0.23 -> 0.34
ms/frame) reading bytes no CPU has touched, and at 1583 fps it is 8.8
GB/s into /dev/shm and the writer's actual cap: the first time in this
demo that MXL itself limits anything. Removing that copy means having the
muxer hand out the grain's address before the packet is built, which is a
change to patch 0010, so it goes to Known gaps next to the read side's
cudaHostRegister note.
Also correct the SM figure on the unpaced --gpu-scale writer row, which
had the paced --gpu-scale both dmon reading (7%) rather than its own; the
case re-measures at 51% on an idle GPU, same 855 fps.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
band_blur_cuda took 0010 on develop, so the MXL patch is 0011 here -- the file, the mkpatch scripts and the series list move with it in the commit that adds them. Regenerate the patch against 0001-0010 instead of leaving its hunks anchored a line off what they now apply to: nothing in the FFmpeg change moves, only the provenance line, the squash's local commit ids and the configure offsets band blur shifts by one. Refresh 8/bases.env for the combined series -- patch_count 11 and the trees it actually produces, which verify.sh reproduces for both bases: FFmpeg n8.1 patch series verified: 93dd33926bd56b56cab82a747ae91c39f31e84ba FFmpeg n8.0 patch series verified: 3f3cacd2a72b5a43a06bbc4cb212307bd7e1a9d6 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
matiaspl
force-pushed
the
mxl-ffmpeg8
branch
from
September 21, 2026 13:24
d343660 to
2927271
Compare
Every number in demos/mxl/README.md came from an ad-hoc wrapper around mxl_demo.py --bench-seconds that lived only on the test host, so the tables were reproducible in principle and not in practice. Collect them under demos/mxl/bench/ as one harness: bench.sh runs a single case (fresh MXL domain, nvidia-smi and cgroup samplers, log triage, ffprobe of the output, one digest instead of a log), cases.sh holds the case lists as named groups, and bench/README.md maps each group to the table it produced. Two checks are not an fps number and stay separate. pack_bitexact.sh publishes the same converted frames with the CPU and the GPU v210 pack, dumps the grains with ffmpeg -c copy and compares them byte for byte at 1920x1080, 1280x720 and 1918x1080, exiting nonzero on a difference so it can gate. pixel_compare.sh covers the rest of "Conversion on the GPU": PSNR/SSIM between the two writer conversions, nvidia-smi dmon for the SM and PCIe columns, and which formats this build's scale_cuda accepts. The scripts reach the demo two ways, because measuring and developing want different things. --image starts a container per case, which is the only mode whose cgroup holds nothing but the run and the only one that can drop the GPU (--no-gpu); --exec uses a container that is already up, for a /build tree you keep rebuilding, where per-node CPU is still exact (the demo reads its own /proc/self/task) but whole-container CPU is not. Anything the harness does not recognise goes to the demo, so its flags work unchanged and beat the group defaults. Verified on the NVIDIA host in both modes: all 54 cases of pack, swsflags, srcfmt, tails, throughput and conversion exit 0 with only the benign "setting edge more than once" triage entry; smoke 11/11; nodecpu, zerocopy, pack_bitexact.sh and all three pixel_compare.sh sections reproduce the documented per-node ms/frame, IDENTICAL grains, and the PSNR/SSIM and dmon figures. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Continuation of #36 by @matiaspl, which GitHub auto-closed when its base branch
ffmpeg8was deleted after #33 merged. Same eight commits rebased onto develop (after the FFmpeg 8 series #34 and the HDR mixer #33), plus one commit resolving the rebase: the MXL patch is now number 0010 of the series (develop carries nine),bases.envre-pinned (both n8.0 and n8.1 verified withverify.sh), the demo importspyplumber.mixer, and the mixer Dockerfile keeps nv-codec-headers 13. Supersedes #32 (older MXL work on the pre-FFmpeg-8 stack).Adds MXL (Media eXchange Layer) shared-memory transport support on top of the FFmpeg 8.x patch series, plus a round-trip demo that exercises it end-to-end.
What's here
8/0010-avformat-libmxl-demuxer-muxer.patch— libmxl demuxer, muxer, URI parser, JSON/diagnostic helpers, FATE coverage and--enable-libmxlglue, squashed fromcbcrc/FFmpegbranchdmf-mxl/8.1(pinned at9eddb90).patch_countin8/bases.envbumped to 10.deps/ffmpeg/Dockerfile.mkpatch,mkpatch-0010-mxl.shandmkpatch-finish.shreplay the fork's history in a container (no compiler, works on Linux and macOS hosts) so the squashed patch can be re-cut against a newer fork branch or FFmpeg base.demos/mixer/Dockerfilebuildsdmf-mxl/mxl@v1.1.0with gcc-13 (the SDK needs C++20; FFmpeg and avplumber keep gcc-11) via vcpkg, and configures FFmpeg with--enable-libmxl --enable-demuxer=mxl --enable-muxer=mxl --enable-protocol=mxl. The build asserts the mxl demuxer and muxer are registered.demos/mxl— writer (lavfi testsrcor a looped file → v210 → MXL) and reader (MXL → v210 unpack → mp4) graphs in one process, using avplumber's genericinput/outputnodes withformat="mxl". No new node types. The reader unpacks straight onto the GPU withv210_to_cuda+ NVENC when a CUDA device is present, and falls back to the CPU v210 decoder otherwise. Grains are taken zero-copy out of the shared-memory ring by default.inputandinput_recnow log a warning for input options libavformat did not consume. Silently dropped options cost real debugging time on this demo (blocking=1was being swallowed); this is the only change undersrc/.Validation
Verified end-to-end on x86_64 Fedora with an RTX 4000 Ada (driver 615.71, Docker + NVIDIA Container Toolkit) against FFmpeg
n8.1-12-g93aafbbfrom this stack: GPU zero-copy, GPU with--no-zero-copy,--gpu-unpack off, from bothlavfi testsrcand a looped file, at 320x240p25 and 640x480p25. All runs started at PTS 0, reported 25.0 fps (25.06 on NVENC) and decoded back to the source pattern.Known gaps (Ctrl-C hanging in teardown, the unpaced reader, the write path still going through swscale and the CPU v210 encoder) are documented in
demos/mxl/README.md.🤖 Generated with Claude Code