perf: speed up ARIS→MP4 convert (~15× small / ~2.6× large) - #23
Conversation
Replace Pillow→ffmpeg CLI with in-process PyAV and a vendored ProcessingPipeline so produce overlaps encode. Bitrate-matched to baseline -q:v 10; optional codec= for HW encoders on other machines. Co-authored-by: Cursor <cursoragent@cursor.com>
Move pixel prep into consumers with numpy I420 (or gray when the codec supports it), split encode/mux timing, and drop unused ProcessingThreadPool from the vendored pipeline module. Co-authored-by: Cursor <cursoragent@cursor.com>
Replace per-pixel np.vectorize/bisect with a shared angle pass and searchsorted so data_import is no longer dominated by LUT build. Co-authored-by: Cursor <cursoragent@cursor.com>
Move remap out of the producer into consumers, and use convertMaps once so each frame remaps via CV_16SC2 maps without per-frame float remap or uint8 cast. Co-authored-by: Cursor <cursoragent@cursor.com>
Chouffe
left a comment
There was a problem hiding this comment.
Thanks for this. The measurement work behind it is careful, and two thirds of the PR is a clear win: I verified the LUT vectorization and the fixed-point convertMaps remap produce byte-identical frames to main on 96 beam data at two geometries (791x1485 and 924x1765), 200 to 300 frames each. That part I would merge as is.
One gap worth naming: every ARIS file I have access to is 96 beam, so only the _beam_breakpoints(96) path is exercised. The 128, 64 and 48 branches are unverified by measurement. They look correct by inspection, since np.searchsorted(..., side="right") - 1 matches bisect_right and the 999 sentinel logic is preserved, but someone with ARIS 3000 data should confirm before this lands.
The speedups reproduce, and if anything you undersold them. End to end (DataImport plus VideoExport, best of 2 to 3 runs, warm page cache, fps=15, vbr=10):
| frames | geometry | main |
this PR | speedup | |
|---|---|---|---|---|---|
| small | 51 | 924x1765 | 4.28 s | 0.09 s | 48x |
| large | 12844 | 791x1485 | 29.22 s | 8.37 s | 3.5x |
Bench environment
| CPU | AMD Ryzen 7 9800X3D, 8 cores / 16 threads, boost 5.27 GHz |
| RAM | 60 GB |
| Storage | NVMe SSD (Samsung 990 PRO), source files in page cache |
| OS | Linux 6.17 |
| Python | 3.13.2 |
| NumPy / OpenCV | 2.3.1 / 4.11.0 |
| PyAV / FFmpeg | 18.0.0 / 6.1.1 |
Single machine, no other load. Absolute times will differ on other hardware; the ratios are what matter, and the encode-bound large-file case will shift most with core count.
Split by stage, the small-file win is almost entirely the vectorized LUT (DataImport 4.03 s to 0.03 s, about 130x) and the large-file win is the overlapped encode (VideoExport 26.3 s to 8.4 s, about 3.2x). Your 4.3 s "before" figure for the 924x1765 geometry and my 4.28 s land in the same place despite different clip lengths (176 frames versus 51), which is itself the point: on main the fixed LUT build dominates and clip length barely matters.
Two caveats on the bench table in the description. First, those numbers are only reachable by passing fps=15 explicitly; on the default CLI path, which derives fps from the file, the run never completes (see comment 1). Second, the large-file comparison is not quality matched: this PR writes 628 MB where main writes 357 MB for the same input, so part of the remaining gap is that the two are not producing the same thing.
The encoder swap needs work before this can land. Three things:
1. It hangs forever on our data. convert_aris_to_video.py against a 96 beam 791x1485 recording (clip D above) exits 124 on a 120 second timeout, produces no file, and reports no error. Root cause is the mpeg4 timebase limit; the reason it hangs rather than fails is that writer thread exceptions deadlock the pipeline. Reproduced on three separate files.
2. The output is a quality regression on every clip I tested. Five clips, 200 to 300 frames each, spanning a 4.3x range in pixel count:
| clip | geometry | pixels | motion |
|---|---|---|---|
| A | 489x786 | 0.38 MP | 8.74 |
| B | 791x1485 | 1.17 MP | 5.02 |
| C | 791x1485 | 1.17 MP | 6.29 |
| D | 791x1485 | 1.17 MP | 4.37 |
| E | 924x1765 | 1.63 MP | 5.85 |
All 96 beam. Motion is the mean absolute difference between consecutive remapped frames, higher meaning busier scene.
Measured against the raw remapped frames, decoded the way a player renders them:
PSNR (peak signal to noise ratio, in dB) compares each encoded frame to the raw remap it came from, so both encoders are scored against the same ground truth rather than against each other. Higher is better, every 6 dB is a halving of RMS pixel error, and differences under about 0.5 dB are not visible. SSIM (structural similarity, 0 to 1, 1 being identical) scores the same comparison but weighs local structure and texture, and largely normalises away a uniform brightness shift. That difference matters here: on clip D this PR scores better on SSIM and worse on PSNR, which is the signature of a systematic intensity error sitting on top of otherwise decent encoding.
| clip | pixels | baseline | PR #23 | PR bitrate | PR PSNR |
|---|---|---|---|---|---|
| A | 0.38 MP | 2.39 Mbps, 30.80 dB | 5.94 Mbps, 30.25 dB | 2.49x | -0.55 dB |
| B | 1.17 MP | 3.68 Mbps, 32.13 dB | 5.96 Mbps, 30.24 dB | 1.62x | -1.89 dB |
| C | 1.17 MP | 4.64 Mbps, 32.16 dB | 5.95 Mbps, 29.74 dB | 1.28x | -2.42 dB |
| D | 1.17 MP | 3.34 Mbps, 32.24 dB | 5.92 Mbps, 31.01 dB | 1.77x | -1.23 dB |
| E | 1.63 MP | 5.88 Mbps, 32.64 dB | 5.91 Mbps, 29.84 dB | 1.005x | -2.80 dB |
More bits spent on all five, worse PSNR on all five. Most of that is the range handling (inline comment below), not the rate control.
3. The 0.81x you compensated for was the right answer. Removing the JPEG stage removed real noise from the source, so the same quantizer legitimately costs fewer bits. Restoring fixed qscale and fixing the range gives:
| clip | baseline | qscale + range fix | ratio | PSNR delta |
|---|---|---|---|---|
| A | 2.39 Mbps | 1.93 Mbps | 0.81x | -0.03 dB |
| C | 4.64 Mbps | 3.82 Mbps | 0.82x | -0.02 dB |
| D | 3.34 Mbps | 2.58 Mbps | 0.77x | -0.01 dB |
| E | 5.88 Mbps | 4.74 Mbps | 0.81x | -0.05 dB |
Every clip within 0.05 dB of baseline at 77 to 82 percent of the bitrate.
I applied all three fixes to a local copy of this branch to check they compose. They do:
main |
this PR | PR + the three fixes | |
|---|---|---|---|
| small file, end to end | 4.28 s | 0.09 s | 0.09 s |
| large file, end to end | 29.22 s | 8.37 s | 8.41 s |
| large file output | 357 MB | 628 MB | 276 MB |
| PSNR vs raw remap | 32.24 dB | 31.01 dB | 32.23 dB |
| pixels crushed to black | 0.001% | 0.055% | 0.001% |
Same picture as main, 23 percent smaller files, and the full speedup is retained.
Visual comparison, error maps and the harness are in pr23-review/ on my machine; happy to attach the error map images here if useful. The short version: in the PR's error map you can see the riverbed, which means the error tracks the signal rather than being quantization noise.
| stream = container.add_stream( | ||
| codec, rate=Fraction(fps).limit_denominator() | ||
| ) |
There was a problem hiding this comment.
Blocking: hangs on every ARIS file we have
limit_denominator() defaults to max_denominator=1000000. Our sonar reports framerate = 15.000149726867676, which becomes 9317058/621131, and mpeg4 caps the timebase denominator at 65535:
[mpeg4] timebase 621131/9317058 not supported by MPEG 4 standard,
the maximum admitted value for the timebase denominator is 65535
avcodec_open2 raises inside the writer thread, the consumers then block forever on writer_queue.put, and run() never returns. The CLI produces no file and no error.
I hit this on every recording I tried, plus the repo's own test fixture. They all report that same framerate, so this is the normal path, not an edge case.
Careful with the fix: limit_denominator(65535) is not enough. It bounds the denominator of the rate, but the timebase is the reciprocal, so the value mpeg4 constrains is the rate numerator. I tried it and still got timebase 40073/601101. Bound the reciprocal instead:
| stream = container.add_stream( | |
| codec, rate=Fraction(fps).limit_denominator() | |
| ) | |
| # mpeg4 caps the timebase denominator at 65535. The timebase is | |
| # 1/rate, so it is the rate *numerator* that must be bounded: | |
| # limit the reciprocal, then invert. | |
| rate = 1 / (1 / Fraction(fps)).limit_denominator(65535) | |
| stream = container.add_stream(codec, rate=rate) |
Verified: 15.000149726867676 becomes 65521/4368 (error 7.9e-05 fps, about 4 ms of drift over a 14 minute file), and exact rates such as 24, 29.97 and 59.94 are preserved unchanged. With this in place convert_aris_to_video.py completes on the real framerate and writes a valid file.
Worth a regression test that exports two frames at fps=15.000149726867676 and asserts a file exists.
| writer(item_to_write) | ||
| writer_queue.task_done() |
There was a problem hiding this comment.
Blocking: writer exceptions deadlock the pipeline
writer(item_to_write) is uncaught. If the writer raises, the thread dies, the writer queue stops draining, and every consumer parks on the bounded writer_queue.put at line 170. run() then blocks in join() forever.
This is what turns the timebase bug above into a silent hang instead of a stack trace. It will do the same for any av.open, stream.encode or container.mux failure: missing output directory, unsupported codec, disk full.
Reproduced in isolation: a pipeline with num_workers=2 and a writer that raises on the first item never returns from run().
| writer(item_to_write) | |
| writer_queue.task_done() | |
| try: | |
| writer(item_to_write) | |
| except Exception as e: | |
| logger.exception("Writer failed, cancelling pipeline") | |
| self._writer_error = e | |
| self._stop_event.set() | |
| finally: | |
| writer_queue.task_done() |
Initialise self._writer_error = None in __init__, and re-raise it at the end of run() so callers actually learn the export failed. Right now process_aris_filepath logs "Successfully converted" for a run that produced nothing.
| if self.writer and result is not None and writer_queue is not None: | ||
| writer_queue.put(result) |
There was a problem hiding this comment.
Blocking: consumers ignore cancellation while enqueuing
writer_queue.put(result) blocks without a timeout and without consulting _stop_event. Once the writer thread has exited (cancellation, or the crash in comment 2), consumers sit here forever and run() cannot join them. The producer already uses the timeout loop pattern at line 139; the consumer should match it.
Reproduced: num_workers=4, a slow writer, and an on_task_done that requests cancellation at item 30. run() never returns, and dumping thread frames shows workers parked at exactly this line.
| if self.writer and result is not None and writer_queue is not None: | |
| writer_queue.put(result) | |
| if self.writer and result is not None and writer_queue is not None: | |
| while not self._stop_event.is_set(): | |
| try: | |
| writer_queue.put(result, timeout=0.1) | |
| break | |
| except queue.Full: | |
| continue |
| # Y = gray, U/V = 128 (I420 = Y plane + half-height chroma block). | ||
| yuv = np.vstack([remap, np.full((h // 2, w), 128, dtype=np.uint8)]) | ||
| video_frame = av.VideoFrame.from_ndarray(yuv, format="yuv420p") |
There was a problem hiding this comment.
Range handling changes every pixel of every video
The raw 0 to 255 grayscale goes straight into a yuv420p Y plane with no full to limited range conversion. Players treat an untagged yuv420p stream as limited range and expand 16 to 235 back out to 0 to 255, so the displayed image is stretched: dark samples crush to black, midtones shift, contrast rises.
On clip D, 300 frames, as rendered by a player:
| mean brightness | PSNR vs raw remap | pixels crushed to black | |
|---|---|---|---|
| raw remap (reference) | 43.31 | - | - |
baseline (main) |
43.33 | 32.24 dB | 0.001% |
| this PR | 40.25 | 31.01 dB | 0.055% |
main tracks the source to within 0.02 of a grey level. This PR is 3.06 low.
The old JPEG to ffmpeg path handled this implicitly, which is why main is correct.
I tried the obvious fix first and it does not work: setting codec_context.color_range = AVCOL_RANGE_JPEG does not survive mpeg4 in mp4, ffprobe still reports color_range=unknown. The pixels have to be converted.
| # Y = gray, U/V = 128 (I420 = Y plane + half-height chroma block). | |
| yuv = np.vstack([remap, np.full((h // 2, w), 128, dtype=np.uint8)]) | |
| video_frame = av.VideoFrame.from_ndarray(yuv, format="yuv420p") | |
| # mpeg4 in mp4 cannot reliably signal full range (color_range does not | |
| # survive the round trip), so convert to limited range the way the old | |
| # JPEG -> ffmpeg path did implicitly. | |
| y = cv2.LUT(np.ascontiguousarray(remap), _FULL_TO_LIMITED) | |
| yuv = np.vstack([y, np.full((h // 2, w), 128, dtype=np.uint8)]) | |
| video_frame = av.VideoFrame.from_ndarray(yuv, format="yuv420p") |
with this next to the other module constants:
# Full-range 0-255 to limited-range 16-235, as the old JPEG -> ffmpeg path did.
_FULL_TO_LIMITED = np.clip(
np.rint(np.arange(256, dtype=np.float32) * (219.0 / 255.0) + 16.0), 0, 255
).astype(np.uint8)Use cv2.LUT rather than the arithmetic inline. I measured the obvious float version (np.clip(np.rint(remap * 219/255 + 16), 0, 255)) at 0.48 ms per frame, which costs about 5 s on a 12844 frame file and eats a chunk of the speedup this PR is for. cv2.LUT is 0.09 ms for a bit-identical result. Plain fancy indexing (_FULL_TO_LIMITED[remap]) is the worst option at 1.07 ms, because remap is a negative-stride flipud view.
With this change alone, keeping your bitrate target, PSNR goes from 31.01 dB to 34.77 dB and black crush drops to 0.001%. It is the single largest quality lever in the PR.
| # Match former JPEG→ffmpeg -q:v <vbr> *output size* (not raw | ||
| # qscale — that under-rates vs the JPEG-pipe path). | ||
| q = max(int(vbr), 1) | ||
| br = max(1, int(_BASELINE_BITRATE * _BASELINE_VBR / q)) | ||
| stream.bit_rate = br | ||
| # Default tolerance (128k) is below our ~5.87 Mbps target. | ||
| stream.codec_context.bit_rate_tolerance = br |
There was a problem hiding this comment.
Fixed bitrate replaces fixed quality
-q:v 10 pinned quality and let bitrate follow the content. This pins bitrate and lets quality float, and the target is a constant with no resolution or frame rate term. Measured output is 5.91 to 5.96 Mbps across a 4.3x range in pixel count (0.38 MP to 1.63 MP): the 489x786 clip gets the same bitrate as the 924x1765 one. It happens to fit the 924x1765 workload it was calibrated on and over provisions everything else, up to 2.49x on our smallest clip.
On the "under-rates at 0.81x" note in the commit message: I think that reading is inverted. Dropping the JPEG stage (Pillow defaults to quality 75) removed real noise from the encoder input, so the same quantizer legitimately needs fewer bits. I measured 0.77x to 0.82x on four clips with PSNR within 0.05 dB of baseline every time. That consistency across content and resolution is what you would expect from removing a fixed noise source, and not what you would expect from an encoder that is under rating.
| # Match former JPEG→ffmpeg -q:v <vbr> *output size* (not raw | |
| # qscale — that under-rates vs the JPEG-pipe path). | |
| q = max(int(vbr), 1) | |
| br = max(1, int(_BASELINE_BITRATE * _BASELINE_VBR / q)) | |
| stream.bit_rate = br | |
| # Default tolerance (128k) is below our ~5.87 Mbps target. | |
| stream.codec_context.bit_rate_tolerance = br | |
| # `-q:v <vbr>` was a fixed quantizer: quality pinned, bitrate free to | |
| # follow content and resolution. Reproduce that rather than targeting | |
| # an absolute bitrate. | |
| ctx = stream.codec_context | |
| ctx.flags |= av.codec.context.Flags.qscale | |
| ctx.global_quality = max(int(vbr), 1) * 118 # FF_QP2LAMBDA | |
| ctx.qmin = ctx.qmax = max(int(vbr), 1) |
_BASELINE_VBR and _BASELINE_BITRATE at lines 32 and 33 become unused.
To be fair to the change: there is a real quality versus size tradeoff here, and spending more bits is a legitimate choice. My objection is not that this PR spends more, it is that an absolute bitrate makes the effective quality setting drift with resolution, so the same config lands somewhere different on every file. vbr does not have that problem. Sweeping it on 300 frames of clip D, with the range fix applied:
vbr |
size | bitrate | PSNR | SSIM |
|---|---|---|---|---|
| 3 | 37.2 MB | 14.87 Mbps | 39.64 dB | 0.974 |
| 5 | 20.0 MB | 8.00 Mbps | 35.99 dB | 0.941 |
| 7 | 12.2 MB | 4.86 Mbps | 33.99 dB | 0.907 |
| 10 | 6.5 MB | 2.58 Mbps | 32.23 dB | 0.857 |
| 14 | 3.4 MB | 1.37 Mbps | 30.91 dB | 0.803 |
main today |
8.4 MB | 3.34 Mbps | 32.24 dB | 0.863 |
vbr=10 reproduces today's output at 77 percent of the size. If we want visibly better video, vbr=7 is the knob to reach for, and it stays adaptive across geometries. Anything above about 14 is not worth it: at vbr=20 mpeg4 hits a floor and the file gets larger than vbr=14 while looking worse.
Worth noting the current PR is not far off vbr=7 in quality per byte once the range is fixed (34.77 dB at 14.8 MB), so if the intent was "make the videos better", say so explicitly and set vbr=7. That gets the same result in a way that survives a change of sonar geometry.
| writer_thread.join() | ||
|
|
||
| # Stop monitor | ||
| self._stop_event.set() |
There was a problem hiding this comment.
Pipeline reports itself cancelled after a clean run
run() sets _stop_event to stop the monitor thread, which conflates "finished" with "cancelled". A separate threading.Event for monitor shutdown keeps the two meanings apart.
Reproduced: a pipeline over 10 items produces all 10 and then reports pipeline.cancelled == True; calling run() again on the same object produces 0 items and raises nothing.
| # TEMP: surface libav / videotoolbox messages on stderr while debugging HW encode. | ||
| av.logging.set_level(av.logging.VERBOSE) |
There was a problem hiding this comment.
Debug instrumentation left in
Self labelled TEMP. This mutates process global libav logging for every caller of the library, not just this function, and floods stderr during batch conversion.
| # TEMP: surface libav / videotoolbox messages on stderr while debugging HW encode. | |
| av.logging.set_level(av.logging.VERBOSE) |
| print(f"queue_stats: {_round_stats(qstats)}", flush=True) | ||
| print(f"timing: {_round_stats(timing)}", flush=True) |
There was a problem hiding this comment.
Debug prints to stdout
Library code printing to stdout on every conversion. video/utils.py and frame.py both go through logging for this kind of thing.
Note pyARIS.py currently has no logging import and no module logger, so swapping in logger.debug needs those added first. If the timing numbers were only for benching the PR, simplest is to drop the two lines:
| print(f"queue_stats: {_round_stats(qstats)}", flush=True) | |
| print(f"timing: {_round_stats(timing)}", flush=True) |
| @@ -1385,14 +1431,16 @@ def VideoExport( | |||
| fontsize : (Int) Size of timestamp font | |||
| ts_pos : (Tuple) (x,y) location of the timestamp | |||
| vbr : (Int) Variable Bitrate of output video (1-31) 1 being highest quality, 31 being lowest quality | |||
There was a problem hiding this comment.
Docstring no longer matches behaviour
With the bitrate change, vbr is no longer a quantizer scale. If you take the qscale suggestion in comment 5 this line stays accurate as written and needs no change; if you keep the bitrate model, it needs rewording to say what the number now does.
| draw.text(ts_pos, ts, font=font, fill="white") | ||
| im.save(pipe.stdin, "JPEG") | ||
| remap = np.asarray(im) | ||
| if use_gray: |
There was a problem hiding this comment.
use_gray skips the padding added in #21
The gray branch skips the even dimension padding and sets pix_fmt = "gray".
I checked which codecs actually take it: _codec_accepts_gray returns True for libx264 and False for both mpeg4 and h264_nvenc. So it never fires on the default path, and the one codec it does fire on is the one you would reach for to produce web video. Two consequences there: the output is 4:0:0 rather than yuv420p, and it loses the odd dimension padding that #21 added specifically to make H.264 output web ready. (4:0:0 H.264 is also widely reported not to decode in Safari and most browsers. I did not verify that here, but the pix_fmt and dimension changes alone are enough to break a downstream format check.)
I confirmed the branch is the only obstacle: with use_gray forced off, codec="libx264" produces correct yuv420p output at the padded dimensions. Worth either restricting use_gray to codecs we have a reason to want it for, or dropping it until a caller needs it.
Bound the reciprocal of fps so the timebase denominator stays <= 65535 (matching ffmpeg's -r reduction). Add unit coverage and an integration export of two frames at each sample file's native fps. Co-authored-by: Cursor <cursoragent@cursor.com>
ProcessingPipeline previously swallowed consumer errors, deadlocked on writer failures, and reported cancelled after clean runs: - Record the first error (producer/consumer/writer) and re-raise it from run(); failures cancel the run instead of hanging or truncating output silently. - Separate cancelled from finished so a clean run leaves cancelled False and the pipeline reusable. - All queue puts go through a timed, stop-aware helper so no thread can block forever on a full queue. - Recreate queues and reset progress/stats each run so a failed or cancelled run cannot leak stale items or sentinels into the next one. - Guard progress counting and first-error recording with a lock; progress callbacks are now serialized and in order. - Raise ValueError for a None producer with workers (was a hang) and remove the unused chained-pipeline path and DEFAULT_PIPELINE_CALLABLE. - VideoExport asserts its reorder buffer is drained after run(). Co-authored-by: Cursor <cursoragent@cursor.com>
Players treat untagged yuv420p as limited range, so full-range remap looked crushed. Map 0–255 to 16–235 before I420, and encode with ffmpeg -q:v semantics (qscale, qmin=qmax=vbr) instead of a fixed bitrate. Always pad odd dims and use yuv420p so the path matches main CLI quality at a smaller file size. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
Thanks for this very completed review. These were good things to address. I just pushed four commits: f3bf3b1 — Fix mpeg4 hang on ARIS fractional frame rates. I think that this fixes everything that you found. The error handling in processing.py needed a bunch of work. I did that and added several new tests there. The range of values being fed to the encoder was also a great catch. I've implemented the remapping you suggested and ran a new benchmark test. The table below has the existing cli vs the new code on all of the test files that I had. Note that I wasn't able to reproduce your 32db psnr but I think that is because my benchmark code only measured over the fan of the sonar and not in the black padding around it.
Note I still haven't been able to verify on ARIS 3000 samples |
Chouffe
left a comment
There was a problem hiding this comment.
Thanks for the follow up work! It works nicely now.
I will merge it.
Summary
Speeds up end-to-end ARIS → MP4 convert by replacing the JPEG→ffmpeg pipe with in-process PyAV, vectorizing cold LUT build, and overlapping remap with encode.
Bench (
fps=15,vbr=10, geom 96×1765→1765×924):2026-04-29_210000_F18_B96_S1765_T06_R0-7_Raw_3975_4150.aris2026-05-05_220000_F18_B96_S1765_T06_R0-7.arisbench/stage3-remap-pipeline)Changes (3 stages)
bench/stage1-pyav-pipeline) —VideoExportvia PyAV mpeg4; newProcessingPipelineoverlaps read/prep/encode; I420 prep (Y=gray,U/V=128) orgraywhen the codec allows;codec=configurable.bench/stage2-numpy-lut) — replacenp.vectorize+ Pythonbisectwith one NumPy pass (searchsorted). Maps verified identical to the old path. Dominates short-file time.bench/stage3-remap-pipeline) — producer reads raw frames (FrameRead(..., remap=False)); consumers callremapARIS;cv2.convertMaps(..., CV_16SC2)once inLUT.Large file is still mostly encode-bound after these changes.
Files
src/aris/video/processing.pysrc/aris/pyARIS/pyARIS.py,src/aris/scripts/convert_aris_to_video.py,pyproject.toml(+ PyAV),uv.lockStage tags are on the fork for bisect/repro:
bench/baseline…bench/stage3-remap-pipeline.Test plan
fps=15,vbr=10and compare wall time to the tableDataImport/FrameReaddefault path still remaps (remap=True)Happy to adjust API / packaging if this should land differently.
Made with Cursor