Tracking issue for the findings from a full audit of the codebase at v2.0.0 (c6c97bb).
How the audit was done
Read every file: module, tests, notebook, Dockerfiles, CI, docs.
Ran ruff (0.11.0, 0.15.7, 0.16.9) and the test suite: 32 passed, 14 GPU tests skipped.
Probed edge cases with synthetic clips.
Ran and profiled the full CPU pipeline on face.mp4 (Python 3.11, numpy 1.26.4, scipy 1.17.1, dtcwt 0.14.0, 4 vCPU).
The GPU path was reviewed by reading only. No CUDA device was available, so no GPU number here was measured.
Headline numbers (face.mp4, CPU)
Peak RSS 8.26 GiB (README: ~1.2 GB)
Wall time 6 m 32 s , single-threaded
59% of output pixels lost to NaN in a test clip with a black half-frame
Phase 1: stop the bleeding (≈ 1 day)
CPU path: exact-zero wavelet coefficients produce NaN, which shows up as black regions in the output #22 CPU path: exact-zero coefficients produce NaN, which shows up as black regions (critical)
save_video says "Output saved" and exits 0 when the video writer fails to open #23 save_video says "Output saved" and exits 0 when the writer fails (critical)
CI will fail on its next run: ruff is unpinned and 0.16 changed its default rules #24 CI will fail on its next run: ruff unpinned, 0.16 changed default rules
CI does not run the unit tests; the pipeline check only confirms a file exists #25 CI does not run the unit tests; the pipeline check only confirms a file exists
GPU Docker image runs on the CPU by default; the README says --gpu is the default entrypoint behaviour #27 GPU Docker image runs on the CPU by default
Colab notebook fails on SciPy >= 1.13 (signal.flattop removed), wraps pixel values, and keeps its own copy of the algorithm #28 Colab notebook fails on SciPy ≥ 1.13, wraps pixels, duplicates the algorithm
Odd frame sizes crash the CPU path at the inverse DTCWT, after the forward pass has already run #30 Odd frame sizes crash the CPU inverse
load_video: corrupt or unreadable input gives a traceback, frames beyond the reported count are dropped, fps=0 isn't handled #31 load_video: tracebacks on bad input, dropped frames, fps=0
Phase 2: correctness baseline (≈ 2–3 days)
CPU FFT temporal filter is one frame out of step with the direct-convolution and GPU paths (even-length windows) #29 CPU FFT temporal filter is one frame out of step with the direct and GPU paths
CPU path peaks at 8.26 GiB on face.mp4 (README says ~1.2 GB); estimate_memory() is never called and underestimates #26 CPU memory ~7× the documented figure; estimate_memory() unused and inaccurate
Tests: some assertions can never fail, and there is no regression test on real output #32 Tests: assertions that can't fail, no regression test on real output
Build reproducibility and supply chain: pin pytorch_wavelets, lock dependencies, fix Docker user handling #33 Build reproducibility and supply chain (pin pytorch_wavelets, lock files, Docker user)
GPU path: add the OOM handling from the design doc, validate --device, and test the GPU code on the CPU #34 GPU: OOM handling, --device validation, test the GPU code on the CPU
Docs don't match the code: GPU padding, VRAM fraction, "bitwise-identical", the 0.2327 constant, versions #35 Docs don't match the code
Phase 3: product and scale (≈ 1–2 weeks)
Package the project (pyproject.toml, entry point) and separate a library API from the CLI; use logging instead of print #41 Packaging plus a library API; logging instead of print
CPU path uses a single core: parallelise the DTCWT per frame and the temporal FFT #36 Parallel CPU path
Add a luma-only mode: magnify Y in YIQ/YCrCb instead of R, G and B separately (~3× faster, no colour fringing) #37 Luma-only mode (~3× faster, no colour fringing)
Let users set the temporal band in Hz (--freq-low/--freq-high) instead of a filter width in frames #38 Set the band in Hz (--freq-low/--freq-high)
Reduce artefacts at high k: amplitude-weighted phase smoothing, limits on large phase shifts, per-level magnification #39 Fewer artefacts at high k: amplitude-weighted smoothing, phase limits, per-level k
Bound memory for HD and long clips by processing overlapping spatial tiles #40 Spatial tiling for bounded memory on HD and long clips
Output: add lossless formats and keep the audio track; move demo GIFs out of git history #42 Lossless output formats, keep audio, move demo GIFs out of history
Suggested order
#24 and #25 first, so CI is green and actually runs the tests. #32 's golden test next, so every later change can be checked against it. Then the Phase 1 bugs. #29 will change the default output slightly, so land it together with the golden test and a CHANGELOG entry.
Tracking issue for the findings from a full audit of the codebase at v2.0.0 (
c6c97bb).How the audit was done
face.mp4(Python 3.11, numpy 1.26.4, scipy 1.17.1, dtcwt 0.14.0, 4 vCPU).Headline numbers (face.mp4, CPU)
Phase 1: stop the bleeding (≈ 1 day)
save_videosays "Output saved" and exits 0 when the writer fails (critical)load_video: tracebacks on bad input, dropped frames, fps=0Phase 2: correctness baseline (≈ 2–3 days)
estimate_memory()unused and inaccurate--devicevalidation, test the GPU code on the CPUPhase 3: product and scale (≈ 1–2 weeks)
--freq-low/--freq-high)Suggested order
#24 and #25 first, so CI is green and actually runs the tests. #32's golden test next, so every later change can be checked against it. Then the Phase 1 bugs. #29 will change the default output slightly, so land it together with the golden test and a CHANGELOG entry.