Skip to content

Bound memory for HD and long clips by processing overlapping spatial tiles #40

Description

@joeljose

Context

Temporal filtering needs the whole time series for each coefficient, which is why the pipeline currently keeps every frame's pyramid in memory. #26 covers constant-factor savings (float32, in-place maths), but memory still grows as frames × pixels: about 55 GiB for 10 s of 1080p30 today, and still around 15 GiB after the #26 fixes.

Suggested approach

All the maths after the DTCWT is independent per spatial position. The only coupling across space is the transform itself, whose support grows with level. So:

  1. Split each frame into tiles of T×T pixels (e.g. 256×256) with an overlap margin M ≥ the coarsest-level filter support (roughly 2**nlevels × filter_len / 2; or cap nlevels so M stays reasonable).
  2. For each tile: decode or read the tile across all frames (from a memmap of the decoded video, or a second decode pass that reads only that region), run forward DTCWT → phase → filter → inverse, then keep only the central, non-overlapping region.
  3. Blend the seams with a small raised-cosine taper in the overlap. With a sufficient margin the result should match whole-frame processing to within about 1e-3.

Peak memory then scales with frames × T² instead of frames × H × W, and tiles can run on separate processes, GPUs or machines.

Related, for very long clips: process the time axis in overlapping windows. The temporal filters have finite support (~344 frames at the default width), so the same overlap-and-crop trick works along time.

Acceptance criteria

  • With --tile 256, 10 s of 1080p30 runs in at most 4 GiB RSS.
  • The tiled result matches the untiled result on face.mp4 (PSNR ≥ 50 dB) with no visible seams at k=5.

Activity

  1. joeljose commented on Sep 27, 2026

    @joeljose
    OwnerAuthor

    The memory arithmetic holds. However, the seam margin needed at level 8 is hundreds of pixels, so 256-pixel tiles would be mostly margin. Suggest deferring: land the #26 float32/in-place fixes first, then consider reusing the GPU two-pass approach on the CPU (keep only phases, recompute amplitudes) before tiling.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions