Skip to content

Performance: the CPU path takes about 78 min for the Chips1 sample; parallelise across frames and vectorise the phase maths #22

Description

@joeljose

Measured

At 704×704 with nlevels=3 and near_sym_b/qshift_b on one core, the forward DTCWT plus phase weighting takes 204 ms per frame.

sample frames estimated CPU time
Chips1-2200Hz-Mary_Had (704×704) 22,859 ≈ 78 min

The README gives a GPU benchmark (3 m 50 s for Chips2) but no CPU figure, and "Future Work" already suggests multiprocessing.

Suggested approach

  1. Parallelise across frames. After the reference is computed, each frame is independent. Have one reader thread decode and convert to grey, and a ProcessPoolExecutor (or threads, since dtcwt's numpy work releases the GIL for large ops) handle blocks of frames and return an (n, nlevels, 6) array per block. This should give close to N× on N cores. Add --jobs.
  2. Use --roi by default where possible. Processing time is proportional to area. Consider an automatic ROI: pick the region with the highest temporal variance in a short preview.
  3. Vectorise the phase maths. Replace amp*amp*angle(c*conj(ref)) with angle(c*conj(ref)) * (c.real**2 + c.imag**2), which avoids the abs and sqrt. Use float32/complex64 on CPU too, since the GPU path already shows the precision is enough.
  4. Postprocessing. sosfiltfilt and the per-band alignment loops are cheap. Filtering with axis=0 on the whole (N, L, 6) array at once is simpler.
  5. GPU. Overlap decoding with GPU compute (a prefetch thread or queue), and use pinned memory for the host-to-device copy.

Acceptance criteria

  • On 4 cores, CPU throughput improves by ≥ 3×, with output identical to atol 1e-6.
  • A CPU benchmark row is added to the README.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions