Measured
At 704×704 with nlevels=3 and near_sym_b/qshift_b on one core, the forward DTCWT plus phase weighting takes 204 ms per frame.
| sample |
frames |
estimated CPU time |
| Chips1-2200Hz-Mary_Had (704×704) |
22,859 |
≈ 78 min |
The README gives a GPU benchmark (3 m 50 s for Chips2) but no CPU figure, and "Future Work" already suggests multiprocessing.
Suggested approach
- Parallelise across frames. After the reference is computed, each frame is independent. Have one reader thread decode and convert to grey, and a
ProcessPoolExecutor (or threads, since dtcwt's numpy work releases the GIL for large ops) handle blocks of frames and return an (n, nlevels, 6) array per block. This should give close to N× on N cores. Add --jobs.
- Use
--roi by default where possible. Processing time is proportional to area. Consider an automatic ROI: pick the region with the highest temporal variance in a short preview.
- Vectorise the phase maths. Replace
amp*amp*angle(c*conj(ref)) with angle(c*conj(ref)) * (c.real**2 + c.imag**2), which avoids the abs and sqrt. Use float32/complex64 on CPU too, since the GPU path already shows the precision is enough.
- Postprocessing.
sosfiltfilt and the per-band alignment loops are cheap. Filtering with axis=0 on the whole (N, L, 6) array at once is simpler.
- GPU. Overlap decoding with GPU compute (a prefetch thread or queue), and use pinned memory for the host-to-device copy.
Acceptance criteria
Measured
At 704×704 with
nlevels=3andnear_sym_b/qshift_bon one core, the forward DTCWT plus phase weighting takes 204 ms per frame.The README gives a GPU benchmark (3 m 50 s for Chips2) but no CPU figure, and "Future Work" already suggests multiprocessing.
Suggested approach
ProcessPoolExecutor(or threads, since dtcwt's numpy work releases the GIL for large ops) handle blocks of frames and return an(n, nlevels, 6)array per block. This should give close to N× on N cores. Add--jobs.--roiby default where possible. Processing time is proportional to area. Consider an automatic ROI: pick the region with the highest temporal variance in a short preview.amp*amp*angle(c*conj(ref))withangle(c*conj(ref)) * (c.real**2 + c.imag**2), which avoids theabsandsqrt. Use float32/complex64 on CPU too, since the GPU path already shows the precision is enough.sosfiltfiltand the per-band alignment loops are cheap. Filtering withaxis=0on the whole(N, L, 6)array at once is simpler.Acceptance criteria