Severity: Medium. Audio is lost silently, and the CPU and GPU results differ in length.
GPU path (extract_audio_gpu)
A batch is processed only when len(batch_frames) == batch_size or fc == frame_count - 1. If cap.read() fails before the reported frame_count (containers often over-report), the loop breaks and the frames already in batch_frames are never processed.
Reproduced by running extract_audio_gpu on the CPU (PyTorch CPU build with pytorch_wavelets; device patched to CPU):
container reports 120 frames, 100 decodable, batch 16:
CPU path returns 100 samples, GPU path returns 96 samples
Up to batch_size - 1 frames are lost at the end of every such video (31 frames at --batch-size 32, about 14 ms of audio at 2200 fps). The progress counter also reports frames_read, not frames processed.
CPU path (extract_audio) and GPU path
Both loop for fc in range(frame_count), so if the container under-reports, all frames after the reported count are silently ignored.
Suggested fix
- Read until
cap.read() fails, not up to frame_count, and use frame_count only for progress and ETA.
- Flush the remaining
batch_frames after the loop. Moving the batch processing into a helper function makes that a single call.
- Warn when the decoded count differs from the metadata.
- Add a test with a fake capture that over- and under-reports, checking that both paths return the true frame count. The GPU path can run on the CPU (see the CI issue).
Acceptance criteria
Severity: Medium. Audio is lost silently, and the CPU and GPU results differ in length.
GPU path (
extract_audio_gpu)A batch is processed only when
len(batch_frames) == batch_size or fc == frame_count - 1. Ifcap.read()fails before the reportedframe_count(containers often over-report), the loopbreaks and the frames already inbatch_framesare never processed.Reproduced by running
extract_audio_gpuon the CPU (PyTorch CPU build withpytorch_wavelets; device patched to CPU):Up to
batch_size - 1frames are lost at the end of every such video (31 frames at--batch-size 32, about 14 ms of audio at 2200 fps). The progress counter also reportsframes_read, not frames processed.CPU path (
extract_audio) and GPU pathBoth loop
for fc in range(frame_count), so if the container under-reports, all frames after the reported count are silently ignored.Suggested fix
cap.read()fails, not up toframe_count, and useframe_countonly for progress and ETA.batch_framesafter the loop. Moving the batch processing into a helper function makes that a single call.Acceptance criteria
CAP_PROP_FRAME_COUNTsays.