Skip to content

Tests: nothing checks that sound is actually recovered; add a synthetic vibrating-surface test with known ground truth #20

Description

@joeljose

Severity: Medium. The current tests check shapes, finiteness and range. The smoke tests feed random-noise frames, where there is no signal to recover, so a change that broke recovery would still pass. That's how #13, #14 and #17 went unnoticed.

Weak or misleading tests

  • test_constant_input_silent asserts only shape and finiteness. It doesn't check that the output is silent.
  • test_passband_preserved asserts std > 0.01 on output that is normalised to [−1, 1], so it passes for any non-constant signal.
  • test_stopband_attenuated compares standard deviations of two outputs, each normalised to its own peak. It passes because of how the 200-sample edges behave, not because the filter works.
  • The input-validation tests pass a file dummy.avi that doesn't exist, so any run that exits non-zero passes, even for an unrelated reason.

Proposed ground-truth test

This generator was used in the audit. It makes exact sub-pixel motion via the Fourier shift theorem, so recovery can be scored against the known signal:

def texture(n=128, seed=0):
    t = ndimage.gaussian_filter(np.random.RandomState(seed).rand(n, n), 2.0)
    return 40 + 170 * (t - t.min()) / (t.max() - t.min())

def shift_img(img, dx, dy):
    H, W = img.shape
    ky, kx = np.fft.fftfreq(H)[:, None], np.fft.fftfreq(W)[None, :]
    return np.real(np.fft.ifft2(np.fft.fft2(img) * np.exp(-2j*np.pi*(kx*dx + ky*dy))))

# frames_k = shift_img(tex, a*s[k]*cos(φ), a*s[k]*sin(φ)) + noise,  s = 440/660/250 Hz tones, fps = 2200
# Feed through a FakeCap(frames) exposing read()/release() into extract_audio / extract_audio_gpu

Suggested assertions:

  1. Horizontal, vertical and 45° motion at 0.05 px: correlation with the truth ≥ 0.99 (vertical currently gives 0.85-0.91, see Sub-band alignment ignores sign: anti-phase orientations are misaligned or cancel (recovery drops to r≈0.85-0.91 for vertical motion) #13).
  2. 0.3 px drift with default settings: correlation ≥ 0.95 (currently 0.26, see With default settings (no filter), slow drift swamps the audio: 0.3 px of drift cuts in-band level by 12 dB #14).
  3. 3 px drift with -fl 100: correlation ≥ 0.99 (currently 0.95, see Measuring phase against frame 0 wraps once the surface drifts; use frame-to-frame differences plus a cumulative sum #17).
  4. CPU vs GPU-path output correlation ≥ 0.999 (GPU run on CPU torch; see the CI issue).
  5. Over- and under-reported frame counts: the output length equals the decodable frames (see GPU path drops the last partial batch when the container over-reports frames; CPU path drops frames when it under-reports #16).

Keep the clips short (1100 frames × 128², about 30 s each on CPU; or 64² for CI speed) and mark them @pytest.mark.slow if needed.

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions