Skip to content

Latest commit

 

History

History
244 lines (207 loc) · 13 KB

File metadata and controls

244 lines (207 loc) · 13 KB

MicFix — architecture and design decisions

1. Requirements analysis (what actually has to be solved)

The microphone symptoms map onto very different DSP problems:

Symptom Nature Right tool
Mains hum + harmonics (50/100/150/200/250 Hz) deterministic, extremely narrow-band, slowly drifting narrow IIR notches at measured f0·k
"Mosquito" buzzing (measured on this mic: ~1 kHz USB-frame whine series up to 16 kHz) deterministic, narrow, but high frequency where a wide notch would damage voice notches with bounded absolute bandwidth (50 Hz)
Hiss / broadband floor stationary random profile-based spectral gain (Wiener)
Rumble low-frequency, partly non-stationary steep high-pass
Thin/harsh tone linear colouration parametric EQ
Inconsistent level dynamics compressor + limiter
Sibilance narrow-band dynamic de-esser

Generic noise suppressors fail here because they treat the tonal hum as part of the "noise" to be estimated statistically: the estimate lags the hum's level, leaves musical residue around it, and over-suppresses the speech bins that share those frequencies. Splitting the problem into deterministic filters for deterministic noise and a statistical estimator only for the random floor is what makes the result clean and natural.

2. Stack decision

Python 3.12 + NumPy/SciPy + Numba + PortAudio (sounddevice) + libsndfile (soundfile) + PyQt6/pyqtgraph.

Why not C++/JUCE: no C/C++ toolchain is available on the target machine, and the workload does not need it. The per-block work is ~0.4 ms for a 5.3 ms block.

How Python is kept out of the hot path:

  • Biquad cascades, compressor, de-esser, limiter, NR gain computation are Numba @njit kernels: LLVM-compiled native loops with explicit float64 state arrays. (Numba was chosen over scipy.signal.sosfilt because the latter's Python wrapper costs ~200 µs per call — significant at 256-sample blocks — and because the dynamics stages need per-sample branching that NumPy cannot express.)
  • FFTs use NumPy's pocketfft (C).
  • Python only orchestrates: ~10 calls per block. Measured: 7 % of real time for the whole chain at 256-sample blocks including Python overhead; ~5 ms of native work per second of audio for the limiter alone.
  • The audio callback runs on PortAudio's thread. sys.setswitchinterval(0.5 ms) keeps GIL hand-off latency far below the block period; a re-entrant lock is held only for the block's processing and for parameter edits (coefficient / state arrays are swapped, so an edit must not race a running kernel).

Libraries and why:

Library Role Why this one
numpy / scipy arrays, FFT, Welch PSD, filter design helpers standard, C-backed
numba native kernels without a compiler toolchain bundles LLVM; cache=True avoids re-JIT
sounddevice PortAudio binding: WASAPI shared mode, duplex callback stream low-latency, WasapiSettings(auto_convert=True) lets us fix 48 kHz
soundfile libsndfile WAV/FLAC I/O robust, 24-bit/float
PyQt6 + pyqtgraph GUI + fast plotting already installed; pyqtgraph redraws 2×2049-point curves at 25 Hz cheaply

No third-party denoiser (RNNoise, noisereduce, etc.) is used. The noise reducer is a from-scratch implementation of the classic decision-directed Wiener estimator because the user explicitly needs controllable suppression with well-understood knobs, and because ML suppressors are exactly the "removes parts of my voice / robotic" tools that failed.

3. Module map

micfix/
  dsp/
    base.py          Processor ABC: prepare/reset/process(block)/latency_samples/describe
    biquad.py        RBJ designs (peak, shelves, notch, BP, HP/LP, Butterworth HP) + SOSChain (stateful)
    kernels.py       Numba: sos_kernel, compressor_kernel, deesser_kernel, limiter_kernel, nr_gain_kernel
    highpass.py      HighPassFilter
    hum_filter.py    HumFilter (harmonic notch bank + extra notches; per-harmonic mask)
    stft.py          StreamingSTFT (fixed-latency overlap-add, any block size)
    noise_reducer.py NoiseProfile, NoiseReducer
    eq.py            EQSettings/EQBand, ParametricEQ, speech_clarity_preset
    deesser.py       DeEsser
    compressor.py    Compressor (+ narration preset)
    limiter.py       Limiter (true-peak, look-ahead) + polyphase FIR design
    chain.py         DSPProcessor: ordered chain, global bypass (latency-aligned), metering, settings I/O
    analysis.py      analyze_noise() -> CalibrationResult; recommend_settings()
  io/
    wavio.py         read_audio / write_wav
    devices.py       device enumeration, virtual-cable detection
    realtime.py      RealtimeEngine: PortAudio duplex stream, recording, calibration capture, analyser history
  gui/
    widgets.py       ParamSlider, LevelMeter
    spectrum_widget.py  live analyser (pyqtgraph)
    main_window.py   wiring only — no DSP
    app.py
  presets.py
tools/micfix_cli.py  offline processing / calibration / synthetic test generation
tests/               signals.py (synthetic generators), test_dsp.py, test_engine.py

Block convention: float32 (frames, channels), full scale ±1. Every stage is stateful and block-size agnostic (tests assert 64-sample and 1000-sample processing give identical output).

4. Processing order and latency

input gain → HPF → hum notches → NR → EQ → de-esser → compressor → limiter → output gain
  • HPF first: rumble would otherwise bias the noise estimate and the notch filters' transient response.
  • Notches before NR: the spectral stage then sees a flat floor instead of tonal peaks 30 dB above it (which cause "musical" residue around them).
  • De-esser before compressor: the compressor would otherwise bring sibilance up.
  • Limiter last, then a static output trim (the limiter ceiling is −1 dBTP by default; positive output gain after it is allowed but defeats the ceiling — the GUI keeps it at 0 by default).
Stage latency
biquad stages (HPF, hum, EQ, de-esser band) 0
noise reducer (frame 1024, hop 256) 768 samples (16 ms)
compressor 0
limiter (1.5 ms look-ahead + 6-tap FIR centre) 78 samples (1.6 ms)
total @48 kHz 846 samples = 17.6 ms

Bypass runs the dry signal through a delay equal to the total latency, so the ORIGINAL/PROCESSED toggle is seamless and offline A/B files are sample-aligned. Bypassed stages with latency (NR, limiter) keep their delay lines running.

5. Stage details

Hum filter

RBJ notch H(z) = (1 − 2cos ω₀ z⁻¹ + z⁻²) / (1 + α − 2cos ω₀ z⁻¹ + (1 − α) z⁻²), α = sin ω₀ / 2Q. For harmonic k the Q is Q·k (capped at 200) so every harmonic notch has the same absolute width (≈1.7 Hz at Q=30): hum harmonics are all a few Hz wide, and a fixed Q would make the 250 Hz notch 5× wider than needed inside the voice fundamental region. Extra (non-mains) notches use Q = max(Q_min, f / 50 Hz): a 12 kHz USB whine gets a 50 Hz-wide notch instead of a 480 Hz hole. Only harmonics the calibration actually detected are enabled (harmonic_mask); the GUI lists the exact frequencies being removed.

Noise reducer

Per frame m, bin k, with noise power N_k from the profile (scaled by the threshold control 10^(thr/10)):

γ  = P / N                                  posterior SNR
ξ  = a · G_prev² · P_prev / N + (1−a) · max(γ−1, 0)   decision-directed a-priori SNR
G  = ξ / (1+ξ)                              Wiener gain
G  = max(G, 10^(−reduction/20))             gain floor (never a hard gate)
G  = triangular smoothing over ±r bins      "artifact control"
Gs = attack/release one-pole on G           fast up (speech onset), slow down
Y  = Gs · X

The decision-directed recursion (a ≈ 0.95) is what suppresses musical noise: random single-frame noise peaks do not survive it, sustained speech energy does. A gain floor rather than a gate is what keeps the result from sounding hollow. Without a profile the reducer learns from the first 160 ms and then updates only bins whose posterior SNR is < 5 dB (a light MCRA).

STFT: frame 1024, hop 256, √Hann analysis and synthesis (perfect reconstruction verified to 1e-5 for block sizes 100…1024). Latency N−H; if the host block size isn't a multiple of the hop, one hop of priming is added so latency is constant.

Compressor

Giannoulis/Massberg/Reiss topology: level (peak-hold with 5 ms decay, or RMS) → static soft-knee curve computed directly as gain reduction in dB → branching attack/release smoothing of the gain reduction → apply with makeup. Smoothing the gain instead of the level is what avoids pumping. Verified: −10 dBFS sine, threshold −20, 4:1 → −17.5 dBFS (±0.3 dB); below threshold bit-identical.

De-esser

Band-pass (0 dB peak, Q from the bandwidth in octaves) feeds a detector; when its envelope exceeds the threshold the band component is subtracted: y = x − (1 − 10^(−GR/20)) · bp(x). At f₀ that is exactly −GR dB, with the shape of a peaking EQ of the same bandwidth; everything outside is untouched (verified: 400 Hz vowel changes < 0.3 dB while the 6–7 kHz band drops > 5 dB).

Limiter

  • True peak: 8× polyphase windowed-sinc interpolation (12 taps/phase, Kaiser β=7; odd prototype length so phases sit at exact k/8 offsets). Worst-case read error 0.04 dB at fs/4 — better than BS.1770's 4× minimum.
  • Detector and delay line are aligned by taking the FIR centre tap as "the sample".
  • Required gain ceiling / tp per sample; the applied gain is the minimum over the look-ahead window, smoothed by a one-pole that reaches 99 % within the window (so gain ramps down before the peak), released exponentially, and clamped so it never exceeds the outgoing sample's own requirement. Final hard clamp at the ceiling only catches the last fraction of a dB.
  • Verified −0.98…−0.99 dBTP output on sines at 300 Hz…12 kHz driven 4–5 dB over, and on heavily clipped speech through the full chain.

Calibration (analysis.py)

Welch PSD (8192 points → 5.9 Hz bins), robust running-median floor per ⅙ octave, prominence = PSD − floor.

  • Mains: score a 50 Hz comb vs a 60 Hz comb (sum of prominence at k·f0); accept only if the winner has ≥ 12 dB summed prominence and the fundamental or 2nd harmonic is itself prominent (≥ 6 dB); refine f0 by parabolic interpolation (detects 59.7 Hz drift; rejects pure white noise).
  • Harmonics: each k·f0 with prominence ≥ 6 dB is DETECTED (only those are notched).
  • Other tonal peaks: prominence ≥ 8 dB, merged, not within 1.5 bins of a mains harmonic; a common fundamental is searched (USB whine ≈ 1 kHz series).
  • Hiss = mean PSD 4–16 kHz (absolute); rumble = <60 Hz floor relative to the 200 Hz–2 kHz floor (a loud flat floor is hiss, not rumble).
  • Speech check: 95th/10th percentile of 100 ms RMS > 15 dB ⇒ warn.
  • Outputs the NR noise profile (same window as the reducer) and an analyser-scale floor curve for the GUI.

Measured on the actual microphone in this session

"Microphone (2- USB PnP Sound Device)": noise floor −49…−54 dBFS RMS, weak 50 Hz component (−70 dB/Hz), a ~999 Hz harmonic series (1992, 4992, 6984, 7986, 8982, 11936, 16002 Hz; 10–40 Hz wide; up to +19 dB above the floor) — the "mosquito" — and strong <60 Hz rumble. The recommended profile notches the prominent whine peaks with 50 Hz-wide notches, sets a 90 Hz 6th-order high-pass and 10–14 dB of profile-based reduction.

6. Real-time routing and the virtual microphone

RealtimeEngine opens one PortAudio duplex stream (WASAPI shared mode, auto_convert=True, latency='low', blocksize 256): input device → chain → output device. Windows offers no user-mode API to create a capture endpoint; that needs a signed kernel driver (what VB-CABLE / Voicemeeter / VoiceMeeter Banana install). MicFix therefore outputs to the cable's playback endpoint ("CABLE Input") and other apps record from the cable's capture endpoint ("CABLE Output"). The device list flags such devices as [virtual].

Limitations:

  • One output device per stream: monitoring on headphones and feeding the cable simultaneously needs Voicemeeter (or a second MicFix instance on a second cable).
  • WASAPI exclusive mode would lower device latency further but locks the mic; shared mode is used deliberately.
  • Sample-rate: the stream is opened at 48 kHz with WASAPI auto-conversion; if a device refuses, choose "device default".

7. Test strategy

Synthetic signals (tests/signals.py): mains hum + harmonics (50/60 Hz, drifting), white and pink noise, rumble, a formant-modulated harmonic "speech-like" signal with syllable envelope and sibilant bursts, clipped speech, silence, and dirty_mic() = speech + hum + hiss + rumble.

Every requirement in the brief has an assertion: hum attenuation, voice-band preservation, no clipping, compressor behaviour, limiter ceiling (sample and true peak), bit-identical bypass, 60 s stability, plus reconstruction, block-size invariance, sample-rate independence, preset application and calibration detection/false-positive tests. tools/micfix_cli.py gen-tests writes the before/after WAVs to test_output/.