Severity: Medium. On the MIT samples, a CPU run takes over an hour (see the performance issue), so a failure at the very end is expensive.
Problems
- The output path is checked last.
save_wav is the final step, so a directory that doesn't exist or can't be written raises FileNotFoundError/PermissionError as a traceback after all frames are processed, and the recovered audio is lost.
- Errors go to stdout. Every
Error: / Warning: message uses plain print(). Scripts and pipelines can't separate diagnostics from output. (The VRAM warning is the only one sent to stderr.)
- GPU device 0 is hard-coded.
torch.device('cuda'), get_device_name(0) and mem_get_info(0) are all fixed, and there's no --device flag, although the sibling projects have one.
estimate_vram ignores its nlevels argument. It's a fixed 15× multiplier, and it only warns, so an OOM still happens later.
- No
torch.no_grad(). The GPU forward pass runs without torch.inference_mode(). It's harmless today because the filters are buffers, but a future pytorch_wavelets could make them parameters and track gradients.
- Timing. Pipeline timing starts before validation, and the CPU path imports
dtcwt twice.
Suggested fix
- Resolve and check the output path at the start of
main(): create the parent directory or fail, and do a test write with os.access.
- Wrap the final write so a failure saves to a fallback path such as
./sound_<timestamp>.wav rather than losing the result.
- Send all diagnostics to
sys.stderr, or use logging. Add -q/--quiet.
- Add
--device cpu|cuda[:N], used for model placement, get_device_name and mem_get_info. It also allows CPU-torch runs in CI.
- Use
with torch.inference_mode(): around the forward and phase maths.
- Either use
nlevels in the estimate or drop the parameter. Catching OOM once and halving the batch automatically would be friendlier than exiting.
Acceptance criteria
Severity: Medium. On the MIT samples, a CPU run takes over an hour (see the performance issue), so a failure at the very end is expensive.
Problems
save_wavis the final step, so a directory that doesn't exist or can't be written raisesFileNotFoundError/PermissionErroras a traceback after all frames are processed, and the recovered audio is lost.Error:/Warning:message uses plainprint(). Scripts and pipelines can't separate diagnostics from output. (The VRAM warning is the only one sent to stderr.)torch.device('cuda'),get_device_name(0)andmem_get_info(0)are all fixed, and there's no--deviceflag, although the sibling projects have one.estimate_vramignores itsnlevelsargument. It's a fixed 15× multiplier, and it only warns, so an OOM still happens later.torch.no_grad(). The GPU forward pass runs withouttorch.inference_mode(). It's harmless today because the filters are buffers, but a futurepytorch_waveletscould make them parameters and track gradients.dtcwttwice.Suggested fix
main(): create the parent directory or fail, and do a test write withos.access../sound_<timestamp>.wavrather than losing the result.sys.stderr, or uselogging. Add-q/--quiet.--device cpu|cuda[:N], used for model placement,get_device_nameandmem_get_info. It also allows CPU-torch runs in CI.with torch.inference_mode():around the forward and phase maths.nlevelsin the estimate or drop the parameter. Catching OOM once and halving the batch automatically would be friendlier than exiting.Acceptance criteria
--deviceis honoured.