English · 简体中文
Frequency-domain GPU acoustic wave simulation.
bornwave solves the 2-D acoustic Helmholtz equation in media with arbitrary
heterogeneous sound speed, density and absorption. It is a matrix-free PyTorch
implementation of the convergent Born series (CBS) solver of Stanziola,
Arridge, Treeby & Cox (JASA, 2026), extended here with joint frequency × shot
batching, CUDA-graph execution, exact band-limited wavefield synthesis, and an
adjoint-state autograd interface. One call produces shot records, full
wavefield movies and FWI gradients.
Time-domain results are obtained by solving the wavelet's frequency components in parallel and superposing them.
Despite the name this is a full-wave solver. "Born" refers to the form of the iterative series, not to the first-order Born approximation: the converged field satisfies the heterogeneous Helmholtz system to the measured true residual, internal multiples and diffractions included.
Two-layer model with a low-velocity lens: pressure wavefield movie and shot
record, both produced by a single call in examples/demo_engine.py.
Requires Python ≥ 3.10 and PyTorch ≥ 2.4 (CUDA optional but strongly
recommended), plus NumPy, SciPy and Matplotlib. ffmpeg is needed for MP4
export; without it, movies fall back to animated GIF.
git clone https://github.com/zzzzswh/bornwave.git
cd bornwave
uv sync # or: pip install -e .Verify the installation — this cross-validates the engine against the reference solver and runs an end-to-end consistency check (~1–2 min on CPU):
uv run tests/test_engine_api.pyThen reproduce the figures above, and the full physics validation suite:
uv run examples/demo_engine.py # records + wavefield movie
uv run examples/two_layer_ricker.py # analytic validation (Hankel, Zoeppritz, moveout)
uv run examples/fwi_gradient.py # single-frequency FWI gradientimport numpy as np
from bornwave import acoustic2d, trace_norm, plot_shot, plot_wavefield_video
nz, nx = 300, 400 # arrays are (nz, nx), i.e. vp[z, x]
dh, dt, nt, f0 = 10.0, 1e-3, 2000, 15.0
vp = np.full((nz, nx), 2500.0); vp[180:] = 3200.0
rho = np.full((nz, nx), 2000.0); rho[180:] = 2300.0
res = acoustic2d(
vp, rho, dh, dt, nt, f0,
sx=[nx // 2], sz=[10], # several shots: pass lists,
rx=np.arange(0, nx, 2), rz=10, # they are solved as one batch
nbc=60, snap_interval=25,
)
res.seis_p, res.seis_vx, res.seis_vz # (nt, nrec) time-domain records
res.snaps # (nsnap, nz, nx) wavefield movie
res.H_p # unit-source transfer functions
res.resynthesize(other_wavelet) # swap wavelets at zero cost
res.stats.kernel_time_s # timing and iteration diagnostics
plot_shot(trace_norm(res.seis_p), "shot.png", dt=dt)
plot_wavefield_video(res.snaps, "wavefield.mp4", fps=12, dh=dh,
snap_times=res.snap_times, adaptive_clims=True, model=vp)The solver works on the first-order acoustic system in the frequency domain,
discretized on a staggered Fourier grid with field ordering
Superscripts
the latter being exactly equivalent to the frequency-dependent
Write the system as
The contraction
pointwise multiply-add → one batched FFT → an unrolled per-$k$
$3\times3$ product → IFFT → pointwise multiply-add
No matrix assembly, no decomposition, no inner solver, and nothing that grows
with the number of shots: all operator tensors are shared across the shot
batch, which is why additional shots cost almost nothing beyond field memory.
The einsum lowers to permute + bmm on CUDA and copies the full
symbol tensor every iteration.
A source wavelet is decomposed with rfft; bins above an amplitude threshold
are retained and solved in chunks of adjacent frequencies. Records follow from
- Arbitrary heterogeneous models — sound speed, density (interpolated onto the staggered half-grid) and absorption in Np/m or constant-$Q$.
- Spectral spatial accuracy — a Fourier pseudospectral discretization, so there is no grid dispersion to accumulate; the validation suite runs at 10 points per wavelength and the Nyquist floor is 2.
-
Frequency × shot batching — the engine iterates a single
(F, B, 3, Nz, Nx)tensor, with converged frequencies finalized and compacted out of the working set on the fly. - CUDA Graphs — at these grid sizes and iteration counts the wall time is dominated by kernel launch latency, not FLOPs. The fixed-point loop is captured and replayed as a single launch; the graph is re-captured after each compaction. Capture failure falls back to eager execution with a warning, and results are unchanged.
-
Wavefield movies without extra solves — snapshots are synthesized from
the same transfer functions as the records, and agree with
np.fft.irfftto machine precision. -
Differentiable —
solve_helmholtzimplements the adjoint-state gradient via the implicit function theorem: exact, and$O(1)$ in memory with respect to the iteration count.
| Argument | Meaning |
|---|---|
vp, rho
|
(nz, nx) velocity [m/s] and density [kg/m³]; strictly positive. rho may be a scalar. |
dh, dt, nt
|
Grid spacing [m], sample interval [s], number of time samples. |
f0 |
Ricker peak frequency [Hz]; ignored if wavelet is supplied. |
sx, sz
|
Shot x/z grid indices — scalars or equal-length sequences, solved as one batch. |
rx, rz
|
Receiver x/z grid indices; rz may be a scalar and is broadcast. |
alpha / Q
|
Absorption [Np/m], or constant-$Q$ quality factor. Mutually exclusive. |
nbc |
Sponge thickness in cells; 40–60 is typical. This is a polynomial |
tol |
Stopping tolerance on the relative increment. 2e-4 gives roughly 0.1–1 % amplitude accuracy (see Validation). |
freq_batch |
Frequencies solved jointly per chunk; the main memory/throughput knob. |
snap_interval |
Store a full pressure wavefield every this many time samples. |
cuda_graph |
True / False / "auto". |
Returns an AcousticResult holding seis_p, seis_vx, seis_vz, snaps,
snap_times, transfer functions H_p, per-frequency iterations and
residuals, a stats namespace, and resynthesize(wavelet).
Note that vx/vz records are the staggered
from bornwave import CBSSolver2D, CBSFreqShotBatch2D, synthesize_shot, solve_helmholtz| Object | Use |
|---|---|
CBSSolver2D |
Single frequency, shot-batched. The reference implementation. |
CBSFreqBatch2D |
Frequency batch with convergence compaction. |
CBSFreqShotBatch2D |
Joint (frequency × shot) batch, CUDA-graph accelerated. Backs acoustic2d. |
synthesize_shot |
Wavelet → frequency band → time-domain gather, without the engine layer. |
solve_helmholtz |
Differentiable single-frequency solve; gradients w.r.t. |
solve_helmholtz is a torch.autograd.Function built on the implicit function
theorem. The backward pass runs one adjoint solve on the same CBS machinery
(examples/fwi_gradient.py shows a three-shot, single-frequency
FWI gradient imaging an interface absent from the starting model.
Every layer is checked against something it did not produce itself — operator identities, analytic Green's functions, plane-wave reflection theory, reciprocity, and cross-validation between implementations.
| Check | Reference | Result | Script |
|---|---|---|---|
|
|
exact identity | 9.0e-16 | test_operator_identity.py |
| Adjoint symbol; skew-Hermitian differential block | exact identity | 1.4e-15 / 0 | test_operator_identity.py |
| Homogeneous medium, 10 ppw | analytic 2-D Hankel Green's function | rel. L2 8.9e-5, amplitude ratio 1.0000, phase 0.00° | test_homogeneous_hankel.py |
| Disk with 2× speed and 2.5× density contrast | acoustic reciprocity | 3.2e-6 | test_heterogeneous_reciprocity.py |
| Two-layer + 15 Hz Ricker, direct-wave window | band-limited Hankel | mean 0.075 %, max 0.18 % | examples/two_layer_ricker.py |
| Zero-offset reflection amplitude | fluid Zoeppritz |
0.5710 (0.08 %) | examples/two_layer_ricker.py |
| AVO curve, offsets ≤ 600 m | fluid Zoeppritz |
mean 0.13 %, max 0.54 % | examples/two_layer_ricker.py |
| Reflection moveout / inverted interface depth | ray theory, |
≤ 1.3 ms (< 1 sample); fitted 596.3 m | examples/two_layer_ricker.py |
| Engine (freq × shot batch) |
CBSFreqBatch2D reference solver |
2.6e-7 | test_engine_api.py |
| Batched shots | the same shots solved sequentially | 0.0 | test_engine_api.py |
| Wavefield snapshots at shared samples | receiver records | 3e-16 | test_engine_api.py |
| Direct-wave lag between receivers | offset / |
exact (50 / 50 samples) | test_engine_api.py |
| Band-limited time-slice synthesis | np.fft.irfft |
machine precision | test_timesynth.py |
| Frequency batch | serial CBSSolver2D
|
shifts/scales exact; |
test_freq_batch.py |
| Autograd through |
torch.autograd.gradcheck + directional finite differences |
pass, rel. diff < 3e-5 | test_autograd.py |
Residuals quoted for the solver are true residuals
A full measured log of these runs is in tests/test-log-260726.txt.
Reference problem — examples/demo_engine.py:
| Grid | 300 × 400, padded to 420 × 525 |
| Time samples | 2000 |
| Frequencies | 95, covering 0.5–47.5 Hz |
| CBS iterations | 76,456 total |
| Kernel time | 27 s, CUDA Graphs enabled |
| Hardware | single CUDA GPU (NVIDIA <model>) |
Additional shots share all operator tensors and cost almost nothing beyond the
extra field memory. The per-chunk working-set size is printed at startup and is
controlled by freq_batch; iteration count grows with frequency, so grouping
adjacent bins keeps compaction waste small.
Four conventions that the source papers leave implicit, or that differ on a staggered grid. Each was determined by experiment and is locked in by a test — worth reading before modifying the internals.
-
Time convention is
$+i\omega t$ (the$H_0^{(2)}$ branch). Implemented literally, it coincides with the NumPy/PyTorch FFT transfer-function convention, so synthesis is$d(t) = \mathrm{irfft}(W \cdot H)$ with no conjugation anywhere. -
The 2-D point-source normalization is amplitude/$\Delta^2$. A single node on a spectral grid represents a band-limited sinc of unit integral; the
$2c_0/\Delta$ correction given in the paper is specific to 1-D. The measured amplitude ratio against the analytic Hankel solution is 1.0000. -
On a staggered grid the per-$k$
$(L+I)^{-1}$ matrix is not symmetric (the stagger phases$e^{\pm ik\Delta/2}$ break it). The adjoint symbol is the per-$k$ conjugate transpose, not the elementwise conjugate — the latter holds only for the non-staggered variant. Both cost the same to apply. -
Grazing-incidence sponge ghosts. Sponge absorption is inefficient for near-horizontally propagating energy. Keep sources and receivers roughly one dominant wavelength away from the absorbing layer; closer than that, residual grazing reflections are not separable from the direct wave and long-offset errors reach several percent. This is the same failure mode as absorbing boundaries in time-domain finite differences.
-
Time wraparound. Frequency sampling
$\Delta f = 1/(n_t \Delta t)$ makes the synthesized response periodic with period$T = n_t \Delta t$ : any coda still ringing at$t = T$ aliases back to$t = 0$ and appears before the source fires. Increasent, or window/damp, when late energy matters. The residual noise floor of the iterative solve is likewise non-causal — uniform in time — and is set bytol. -
No free surface yet. Vacuum cells (
vp = 0) lie outside the CBS convergence domain by construction, since the contraction requires bounded contrast. All four boundaries are absorbing sponges, so there are no surface multiples; internal multiples are fully present. - 2-D only, with a single uniform grid spacing.
bornwave/
solver.py CBSSolver2D — single frequency, shot-batched
multifreq.py CBSFreqBatch2D — frequency batch + convergence compaction
engine.py CBSFreqShotBatch2D — (freq × shot) batch + CUDA Graphs
api.py acoustic2d — one-call forward-modeling entry point
autograd.py differentiable solve (implicit function theorem / adjoint)
operators.py per-k Fourier symbols of (L+I)^-1 and (L+I), staggered
grid.py FFT-friendly sizes, sponge profiles, staggered averaging
synthesis.py Ricker wavelet, band selection, per-frequency synthesis
timesynth.py band spectra → time slices (torch-free, machine precision)
viz.py shot plots, wavefield movies (torch-free)
analytic.py Hankel Green's function, fluid Zoeppritz (validation only)
examples/ demo_engine.py, two_layer_ricker.py, fwi_gradient.py
tests/ validation suite + measured log
- Free surface via the image method
- Complex-frequency damping to suppress time wraparound
- Anderson fluid-cylinder analytic benchmark for the variable-density path
- True-residual, frequency-adaptive stopping criterion
- Frequency-bucket scheduling
- Osnabrugge (2021) ultra-thin absorbing boundary layer
- 3-D (four fields, eight FFTs per iteration)
This repository is an independent implementation. All algorithmic credit belongs to the following papers; the implementation, the engine layer and any bugs are ours.
- A. Stanziola, S. R. Arridge, B. E. Treeby, B. T. Cox, Iterative Born solver for the acoustic Helmholtz equation with heterogeneous sound speed and density, J. Acoust. Soc. Am. 159 (2026) 1457–1470. doi:10.1121/10.0042259 · arXiv:2507.16087
- T. Vettenburg, I. M. Vellekoop, A universal matrix-free split preconditioner for the fixed-point iterative solution of non-symmetric linear systems, arXiv:2207.14222 (2022).
- G. Osnabrugge, S. Leedumrongwatthanakun, I. M. Vellekoop, A convergent Born series for solving the inhomogeneous Helmholtz equation in arbitrarily large media, J. Comput. Phys. 322 (2016) 113–124. doi:10.1016/j.jcp.2016.06.034
BibTeX
@article{stanziola2026iterativeborn,
title = {Iterative {Born} solver for the acoustic {Helmholtz} equation
with heterogeneous sound speed and density},
author = {Stanziola, Antonio and Arridge, Simon R. and
Treeby, Bradley E. and Cox, Benjamin T.},
journal = {The Journal of the Acoustical Society of America},
volume = {159},
number = {2},
pages = {1457--1470},
year = {2026},
doi = {10.1121/10.0042259}
}
@misc{vettenburg2022universal,
title = {A universal matrix-free split preconditioner for the
fixed-point iterative solution of non-symmetric linear systems},
author = {Vettenburg, Tom and Vellekoop, Ivo M.},
year = {2022},
eprint = {2207.14222},
archivePrefix = {arXiv},
primaryClass = {math.NA}
}
@article{osnabrugge2016convergent,
title = {A convergent {Born} series for solving the inhomogeneous
{Helmholtz} equation in arbitrarily large media},
author = {Osnabrugge, Gerwin and Leedumrongwatthanakun, Saroch and
Vellekoop, Ivo M.},
journal = {Journal of Computational Physics},
volume = {322},
pages = {113--124},
year = {2016},
doi = {10.1016/j.jcp.2016.06.034}
}Not yet declared. Please open an issue if you need a specific license for your use case.
