Drop-in speed patches for PGSR training — same geometry, a lower GPU bill.
pgsr-fast is an idempotent patcher that strips avoidable per-iteration overhead
out of the PGSR training loop without changing what the training does. PGSR
(planar-based Gaussian splatting for surface reconstruction, a fork of 3D
Gaussian Splatting) repeats a handful of small wastes on every single step:
buffers born on the CPU and copied over PCIe, a multi-megapixel pixel grid
rebuilt from scratch each iteration, five all-ones convolutions that are
mathematically just sums, and GPU→CPU syncs issued purely for logging. It also
ships a real operator-precedence bug in three of its logging EMAs. This tool
rewrites those spots in place, proves the rewrites are numerically equivalent on
CPU, and leaves the geometry untouched. The mesh-extraction path in render.py
gets the same treatment: one real crash fixed, and four accuracy traps exposed as
flags that default to exactly what upstream hardcoded.
You run one command against your own PGSR checkout. Every change is anchored to
exact source text, tagged with a # pgsr-fast marker, and safe to re-run.
| Speedup | 1.095× (+9.5%) (training loop, same recipe) |
| Loss parity | within_noise — within the baseline-vs-baseline noise floor |
| Hardware | RTX 4090 |
| Scene | real 469-photo, 12 MP capture |
| Geometry | unchanged — parity is the whole point |
Thirteen patches, P1–P13. Everything applies by default except P7
(--fused-ssim) and P8 (--gpu-tsdf), which are opt-in. P1–P8 are about
the training loop; P9–P13 are about the mesh-extraction path in render.py.
P9 is a bug fix — it ships on by default because it's identical on
single-camera captures and corrects a real crash on multi-camera ones. P10–P13
add render-time flags whose defaults are the constants they replace, so
patching changes nothing until you pass one. Each row is one waste (or bug), one
fix, and why it is safe.
| # | File | The waste | The fix | Why it's safe |
|---|---|---|---|---|
| P1 ‡ | gaussian_renderer/__init__.py |
input_all_map (~28 MB) allocated in pageable host memory, then a synchronous H2D copy over PCIe — ~2×/iter |
Allocate the buffer directly on the GPU (device="cuda") |
Same zeros, same dtype — only the allocation device changes. Numerically identical. |
| P2 † | train.py |
The multi-view pixel grid (12 MP → ~96 MB) rebuilt on the CPU with meshgrid and copied to the GPU every multi-view iteration |
Build it once per (H, W, device) and cache it |
The grid depends only on resolution and is read-only downstream. Same values, no mutation. |
| P3 † | utils/loss_utils.py |
lncc runs five 11×11 all-ones conv2ds but reads only the central pixel — 121× the FLOPs of what's used |
Replace each with a direct patch sum | Mathematical identity: an all-ones convolution at its fully-overlapping centre is the sum over the patch (holds for odd patch_size). Verified numerically in tests. |
| P4 † | train.py |
Operator-precedence bug: a if c else 0.0 + 0.6*ema parses as a if c else (0.6*ema), so the Single/Geo/Pho tqdm columns dropped the running average entirely — they were never EMAs |
Parenthesize correctly; leave an EMA untouched when its loss is absent | Logging-only. Zero effect on training; it makes the health telemetry read true. |
| P5 | train.py |
Per-iteration .item() / TensorBoard scalar writes force a GPU→CPU sync that stalls the pipeline every step |
Keep the log EMAs as GPU scalars; flush .item() and TB scalars every 10th iter (the tqdm cadence) |
Logging-only. Identical numbers, read on a coarser schedule. |
| P6 | train.py |
A heavy debug-visualization JPEG dump runs on a fixed interval whether or not anyone wants it | Gate it behind PGSR_DEBUG_VIZ=1 (off by default) |
Debug-only. Re-enable any time with the env var. |
| P7 | train.py |
Reference SSIM runs as six separate convolution passes in PyTorch | (opt-in, --fused-ssim) Wire in fused_ssim when installed, with a silent fallback to reference SSIM |
Not bit-identical — a fused CUDA kernel, off by default. A/B it before relying on it. |
| P8 | render.py |
Mesh TSDF fusion runs on the CPU via Open3D ScalableTSDFVolume — single-threaded integration + marching cubes, minutes on a large capture |
(opt-in, --gpu-tsdf, runtime PGSR_GPU_TSDF=1) Swap in open3d.t.geometry.VoxelBlockGrid CUDA fusion, with a clean fallback to the CPU volume if Open3D-CUDA is unavailable |
Not bit-identical, opt-in — A/B your mesh first. A different fusion + marching-cubes implementation, off by default; the default CPU path stays byte-identical to upstream. |
| P9 ★ | render.py |
Real upstream bug (not a waste). TSDF fusion torch.stacks the per-view depth maps and reuses one global W,H — the last view's dims — for every camera's PinholeCameraIntrinsic. On any capture with more than one camera group (mixed image sizes) the stack crashes, and even when it doesn't, other groups integrate at the wrong intrinsics |
Keep the depths as a list (indexed per view, exactly as before); build each view's intrinsic from that view's own ref_depth.shape |
Bug fix — identical on single-camera datasets, correct on multi-camera. The list indexes exactly as the stack did and per-view dims equal the globals when there's one camera; default-on because it's a correctness fix, not a tradeoff. |
| P10 ◆ | render.py |
Depth is quantized to uint16 at 1000 counts/m before Open3D fuses it — 1 mm steps, coarser than the default 2 mm voxel and a quarter of the truncation band. The uint16 image is what Open3D's integrators take, so the quantization is not avoidable by "keeping float32"; the scale is the only lever | --depth_scale (default 1000 = vendor). 10000 quantizes at 0.1 mm, an order of magnitude under the voxel |
Default reproduces the vendor value exactly. A uint16 holds 65535 counts, so the scale is clamped to 65535/max_depth with a warning — an overflow wraps far geometry to near depth silently. |
| P11 ◆ | render.py |
sdf_trunc = 4.0 * voxel_size is hardcoded. The truncation band is the distance each view's signed distance is blended over, so it doubles as a low-pass filter: a four-voxel band averages away relief shallower than itself |
--sdf_trunc_mult (default 4.0 = vendor). A narrower band (2.0) keeps shallow relief the wide one flattens, with less cross-view smoothing |
Default is the hardcoded constant, so an unflagged run is unchanged. Honours --gpu-tsdf too, so the knob can't silently do nothing on the CUDA path. |
| P12 ◆ | render.py |
ref_depth[ref_depth>max_depth]=0 is a guillotine, not a clip: every view whose camera sits beyond max_depth loses all of its depth and fuses nothing — no surface, and no free-space carving either. On a real 468-photo capture it silently emptied 149 views, and nothing in the logs said so |
Count the views the cut empties and print one warning at the end of fusion, naming the count, the max_depth in force, and the symptom |
Diagnostic only — zero behaviour change. The cut is left exactly as it was; the patch adds a probe and a print. |
| P13 ◆ | render.py, gaussian_renderer/__init__.py |
PGSR's depth is an alpha-weighted expectation with no coverage test — the rasterizer emits a depth for every pixel of the frustum, so a pixel holding a few percent of accumulated alpha still votes a plausible depth into empty space. Those votes land a voxel or two off, disagree between views, and the TSDF fuses the disagreement into thin sheets of dust around the object. The rasterizer already computes rendered_alpha; the return dict drops it |
Return rendered_alpha (no extra kernel — it was being discarded), and zero depth samples below --alpha_gate before fusion |
Default 0.0 = no gate, vanilla path. At 0.5 — what 2DGS and RaDe-GS mask at — the dust goes and the real surface is preserved: real geometry renders at alpha ≈ 1, so the gate has nothing to take from it. |
Not just "known tricks." These came from a line-by-line read of the training loop, and the audit that followed is part of why we trust them:
- P2, P3 and P4 have no public record — no issue, PR, or fork on PGSR or 3DGS addresses them, as of a July 2026 sweep of the ecosystem.
- P1 has been fixed independently in at least three forks (
DP-GS,ibgs,Endo-MCGS) — three teams converged on the exact same allocation fix — yet it was never reported or named as an optimization anywhere, and ~60 other forks still carry the waste. Independent convergence is the strongest evidence the fix is correct. - P9 was found live — it crashed our own multi-camera reconstruction mid-run
(two camera groups from a re-run extraction). The same stale-
W,Hpath was also integrating every view of the second group at the wrong intrinsics even in the runs where it didn't crash. It's a real upstream bug, shipped default-on because the fix is identical on single-camera captures and correct on multi-camera ones. - P10–P13 came out of a mesh-dust post-mortem. A finished 468-photo reconstruction produced a mesh wrapped in thin sheets of surface dust and ghost shells. Tracing it back through the extraction path turned up four independent causes — a coarse depth quantization, a hardcoded truncation band, a depth cut that silently emptied a third of the views, and depth fused with no coverage test at all. P13 is the one that mattered: an alpha gate at 0.5 removed the dust and left the surface intact. All four are flags because they are trade-offs and scene-dependent, not universal wins — the defaults keep upstream's behaviour.
† no public record in PGSR/3DGS issues, PRs, or forks · ‡ independently fixed in ≥3 forks, never reported · ★ real upstream bug, found live when it crashed our own multi-camera reconstruction · ◆ found while diagnosing surface dust and ghost shells on a real 468-photo capture
Optimizations that quietly degrade quality are the classic sin of this niche, so parity is treated as the deliverable, not a footnote.
Numerical-equivalence tests on CPU. Before any GPU ever runs, the suite copies
the real vendor PGSR files, patches them, and compares the patched behaviour to a
reference copied from the original code — lncc value equivalence, the pixel-grid
cache (value and object identity on a cache hit), the corrected EMA formula
(and proof the buggy one really differed), idempotency (byte-identical on re-run),
and that every anchor still matches the current vendor source.
A/B parity judged against a noise floor, not against zero. PGSR fixes its seed
(setup_seed(22), cudnn.deterministic=True), yet training is still not
bit-reproducible: the custom rasterizer's backward pass accumulates gradients with
floating-point atomicAdd, whose summation order depends on nondeterministic GPU
thread scheduling. Floating-point addition isn't associative, so two runs of the
exact same code drift apart over a measurement window. There is no "identical
loss" target to hit — the honest baseline is the same-code, run-to-run noise
envelope. The patched arm passes when its Loss curve stays inside that
baseline-vs-baseline envelope and fails only if it diverges beyond it. The harness
interleaves runs with an alternating lead to cancel warm-up and neighbour noise.
The render-time knobs default to the vendor's own constants. P10–P13 turn four
hardcoded values in the extraction path into flags, and a run that passes none of
them fuses exactly what upstream fused. That claim is tested rather than asserted:
the suite reads the module defaults and the argparse defaults back out of the
patched file's AST (1000, 4.0, 0.0), checks the gate is behind an
if alpha_gate > 0.0, and checks P12 left the depth cut it warns about byte-for-byte
alone. What each flag does when you do pass it is proved separately, on the
arithmetic in isolation.
fused-ssim is kept separate on purpose. P7 is a fused CUDA kernel, not bit-identical to the reference SSIM. It is off by default, gated behind a flag, and benchmarked as its own arm so its (larger, expected) loss delta never contaminates the parity claim for P1–P6.
Apply the patches to your own PGSR checkout (run it on a clean git checkout so
git diff shows exactly what changed):
python patch_fast_train.py /path/to/PGSR # P1–P6 + P9–P13 (default-on)
python patch_fast_train.py /path/to/PGSR --fused-ssim # also wire in P7
python patch_fast_train.py /path/to/PGSR --gpu-tsdf # also wire in P8 (render.py CUDA TSDF)--fused-ssim and --gpu-tsdf are independent and combinable. --gpu-tsdf only
wires the code in; the CUDA mesh-fusion path activates at runtime when you set
PGSR_GPU_TSDF=1 for the render.py extraction run (default stays on the CPU
volume, byte-identical to upstream — A/B the mesh before trusting the GPU one).
P10–P13 add flags to your render.py and change nothing until you use them. If
your mesh shows dust or ghost shells, this is the recipe they came from:
python render.py -m <model> --max_depth 8.0 --alpha_gate 0.5 --depth_scale 10000 --sdf_trunc_mult 2.0--alpha_gate 0.5(P13) — drop depth from pixels almost nothing was rendered into. This is the one that removes the dust; try it alone first.--depth_scale 10000(P10) — quantize at 0.1 mm instead of 1 mm.--sdf_trunc_mult 2.0(P11) — narrower truncation band; keeps shallow relief.--max_depthpast your furthest camera — P12 prints one warning at the end of fusion if the cut emptied whole views.
The tool is idempotent: re-running skips already-patched spots and writes nothing
if a required anchor is missing (it tells you which one). Re-enable the debug
visualization after patching with PGSR_DEBUG_VIZ=1 in the training environment.
Run the equivalence tests (CPU-only, no GPU needed):
python -m venv .venv && . .venv/bin/activate # Windows: .venv\Scripts\Activate.ps1
pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install -r requirements.txt
PGSR_VENDOR=/path/to/PGSR pytest tests/ -v
# Windows PowerShell: $env:PGSR_VENDOR="C:\path\to\PGSR"; pytest tests/ -vPGSR_VENDOR points the tests at a real PGSR checkout to diff against; if it is
unset the vendor-dependent tests skip cleanly.
Run the A/B benchmark yourself (pod-side, resume-from-checkpoint window):
bash bench/bench_ab.sh # first call seeds the shared flags; second call runs A/BSee bench/README-bench.md for the full harness,
interleaving scheme, and the parity verdict semantics.
Measured on an RTX 4090 against a real 469-photo, 12 MP capture, resuming from a shared checkpoint and training an identical 1,500-iteration window (resume 7000 → 8500) on both arms.
| Arm | Run | Loop rate (it/s) | Wall (s) |
|---|---|---|---|
| baseline | 1 | 1.122 | 1628.5 |
| baseline | 2 | 1.109 | 1644.1 |
| baseline | mean | 1.116 | 1636.3 |
| patched | 1 | 1.237 | 1481.1 |
| patched | 2 | 1.207 | 1520.8 |
| patched | mean | 1.222 | 1501.0 |
Methodology: 4 runs interleaved with an alternating lead (round 1 baseline, patched; round 2 patched, baseline) to cancel warm-up and neighbour noise; parity judged against the same-code run-to-run noise floor, not against zero. Loop rate isolates the training loop; wall time includes the one-off checkpoint load/save.
Speedup: 1.095× (+9.5%) on the training-loop rate (1.090× on wall time).
Our initial estimate from reading the loop was 12–25%; the measured gain on this scene is 9.5% — measured numbers rule.
Parity verdict for the pure patched arm (P1–P6): within_noise — the
cross-arm Loss difference (max 0.00472) is at or below the same-code noise
floor (0.00477). The Single / Geo / Pho columns are expected to differ:
P4 fixes the EMA bug, so those numbers change on purpose.
Beyond the interleaved A/B window, the patched tree has since carried a complete
production training end-to-end on the same 469-photo scene: 23,000 iterations
(checkpoint resume 7k → 30k) at a sustained ~1.22 it/s with zero incidents,
finishing at 2.80M gaussians, followed by TSDF mesh extraction through the P9
multi-camera path on the resulting model. That run is what gates the v0.1.0
release tag.
- DISTWAR port — cut the backward-pass
atomicAddcontention in the diff-plane rasterizer (the dominant per-iteration cost). - Coarse-to-fine resolution schedule — Trick-GS-style, validated against Chamfer / F-score so geometry is provably preserved.
- Wider-pipeline notes — GLOMAP in place of the COLMAP mapper, for anyone
running PGSR end-to-end to a mesh. (CUDA TSDF mesh fusion via
VoxelBlockGridnow ships as opt-in P8 above —--gpu-tsdf; A/B the mesh before trusting it.)
The pgsr-fast patcher and benchmark harness (this repository's own code) are
released under the MIT license.
They do not relicense PGSR. The patches apply to your PGSR checkout, in
place. PGSR carries its own license from ZJU — educational, research, and
non-profit use only; any derivative must stay open-source; commercial use is
prohibited; all copyright, patent, trademark, and attribution notices must be
retained. Running these patches produces a modified PGSR checkout, and that
result is governed by PGSR's license, not ours. If your use is anything other
than non-commercial research, that is between you and ZJU — this tool changes
nothing about it. When in doubt, read LICENSE.md in the PGSR repo and contact
ZJU for commercial terms.
If pgsr-fast is useful, please cite the work it builds on:
- PGSR — Chen et al., PGSR: Planar-based Gaussian Splatting for Efficient and High-Fidelity Surface Reconstruction. https://github.com/zju3dv/PGSR
- 3D Gaussian Splatting — Kerbl et al., 3D Gaussian Splatting for Real-Time Radiance Field Rendering, SIGGRAPH 2023, Inria / MPII. https://github.com/graphdeco-inria/gaussian-splatting
- fused-ssim (P7, optional) — https://github.com/rahul-goel/fused-ssim (MIT).
Found via a line-by-line read of the training loop. Issues and PRs welcome.

