Reconstructing a real room as a 3D Gaussian splat from a single handheld phone video, on a consumer GPU, in about 20 minutes.
This is a field log, not a library. It records what actually went wrong across three capture attempts and what fixed it — including a run where COLMAP registered 2 of 353 images and the two-line change that took it to 351.
Live result: https://room-sepia.vercel.app/scan.html · 中文完整版
Same machine, same software, same steps. Only the source footage changed.
| v1 living room | v2 bedroom | v3 reshoot | |
|---|---|---|---|
| Source | 1440×810, 2.5 Mbps, 48 s | 3840×2160, 44.9 Mbps, 80 s | 3840×2160, 44.9 Mbps, 80 s |
| Frames used | 240 | 402 | 353 |
| COLMAP registered | 240 / 240 | 361 / 402 | 351 / 353 |
| Sparse points | 25,090 | 80,658 | 72,400 |
| Mean track length | 10.36 | 12.40 | — |
| Mean reproj. error | 0.653 px | 0.624 px | — |
| Final splats | 735,407 | 612,803 | — |
| Opacity p50 | 0.057 | 0.10 | — |
| Needle ratio p99 | 1,701.74 | 53.57 | 39.70 |
| Train time | 14 min (15k steps) | 8.7 min (20k steps) | — |
| Delivered | — | — | 5.79 MB SOG |
Needle ratio p99 is the direct measure of the smeared, ink-wash floaters you see when you orbit away from the capture path. v1 → v2 dropped it 32×, and nothing in the software changed.
Two things fall out of that table:
- The footage decides the result. The software only moves information around. Reshooting beat every post-process I tried.
- Better data trains faster. v2 ran more steps in less time than v1, because the optimizer stopped fighting contradictory pixels.
v3 was the best footage of the three by every input metric. Its first COLMAP run registered 2 images out of 353. Two changes, in this order of importance:
COLMAP's default focal guess is 1.2 × max(width, height). For a portrait 4K iPhone frame that guesses 2592 px. The true focal length is 945 px — off by 2.7×. A bad initial guess triangulates the first pairs at the wrong depth, and the model grows crooked from image three onward.
colmap feature_extractor --database_path db.db --image_path images \
--ImageReader.single_camera 1 --ImageReader.camera_model OPENCV \
--ImageReader.camera_params "945.36,944.21,608,1080,-0.00033,0.00072,0.00009,-0.00007"
# and pin it during mapping:
colmap mapper ... --Mapper.ba_refine_focal_length 0Those are the values for an iPhone 16 shooting 4K portrait, resized to 1216×2160. Get your own from a checkerboard calibration, or from an EXIF-derived estimate — anything is better than the 1.2× default.
The distortion coefficients turned out to be tiny (k1 = −0.00033), which means the phone already corrects the lens internally. SIMPLE_RADIAL would have been fine. It was the focal length that mattered, not the model.
Every tutorial says sequential_matcher for video. With correct intrinsics but sequential matching, v3 still shattered into five disconnected components (221 / 87 / 54 / 35 / 10) — the breaks are wherever the camera turned quickly. Exhaustive matching registered 351 / 353 in one pass, and took 8 minutes on an RTX 5070.
Under ~600 images, always use exhaustive_matcher. Sequential only starts paying for itself in the thousands.
| Setup | Registered |
|---|---|
| No intrinsics prior + sequential | 2 / 353 |
| Intrinsics prior + sequential | 54 / 353 |
| Intrinsics prior + exhaustive | 351 / 353 |
v2 succeeding without a prior was luck. It happened to hit an initial pair that survived the bad guess.
The first complaint about v2 was that the black door, the dark wardrobe and the ceiling shadows smeared worse than everything else. That is measurable:
| Brightness | Share of splats | Needle p50 | Needle p90 |
|---|---|---|---|
| very dark <0.10 | 16.6% | 5.17 | 20.1 |
| dark 0.10–0.20 | 6.9% | 3.95 | 15.2 |
| mid 0.20–0.35 | 11.6% | 3.64 | 12.9 |
| bright 0.35–0.55 | 21.7% | 3.24 | 11.2 |
| very bright >0.55 | 43.3% | 3.28 | 11.8 |
Monotonic. Dark splats are ~60% more needle-shaped than bright ones.
The decisive evidence is elsewhere. v2 had 41 unregistered frames, and they were contiguous: f0001–f0041, the opening pan across the black door. Mean luminance of that block was about 24 / 255; from f0060 onward it was about 91 / 255. Nearly 4× darker. COLMAP found no reliable features there and dropped the whole segment — the door is not in the model at all. What you see is not a bad reconstruction of the door. It is Gaussians guessing at a place with no data.
Why dark regions are structurally hard:
- H.264 deliberately starves them of bits. Encoders know the eye is insensitive there. 45 Mbps does not change this.
- Sensor noise beats signal. SIFT latches onto noise, not structure. Shooting the same dark corner ten times gives you ten different noise patterns, so repetition does not help.
- The loss is an RGB difference. Near-zero pixels produce near-zero gradients, so those Gaussians are barely optimized and keep their large, elongated initial shape.
- Darkness is genuinely ambiguous. "Dark opaque surface" and "empty space in front of a black background" render identically. The optimizer cannot tell them apart and hedges with semi-transparent dark Gaussians.
prune.mjs has a brightness-aware rule (splats below a luminance threshold get a stricter needle limit). It removed 79% of the dark splats and the picture barely improved — because the problem is missing data, not excess noise.
- Lock exposure. This is the big one. iPhone auto-exposure keeps adjusting as you sweep between bright and dark, so the same surface has different brightness in different frames. That directly violates the photometric consistency 3DGS assumes, and the optimizer reconciles the contradiction with semi-transparent haze. Long-press to lock AE/AF before you start recording.
- After locking, deliberately overexpose. Push the slider up, blow out the highlights, lift the shadows. The video will look bad and the reconstruction will be better. You want information, not a nice-looking clip.
- Add light. Turn everything on, open the door, point a second phone's torch into the dark corner.
- Do not open on the darkest thing in the room. That segment will probably be discarded in full.
An untested fourth option is in this repo: train on gamma-lifted frames (eq=gamma=), then undo the lift on the spherical-harmonic DC term afterwards with degamma.mjs, keeping the improved geometry and the original colors. The delivered scan is still the plain v3 output; this variant has not been shipped.
Orbit away from the capture path and you get translucent brush-stroke smears. Two causes, measured on v1:
| Metric | p50 | p90 | p99 |
|---|---|---|---|
| Opacity | 0.057 | 0.241 | 0.643 |
| Needle s0/s1 | 4.96 | 39.96 | 1,701.74 |
| Disc s1/s2 | 3.72 | 69.47 | 3,331.56 |
Half the splats have a longest axis 5× their middle axis; the 99th percentile is 1,700×. Those thin needles look correct from the training view and become streaks from anywhere else. Separately, a median opacity of 0.057 means a very large number of near-invisible splats stacking into haze.
Do not use max/min extent as a single anisotropy score. Sort the three scales s0 ≥ s1 ≥ s2 and read them as two different shapes: a large s0/s1 is a needle and should go, a large s1/s2 is a disc, which is the normal shape of a flat surface. Cut discs and your walls develop holes. My first version conflated them and wanted to delete 77% of the model.
node prune.mjs splats/export_15000.ply splats/pruned_mid.ply 4 0.05 10| Preset | Needle | Opacity | Distance | Kept | Size | Result |
|---|---|---|---|---|---|---|
| soft | <6 | >0.03 | <12 | 42.8% | 71 MB | still hazy |
| mid | <4 | >0.05 | <10 | 26.3% | 44 MB | best |
| hard | <3 | >0.08 | <10 | 15.3% | 25 MB | worse |
More pruning is not better. The hard preset strips the low-opacity splats that make up soft surfaces; the sofa collapses into a dark smear. mid is the sweet spot.
One free improvement that beats any pruning: constrain the viewer camera to stay near the original capture path. Splats are only correct near the views they were trained on.
The wardrobe mirror put a recognizable person into the scan. It also produced geometry that does not exist: photogrammetry treats a mirror as "another room behind the wall," and with a reflective surface opposite, the reflection bounces. The depth histogram had three peaks (roughly 4 m / 8.5 m / 13 m) and fake geometry extended to 15 m out — world-space X reached 19.79 in a room that is 4.09 across.
demirror.mjs handles it by projection, not by bounding box:
node demirror.mjs in.ply out.ply mirror-cam.json 2.5Take the mirror's screen-space rectangle in one frame, extend it into the scene as a frustum, and delete every splat inside that frustum beyond a depth threshold. It removed 56,370 splats (7.8%).
A world-space bounding box does not work here. The first reflection lands inside the real room's extent, so no box can separate it from real furniture. Verified by re-rendering from the exact camera that had captured the person: they are gone.
If you are scanning a room with a mirror and plan to publish the result, this step is not optional.
- RTX 5070 12 GB, compute capability 12.0 (sm_120), driver 610.62
- COLMAP 4.1.1, official prebuilt Windows CUDA binary
- Brush v0.3.0, single Windows binary, Apache-2.0
sm_120 is why most trainers were unusable:
- OpenSplat — the documented Windows recipe is CUDA 11.8 + libtorch 2.1.2, a 2023 combination. Does not build.
- nerfstudio — effectively dead. Last release 2024-11, last commit 2025-07-29.
- Original 3DGS (Inria) — Windows CUDA build pain, plus a non-commercial license.
Brush is Rust + WGPU. It never touches the CUDA toolkit, so a new architecture is not its problem, and it is one binary with no Python environment.
# --max-splats is mandatory; the 10M default will exhaust 12 GB
brush_app.exe dataset --total-steps 15000 --max-splats 2000000 \
--export-path out --export-every 5000Measured peak VRAM 4.6 / 12.2 GB, GPU 98%, 67 °C.
COLMAP 4.1.1 renamed options. --SiftExtraction.use_gpu → --FeatureExtraction.use_gpu, --SiftExtraction.max_image_size → --FeatureExtraction.max_image_size, --SiftMatching.use_gpu → --FeatureMatching.use_gpu. --SiftExtraction.max_num_features still exists.
ffmpeg's select expression has a length limit. 402 chained eq(n,X) terms (4,236 characters) fails at filter-graph init with a misleading message:
[AVFilterGraph] Error initializing filters
Error opening output files: Cannot allocate memory
Split into batches of 60 and use an incrementing -start_number.
PLY field order. Brush writes a standard 3DGS PLY, 59 float32 per splat: f_dc_0..2 (3) + f_rest_0..44 (45, SH degree 3) + opacity (1) + rot_0..3 (4) + scale_0..2 (3) + x,y,z (3). Check: 735407 × 59 × 4 + 1551 header = 173,557,603 bytes, which matches the file exactly. Note that x,y,z are fields 56–58, at the end, not the beginning. Any converter that assumes position comes first will silently produce garbage.
Frame selection. Best-of-window: take the sharpest frame from every group of 4, using Laplacian edge energy. Keeps sharpness without sacrificing temporal coverage.
Do not average the camera forward vectors to pick a default view. Walking around produces many downward-tilted frames, so the mean points at the floor and the viewer opens looking at the bed. Instead pick the real pose whose forward vector is most perpendicular to the scene up axis (v2: f0096, tilt 0.086). That guarantees the default view sits somewhere with actual data, and sidesteps roll conversion entirely.
Viewer, with Spark:
SplatMeshdoes not draw itself. You must also add aSparkRendererto the scene. Without it the canvas is black with no error at all; the only clue isrenderer.info.render.calls === 0.- Derive
camera.upfrom the COLMAP poses, not from convention. For v1 the room's up was(0.023, −0.993, −0.120)— that is −Y. Use the camera path centroid, not the point cloud centroid, as the orbit target; points outside the window drag the cloud centroid far off (I got(1.58, 0.11, 3.26)instead of(−0.29, −0.01, 0.37)). - When the browser panel is hidden,
innerWidthreads 0, the canvas becomes 0×0 and never recovers. Compare sizes every frame and self-heal; do not trust theresizeevent. - Do not use mkkellogg's three.js splat viewer — unmaintained since 2025-01.
Delivery. Convert to SOG with @playcanvas/splat-transform; keep the PLY as the master. A 165 MB PLY will not go onto Cloudflare Pages (25 MiB per-file limit); SOG is sharded, so it slips under. On iPhone, disable MSAA and clamp devicePixelRatio to 2 or you will see 30–40 fps.
bash pipeline.sh <video> <workdir> 6 20000 3000000pipeline.sh picks the matcher by image count, applies the intrinsics prior, and prints any disconnected components so you know whether the model fragmented.
pipeline.sh frame selection → COLMAP → Brush, end to end
prune.mjs needle/disc/opacity/distance pruning, brightness-aware
demirror.mjs delete mirror-image geometry by camera frustum
degamma.mjs undo a training-time gamma lift on the SH DC term
web/ minimal three.js + Spark viewer, `node web/server.mjs`
Splats, COLMAP databases, extracted frames and source video are not in this repo — about 3 GB, and the room is my bedroom.
- No metric scale. COLMAP units are arbitrary. To fix it, measure any identifiable object with a tape, measure the same span in the model, and divide. The iPhone 16 (non-Pro) has no LiDAR, so scale cannot be captured directly.
- Dark residue persists, even in v3. It is inherent to photogrammetry, not a parameter you can tune.
- Screens ruin their surroundings. Anything self-illuminated with changing content is noise. Turn monitors off; I did not, three times.
- Pose alignment is unresolved. Placing the viewer camera at the COLMAP pose of
f0000.jpgand comparing against the source frame gives SSIM 0.22–0.33 and PSNR ~8 dB across all four roll options, which means the conversion is wrong in more than just roll. It does not affect the viewer, but any "overlay the video on the splat" effect would need this solved first.
This was built by one person working with an AI coding assistant (Claude, via Claude Code). A repo that reads like a solo engineering log should say who did what, so:
Written by the AI. Every script here — pipeline.sh, prune.mjs, demirror.mjs, degamma.mjs, the web/ viewer — plus this README and the Chinese log it was translated from. Also the technical diagnoses: that a wrong focal prior was what broke COLMAP, that the matcher should be chosen by image count, that needles and discs have to be scored separately, the frustum method for cutting mirror geometry, and the luminance analysis of the 41 dropped frames. Where this README says "I" inside a debugging story, read it as that collaboration.
Mine. The parts that decided how it went:
- The premise. I wanted to record the things around me with a phone and get 3D back, and went looking for a way.
- v1's source. Instead of shooting anything, use the opening video already running on my personal site. That is the only reason the first result took 20 minutes instead of a weekend.
- Spotting the failure. I described the v1 output as ink-wash smearing. Nobody measures needle ratio until there is a complaint that needs explaining — that observation is upstream of every number in this repo.
- Asking which single fix had the highest payoff, rather than running every option and seeing what stuck.
- Rejecting v2. "Better, but the shadows still smear." v3 exists because I did not accept v2.
- All three videos are mine, shot on my phone. v3 is 80 seconds because I move footage through Google Drive and long files take too long to upload.
- The hardware, the runs, and the room.
Every number here comes from an actual run on this machine. None of them are estimated, or carried over from a paper.