Important
This repository is a research extension of Lukas Höllein and Matthias Nießner's video_to_world project and the paper World Reconstruction From Inconsistent Views. The upstream authorship and MIT license are preserved. The extensions in this repository focus on a deployable 3DGS path, transient-aware refinement, evaluation, and an upload-to-world demo.
Video to Gaussian World turns a 5-20 second monocular camera-motion video into one explicit 3D Gaussian scene:
video → depth & cameras → non-rigid alignment → canonical world
→ 3DGS-MCMC → transient refinement → MP4 + 3DGS PLY
It is designed for views that are not perfectly consistent, including videos produced by generative video models and casually captured real-world footage. Instead of assuming that every frame depicts exactly the same geometry and appearance, the system separates:
- persistent structure that belongs to the canonical 3D world; and
- transient appearance caused by lighting changes, occlusion, dynamic content, or generative flicker.
The result can be rendered along a camera trajectory, downloaded as a standard 3DGS PLY, and inspected in tools such as SuperSplat.
| Stage | Component | Output |
|---|---|---|
| Lift | Depth Anything 3 | depth, confidence, intrinsics, extrinsics, RGB |
| Align | RoMa + confidence-aware non-rigid ICP | canonical point cloud and per-frame deformation |
| Refine | global optimization | sharper and more coherent canonical geometry |
| Invert | inverse deformation network | canonical-to-frame warp for differentiable rendering |
| Optimize | budgeted 3DGS-MCMC | persistent 3D Gaussian topology, capped at 1.5M primitives |
| Factor | frozen-core rank-8 transient branch | time-conditioned color and opacity residuals |
| Export | H.264 + PLY exporters | rendered MP4, standard 3DGS PLY, neural checkpoint |
The current showcase export contains 1,500,000 Gaussians and is approximately 338 MB as a standard PLY. It has been verified in SuperSplat:
The reconstruction is a 3D Gaussian scene, not a triangle mesh. Unobserved regions may remain empty, and quality falls when the viewer moves far beyond the training camera distribution.
After a 10,000-iteration 3DGS-MCMC reconstruction, the system freezes canonical positions, scales, rotations, camera parameters, and static spherical harmonics. A 2,000-iteration rank-8 branch then learns sparse, frame-conditioned residuals over SH-DC color and opacity.
For Gaussian i at time t:
color(i,t) = color_static(i) + gate(i) · color_basis(i) · code(t)
opacity(i,t) = opacity_static(i) + gate(i) · opacity_basis(i) · code(t)
Sparsity and temporal smoothness regularization prevent this branch from replacing the persistent scene. Time codes for appearance-held-out frames are interpolated from neighboring supervised timestamps.
- ✅ Budgeted
default,AbsGrad, andMCMCdensity strategies - ✅ Confidence-weighted ICP and independent GS confidence ablations
- ✅ PSNR, SSIM, LPIPS, and flow-guided temporal diagnostics
- ✅ Classic/antialiased rasterization and Mip-Splatting-style 3D filtering
- ✅ DSSIM, multiscale L1, edge, dynamic-mask, and appearance controls
- ✅
local_se3, polar, partial-Jacobian, and full-Jacobian covariance transport - ✅ Standard 3DGS PLY export and external-viewer validation
- 🧪 Reliability-gated topology growth remains an active research direction
- 🧪 Continuous position/rotation dynamics remain future work
Results below use the locked Voyager 2 / 24 frames / 768 × 512 single-scene protocol. These are reproducible preliminary results, not a multi-dataset SOTA claim.
| Protocol | Model | PSNR ↑ | SSIM ↑ | LPIPS ↓ | Gaussians |
|---|---|---|---|---|---|
| All-frame input fitting | 3DGS-MCMC, 10k | 27.3264 | 0.8993 | 0.0391 | 1,500,000 |
| All-frame input fitting | + rank-8 transient, 12k | 28.8959 | 0.9206 | 0.0308 | 1,500,000 |
| Stage-3 appearance holdout | 3DGS-MCMC, 3k | 23.4193 | 0.7873 | 0.1880 | 1,016,200 |
| Stage-3 appearance holdout | + interpolated transient, 4k | 24.0226 | 0.8001 | 0.1806 | 1,016,200 |
Key observations:
- +1.57 dB PSNR on all-frame fitting with no increase in Gaussian count.
- +0.60 dB PSNR on Stage-3 appearance-held-out frames.
- The holdout still uses all frames in upstream geometry and inverse-deformation stages; it is not a fully unseen-camera benchmark.
- The upload shown in the demo reports 31.09 dB / 0.9355 / 0.0225 under its task-local fitting protocol and should not be compared directly with the locked table above.
Full notes are available in:
experiments/voyager2_hq24_baseline.mdexperiments/phase4_transient_results.mdexperiments/deep_research_3dgs_quality_2026-09-03.md
git clone https://github.com/BoPythonAI/video-to-gaussian-world.git
cd video-to-gaussian-world
conda create -n video_to_world python=3.10
conda activate video_to_world
pip install "numpy<2" "opencv-python<4.12"mkdir -p third_party
git clone https://github.com/ByteDance-Seed/depth-anything-3 third_party/depth-anything-3
git -C third_party/depth-anything-3 checkout 2c21ea849ceec7b469a3e62ea0c0e270afc3281a
pip install "torch>=2" torchvision xformers
pip install -e third_party/depth-anything-3
git -C third_party/depth-anything-3 apply ../../patches/da3-export-trajectory.patchpip install --no-build-isolation \
"git+https://github.com/nerfstudio-project/gsplat.git@v1.5.3"
pip install setuptools==81.0.0
pip install "git+https://github.com/NVlabs/tiny-cuda-nn/#subdirectory=bindings/torch" \
--no-build-isolation
pip install open3d scipy tyro tqdm tensorboard lpips viser nerfview romatch
git clone https://github.com/Parskatt/RoMaV2 third_party/RoMaV2
git -C third_party/RoMaV2 apply ../../patches/romav2-dataclasses.patch
pip install -e "third_party/RoMaV2[fused-local-corr]"python run_reconstruction.py \
--config.input-video /path/to/video.mp4 \
--config.renderer 3dgsUse --config.mode extensive for the full pipeline. Advanced options are defined in configs/ and exposed through Tyro CLI arguments.
Launch the web interface against an existing evaluated scene:
python -m demo.server \
--host 0.0.0.0 \
--port 7860 \
--scene-root /path/to/evaluated_scene \
--workspace /path/to/demo_workspaceOpen http://127.0.0.1:7860. For a remote GPU host:
ssh -L 7860:127.0.0.1:7860 -p <SSH_PORT> <USER>@<HOST>Every public upload runs the strongest showcase recipe:
- 24 temporally distributed frames
- 60 non-rigid ICP iterations
- 30 inverse-deformation epochs
- 10,000 3DGS-MCMC iterations
- 2,000 frozen-core rank-8 transient iterations
- maximum 1.5M Gaussians
Only one GPU job is admitted at a time. The embedded viewer loads GaussianSplats3D from jsDelivr, so the browser needs internet access even though reconstruction runs locally.
| Property | Recommendation |
|---|---|
| Duration | 5-20 seconds |
| Resolution | 720p or 1080p |
| Camera motion | smooth translation around a mostly static subject |
| Formats | MP4, MOV, WEBM, MKV |
| Maximum size | 256 MB in the upload demo |
| Avoid | cuts, zooms, heavy blur, exposure jumps, large moving foreground objects |
Good parallax matters more than pure in-place rotation. Unobserved surfaces cannot be reconstructed reliably from a single video.
<scene_root>/
├── exports/npz/results.npz # depth, camera, confidence, RGB
├── frame_to_model_icp_*/
│ ├── after_non_rigid_icp/ # aligned canonical geometry
│ ├── after_global_optimization/ # globally refined geometry
│ ├── inverse_deformation/ # canonical-to-frame warp
│ ├── gs_3dgs/ # persistent 3DGS-MCMC core
│ └── gs_3dgs_transient/
│ ├── model_final.pt # full neural checkpoint
│ ├── splats_3dgs.ply # canonical 3D Gaussian scene
│ └── gs_video_eval/render_gs_video.mp4
└── frames_subsampled/
The exported PLY contains the persistent canonical scene. Time-dependent transient parameters remain in the PyTorch checkpoint.
Run the CPU-side test suite without loading unrelated global pytest plugins:
PYTEST_DISABLE_PLUGIN_AUTOLOAD=1 python -m pytest -qCurrent local verification: 65 passed.
| Path | Purpose |
|---|---|
run_reconstruction.py |
end-to-end pipeline orchestration |
preprocess_video.py |
frame sampling and DA3 preprocessing |
frame_to_model_icp.py |
confidence-aware non-rigid alignment |
global_optimization.py |
joint deformation refinement |
train_inverse_deformation.py |
inverse warp training |
train_gs.py |
2DGS/3DGS training and transient continuation |
models/canonical_gs_model.py |
Gaussian, appearance, transient, and deformation model |
utils/density_control.py |
adaptive 3DGS topology strategies |
utils/covariance_transport.py |
deformation-aware Gaussian shape transport |
utils/eval_metrics.py |
PSNR, SSIM, LPIPS, and temporal diagnostics |
demo/ |
upload server and presentation UI |
scripts/ |
reproducible experiment launchers |
tests/ |
CPU-side regression suite |
- Results are strongest near the observed camera trajectory; this is not closed 360° reconstruction.
- The promoted transient branch changes color and opacity, not Gaussian position or rotation.
- The PLY is a Gaussian representation rather than a collision-ready mesh.
- Current quantitative evidence is primarily from one 24-frame scene.
- Multi-scene, multi-seed, strict held-out-camera evaluation remains required before paper-level claims.
- A 1.5M-Gaussian PLY is large; compression, quantization, and level-of-detail streaming are future work.
- Multi-scene evaluation — 5-10 scenes, three seeds, mean ± standard deviation.
- Reliable growth — use depth, visibility, motion, residual, and feature consistency to gate topology updates.
- Jacobian-consistent shape transport — compare polar, partial-stretch, and full covariance transport.
- Continuous dynamic Gaussians — learn time-continuous position/rotation residuals with local rigidity.
- Deployment — compressed splat export, stable browser streaming, shareable hosted examples.
- Physical AI bridge — mesh/TSDF extraction, semantics, collision geometry, and downstream robot interfaces.
This project builds on:
If you use the upstream method, please cite the original authors:
@misc{hoellein2026worldreconstructioninconsistentviews,
title = {World Reconstruction From Inconsistent Views},
author = {Lukas H{\"o}llein and Matthias Nie{\ss}ner},
year = {2026},
eprint = {2603.16736},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2603.16736}
}Released under the MIT License. The original copyright notice is retained.

