Skip to content

Repository files navigation

Video to Gaussian World

Reconstruct a persistent, transient-aware 3D Gaussian world from one short video.

视频重建三维高斯世界

Python PyTorch CUDA 3DGS Tests License Stars


Video to Gaussian World web demo result

Important

This repository is a research extension of Lukas Höllein and Matthias Nießner's video_to_world project and the paper World Reconstruction From Inconsistent Views. The upstream authorship and MIT license are preserved. The extensions in this repository focus on a deployable 3DGS path, transient-aware refinement, evaluation, and an upload-to-world demo.

✦ What this project does

Video to Gaussian World turns a 5-20 second monocular camera-motion video into one explicit 3D Gaussian scene:

video → depth & cameras → non-rigid alignment → canonical world
      → 3DGS-MCMC → transient refinement → MP4 + 3DGS PLY

It is designed for views that are not perfectly consistent, including videos produced by generative video models and casually captured real-world footage. Instead of assuming that every frame depicts exactly the same geometry and appearance, the system separates:

  • persistent structure that belongs to the canonical 3D world; and
  • transient appearance caused by lighting changes, occlusion, dynamic content, or generative flicker.

The result can be rendered along a camera trajectory, downloaded as a standard 3DGS PLY, and inspected in tools such as SuperSplat.

🧭 System architecture

Video to Gaussian World system architecture

Stage Component Output
Lift Depth Anything 3 depth, confidence, intrinsics, extrinsics, RGB
Align RoMa + confidence-aware non-rigid ICP canonical point cloud and per-frame deformation
Refine global optimization sharper and more coherent canonical geometry
Invert inverse deformation network canonical-to-frame warp for differentiable rendering
Optimize budgeted 3DGS-MCMC persistent 3D Gaussian topology, capped at 1.5M primitives
Factor frozen-core rank-8 transient branch time-conditioned color and opacity residuals
Export H.264 + PLY exporters rendered MP4, standard 3DGS PLY, neural checkpoint

✨ Interactive Gaussian result

The current showcase export contains 1,500,000 Gaussians and is approximately 338 MB as a standard PLY. It has been verified in SuperSplat:

1.5M Gaussian scene opened in SuperSplat

The reconstruction is a 3D Gaussian scene, not a triangle mesh. Unobserved regions may remain empty, and quality falls when the viewer moves far beyond the training camera distribution.

🧠 Main research extension

Frozen-core static/transient factorization

After a 10,000-iteration 3DGS-MCMC reconstruction, the system freezes canonical positions, scales, rotations, camera parameters, and static spherical harmonics. A 2,000-iteration rank-8 branch then learns sparse, frame-conditioned residuals over SH-DC color and opacity.

For Gaussian i at time t:

color(i,t)   = color_static(i) + gate(i) · color_basis(i) · code(t)
opacity(i,t) = opacity_static(i) + gate(i) · opacity_basis(i) · code(t)

Sparsity and temporal smoothness regularization prevent this branch from replacing the persistent scene. Time codes for appearance-held-out frames are interpolated from neighboring supervised timestamps.

Implemented research controls

  • ✅ Budgeted default, AbsGrad, and MCMC density strategies
  • ✅ Confidence-weighted ICP and independent GS confidence ablations
  • ✅ PSNR, SSIM, LPIPS, and flow-guided temporal diagnostics
  • ✅ Classic/antialiased rasterization and Mip-Splatting-style 3D filtering
  • ✅ DSSIM, multiscale L1, edge, dynamic-mask, and appearance controls
  • ✅ local_se3, polar, partial-Jacobian, and full-Jacobian covariance transport
  • ✅ Standard 3DGS PLY export and external-viewer validation
  • 🧪 Reliability-gated topology growth remains an active research direction
  • 🧪 Continuous position/rotation dynamics remain future work

📊 Quantitative results

Results below use the locked Voyager 2 / 24 frames / 768 × 512 single-scene protocol. These are reproducible preliminary results, not a multi-dataset SOTA claim.

Protocol Model PSNR ↑ SSIM ↑ LPIPS ↓ Gaussians
All-frame input fitting 3DGS-MCMC, 10k 27.3264 0.8993 0.0391 1,500,000
All-frame input fitting + rank-8 transient, 12k 28.8959 0.9206 0.0308 1,500,000
Stage-3 appearance holdout 3DGS-MCMC, 3k 23.4193 0.7873 0.1880 1,016,200
Stage-3 appearance holdout + interpolated transient, 4k 24.0226 0.8001 0.1806 1,016,200

Key observations:

  • +1.57 dB PSNR on all-frame fitting with no increase in Gaussian count.
  • +0.60 dB PSNR on Stage-3 appearance-held-out frames.
  • The holdout still uses all frames in upstream geometry and inverse-deformation stages; it is not a fully unseen-camera benchmark.
  • The upload shown in the demo reports 31.09 dB / 0.9355 / 0.0225 under its task-local fitting protocol and should not be compared directly with the locked table above.

Full notes are available in:

🚀 Quick start

1. Create the environment

git clone https://github.com/BoPythonAI/video-to-gaussian-world.git
cd video-to-gaussian-world

conda create -n video_to_world python=3.10
conda activate video_to_world
pip install "numpy<2" "opencv-python<4.12"

2. Install Depth Anything 3

mkdir -p third_party
git clone https://github.com/ByteDance-Seed/depth-anything-3 third_party/depth-anything-3
git -C third_party/depth-anything-3 checkout 2c21ea849ceec7b469a3e62ea0c0e270afc3281a
pip install "torch>=2" torchvision xformers
pip install -e third_party/depth-anything-3
git -C third_party/depth-anything-3 apply ../../patches/da3-export-trajectory.patch

3. Install the 3D stack

pip install --no-build-isolation \
  "git+https://github.com/nerfstudio-project/gsplat.git@v1.5.3"

pip install setuptools==81.0.0
pip install "git+https://github.com/NVlabs/tiny-cuda-nn/#subdirectory=bindings/torch" \
  --no-build-isolation

pip install open3d scipy tyro tqdm tensorboard lpips viser nerfview romatch

git clone https://github.com/Parskatt/RoMaV2 third_party/RoMaV2
git -C third_party/RoMaV2 apply ../../patches/romav2-dataclasses.patch
pip install -e "third_party/RoMaV2[fused-local-corr]"

4. Reconstruct a Gaussian world

python run_reconstruction.py \
  --config.input-video /path/to/video.mp4 \
  --config.renderer 3dgs

Use --config.mode extensive for the full pipeline. Advanced options are defined in configs/ and exposed through Tyro CLI arguments.

🖥️ Upload demo

Launch the web interface against an existing evaluated scene:

python -m demo.server \
  --host 0.0.0.0 \
  --port 7860 \
  --scene-root /path/to/evaluated_scene \
  --workspace /path/to/demo_workspace

Open http://127.0.0.1:7860. For a remote GPU host:

ssh -L 7860:127.0.0.1:7860 -p <SSH_PORT> <USER>@<HOST>

Every public upload runs the strongest showcase recipe:

  • 24 temporally distributed frames
  • 60 non-rigid ICP iterations
  • 30 inverse-deformation epochs
  • 10,000 3DGS-MCMC iterations
  • 2,000 frozen-core rank-8 transient iterations
  • maximum 1.5M Gaussians

Only one GPU job is admitted at a time. The embedded viewer loads GaussianSplats3D from jsDelivr, so the browser needs internet access even though reconstruction runs locally.

🎥 Recommended input

Property Recommendation
Duration 5-20 seconds
Resolution 720p or 1080p
Camera motion smooth translation around a mostly static subject
Formats MP4, MOV, WEBM, MKV
Maximum size 256 MB in the upload demo
Avoid cuts, zooms, heavy blur, exposure jumps, large moving foreground objects

Good parallax matters more than pure in-place rotation. Unobserved surfaces cannot be reconstructed reliably from a single video.

📦 Outputs

<scene_root>/
├── exports/npz/results.npz                 # depth, camera, confidence, RGB
├── frame_to_model_icp_*/
│   ├── after_non_rigid_icp/                # aligned canonical geometry
│   ├── after_global_optimization/          # globally refined geometry
│   ├── inverse_deformation/                # canonical-to-frame warp
│   ├── gs_3dgs/                            # persistent 3DGS-MCMC core
│   └── gs_3dgs_transient/
│       ├── model_final.pt                  # full neural checkpoint
│       ├── splats_3dgs.ply                 # canonical 3D Gaussian scene
│       └── gs_video_eval/render_gs_video.mp4
└── frames_subsampled/

The exported PLY contains the persistent canonical scene. Time-dependent transient parameters remain in the PyTorch checkpoint.

🧪 Testing

Run the CPU-side test suite without loading unrelated global pytest plugins:

PYTEST_DISABLE_PLUGIN_AUTOLOAD=1 python -m pytest -q

Current local verification: 65 passed.

🗺️ Repository map

Path Purpose
run_reconstruction.py end-to-end pipeline orchestration
preprocess_video.py frame sampling and DA3 preprocessing
frame_to_model_icp.py confidence-aware non-rigid alignment
global_optimization.py joint deformation refinement
train_inverse_deformation.py inverse warp training
train_gs.py 2DGS/3DGS training and transient continuation
models/canonical_gs_model.py Gaussian, appearance, transient, and deformation model
utils/density_control.py adaptive 3DGS topology strategies
utils/covariance_transport.py deformation-aware Gaussian shape transport
utils/eval_metrics.py PSNR, SSIM, LPIPS, and temporal diagnostics
demo/ upload server and presentation UI
scripts/ reproducible experiment launchers
tests/ CPU-side regression suite

⚠️ Current limitations

  • Results are strongest near the observed camera trajectory; this is not closed 360° reconstruction.
  • The promoted transient branch changes color and opacity, not Gaussian position or rotation.
  • The PLY is a Gaussian representation rather than a collision-ready mesh.
  • Current quantitative evidence is primarily from one 24-frame scene.
  • Multi-scene, multi-seed, strict held-out-camera evaluation remains required before paper-level claims.
  • A 1.5M-Gaussian PLY is large; compression, quantization, and level-of-detail streaming are future work.

🔭 Research roadmap

  1. Multi-scene evaluation — 5-10 scenes, three seeds, mean ± standard deviation.
  2. Reliable growth — use depth, visibility, motion, residual, and feature consistency to gate topology updates.
  3. Jacobian-consistent shape transport — compare polar, partial-stretch, and full covariance transport.
  4. Continuous dynamic Gaussians — learn time-continuous position/rotation residuals with local rigidity.
  5. Deployment — compressed splat export, stable browser streaming, shareable hosted examples.
  6. Physical AI bridge — mesh/TSDF extraction, semantics, collision geometry, and downstream robot interfaces.

🙏 Acknowledgements and citation

This project builds on:

If you use the upstream method, please cite the original authors:

@misc{hoellein2026worldreconstructioninconsistentviews,
  title         = {World Reconstruction From Inconsistent Views},
  author        = {Lukas H{\"o}llein and Matthias Nie{\ss}ner},
  year          = {2026},
  eprint        = {2603.16736},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2603.16736}
}

License

Released under the MIT License. The original copyright notice is retained.

Releases

Packages

Contributors

Languages