Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Photorealistic Augmented Reality

A real-time, markerless AR pipeline that makes virtual objects look like they belong in the scene — tracking without a marker, relighting to match the room, and occluding behind real geometry — at ~49 fps on a single consumer GPU with no depth sensor.

Built in C++17 with CUDA-accelerated OpenCV, OpenGL, and ONNX Runtime. This project takes a conventional checkerboard-marker AR system and rebuilds the three subsystems that most break the illusion of a virtual object sitting in the real world:

Problem Baseline This project
Tracking Checkerboard must stay fully visible every frame Markerless: ORB keyframe anchoring + Lucas–Kanade optical flow
Lighting Single hand-tuned point light Per-frame spherical-harmonic irradiance estimated from the camera image
Occlusion Virtual object always drawn on top Per-fragment depth test against monocular MiDaS depth

All three run together and still sustain 20.2 ms/frame (49.5 fps) — within 3 ms of the marker-only baseline.

The same virtual backpack relit by the pipeline under four real lighting conditions

The same virtual backpack, composited live into the camera feed and relit by the spherical-harmonic estimator under four different room lighting conditions — no manual tuning between shots.

📄 Full technical writeup: paper.pdf  ·  🎦 Narrated presentation (13 min walkthrough of the design and results)


Why it's interesting

Perceptual coherence in AR depends on solving three interdependent problems at once — get any one wrong and the illusion breaks. Prior work tends to address them in isolation. This is an integrated pipeline where a shared per-frame MiDaS depth pass feeds both the occlusion renderer and the tracker's bootstrap, and where each subsystem was chosen for its quality-to-cost ratio so the three can coexist in a real-time budget:

  • Binary ORB over SIFT/SURF for keyframe features
  • Nine-coefficient spherical harmonics over a full environment probe for lighting
  • MiDaS Small over larger depth variants

The result is photorealistic markerless AR on commodity hardware with a single camera — no LiDAR, no structured light, no time-of-flight.

Results

Tracking robustness

Evaluated across three motion conditions against the checkerboard baseline. The relevant metric for coherence is the worst-case pose-loss burst (how long the object can vanish), not just the average:

Tracking comparison

  • When the camera pans the board out of view, the baseline produces no pose for up to 297 consecutive frames (~5 s); the markerless tracker bounds its worst case to 39 frames.
  • When the board itself moves — a condition the baseline cannot handle by construction — the baseline loses 30.5% of frames; the markerless system holds pose on 100% of frames with no loss bursts.

Lighting estimation

Spherical-harmonic irradiance recovers the scene's dominant chromaticity and intensity per frame with no manual tuning, where the prior point light was correct only in the one configuration it was tuned for (see the four-panel comparison above). Estimator cost: < 1 ms/frame.

Depth occlusion

MiDaS relative depth is fit to metric scale from the tracker's anchor cloud each frame. The fit tracks real geometry — estimated origin depth scales monotonically with physical distance — and is cleanest (≈9% median residual) when the checkerboard bootstrap gives the anchor cloud real spread in 1/Z:

Depth occlusion characterization across distance and board orientation

End-to-end performance

Per-stage timing across configurations. SH lighting is essentially free; depth occlusion (MiDaS inference) is the only meaningful new cost at ~8 ms:

Performance breakdown

With all three components active: 20.2 ms mean, 24.5 ms p95, sustaining 49.5 fps.


Architecture

The pipeline is three loosely-coupled components sharing a common per-frame state:

camera frame ─┬─► MarkerlessTracker ──► camera pose + 3D↔2D anchors ─┐
              │       (ORB keyframes + pyramidal LK optical flow)     │
              ├─► ShLighting ─────────► 9 SH irradiance coefficients ─┤
              │       (per-pixel sample @ stride 8, EMA smoothed)     ├─► ArRenderer (OpenGL)
              └─► DepthEstimator ─────► metric-fit inverse depth map ─┘      Lambertian shade
                      (MiDaS v2.1 Small via ONNX Runtime + CUDA)             + per-fragment
                                                                             depth occlusion
  • Markerless tracking (markerless_tracker.cpp) — Uses the checkerboard once at startup as a bootstrap coordinate frame, then tracks a map of natural ORB features. Anchor frames match descriptors with a Hamming brute-force matcher + Lowe's ratio test and solve pose via solvePnPRansac; intermediate frames track the surviving inliers with CUDA pyramidal Lucas–Kanade (4/5 frames take this cheap path). Pose is smoothed with an exponential moving average.
  • Spherical-harmonic lighting (sh_lighting.cpp) — Samples the frame at an 8-pixel stride, linearizes sRGB, reconstructs each sample's camera-space ray, and accumulates the 9 SH irradiance coefficients (Ramamoorthi & Hanrahan). Coefficients are uploaded as a vec3[9] uniform and evaluated per-fragment for Lambertian shading.
  • Depth occlusion (depth_estimator.cpp, midas_fit.cpp) — Runs MiDaS Small through ONNX Runtime on the CUDA provider, fits its relative inverse depth to metric scale using the tracker's anchor cloud (closed-form 1/Z regression), and uploads the map as an R32F texture. The fragment shader discards virtual fragments that fall behind real geometry.

Most per-frame work is on the GPU: OpenCV's CUDA modules handle ORB, matching, and optical flow; ONNX Runtime runs MiDaS on CUDA; lighting and occlusion are evaluated in GLSL alongside rasterization.

Code layout

src/            include/ar/
├── main.cpp            app dispatch (calibrate / ar / orb modes)
├── app.cpp             per-frame loop, wiring the three components
├── markerless_tracker  ORB keyframe + LK optical-flow tracking
├── tracker             checkerboard baseline tracker
├── sh_lighting         spherical-harmonic irradiance estimation
├── depth_estimator     MiDaS inference (ONNX Runtime + CUDA)
├── midas_fit           relative → metric depth fitting
├── ar_renderer         OpenGL model rendering, SH shading, occlusion
├── background_quad     camera-feed background
├── calibration         checkerboard intrinsics calibration
├── cv_gl_math          OpenCV ↔ OpenGL matrix conversion
├── metrics_logger      per-frame CSV (timings + drift) for evaluation
└── frame_stats         timing instrumentation

data/
├── shaders/            GLSL (model + background)
├── models/backpack/    textured OBJ test model
├── models/midas/       bundled MiDaS Small depth model (ONNX)
├── results/            evaluation CSVs (source data for figures)
└── figures/            plots used above
experiments/
└── make_figures.py     regenerates figures from data/results/
setup.sh                one-shot dependency install + build
scripts/
└── install_opencv_cuda.sh   builds OpenCV 4.10 with CUDA

Quickstart

This is a live, GPU-accelerated AR application, so it needs real hardware: an NVIDIA GPU + CUDA toolkit, a webcam, and a display. It targets Ubuntu 22.04/24.04 (Linux is where CUDA-enabled OpenCV lives). The MiDaS depth model is already bundled in data/, so there's nothing to download.

git clone https://github.com/chopndolphy/photorealistic-ar.git
cd photorealistic-ar
./setup.sh --build-opencv      # installs all deps + builds the app (see notes below)

Then run it:

./build/ar_project calibrate                     # capture checkerboard views  (S save · C compute)
./build/ar_project ar --markerless --sh --depth  # the full pipeline

That's it. Each feature is an independent flag — run ar with any subset (--markerless, --sh, --depth) to A/B a component against the marker-based baseline. Add --metrics run.csv to log per-frame timings and drift. ar_project orb visualizes GPU ORB keypoints; run with no arguments for full usage.

In-app keys: Q quit · R reset tracker/occlusion fit · D toggle occlusion debug · A/U snapshot/clear intended pose for drift measurement.

What setup.sh does

The one script handles everything (setup.sh, scripts/install_opencv_cuda.sh):

  1. apt installs the system libraries — GLFW, Assimp, GLM, OpenGL, build tools.
  2. Downloads and installs ONNX Runtime (GPU) to /usr/local/onnxruntime.
  3. Builds OpenCV 4.10 with CUDA from source, auto-detecting your GPU's compute capability (--build-opencv; ~30–60 min, the long pole — omit the flag if you already have a CUDA-enabled OpenCV 4.10).
  4. Configures and builds the project with CMake.

Re-run ./setup.sh --skip-deps to just rebuild the app after code changes. Prefer to drive CMake yourself? The manual path is unchanged:

cmake -B build -DCMAKE_BUILD_TYPE=Release && cmake --build build -j

Dependency summary: C++17 · CMake ≥ 3.18 · CUDA-capable NVIDIA GPU · OpenCV 4.10 with CUDA (core imgproc imgcodecs videoio highgui calib3d features2d dnn cuda{arithm,imgproc,features2d,warping,optflow}) · OpenGL · GLFW 3.3 · Assimp · GLM · ONNX Runtime (CUDA EP). GLAD and stb are vendored under third_party/; the MiDaS model is bundled under data/models/midas/ (override with the AR_DEPTH_MODEL env var).

Reproducing the figures

python experiments/make_figures.py   # reads data/results/*.csv → data/figures/

Limitations & future work

  • Tracking still needs the checkerboard at session startup and can drift under sustained fast motion; loop closure or inertial measurements would help.
  • SH lighting captures low-frequency diffuse illumination only — no cast shadows, specular highlights, or sharp directional light (the standard tradeoff of the SH irradiance approach; fine for Lambertian materials).
  • The occlusion boundary shows a faint silhouette halo from upsampling MiDaS's 256×256 output; a higher-resolution or learned metric depth model would sharpen it.

Credits

Test model — Survival Guitar Backpack by Berk Gedik (material assignment adapted by Joey de Vries).

Author: Christopher John Allison · Khoury College of Computer Sciences, Northeastern University.

About

Real-time photorealistic markerless AR: ORB+optical-flow tracking, spherical-harmonic lighting estimation, and MiDaS monocular depth occlusion in C++/CUDA/OpenGL (~49 fps).

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages