A real-time, markerless AR pipeline that makes virtual objects look like they belong in the scene — tracking without a marker, relighting to match the room, and occluding behind real geometry — at ~49 fps on a single consumer GPU with no depth sensor.
Built in C++17 with CUDA-accelerated OpenCV, OpenGL, and ONNX Runtime. This project takes a conventional checkerboard-marker AR system and rebuilds the three subsystems that most break the illusion of a virtual object sitting in the real world:
| Problem | Baseline | This project |
|---|---|---|
| Tracking | Checkerboard must stay fully visible every frame | Markerless: ORB keyframe anchoring + Lucas–Kanade optical flow |
| Lighting | Single hand-tuned point light | Per-frame spherical-harmonic irradiance estimated from the camera image |
| Occlusion | Virtual object always drawn on top | Per-fragment depth test against monocular MiDaS depth |
All three run together and still sustain 20.2 ms/frame (49.5 fps) — within 3 ms of the marker-only baseline.
The same virtual backpack, composited live into the camera feed and relit by the spherical-harmonic estimator under four different room lighting conditions — no manual tuning between shots.
📄 Full technical writeup:
paper.pdf· 🎦 Narrated presentation (13 min walkthrough of the design and results)
Perceptual coherence in AR depends on solving three interdependent problems at once — get any one wrong and the illusion breaks. Prior work tends to address them in isolation. This is an integrated pipeline where a shared per-frame MiDaS depth pass feeds both the occlusion renderer and the tracker's bootstrap, and where each subsystem was chosen for its quality-to-cost ratio so the three can coexist in a real-time budget:
- Binary ORB over SIFT/SURF for keyframe features
- Nine-coefficient spherical harmonics over a full environment probe for lighting
- MiDaS Small over larger depth variants
The result is photorealistic markerless AR on commodity hardware with a single camera — no LiDAR, no structured light, no time-of-flight.
Evaluated across three motion conditions against the checkerboard baseline. The relevant metric for coherence is the worst-case pose-loss burst (how long the object can vanish), not just the average:
- When the camera pans the board out of view, the baseline produces no pose for up to 297 consecutive frames (~5 s); the markerless tracker bounds its worst case to 39 frames.
- When the board itself moves — a condition the baseline cannot handle by construction — the baseline loses 30.5% of frames; the markerless system holds pose on 100% of frames with no loss bursts.
Spherical-harmonic irradiance recovers the scene's dominant chromaticity and intensity per frame with no manual tuning, where the prior point light was correct only in the one configuration it was tuned for (see the four-panel comparison above). Estimator cost: < 1 ms/frame.
MiDaS relative depth is fit to metric scale from the tracker's anchor cloud each frame. The fit tracks real geometry — estimated origin depth scales monotonically with physical distance — and is cleanest (≈9% median residual) when the checkerboard bootstrap gives the anchor cloud real spread in 1/Z:
Per-stage timing across configurations. SH lighting is essentially free; depth occlusion (MiDaS inference) is the only meaningful new cost at ~8 ms:
With all three components active: 20.2 ms mean, 24.5 ms p95, sustaining 49.5 fps.
The pipeline is three loosely-coupled components sharing a common per-frame state:
camera frame ─┬─► MarkerlessTracker ──► camera pose + 3D↔2D anchors ─┐
│ (ORB keyframes + pyramidal LK optical flow) │
├─► ShLighting ─────────► 9 SH irradiance coefficients ─┤
│ (per-pixel sample @ stride 8, EMA smoothed) ├─► ArRenderer (OpenGL)
└─► DepthEstimator ─────► metric-fit inverse depth map ─┘ Lambertian shade
(MiDaS v2.1 Small via ONNX Runtime + CUDA) + per-fragment
depth occlusion
- Markerless tracking (
markerless_tracker.cpp) — Uses the checkerboard once at startup as a bootstrap coordinate frame, then tracks a map of natural ORB features. Anchor frames match descriptors with a Hamming brute-force matcher + Lowe's ratio test and solve pose viasolvePnPRansac; intermediate frames track the surviving inliers with CUDA pyramidal Lucas–Kanade (4/5 frames take this cheap path). Pose is smoothed with an exponential moving average. - Spherical-harmonic lighting (
sh_lighting.cpp) — Samples the frame at an 8-pixel stride, linearizes sRGB, reconstructs each sample's camera-space ray, and accumulates the 9 SH irradiance coefficients (Ramamoorthi & Hanrahan). Coefficients are uploaded as avec3[9]uniform and evaluated per-fragment for Lambertian shading. - Depth occlusion (
depth_estimator.cpp,midas_fit.cpp) — Runs MiDaS Small through ONNX Runtime on the CUDA provider, fits its relative inverse depth to metric scale using the tracker's anchor cloud (closed-form1/Zregression), and uploads the map as anR32Ftexture. The fragment shader discards virtual fragments that fall behind real geometry.
Most per-frame work is on the GPU: OpenCV's CUDA modules handle ORB, matching, and optical flow; ONNX Runtime runs MiDaS on CUDA; lighting and occlusion are evaluated in GLSL alongside rasterization.
src/ include/ar/
├── main.cpp app dispatch (calibrate / ar / orb modes)
├── app.cpp per-frame loop, wiring the three components
├── markerless_tracker ORB keyframe + LK optical-flow tracking
├── tracker checkerboard baseline tracker
├── sh_lighting spherical-harmonic irradiance estimation
├── depth_estimator MiDaS inference (ONNX Runtime + CUDA)
├── midas_fit relative → metric depth fitting
├── ar_renderer OpenGL model rendering, SH shading, occlusion
├── background_quad camera-feed background
├── calibration checkerboard intrinsics calibration
├── cv_gl_math OpenCV ↔ OpenGL matrix conversion
├── metrics_logger per-frame CSV (timings + drift) for evaluation
└── frame_stats timing instrumentation
data/
├── shaders/ GLSL (model + background)
├── models/backpack/ textured OBJ test model
├── models/midas/ bundled MiDaS Small depth model (ONNX)
├── results/ evaluation CSVs (source data for figures)
└── figures/ plots used above
experiments/
└── make_figures.py regenerates figures from data/results/
setup.sh one-shot dependency install + build
scripts/
└── install_opencv_cuda.sh builds OpenCV 4.10 with CUDA
This is a live, GPU-accelerated AR application, so it needs real hardware: an NVIDIA GPU + CUDA toolkit, a webcam, and a display. It targets Ubuntu 22.04/24.04 (Linux is where CUDA-enabled OpenCV lives). The MiDaS depth model is already bundled in data/, so there's nothing to download.
git clone https://github.com/chopndolphy/photorealistic-ar.git
cd photorealistic-ar
./setup.sh --build-opencv # installs all deps + builds the app (see notes below)Then run it:
./build/ar_project calibrate # capture checkerboard views (S save · C compute)
./build/ar_project ar --markerless --sh --depth # the full pipelineThat's it. Each feature is an independent flag — run ar with any subset (--markerless, --sh, --depth) to A/B a component against the marker-based baseline. Add --metrics run.csv to log per-frame timings and drift. ar_project orb visualizes GPU ORB keypoints; run with no arguments for full usage.
In-app keys: Q quit · R reset tracker/occlusion fit · D toggle occlusion debug · A/U snapshot/clear intended pose for drift measurement.
The one script handles everything (setup.sh, scripts/install_opencv_cuda.sh):
aptinstalls the system libraries — GLFW, Assimp, GLM, OpenGL, build tools.- Downloads and installs ONNX Runtime (GPU) to
/usr/local/onnxruntime. - Builds OpenCV 4.10 with CUDA from source, auto-detecting your GPU's compute capability (
--build-opencv; ~30–60 min, the long pole — omit the flag if you already have a CUDA-enabled OpenCV 4.10). - Configures and builds the project with CMake.
Re-run ./setup.sh --skip-deps to just rebuild the app after code changes. Prefer to drive CMake yourself? The manual path is unchanged:
cmake -B build -DCMAKE_BUILD_TYPE=Release && cmake --build build -jDependency summary: C++17 · CMake ≥ 3.18 · CUDA-capable NVIDIA GPU · OpenCV 4.10 with CUDA (core imgproc imgcodecs videoio highgui calib3d features2d dnn cuda{arithm,imgproc,features2d,warping,optflow}) · OpenGL · GLFW 3.3 · Assimp · GLM · ONNX Runtime (CUDA EP). GLAD and stb are vendored under third_party/; the MiDaS model is bundled under data/models/midas/ (override with the AR_DEPTH_MODEL env var).
python experiments/make_figures.py # reads data/results/*.csv → data/figures/- Tracking still needs the checkerboard at session startup and can drift under sustained fast motion; loop closure or inertial measurements would help.
- SH lighting captures low-frequency diffuse illumination only — no cast shadows, specular highlights, or sharp directional light (the standard tradeoff of the SH irradiance approach; fine for Lambertian materials).
- The occlusion boundary shows a faint silhouette halo from upsampling MiDaS's 256×256 output; a higher-resolution or learned metric depth model would sharpen it.
Test model — Survival Guitar Backpack by Berk Gedik (material assignment adapted by Joey de Vries).
Author: Christopher John Allison · Khoury College of Computer Sciences, Northeastern University.


