One forward camera in, metric lane geometry and a metric obstacle list out, with every stage scored against ground truth that was written down before the pixels existed.
That strip is the whole repository. A frame arrives bent by the lens, is rectified, is reduced to the pixels that could be paint, and is mapped onto a grid of the road whose scale is a consequence of the calibration rather than a constant somebody typed in. What comes out is a lane in metres and a list of things standing on the road, in metres.
Everything the pipeline needs is generated at run time. A synthetic chessboard dataset is projected through known intrinsics and known distortion, a target of known geometry is placed on the road at a known position, and a synthetic road scene generator produces frames with exact lane curvature, lateral offset, and obstacle footprints. Nothing is downloaded, nothing is stored in the repository, and the whole thing runs offline on a CPU. That is the point rather than a convenience: on recorded data the only available score is self consistency, and self consistency is exactly what a systematically wrong calibration preserves.
The rest of this page walks down the pipeline in the order a frame does. Each section says what the stage decides, what it was measured against, and what the measurement came out at.
Requires Python 3.12 or later. Continuous integration runs the whole suite on
3.12 and 3.13, on Linux and on Windows, so the version floor in pyproject.toml
is a tested claim rather than a declared one.
git clone https://github.com/Eelis03/av-perception-pipeline.git
cd av-perception-pipeline
uv syncUsing pip instead of uv:
python -m venv .venv
.venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e ".[dev]"Five scripts produce every number and every figure on this page:
uv run python examples/run_calibration.py
uv run python examples/run_birdseye.py
uv run python examples/run_lane_detection.py
uv run python examples/run_obstacle_extraction.py
uv run python examples/make_docs_figures.pyThe first four accept --output and --no-figures; the two sequence scripts
also accept --frames and --seed. To run on recorded images instead of
generated ones, pass a directory to open_frame_source. If the directory is
missing or holds no readable images the synthetic sequence is used instead,
which is why every command above works on a clean checkout with no data.
As a library:
from av_perception import (
PipelineConfig,
SequenceConfig,
SyntheticFrameSource,
evaluate_lane,
reference_camera,
run_sequence,
)
camera = reference_camera()
source = SyntheticFrameSource(camera, sequence=SequenceConfig(frames=40, seed=20260731))
trace = run_sequence(camera, source, PipelineConfig())
accuracy = evaluate_lane(trace)
geometry = trace.frames[0].geometry
print(f"detection rate {accuracy.detection_rate:.3f}")
print(f"lateral offset RMSE {accuracy.offset_rmse_cm:.2f} cm")
print(f"curvature RMSE {accuracy.curvature_rmse_per_km:.3f} per km")
print(f"mean lane width {accuracy.width_mean:.3f} m")
print(f"frame 0 offset {geometry.lateral_offset:+.3f} m")
print(f"frame 0 radius {geometry.curvature_radius:.0f} m")detection rate 0.975
lateral offset RMSE 1.07 cm
curvature RMSE 0.091 per km
mean lane width 3.697 m
frame 0 offset +0.224 m
frame 0 radius 24054 m
One line per stage, each expanded in the section that follows. Every figure in this table is the output of the command beside it, on Python 3.12.10 with numpy 2.5.1, opencv-python-headless 5.0.0, and matplotlib 3.11.1, on Windows 11. Everything except the per frame timing is deterministic given the seed.
| Stage | Measured against | Result | Command |
|---|---|---|---|
| Lens | The intrinsics and distortion the board was projected through | focal length to 3.3e-04 relative, two rectifications 1.33 px apart at worst | run_calibration.py |
| Placement | The height and angles the ground target was drawn from | height to 2.9 mm, pitch to 0.99 mrad, yaw to 0.28 mrad, roll to 0.22 mrad | run_calibration.py |
| Road plane | A rectangle of known metric size, and a second construction of the same homography | 3.70 m by 18.00 m becomes 74.0000 by 360.0000 px; the two constructions agree to 3.8e-09 | run_birdseye.py |
| Lane | The exact arc each frame was drawn from | lateral offset 1.07 cm RMSE, heading 1.46 mrad, curvature 0.091 per km, 39 of 40 frames accepted | run_lane_detection.py |
| Free space | The exact box footprints each frame was drawn from | recall 1.0000, precision 1.0000, near edge range 7.27 cm RMSE over 80 instances | run_obstacle_extraction.py |
| Whole chain | The true rig, over identical frames | recovering both the lens and the placement moves the reported lateral offset by 1.03 cm on average and changes no accept or reject decision | run_calibration.py |
The rig those numbers were measured on: 1280 by 720 at a focal length of 900
pixels, a horizontal field of view of 70.83 degrees, principal point at the
sensor centre. A mild barrel lens, k1 = -0.28, k2 = 0.10, p1 = 0.0008,
p2 = -0.0006, k3 = -0.02, whose radial map stays strictly increasing out to a
normalised radius of 1.2 against the 0.82 reached by the image corner. The camera
sits 1.35 m above the road on the vehicle centreline, pitched 3 degrees down,
which puts the horizon at row 312.33. The road is a 3.7 m lane, the design width
for a United States freeway, bounded by a solid yellow line on the left and a
broken white line on the right with 3 m marks on a 12 m cycle, which is the
standard broken line of the MUTCD. The sequence is 40 frames at 10 Hz at 25 m/s,
with curvature, lateral offset, and heading varying sinusoidally at three
incommensurate periods, a minimum curvature radius of 450 m, a lateral offset
amplitude of 0.35 m, and traffic in both adjacent lanes. Frame 11 is drawn with
the lane markings removed, standing in for a stretch of worn or covered road.
The three figures on this page are snapshots of one run of
uv run python examples/make_docs_figures.pywhich rewrites all three into docs/figures and fails if they exceed 250
kilobytes between them. They are not test fixtures and nothing compares them byte
for byte, in CI or anywhere else: matplotlib does not rasterise identically
across platforms, versions, or freetype builds, so a byte comparison would fail
on a correct change and would say nothing about whether the figure is right. What
CI does check is that the command runs and that the tracked figures stay inside
the budget.
Coverage is measured by
uv run pytest --cov=src/av_perception --cov-report=term-missingand stands at 90 percent of 2145 statements. CI runs the same command with
--cov-fail-under=88.
src/av_perception/algorithm/calibration.py, run_calibration.py
A pinhole projection with Brown-Conrady distortion: a radial series in
k1, k2, k3 and the two tangential terms p1, p2 of Brown (1966) and Conrady
(1919). A 1280 by 720 camera with a mild barrel lens moves a point at the corner
of the frame by more than a hundred pixels, and every geometric step downstream
assumes that straight lines in the world are straight in the image, so this is
the first thing that has to be right.
The inverse of that model has no closed form. The usual scheme, and the one OpenCV uses, is a fixed point iteration with a fixed step count; it converges linearly and leaves a residual of up to 0.37 pixels at the corner of this lens. Newton on the same equation, with the analytic Jacobian, reaches below 1e-12 pixels in about five steps, and it is what this repository uses. A solution at a negative radius is refused rather than returned, because past the turning point of the radial polynomial the equation has a root that reflects the image through the principal point, and that root has zero residual and is completely wrong.
The estimator is Zhang's method (2000) on planar chessboard views, refined by
Levenberg-Marquardt, through cv2.calibrateCamera. Calibration is where
correctness is hardest to check: on recorded images the only available score is
reprojection error, and reprojection error is exactly what a wrong calibration
preserves, because focal length and radial distortion trade against one another.
That is why the target here is synthetic, and why two observation paths are
reported rather than one.
projected corners
source projected corners
views 12
corners 648
image area covered 0.969
RMS reprojection error [px] 0.00002
worst view error [px] 0.00003
fx error [px] +0.00002
fy error [px] +0.00001
cx error [px] -0.00000
cy error [px] +0.00005
focal relative error 1.396e-08
k1 error -1.490e-07
k2 error +8.640e-08
p1 error -2.783e-08
p2 error +4.381e-08
k3 error +1.935e-07
rectification disagreement [px] 0.00064
radial map invertible True
rendered images with corner detection
source rendered images, 12 of 12 views detected
views 12
corners 648
image area covered 0.969
RMS reprojection error [px] 0.12916
worst view error [px] 0.26396
fx error [px] -0.26570
fy error [px] -0.32468
cx error [px] -0.08707
cy error [px] +0.25631
focal relative error 3.280e-04
k1 error +4.233e-04
k2 error -2.087e-03
p1 error +4.914e-05
p2 error -3.540e-05
k3 error +2.113e-03
rectification disagreement [px] 1.33243
radial map invertible True
The first path computes the corners with the forward model, which isolates the estimator: it returns the parameters it was given, the focal length to fourteen parts in a thousand million and every distortion coefficient to better than 2e-07. That is a statement about the implementation and not about calibration in general, and it is what makes any error here a bug rather than a measurement.
The second path renders images of the same board at the same poses and finds the
corners with cv2.findChessboardCorners and cv2.cornerSubPix. The detector
contributes about 0.12 pixels RMS, the recovered focal length moves by 0.27
pixels, three parts in ten thousand, and k2 and k3 move by two parts in a
thousand. Those five coefficient errors are not individually interpretable, since
the terms trade against each other, so they are folded into one number that a
downstream stage would actually feel: rectifying a frame with the true
coefficients and with the recovered ones puts the same pixel in two places 1.33
pixels apart at the worst point in the frame.
Alongside the errors the implementation reports the fraction of the frame the corners covered, 96.9 percent of an eight by eight grid here, and whether the recovered radial polynomial is still invertible out to the image corner. A calibration built entirely from the middle of the frame reports a small reprojection error and is still wrong at the edges, where distortion lives.
src/av_perception/algorithm/extrinsics.py,
src/av_perception/pipeline/ground_target.py, run_calibration.py
The lens is half of a calibration. The ground plane homography is
K [r1 | r2 | t], so the camera height and its three angles set the scale and
the shape of the bird's-eye view just as firmly as the focal length does. Until
recently this repository assumed them, and said so in its limitations. It now
measures them.
A chessboard lies on the road at a measured position and is read by the same
corner detector the intrinsics use. The estimator is the planar pose problem run
in the road frame: rectify the corners, solve the road to image homography by the
normalised direct linear transform, decompose K^-1 H into [r1 | r2 | t] with
the scale fixed by the unit length of the rotation columns and the sign fixed by
requiring the target to be in front of the camera, project the result onto the
rotation group by a singular value decomposition, and refine over all corners by
Levenberg-Marquardt through the full distortion model.
placement from a rendered ground target, using the recovered lens
source rendered image with corner detection
corners 45
target span [px] 533 across, 73 along
RMS reprojection error [px] 0.39089
height [m] 1.35293
height error [mm] +2.927
pitch [deg] 3.05667
pitch error [mrad] +0.9891
yaw error [mrad] -0.2819
roll error [mrad] +0.2234
lateral error [mm] -0.426
longitudinal residual [mm] +2.283
road displacement [cm] mean 21.99, max 62.17
From projected corners rather than a rendered image the same estimator returns the placement it was drawn from to solver precision, which is the same bug check the intrinsics get.
The target has to be large, and the target span line is why: 533 pixels across
the road, 73 along it. Foreshortening compresses range, so at 4 m ahead a metre
along the line of sight occupies about a fifth of the pixels a metre across the
road does. A hand held board of the kind used for the intrinsics would be a few
pixels deep on the ground and would constrain the height not at all. The
published target is a 3.0 m by 1.8 m printed pattern with 0.30 m squares, which
is a floor mounted target in an alignment bay rather than something carried in a
boot.
longitudinal residual is not an accuracy figure. The road frame origin is
defined as the point directly under the camera, so the recovered camera centre
must return to it, and 2.3 mm is how close it came. That residual exists because
nothing else can catch a particular failure: the corner detector reads a
rectangular grid in one of two rigid orders, and reading the target through half
a turn describes a real pose of a real board, reprojects perfectly, and is
completely wrong. Goodness of fit cannot see it. The distance from the recovered
centre to the origin can, and a test asserts exactly that.
road displacement folds the height and the three angle errors into one metric
number the same way the rectification disagreement folds the five distortion
coefficients: a road point projected through the true placement and read back
through the recovered one lands 22 cm away on average and 62 cm away at worst
over the 4 m to 30 m window, most of that at the far edge where a milliradian of
pitch is worth decimetres.
The last block runs the lane pipeline three times over identical frames, so that the two calibrations can be separated and then charged together:
effect on the lane pipeline, recovered camera against true camera
frames 12
accepted with the true rig 11
recovered lens only
accepted 11
lateral offset change [cm] mean 0.602, max 1.165
lane width change [cm] mean 1.405, max 1.930
recovered lens and placement
accepted 11
lateral offset change [cm] mean 1.028, max 2.120
lane width change [cm] mean 2.528, max 3.155
A rig calibrated entirely from its own targets, lens and placement alike, reports a lateral offset 1.03 cm from the one a perfectly known rig reports, and changes no accept or reject decision. What remains uncalibrated is stated under what it does not do.
src/av_perception/model/homography.py,
src/av_perception/pipeline/runner.py, run_birdseye.py
Panels one and two of the figure at the top of this page are this stage. Every frame is rectified onto its own camera matrix, so the ground plane homography, which is built from that matrix, applies to the rectified image with no further adjustment.
A road point lies on the plane z = 0, so its camera frame coordinate is
[r1 r2 t] [x, y, 1] and the image of the road plane is the homography
H = K [r1 | r2 | t], fixed entirely by the intrinsics and the placement, both
of which are now measured. Composing its inverse with a metric to raster
similarity gives the inverse perspective mapping of Mallot et al. (1991). The
window is defined in metres first, 4 m to 30 m ahead and 6 m either side at
0.05 m per pixel, and that definition is the pixel to metre scale, so nothing
downstream needs a separate scaling constant.
image size 1280 x 720
horizontal field of view [deg] 70.83
camera height [m] 1.350
camera pitch [deg] 3.000
horizon row 312.33
bird's-eye window [m] x 4.0 to 30.0, y -6.0 to 6.0
resolution [m/px] 0.050
bird's-eye raster 241 x 521
four point against analytic 3.798e-09
rectangle 3.70 m by 18.00 m maps to:
width [px] 74.0000 expected 74.0000
length [px] 360.0000 expected 360.0000
corner angle [deg] 90.000000
observed cells 119965 of 125561
observed fraction 0.9554
ego lane fully observed from [m] 4.00
A rectangle 3.70 m by 18.00 m on the road becomes a rectangle 74.0000 by 360.0000 pixels with corners at 90.000000 degrees. That is what metric by construction means: 74 pixels at 0.05 m per pixel is 3.70 m exactly, with no constant fitted anywhere.
The four point construction usually seen in lane detection code, where a trapezoid is picked by hand on a straight road and asserted to be some size, is also implemented, by the normalised direct linear transform of Hartley and Zisserman. It agrees with the analytic homography to 3.8e-09 and exists so the two can be compared, not because it is used: in the four point route every metre the pipeline later reports inherits the error in that assertion, and nothing downstream can detect it.
The near and far edges of the window are set by the camera rather than chosen. Closer than about 3.5 m the 2.5 m corridor either side of the vehicle has already left the frame, which is also why 95.5 percent of the window is observed and the missing 4.5 percent is the two near corners. Beyond 30 m a 0.12 m marking subtends fewer than 3.6 pixels and the thresholding stage loses it.
src/av_perception/algorithm/threshold.py,
src/av_perception/algorithm/lane_search.py,
src/av_perception/algorithm/fitting.py, run_lane_detection.py
Panels three and four of the figure at the top are this stage. Thresholding runs
on the rectified perspective image rather than on the warped one, because the
warp resamples and resampling a 3.6 pixel wide marking at 30 m before
differentiating it throws away the resolution the gradient operator needs. Three
channels: high lightness for white paint, high saturation for yellow paint, and a
Sobel gradient in the image x direction restricted to steep edges, which keeps
marking edges and discards the horizon and shadows lying across the road. The
binary result is warped into the bird's-eye grid.
Pixels are associated either by a stack of sliding windows started from a column
histogram, in the acquisition mode, or by a corridor around the previous accepted
fit, in the tracking mode. The fit is a second order polynomial of lateral
position against longitudinal distance, in metres, so the coefficients need no
rescaling afterwards: c is the lateral position at the vehicle, b is the
tangent of the heading error, and 2a is the curvature. The common alternative,
fitting in pixels and multiplying by a metres per pixel ratio afterwards, is
algebraically identical and was rejected because the rescaling constants end up
in a different file from the transform that defines them.
The two boundaries are fitted jointly, sharing one quadratic shape and carrying one offset each. That constraint is the definition of a lane, and it matters because support is asymmetric: the left boundary is solid and imaged along the whole window, the right is broken, and on about a quarter of frames its nearest mark is beyond ten metres. Both models are implemented and both are reported.
parallel model with tracking
frames 40
accepted 39
detection rate 0.9750
lateral offset RMSE [cm] 1.07
lateral offset max [cm] 3.65
lateral offset bias [cm] +0.13
heading RMSE [mrad] 1.456
curvature RMSE [1/km] 0.0911
curvature max error [1/km] 0.2392
radius relative RMSE 0.0927
mean lane width [m] 3.6968
lane width RMSE [cm] 0.63
mean fit residual [cm] 5.60
window searches 3
corridor searches 37
tracking fallbacks 1
mean frame time [ms] 36.67
independent model with tracking
frames 40
accepted 39
detection rate 0.9750
lateral offset RMSE [cm] 10.13
lateral offset max [cm] 22.47
lateral offset bias [cm] +6.64
heading RMSE [mrad] 12.865
curvature RMSE [1/km] 0.7388
curvature max error [1/km] 1.4917
radius relative RMSE 0.9146
mean lane width [m] 3.6857
lane width RMSE [cm] 1.99
mean fit residual [cm] 5.58
parallel model, full search every frame
frames 40
accepted 39
detection rate 0.9750
lateral offset RMSE [cm] 1.13
lateral offset max [cm] 3.66
lateral offset bias [cm] +0.08
heading RMSE [mrad] 1.747
curvature RMSE [1/km] 0.1033
curvature max error [1/km] 0.3628
radius relative RMSE 0.0975
mean lane width [m] 3.6996
lane width RMSE [cm] 0.95
window searches 40
corridor searches 0
tracking fallbacks 0
rejected frames 1
left support 0 below 200 pixels
right support 0 below 200 pixels
left span 0.0 m below 6.0 m
right span 0.0 m below 6.0 m
tracking against full search max offset difference 1.13 cm over 39 frames
The lateral offset is recovered to 1.07 cm RMS with a bias of 1.3 mm, over a true offset that sweeps 0.70 m. Lane width comes back as 3.6968 m against a true 3.7000 m, a systematic 3 mm narrow, which is the separation of the fitted line centres rather than of the paint. Heading is recovered to 1.46 mrad, which is 0.08 degrees.
Curvature is the weakest output and is reported as curvature rather than as
radius. The RMS error is 0.091 per kilometre against a true amplitude of 2.22 per
kilometre, and the relative radius error is 9.3 percent. That is a property of
the measurement rather than of this implementation: a second order fit over a
26 m window is estimating the coefficient of s^2, and at a 450 m radius the
whole lateral excursion the fit has to work with is 0.75 m. Any single frame
monocular curvature estimate over this range behaves this way, which is why
production systems filter it across frames and why this one does not, since a
filtered number would report the filter's ability to average rather than the
detector's ability to measure.
The three variants differ in exactly one respect each. The first two isolate the fitting model: sharing one quadratic shape between the boundaries reduces the lateral offset error by a factor of 9.5 and the relative radius error by a factor of 9.9, because the broken right boundary no longer has to estimate its own shape from two distant marks. The first and the third isolate the search: the corridor and the full histogram agree to 1.13 cm at worst over the 39 frames both accept, so tracking changes the cost of the answer and not the answer.
The vertical band in the figure is the frame drawn without markings, and it is rejected for the right reason: no support on either side. The tracking corridor found nothing, the full search was rerun on the same frame as a fallback and also found nothing, the tracking state was discarded, and the next frame reacquired from the histogram and was accurate again. Acceptance needs enough pixels over a long enough span on each side, a lane width between 2.6 m and 4.6 m, a small fit residual, and a curvature radius above 60 m. The mean frame time is the only machine dependent number on this page and excludes scene generation.
src/av_perception/algorithm/obstacles.py, run_obstacle_extraction.py
The lane says where to go. It does not say whether something is standing in the way. The bird's-eye view assumes everything it shows lies on the road plane, and anything standing above the plane breaks that assumption in a specific and predictable way: the ray through the top of an object strikes the ground further away than the object stands, so the object smears outwards along the line of sight, beginning exactly at its ground contact edge nearest the camera.
The left panel is the bird's-eye colour image and the right panel is the decision
made from it: unknown, free, obstacle, with the detected near edge in red and the
true footprint in dashed yellow. The two long yellow wedges running to the top of
the frame are the smear, and they are not an error to be fixed. An object of
height h seen at range r from a camera at height H occupies the mapped
ground out to r H / (H - h). The extractor therefore reports the near edge
separately from the full component, scores accuracy against the near edge only,
and reports the smear without scoring it.
frames 40
bird's-eye window [m2] 312.0
observed area [m2] 299.9
free area [m2] 270.1
obstacle area [m2] 29.8
free fraction of observed 0.9008
true obstacle instances 80
detections 80
matched 80
missed 0
spurious 0
recall 1.0000
precision 1.0000
near edge range RMSE [cm] 7.27
near edge range bias [cm] -1.94
near edge range max [cm] 35.33
lateral centre RMSE [cm] 8.77
lateral width bias [cm] +13.06
predicted smear beyond the near edge, from the flat world assumption:
mean [m] 7.08
max [m] 12.08
obstacle height [m] 0.55, 0.45
camera height [m] 1.35
Across 80 obstacle instances there is no miss and no spurious detection. The near edge range is recovered to 7.3 cm RMS with a bias of 1.9 cm short, and the lateral centre of the near edge to 8.8 cm. The near edge width is overestimated by 13.1 cm, about two and a half bird's-eye cells, which is the morphological closing and the antialiased silhouette of the box adding a cell on each side.
The obstacle area of 29.8 m2 is more than eight times the 3.6 m2 the two footprints actually occupy, and the predicted smear of 7.08 m on average and 12.08 m at worst is exactly that difference. An object taller than the camera has no ray through its top that ever meets the road, and its smear runs to the horizon. The remaining 10 percent of the observed window that is not called free is entirely this smear: on a frame with no obstacles the segmentation returns no components at all.
The code is laid out in the order this page reads. model holds pure functions
and dataclasses with no I/O and no state, algorithm holds the decisions and
reads no files and draws no random numbers, pipeline is the only place a random
number is drawn or a file is read, analysis reads traces and produces numbers
and figures, and the example scripts contain wiring and printing and no logic
that is not tested elsewhere. The dependency direction never runs backwards.
Both of these were found by running the pipeline against ground truth, not by reading the code, and both are recorded because they are the evidence that the evaluation was real.
The yellow lane line was a 26 metre obstacle. The free space test asks whether a pixel is achromatic and not dark. Yellow paint at high saturation is neither, so the solid left boundary came back as a connected component 26 m long lying down the middle of the drivable area, blocking the lane it was defining. The fix is the second clause of the appearance model, which admits yellow at high saturation and moderate lightness as road surface. It is not decoration and it is the reason the model has two clauses instead of one.
An adjacent lane vehicle biased the lateral offset by 17 cm. During development an obstacle whose silhouette came within 0.13 m of the left boundary at long range was claimed by the sliding window and dragged the fit with it. The sliding window half width of 0.55 m is what sets the distance at which this happens, and the parallelism and residual tests are what catch it when it does. The generated scenes now place traffic in the adjacent lanes at no less than 0.90 m clearance from the ego lane markings on every frame regardless of curvature, so that the lane accuracy figures above measure the lane pipeline rather than an interaction. The interaction itself is real and is documented rather than removed. A vehicle in the ego lane, which is the case that matters most, would occlude the boundaries outright and the frame would be rejected for lack of support.
uv run pytest
uv run ruff check .
uv run ruff format --check .
uv run mypy195 tests run in about 36 seconds, dominated by the frames the scene generator has to render, and in about 46 seconds with coverage measurement on. They come in three tiers.
The first tier is property and invariant tests over the mathematics. Projecting a
camera frame point and unprojecting it returns the point to 1e-09. Distortion
applied and then removed is the identity to 1e-12 in normalised units. The
analytic distortion Jacobian matches a central difference. The forward model
agrees with cv2.projectPoints and the inverse leaves a smaller residual than
cv2.undistortPoints. A non-invertible lens raises rather than returning a
reflected root. The direct linear transform matches cv2.getPerspectiveTransform
and rejects collinear correspondences. The bird's-eye homography maps a 3.7 m by
18 m road rectangle to a 74 by 360 pixel axis aligned rectangle. Calibration on
the synthetic board recovers the known focal length to 0.01 pixels and every
distortion coefficient to 1e-04, and degrades in proportion to added corner noise.
The extrinsic estimator returns the placement it was drawn from to 1e-09 for
yawed and rolled rigs as well as level ones, inverts the rig parameterisation
exactly, repairs a corner reading that is half a turn out, refuses one that is
mirrored, and recovers a rendered target to within a centimetre of height and
3 mrad of every angle. A straight lane yields a curvature radius above 10 000 m
and a known lateral offset within 2 cm. A lane of known constant radius yields
that radius within 5 percent. The corridor search and the window search claim the
same pixels and produce lateral offsets within 3 cm on the same frames. The
parallel model beats the independent one on a boundary visible only in two distant
marks. A frame drawn without markings is rejected for lack of support while the
next frame reacquires and is accurate again. And the loop between the two mappings
closes: the yellow line drawn by forward projection lands on the true arc after
distortion correction and inverse perspective mapping, to 0.5 cm mean and 5 cm
worst case.
The second tier replays a recorded 14 frame run against
tests/data/reference_run.json; regenerate it with
uv run python tests/generate_reference.py when a change to the algorithms is
intended. What it pins, and what it does not, is deliberate. Search modes, accept
and reject decisions, obstacle counts, and detection counts are pinned exactly,
because they are decisions rather than measurements. Lane offset and lane width
are pinned in centimetres to half a bird's-eye cell, obstacle near edge range to
five cells, and the calibration to 1e-03 pixels, because those come out of
resampled images and cv2.remap and cv2.warpPerspective are entitled to differ
in the last bit between OpenCV builds and platforms, which can move a lane pixel
across a threshold. The count of thresholded pixels, the raw polynomial
coefficients, and the curvature radius are not pinned at all: the first two are
image statistics that move with the last bit of an interpolation, and the third
is the reciprocal of the estimated quantity and diverges on a straight road, so
the signed curvature is pinned instead and a separate test asserts the
qualitative property that a near straight frame reports a radius above 2000 m.
The third tier runs every script in examples/ as a subprocess under reduced
frame and view counts, including the figure writing paths, the fallback from an
absent frame directory to the synthetic sequence, and the figure budget. A fourth
handful of tests covers what the repository ships rather than what it computes:
that py.typed is present, empty, and inside the importable package, that the
wheel is built from the directory holding it, that the three published figures
exist and fit the budget, and that the README shows all of them.
CI runs the suite, the linter, and the type checker on Ubuntu and on Windows,
with --cov-fail-under=88.
The full list, with the reasoning, is in docs/design-notes.md, which also records the alternatives that were rejected, including the Hough transform, RANSAC, clothoid road models, recursive filtering across frames, stereo free space, and learned lane detectors, and one limitation that has since been closed. The four that matter most here:
- Fixed thresholds are brittle to lighting. Lightness at least 170, saturation at least 90, gradient at least 40 after scaling. A hard shadow drops the paint under it below threshold, a wet surface in low sun fills the mask, and worn paint on light concrete never reaches either colour threshold. This is the gap between the accuracy figures above and a real road.
- Free space is decided by appearance. A shadow, a patch of new asphalt, or a wet stain will be called an obstacle. From one monocular frame a black object standing on the road and a black shadow lying on it produce the same pixels, so there is no fix here that is not a different sensor.
- The placement is calibrated once, not continuously. A vehicle pitches under braking, the suspension moves with load, and the road has crest and sag curves. The calibration answers where the camera was on the day, not where it is on this frame.
- The scene generator does not produce the failures that matter. Clean paint, uniform lighting, no shadows, no weather, no occlusion of the ego lane boundaries. The geometric results carry over to a real road because they do not depend on appearance; the detection rate does not.
- Brown, D. C. "Decentering Distortion of Lenses". Photogrammetric Engineering, 32(3):444 to 462, 1966. https://www.asprs.org/wp-content/uploads/pers/1966journal/may/1966_may_444-462.pdf. The radial and decentering distortion model implemented here.
- Conrady, A. E. "Decentred Lens-Systems". Monthly Notices of the Royal Astronomical Society, 79(5):384 to 390, 1919. DOI: 10.1093/mnras/79.5.384. The tangential terms.
- Zhang, Z. "A Flexible New Technique for Camera Calibration". IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(11):1330 to 1334, 2000. DOI: 10.1109/34.888718. The calibration method: planar homographies, the closed form intrinsics, and the nonlinear refinement. The same planar pose decomposition, run in the road frame, is the extrinsic estimator.
- Levenberg, K. "A Method for the Solution of Certain Non-Linear Problems in Least Squares". Quarterly of Applied Mathematics, 2(2):164 to 168, 1944. DOI: 10.1090/qam/10666.
- Marquardt, D. W. "An Algorithm for Least-Squares Estimation of Nonlinear Parameters". Journal of the Society for Industrial and Applied Mathematics, 11(2):431 to 441, 1963. DOI: 10.1137/0111030. The refinement both calibration solvers run.
- Hartley, R. and Zisserman, A. "Multiple View Geometry in Computer Vision", second edition, Cambridge University Press, 2004. DOI: 10.1017/CBO9780511811685. Chapter 4 for the direct linear transform and the plane to plane homography.
- Hartley, R. "In Defense of the Eight-Point Algorithm". IEEE Transactions on Pattern Analysis and Machine Intelligence, 19(6):580 to 593, 1997. DOI: 10.1109/34.601246. The isotropic normalisation applied before the direct linear transform.
- Rodrigues, O. "Des lois geometriques qui regissent les deplacements d'un systeme solide dans l'espace". Journal de Mathematiques Pures et Appliquees, 5:380 to 440, 1840. https://eudml.org/doc/234443. The axis-angle to rotation matrix formula used for board poses.
- Mallot, H. A., Bulthoff, H. H., Little, J. J., and Bohrer, S. "Inverse Perspective Mapping Simplifies Optical Flow Computation and Obstacle Detection". Biological Cybernetics, 64(3):177 to 185, 1991. DOI: 10.1007/BF00201978. The inverse perspective mapping, and the observation that objects above the road plane smear along the line of sight.
- Bertozzi, M. and Broggi, A. "GOLD: A Parallel Real-Time Stereo Vision System for Generic Obstacle and Lane Detection". IEEE Transactions on Image Processing, 7(1):62 to 81, 1998. DOI: 10.1109/83.650851. Lane and obstacle detection in a remapped bird's-eye view.
- Bertozzi, M., Broggi, A., and Fascioli, A. "Stereo Inverse Perspective Mapping: Theory and Applications". Image and Vision Computing, 16(8):585 to 590, 1998. DOI: 10.1016/S0262-8856(97)00093-0.
- Aly, M. "Real Time Detection of Lane Markers in Urban Streets". In IEEE Intelligent Vehicles Symposium, 2008, pp. 7 to 12. DOI: 10.1109/IVS.2008.4621152. Thresholding, grouping, and fitting of lane candidates in a bird's-eye view.
- Dickmanns, E. D. and Mysliwetz, B. D. "Recursive 3-D Road and Relative Ego-State Recognition". IEEE Transactions on Pattern Analysis and Machine Intelligence, 14(2):199 to 213, 1992. DOI: 10.1109/34.121789. The parallel boundary road model and its recursive estimation.
- McCall, J. C. and Trivedi, M. M. "Video-Based Lane Estimation and Tracking for Driver Assistance: Survey, System, and Evaluation". IEEE Transactions on Intelligent Transportation Systems, 7(1):20 to 37, 2006. DOI: 10.1109/TITS.2006.869595. Survey of the classical stack and of how lane estimates are evaluated.
- Hillel, A. B., Lerner, R., Levi, D., and Raz, G. "Recent Progress in Road and Lane Detection: A Survey". Machine Vision and Applications, 25(3):727 to 745, 2014. DOI: 10.1007/s00138-011-0404-2.
- Sobel, I. and Feldman, G. "A 3x3 Isotropic Gradient Operator for Image Processing". Presented at the Stanford Artificial Intelligence Project, 1968, and reproduced in Duda, R. O. and Hart, P. E. "Pattern Classification and Scene Analysis", Wiley, 1973. https://www.researchgate.net/publication/239398674. The gradient operator used by the thresholding stage.
- Rosenfeld, A. and Pfaltz, J. L. "Sequential Operations in Digital Picture Processing". Journal of the ACM, 13(4):471 to 494, 1966. DOI: 10.1145/321356.321357. Connected component labelling.
- Bolelli, F., Allegretti, S., Baraldi, L., and Grana, C. "Spaghetti Labeling:
Directed Acyclic Graphs for Block-Based Connected Components Labeling". IEEE
Transactions on Image Processing, 29:1999 to 2012, 2020.
DOI: 10.1109/TIP.2019.2946979. The
algorithm behind
cv2.connectedComponentsWithStats. - Serra, J. "Image Analysis and Mathematical Morphology". Academic Press, 1982. https://shop.elsevier.com/books/image-analysis-and-mathematical-morphology/serra/978-0-12-637240-3. Opening and closing, used to clean the obstacle mask.
- Smith, A. R. "Color Gamut Transform Pairs". ACM SIGGRAPH Computer Graphics, 12(3):12 to 19, 1978. DOI: 10.1145/965139.807361. The hue, lightness, saturation space the colour thresholds are defined in.
- Federal Highway Administration. "Manual on Uniform Traffic Control Devices for Streets and Highways", 11th edition, 2023. https://mutcd.fhwa.dot.gov/. Part 3 for the yellow line on the left of a one-way carriageway, the white line on the right, and the broken line pattern of a 3 m mark on a 12 m cycle used by the scene generator.
- American Association of State Highway and Transportation Officials. "A Policy on Geometric Design of Highways and Streets", 7th edition, 2018. https://store.transportation.org/item/collectiondetail/180. Lane widths of 2.7 m to 3.6 m and the minimum curve radii that set the acceptance band on curvature.
- Duda, R. O. and Hart, P. E. "Use of the Hough Transformation to Detect Lines and Curves in Pictures". Communications of the ACM, 15(1):11 to 15, 1972. DOI: 10.1145/361237.361242.
- Fischler, M. A. and Bolles, R. C. "Random Sample Consensus: A Paradigm for Model Fitting with Applications to Image Analysis and Automated Cartography". Communications of the ACM, 24(6):381 to 395, 1981. DOI: 10.1145/358669.358692.
- Southall, B. and Taylor, C. J. "Stochastic Road Shape Estimation". In IEEE International Conference on Computer Vision, 2001, pp. 205 to 212. DOI: 10.1109/ICCV.2001.937519.
- Badino, H., Franke, U., and Pfeiffer, D. "The Stixel World: A Compact Medium Level Representation of the 3D-World". In Pattern Recognition, DAGM 2009, Lecture Notes in Computer Science volume 5748, pp. 51 to 60. DOI: 10.1007/978-3-642-03798-6_6.
- Pan, X., Shi, J., Luo, P., Wang, X., and Tang, X. "Spatial As Deep: Spatial CNN for Traffic Scene Understanding". In AAAI Conference on Artificial Intelligence, 32(1), 2018. DOI: 10.1609/aaai.v32i1.12301.
- Neven, D., De Brabandere, B., Georgoulis, S., Proesmans, M., and Van Gool, L. "Towards End-to-End Lane Detection: an Instance Segmentation Approach". In IEEE Intelligent Vehicles Symposium, 2018, pp. 286 to 291. DOI: 10.1109/IVS.2018.8500547.
- Tabelini, L., Berriel, R., Paixao, T. M., Badue, C., De Souza, A. F., and Oliveira-Santos, T. "Keep Your Eyes on the Lane: Real-Time Attention-Guided Lane Detection". In IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp. 294 to 302. DOI: 10.1109/CVPR46437.2021.00036.
| Package | Version | Purpose | Licence |
|---|---|---|---|
| numpy | >= 2.0 | Array storage, least squares and singular value decomposition, seeded random number generation | BSD-3-Clause |
| opencv-python-headless | >= 4.10 | Calibration solver, chessboard corner detection and subpixel refinement, pose refinement, image remapping and perspective warping, colour conversion, Sobel filtering, morphology, connected components | Apache-2.0 for the OpenCV library, MIT for the Python packaging |
| matplotlib | >= 3.9 | Calibration, pipeline stage, lane geometry, and obstacle figures | Matplotlib licence, a BSD-compatible licence derived from the Python Software Foundation licence |
| pytest | >= 8.3 | Test runner, development only | MIT |
| pytest-cov | >= 6.0 | Coverage measurement, development only | MIT |
| ruff | >= 0.8 | Linter, development only | MIT |
| mypy | >= 1.13 | Static type checker, development only | MIT |
Citations for the runtime dependencies:
- Harris, C. R. et al. "Array Programming with NumPy". Nature, 585:357 to 362, 2020. DOI: 10.1038/s41586-020-2649-2.
- Bradski, G. "The OpenCV Library". Dr. Dobb's Journal of Software Tools, 2000. https://opencv.org/.
- Hunter, J. D. "Matplotlib: A 2D Graphics Environment". Computing in Science and Engineering, 9(3):90 to 95, 2007. DOI: 10.1109/MCSE.2007.55.
Released under the MIT license. See LICENSE.


