Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Av Perception Pipeline

One forward camera in, metric lane geometry and a metric obstacle list out, with every stage scored against ground truth that was written down before the pixels existed.

CI Python License

Four stages of one frame side by side: the raw camera image, the same frame after distortion correction has straightened the road edges, the colour and gradient mask holding the paint and dropping the road, and the bird's-eye view with the two lane boundaries and the lane centre fitted in metres

That strip is the whole repository. A frame arrives bent by the lens, is rectified, is reduced to the pixels that could be paint, and is mapped onto a grid of the road whose scale is a consequence of the calibration rather than a constant somebody typed in. What comes out is a lane in metres and a list of things standing on the road, in metres.

Everything the pipeline needs is generated at run time. A synthetic chessboard dataset is projected through known intrinsics and known distortion, a target of known geometry is placed on the road at a known position, and a synthetic road scene generator produces frames with exact lane curvature, lateral offset, and obstacle footprints. Nothing is downloaded, nothing is stored in the repository, and the whole thing runs offline on a CPU. That is the point rather than a convenience: on recorded data the only available score is self consistency, and self consistency is exactly what a systematically wrong calibration preserves.

The rest of this page walks down the pipeline in the order a frame does. Each section says what the stage decides, what it was measured against, and what the measurement came out at.

Installation

Requires Python 3.12 or later. Continuous integration runs the whole suite on 3.12 and 3.13, on Linux and on Windows, so the version floor in pyproject.toml is a tested claim rather than a declared one.

git clone https://github.com/Eelis03/av-perception-pipeline.git
cd av-perception-pipeline
uv sync

Using pip instead of uv:

python -m venv .venv
.venv/bin/activate      # Windows: .venv\Scripts\activate
pip install -e ".[dev]"

Five scripts produce every number and every figure on this page:

uv run python examples/run_calibration.py
uv run python examples/run_birdseye.py
uv run python examples/run_lane_detection.py
uv run python examples/run_obstacle_extraction.py
uv run python examples/make_docs_figures.py

The first four accept --output and --no-figures; the two sequence scripts also accept --frames and --seed. To run on recorded images instead of generated ones, pass a directory to open_frame_source. If the directory is missing or holds no readable images the synthetic sequence is used instead, which is why every command above works on a clean checkout with no data.

As a library:

from av_perception import (
    PipelineConfig,
    SequenceConfig,
    SyntheticFrameSource,
    evaluate_lane,
    reference_camera,
    run_sequence,
)

camera = reference_camera()
source = SyntheticFrameSource(camera, sequence=SequenceConfig(frames=40, seed=20260731))
trace = run_sequence(camera, source, PipelineConfig())
accuracy = evaluate_lane(trace)
geometry = trace.frames[0].geometry

print(f"detection rate        {accuracy.detection_rate:.3f}")
print(f"lateral offset RMSE   {accuracy.offset_rmse_cm:.2f} cm")
print(f"curvature RMSE        {accuracy.curvature_rmse_per_km:.3f} per km")
print(f"mean lane width       {accuracy.width_mean:.3f} m")
print(f"frame 0 offset        {geometry.lateral_offset:+.3f} m")
print(f"frame 0 radius        {geometry.curvature_radius:.0f} m")
detection rate        0.975
lateral offset RMSE   1.07 cm
curvature RMSE        0.091 per km
mean lane width       3.697 m
frame 0 offset        +0.224 m
frame 0 radius        24054 m

Results

One line per stage, each expanded in the section that follows. Every figure in this table is the output of the command beside it, on Python 3.12.10 with numpy 2.5.1, opencv-python-headless 5.0.0, and matplotlib 3.11.1, on Windows 11. Everything except the per frame timing is deterministic given the seed.

Stage Measured against Result Command
Lens The intrinsics and distortion the board was projected through focal length to 3.3e-04 relative, two rectifications 1.33 px apart at worst run_calibration.py
Placement The height and angles the ground target was drawn from height to 2.9 mm, pitch to 0.99 mrad, yaw to 0.28 mrad, roll to 0.22 mrad run_calibration.py
Road plane A rectangle of known metric size, and a second construction of the same homography 3.70 m by 18.00 m becomes 74.0000 by 360.0000 px; the two constructions agree to 3.8e-09 run_birdseye.py
Lane The exact arc each frame was drawn from lateral offset 1.07 cm RMSE, heading 1.46 mrad, curvature 0.091 per km, 39 of 40 frames accepted run_lane_detection.py
Free space The exact box footprints each frame was drawn from recall 1.0000, precision 1.0000, near edge range 7.27 cm RMSE over 80 instances run_obstacle_extraction.py
Whole chain The true rig, over identical frames recovering both the lens and the placement moves the reported lateral offset by 1.03 cm on average and changes no accept or reject decision run_calibration.py

The rig those numbers were measured on: 1280 by 720 at a focal length of 900 pixels, a horizontal field of view of 70.83 degrees, principal point at the sensor centre. A mild barrel lens, k1 = -0.28, k2 = 0.10, p1 = 0.0008, p2 = -0.0006, k3 = -0.02, whose radial map stays strictly increasing out to a normalised radius of 1.2 against the 0.82 reached by the image corner. The camera sits 1.35 m above the road on the vehicle centreline, pitched 3 degrees down, which puts the horizon at row 312.33. The road is a 3.7 m lane, the design width for a United States freeway, bounded by a solid yellow line on the left and a broken white line on the right with 3 m marks on a 12 m cycle, which is the standard broken line of the MUTCD. The sequence is 40 frames at 10 Hz at 25 m/s, with curvature, lateral offset, and heading varying sinusoidally at three incommensurate periods, a minimum curvature radius of 450 m, a lateral offset amplitude of 0.35 m, and traffic in both adjacent lanes. Frame 11 is drawn with the lane markings removed, standing in for a stretch of worn or covered road.

The three figures on this page are snapshots of one run of

uv run python examples/make_docs_figures.py

which rewrites all three into docs/figures and fails if they exceed 250 kilobytes between them. They are not test fixtures and nothing compares them byte for byte, in CI or anywhere else: matplotlib does not rasterise identically across platforms, versions, or freetype builds, so a byte comparison would fail on a correct change and would say nothing about whether the figure is right. What CI does check is that the command runs and that the tracked figures stay inside the budget.

Coverage is measured by

uv run pytest --cov=src/av_perception --cov-report=term-missing

and stands at 90 percent of 2145 statements. CI runs the same command with --cov-fail-under=88.

Calibrating the lens

src/av_perception/algorithm/calibration.py, run_calibration.py

A pinhole projection with Brown-Conrady distortion: a radial series in k1, k2, k3 and the two tangential terms p1, p2 of Brown (1966) and Conrady (1919). A 1280 by 720 camera with a mild barrel lens moves a point at the corner of the frame by more than a hundred pixels, and every geometric step downstream assumes that straight lines in the world are straight in the image, so this is the first thing that has to be right.

The inverse of that model has no closed form. The usual scheme, and the one OpenCV uses, is a fixed point iteration with a fixed step count; it converges linearly and leaves a residual of up to 0.37 pixels at the corner of this lens. Newton on the same equation, with the analytic Jacobian, reaches below 1e-12 pixels in about five steps, and it is what this repository uses. A solution at a negative radius is refused rather than returned, because past the turning point of the radial polynomial the equation has a root that reflects the image through the principal point, and that root has zero residual and is completely wrong.

The estimator is Zhang's method (2000) on planar chessboard views, refined by Levenberg-Marquardt, through cv2.calibrateCamera. Calibration is where correctness is hardest to check: on recorded images the only available score is reprojection error, and reprojection error is exactly what a wrong calibration preserves, because focal length and radial distortion trade against one another. That is why the target here is synthetic, and why two observation paths are reported rather than one.

projected corners
  source                       projected corners
  views                        12
  corners                      648
  image area covered           0.969
  RMS reprojection error [px]  0.00002
  worst view error [px]        0.00003
  fx error [px]                +0.00002
  fy error [px]                +0.00001
  cx error [px]                -0.00000
  cy error [px]                +0.00005
  focal relative error         1.396e-08
  k1 error                     -1.490e-07
  k2 error                     +8.640e-08
  p1 error                     -2.783e-08
  p2 error                     +4.381e-08
  k3 error                     +1.935e-07
  rectification disagreement [px] 0.00064
  radial map invertible        True

rendered images with corner detection
  source                       rendered images, 12 of 12 views detected
  views                        12
  corners                      648
  image area covered           0.969
  RMS reprojection error [px]  0.12916
  worst view error [px]        0.26396
  fx error [px]                -0.26570
  fy error [px]                -0.32468
  cx error [px]                -0.08707
  cy error [px]                +0.25631
  focal relative error         3.280e-04
  k1 error                     +4.233e-04
  k2 error                     -2.087e-03
  p1 error                     +4.914e-05
  p2 error                     -3.540e-05
  k3 error                     +2.113e-03
  rectification disagreement [px] 1.33243
  radial map invertible        True

The first path computes the corners with the forward model, which isolates the estimator: it returns the parameters it was given, the focal length to fourteen parts in a thousand million and every distortion coefficient to better than 2e-07. That is a statement about the implementation and not about calibration in general, and it is what makes any error here a bug rather than a measurement.

The second path renders images of the same board at the same poses and finds the corners with cv2.findChessboardCorners and cv2.cornerSubPix. The detector contributes about 0.12 pixels RMS, the recovered focal length moves by 0.27 pixels, three parts in ten thousand, and k2 and k3 move by two parts in a thousand. Those five coefficient errors are not individually interpretable, since the terms trade against each other, so they are folded into one number that a downstream stage would actually feel: rectifying a frame with the true coefficients and with the recovered ones puts the same pixel in two places 1.33 pixels apart at the worst point in the frame.

Alongside the errors the implementation reports the fraction of the frame the corners covered, 96.9 percent of an eight by eight grid here, and whether the recovered radial polynomial is still invertible out to the image corner. A calibration built entirely from the middle of the frame reports a small reprojection error and is still wrong at the edges, where distortion lives.

Calibrating where the camera is

src/av_perception/algorithm/extrinsics.py, src/av_perception/pipeline/ground_target.py, run_calibration.py

The lens is half of a calibration. The ground plane homography is K [r1 | r2 | t], so the camera height and its three angles set the scale and the shape of the bird's-eye view just as firmly as the focal length does. Until recently this repository assumed them, and said so in its limitations. It now measures them.

A chessboard lies on the road at a measured position and is read by the same corner detector the intrinsics use. The estimator is the planar pose problem run in the road frame: rectify the corners, solve the road to image homography by the normalised direct linear transform, decompose K^-1 H into [r1 | r2 | t] with the scale fixed by the unit length of the rotation columns and the sign fixed by requiring the target to be in front of the camera, project the result onto the rotation group by a singular value decomposition, and refine over all corners by Levenberg-Marquardt through the full distortion model.

placement from a rendered ground target, using the recovered lens
  source                       rendered image with corner detection
  corners                      45
  target span [px]             533 across, 73 along
  RMS reprojection error [px]  0.39089
  height [m]                   1.35293
  height error [mm]            +2.927
  pitch [deg]                  3.05667
  pitch error [mrad]           +0.9891
  yaw error [mrad]             -0.2819
  roll error [mrad]            +0.2234
  lateral error [mm]           -0.426
  longitudinal residual [mm]   +2.283
  road displacement [cm]       mean 21.99, max 62.17

From projected corners rather than a rendered image the same estimator returns the placement it was drawn from to solver precision, which is the same bug check the intrinsics get.

The target has to be large, and the target span line is why: 533 pixels across the road, 73 along it. Foreshortening compresses range, so at 4 m ahead a metre along the line of sight occupies about a fifth of the pixels a metre across the road does. A hand held board of the kind used for the intrinsics would be a few pixels deep on the ground and would constrain the height not at all. The published target is a 3.0 m by 1.8 m printed pattern with 0.30 m squares, which is a floor mounted target in an alignment bay rather than something carried in a boot.

longitudinal residual is not an accuracy figure. The road frame origin is defined as the point directly under the camera, so the recovered camera centre must return to it, and 2.3 mm is how close it came. That residual exists because nothing else can catch a particular failure: the corner detector reads a rectangular grid in one of two rigid orders, and reading the target through half a turn describes a real pose of a real board, reprojects perfectly, and is completely wrong. Goodness of fit cannot see it. The distance from the recovered centre to the origin can, and a test asserts exactly that.

road displacement folds the height and the three angle errors into one metric number the same way the rectification disagreement folds the five distortion coefficients: a road point projected through the true placement and read back through the recovered one lands 22 cm away on average and 62 cm away at worst over the 4 m to 30 m window, most of that at the far edge where a milliradian of pitch is worth decimetres.

The last block runs the lane pipeline three times over identical frames, so that the two calibrations can be separated and then charged together:

effect on the lane pipeline, recovered camera against true camera
  frames                       12
  accepted with the true rig   11
  recovered lens only
    accepted                   11
    lateral offset change [cm] mean 0.602, max 1.165
    lane width change [cm]     mean 1.405, max 1.930
  recovered lens and placement
    accepted                   11
    lateral offset change [cm] mean 1.028, max 2.120
    lane width change [cm]     mean 2.528, max 3.155

A rig calibrated entirely from its own targets, lens and placement alike, reports a lateral offset 1.03 cm from the one a perfectly known rig reports, and changes no accept or reject decision. What remains uncalibrated is stated under what it does not do.

Rectifying and mapping the road plane

src/av_perception/model/homography.py, src/av_perception/pipeline/runner.py, run_birdseye.py

Panels one and two of the figure at the top of this page are this stage. Every frame is rectified onto its own camera matrix, so the ground plane homography, which is built from that matrix, applies to the rectified image with no further adjustment.

A road point lies on the plane z = 0, so its camera frame coordinate is [r1 r2 t] [x, y, 1] and the image of the road plane is the homography H = K [r1 | r2 | t], fixed entirely by the intrinsics and the placement, both of which are now measured. Composing its inverse with a metric to raster similarity gives the inverse perspective mapping of Mallot et al. (1991). The window is defined in metres first, 4 m to 30 m ahead and 6 m either side at 0.05 m per pixel, and that definition is the pixel to metre scale, so nothing downstream needs a separate scaling constant.

image size                     1280 x 720
horizontal field of view [deg] 70.83
camera height [m]              1.350
camera pitch [deg]             3.000
horizon row                    312.33
bird's-eye window [m]          x 4.0 to 30.0, y -6.0 to 6.0
resolution [m/px]              0.050
bird's-eye raster              241 x 521
four point against analytic    3.798e-09
rectangle 3.70 m by 18.00 m maps to:
  width [px]                   74.0000  expected 74.0000
  length [px]                  360.0000  expected 360.0000
  corner angle [deg]           90.000000
observed cells                 119965 of 125561
observed fraction              0.9554
ego lane fully observed from [m] 4.00

A rectangle 3.70 m by 18.00 m on the road becomes a rectangle 74.0000 by 360.0000 pixels with corners at 90.000000 degrees. That is what metric by construction means: 74 pixels at 0.05 m per pixel is 3.70 m exactly, with no constant fitted anywhere.

The four point construction usually seen in lane detection code, where a trapezoid is picked by hand on a straight road and asserted to be some size, is also implemented, by the normalised direct linear transform of Hartley and Zisserman. It agrees with the analytic homography to 3.8e-09 and exists so the two can be compared, not because it is used: in the four point route every metre the pipeline later reports inherits the error in that assertion, and nothing downstream can detect it.

The near and far edges of the window are set by the camera rather than chosen. Closer than about 3.5 m the 2.5 m corridor either side of the vehicle has already left the frame, which is also why 95.5 percent of the window is observed and the missing 4.5 percent is the two near corners. Beyond 30 m a 0.12 m marking subtends fewer than 3.6 pixels and the thresholding stage loses it.

Finding the lane

src/av_perception/algorithm/threshold.py, src/av_perception/algorithm/lane_search.py, src/av_perception/algorithm/fitting.py, run_lane_detection.py

Panels three and four of the figure at the top are this stage. Thresholding runs on the rectified perspective image rather than on the warped one, because the warp resamples and resampling a 3.6 pixel wide marking at 30 m before differentiating it throws away the resolution the gradient operator needs. Three channels: high lightness for white paint, high saturation for yellow paint, and a Sobel gradient in the image x direction restricted to steep edges, which keeps marking edges and discards the horizon and shadows lying across the road. The binary result is warped into the bird's-eye grid.

Pixels are associated either by a stack of sliding windows started from a column histogram, in the acquisition mode, or by a corridor around the previous accepted fit, in the tracking mode. The fit is a second order polynomial of lateral position against longitudinal distance, in metres, so the coefficients need no rescaling afterwards: c is the lateral position at the vehicle, b is the tangent of the heading error, and 2a is the curvature. The common alternative, fitting in pixels and multiplying by a metres per pixel ratio afterwards, is algebraically identical and was rejected because the rescaling constants end up in a different file from the transform that defines them.

The two boundaries are fitted jointly, sharing one quadratic shape and carrying one offset each. That constraint is the definition of a lane, and it matters because support is asymmetric: the left boundary is solid and imaged along the whole window, the right is broken, and on about a quarter of frames its nearest mark is beyond ten metres. Both models are implemented and both are reported.

Lateral offset, heading and curvature over 40 frames with the estimate following the true value throughout, and a vertical band on the one frame drawn without markings, which is rejected and reacquired on the next frame

parallel model with tracking
  frames                       40
  accepted                     39
  detection rate               0.9750
  lateral offset RMSE [cm]     1.07
  lateral offset max [cm]      3.65
  lateral offset bias [cm]     +0.13
  heading RMSE [mrad]          1.456
  curvature RMSE [1/km]        0.0911
  curvature max error [1/km]   0.2392
  radius relative RMSE         0.0927
  mean lane width [m]          3.6968
  lane width RMSE [cm]         0.63
  mean fit residual [cm]       5.60
  window searches              3
  corridor searches            37
  tracking fallbacks           1
  mean frame time [ms]         36.67

independent model with tracking
  frames                       40
  accepted                     39
  detection rate               0.9750
  lateral offset RMSE [cm]     10.13
  lateral offset max [cm]      22.47
  lateral offset bias [cm]     +6.64
  heading RMSE [mrad]          12.865
  curvature RMSE [1/km]        0.7388
  curvature max error [1/km]   1.4917
  radius relative RMSE         0.9146
  mean lane width [m]          3.6857
  lane width RMSE [cm]         1.99
  mean fit residual [cm]       5.58

parallel model, full search every frame
  frames                       40
  accepted                     39
  detection rate               0.9750
  lateral offset RMSE [cm]     1.13
  lateral offset max [cm]      3.66
  lateral offset bias [cm]     +0.08
  heading RMSE [mrad]          1.747
  curvature RMSE [1/km]        0.1033
  curvature max error [1/km]   0.3628
  radius relative RMSE         0.0975
  mean lane width [m]          3.6996
  lane width RMSE [cm]         0.95
  window searches              40
  corridor searches            0
  tracking fallbacks           0

rejected frames                1
  left support 0 below 200 pixels
  right support 0 below 200 pixels
  left span 0.0 m below 6.0 m
  right span 0.0 m below 6.0 m
tracking against full search   max offset difference 1.13 cm over 39 frames

The lateral offset is recovered to 1.07 cm RMS with a bias of 1.3 mm, over a true offset that sweeps 0.70 m. Lane width comes back as 3.6968 m against a true 3.7000 m, a systematic 3 mm narrow, which is the separation of the fitted line centres rather than of the paint. Heading is recovered to 1.46 mrad, which is 0.08 degrees.

Curvature is the weakest output and is reported as curvature rather than as radius. The RMS error is 0.091 per kilometre against a true amplitude of 2.22 per kilometre, and the relative radius error is 9.3 percent. That is a property of the measurement rather than of this implementation: a second order fit over a 26 m window is estimating the coefficient of s^2, and at a 450 m radius the whole lateral excursion the fit has to work with is 0.75 m. Any single frame monocular curvature estimate over this range behaves this way, which is why production systems filter it across frames and why this one does not, since a filtered number would report the filter's ability to average rather than the detector's ability to measure.

The three variants differ in exactly one respect each. The first two isolate the fitting model: sharing one quadratic shape between the boundaries reduces the lateral offset error by a factor of 9.5 and the relative radius error by a factor of 9.9, because the broken right boundary no longer has to estimate its own shape from two distant marks. The first and the third isolate the search: the corridor and the full histogram agree to 1.13 cm at worst over the 39 frames both accept, so tracking changes the cost of the answer and not the answer.

The vertical band in the figure is the frame drawn without markings, and it is rejected for the right reason: no support on either side. The tracking corridor found nothing, the full search was rerun on the same frame as a fallback and also found nothing, the tracking state was discarded, and the next frame reacquired from the histogram and was accurate again. Acceptance needs enough pixels over a long enough span on each side, a lane width between 2.6 m and 4.6 m, a small fit residual, and a curvature radius above 60 m. The mean frame time is the only machine dependent number on this page and excludes scene generation.

Finding the free space

src/av_perception/algorithm/obstacles.py, run_obstacle_extraction.py

The lane says where to go. It does not say whether something is standing in the way. The bird's-eye view assumes everything it shows lies on the road plane, and anything standing above the plane breaks that assumption in a specific and predictable way: the ray through the top of an object strikes the ground further away than the object stands, so the object smears outwards along the line of sight, beginning exactly at its ground contact edge nearest the camera.

Bird's-eye free space decision in which every obstacle is smeared along the line of sight to the far edge of the window, while the box drawn at its near edge sits on the true footprint

The left panel is the bird's-eye colour image and the right panel is the decision made from it: unknown, free, obstacle, with the detected near edge in red and the true footprint in dashed yellow. The two long yellow wedges running to the top of the frame are the smear, and they are not an error to be fixed. An object of height h seen at range r from a camera at height H occupies the mapped ground out to r H / (H - h). The extractor therefore reports the near edge separately from the full component, scores accuracy against the near edge only, and reports the smear without scoring it.

frames                         40
bird's-eye window [m2]         312.0
observed area [m2]             299.9
free area [m2]                 270.1
obstacle area [m2]             29.8
free fraction of observed      0.9008

true obstacle instances        80
detections                     80
matched                        80
missed                         0
spurious                       0
recall                         1.0000
precision                      1.0000

near edge range RMSE [cm]      7.27
near edge range bias [cm]      -1.94
near edge range max [cm]       35.33
lateral centre RMSE [cm]       8.77
lateral width bias [cm]        +13.06

predicted smear beyond the near edge, from the flat world assumption:
  mean [m]                     7.08
  max [m]                      12.08
obstacle height [m]            0.55, 0.45
camera height [m]              1.35

Across 80 obstacle instances there is no miss and no spurious detection. The near edge range is recovered to 7.3 cm RMS with a bias of 1.9 cm short, and the lateral centre of the near edge to 8.8 cm. The near edge width is overestimated by 13.1 cm, about two and a half bird's-eye cells, which is the morphological closing and the antialiased silhouette of the box adding a cell on each side.

The obstacle area of 29.8 m2 is more than eight times the 3.6 m2 the two footprints actually occupy, and the predicted smear of 7.08 m on average and 12.08 m at worst is exactly that difference. An object taller than the camera has no ray through its top that ever meets the road, and its smear runs to the horizon. The remaining 10 percent of the observed window that is not called free is entirely this smear: on a frame with no obstacles the segmentation returns no components at all.

The code is laid out in the order this page reads. model holds pure functions and dataclasses with no I/O and no state, algorithm holds the decisions and reads no files and draws no random numbers, pipeline is the only place a random number is drawn or a file is read, analysis reads traces and produces numbers and figures, and the example scripts contain wiring and printing and no logic that is not tested elsewhere. The dependency direction never runs backwards.

Two defects the evaluation exposed

Both of these were found by running the pipeline against ground truth, not by reading the code, and both are recorded because they are the evidence that the evaluation was real.

The yellow lane line was a 26 metre obstacle. The free space test asks whether a pixel is achromatic and not dark. Yellow paint at high saturation is neither, so the solid left boundary came back as a connected component 26 m long lying down the middle of the drivable area, blocking the lane it was defining. The fix is the second clause of the appearance model, which admits yellow at high saturation and moderate lightness as road surface. It is not decoration and it is the reason the model has two clauses instead of one.

An adjacent lane vehicle biased the lateral offset by 17 cm. During development an obstacle whose silhouette came within 0.13 m of the left boundary at long range was claimed by the sliding window and dragged the fit with it. The sliding window half width of 0.55 m is what sets the distance at which this happens, and the parallelism and residual tests are what catch it when it does. The generated scenes now place traffic in the adjacent lanes at no less than 0.90 m clearance from the ego lane markings on every frame regardless of curvature, so that the lane accuracy figures above measure the lane pipeline rather than an interaction. The interaction itself is real and is documented rather than removed. A vehicle in the ego lane, which is the case that matters most, would occlude the boundaries outright and the frame would be rejected for lack of support.

How it is checked

uv run pytest
uv run ruff check .
uv run ruff format --check .
uv run mypy

195 tests run in about 36 seconds, dominated by the frames the scene generator has to render, and in about 46 seconds with coverage measurement on. They come in three tiers.

The first tier is property and invariant tests over the mathematics. Projecting a camera frame point and unprojecting it returns the point to 1e-09. Distortion applied and then removed is the identity to 1e-12 in normalised units. The analytic distortion Jacobian matches a central difference. The forward model agrees with cv2.projectPoints and the inverse leaves a smaller residual than cv2.undistortPoints. A non-invertible lens raises rather than returning a reflected root. The direct linear transform matches cv2.getPerspectiveTransform and rejects collinear correspondences. The bird's-eye homography maps a 3.7 m by 18 m road rectangle to a 74 by 360 pixel axis aligned rectangle. Calibration on the synthetic board recovers the known focal length to 0.01 pixels and every distortion coefficient to 1e-04, and degrades in proportion to added corner noise. The extrinsic estimator returns the placement it was drawn from to 1e-09 for yawed and rolled rigs as well as level ones, inverts the rig parameterisation exactly, repairs a corner reading that is half a turn out, refuses one that is mirrored, and recovers a rendered target to within a centimetre of height and 3 mrad of every angle. A straight lane yields a curvature radius above 10 000 m and a known lateral offset within 2 cm. A lane of known constant radius yields that radius within 5 percent. The corridor search and the window search claim the same pixels and produce lateral offsets within 3 cm on the same frames. The parallel model beats the independent one on a boundary visible only in two distant marks. A frame drawn without markings is rejected for lack of support while the next frame reacquires and is accurate again. And the loop between the two mappings closes: the yellow line drawn by forward projection lands on the true arc after distortion correction and inverse perspective mapping, to 0.5 cm mean and 5 cm worst case.

The second tier replays a recorded 14 frame run against tests/data/reference_run.json; regenerate it with uv run python tests/generate_reference.py when a change to the algorithms is intended. What it pins, and what it does not, is deliberate. Search modes, accept and reject decisions, obstacle counts, and detection counts are pinned exactly, because they are decisions rather than measurements. Lane offset and lane width are pinned in centimetres to half a bird's-eye cell, obstacle near edge range to five cells, and the calibration to 1e-03 pixels, because those come out of resampled images and cv2.remap and cv2.warpPerspective are entitled to differ in the last bit between OpenCV builds and platforms, which can move a lane pixel across a threshold. The count of thresholded pixels, the raw polynomial coefficients, and the curvature radius are not pinned at all: the first two are image statistics that move with the last bit of an interpolation, and the third is the reciprocal of the estimated quantity and diverges on a straight road, so the signed curvature is pinned instead and a separate test asserts the qualitative property that a near straight frame reports a radius above 2000 m.

The third tier runs every script in examples/ as a subprocess under reduced frame and view counts, including the figure writing paths, the fallback from an absent frame directory to the synthetic sequence, and the figure budget. A fourth handful of tests covers what the repository ships rather than what it computes: that py.typed is present, empty, and inside the importable package, that the wheel is built from the directory holding it, that the three published figures exist and fit the budget, and that the README shows all of them.

CI runs the suite, the linter, and the type checker on Ubuntu and on Windows, with --cov-fail-under=88.

What it does not do

The full list, with the reasoning, is in docs/design-notes.md, which also records the alternatives that were rejected, including the Hough transform, RANSAC, clothoid road models, recursive filtering across frames, stereo free space, and learned lane detectors, and one limitation that has since been closed. The four that matter most here:

  • Fixed thresholds are brittle to lighting. Lightness at least 170, saturation at least 90, gradient at least 40 after scaling. A hard shadow drops the paint under it below threshold, a wet surface in low sun fills the mask, and worn paint on light concrete never reaches either colour threshold. This is the gap between the accuracy figures above and a real road.
  • Free space is decided by appearance. A shadow, a patch of new asphalt, or a wet stain will be called an obstacle. From one monocular frame a black object standing on the road and a black shadow lying on it produce the same pixels, so there is no fix here that is not a different sensor.
  • The placement is calibrated once, not continuously. A vehicle pitches under braking, the suspension moves with load, and the road has crest and sag curves. The calibration answers where the camera was on the day, not where it is on this frame.
  • The scene generator does not produce the failures that matter. Clean paint, uniform lighting, no shadows, no weather, no occlusion of the ego lane boundaries. The geometric results carry over to a real road because they do not depend on appearance; the detection rate does not.

References

Camera model and calibration

  • Brown, D. C. "Decentering Distortion of Lenses". Photogrammetric Engineering, 32(3):444 to 462, 1966. https://www.asprs.org/wp-content/uploads/pers/1966journal/may/1966_may_444-462.pdf. The radial and decentering distortion model implemented here.
  • Conrady, A. E. "Decentred Lens-Systems". Monthly Notices of the Royal Astronomical Society, 79(5):384 to 390, 1919. DOI: 10.1093/mnras/79.5.384. The tangential terms.
  • Zhang, Z. "A Flexible New Technique for Camera Calibration". IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(11):1330 to 1334, 2000. DOI: 10.1109/34.888718. The calibration method: planar homographies, the closed form intrinsics, and the nonlinear refinement. The same planar pose decomposition, run in the road frame, is the extrinsic estimator.
  • Levenberg, K. "A Method for the Solution of Certain Non-Linear Problems in Least Squares". Quarterly of Applied Mathematics, 2(2):164 to 168, 1944. DOI: 10.1090/qam/10666.
  • Marquardt, D. W. "An Algorithm for Least-Squares Estimation of Nonlinear Parameters". Journal of the Society for Industrial and Applied Mathematics, 11(2):431 to 441, 1963. DOI: 10.1137/0111030. The refinement both calibration solvers run.
  • Hartley, R. and Zisserman, A. "Multiple View Geometry in Computer Vision", second edition, Cambridge University Press, 2004. DOI: 10.1017/CBO9780511811685. Chapter 4 for the direct linear transform and the plane to plane homography.
  • Hartley, R. "In Defense of the Eight-Point Algorithm". IEEE Transactions on Pattern Analysis and Machine Intelligence, 19(6):580 to 593, 1997. DOI: 10.1109/34.601246. The isotropic normalisation applied before the direct linear transform.
  • Rodrigues, O. "Des lois geometriques qui regissent les deplacements d'un systeme solide dans l'espace". Journal de Mathematiques Pures et Appliquees, 5:380 to 440, 1840. https://eudml.org/doc/234443. The axis-angle to rotation matrix formula used for board poses.

Inverse perspective mapping and lane detection

  • Mallot, H. A., Bulthoff, H. H., Little, J. J., and Bohrer, S. "Inverse Perspective Mapping Simplifies Optical Flow Computation and Obstacle Detection". Biological Cybernetics, 64(3):177 to 185, 1991. DOI: 10.1007/BF00201978. The inverse perspective mapping, and the observation that objects above the road plane smear along the line of sight.
  • Bertozzi, M. and Broggi, A. "GOLD: A Parallel Real-Time Stereo Vision System for Generic Obstacle and Lane Detection". IEEE Transactions on Image Processing, 7(1):62 to 81, 1998. DOI: 10.1109/83.650851. Lane and obstacle detection in a remapped bird's-eye view.
  • Bertozzi, M., Broggi, A., and Fascioli, A. "Stereo Inverse Perspective Mapping: Theory and Applications". Image and Vision Computing, 16(8):585 to 590, 1998. DOI: 10.1016/S0262-8856(97)00093-0.
  • Aly, M. "Real Time Detection of Lane Markers in Urban Streets". In IEEE Intelligent Vehicles Symposium, 2008, pp. 7 to 12. DOI: 10.1109/IVS.2008.4621152. Thresholding, grouping, and fitting of lane candidates in a bird's-eye view.
  • Dickmanns, E. D. and Mysliwetz, B. D. "Recursive 3-D Road and Relative Ego-State Recognition". IEEE Transactions on Pattern Analysis and Machine Intelligence, 14(2):199 to 213, 1992. DOI: 10.1109/34.121789. The parallel boundary road model and its recursive estimation.
  • McCall, J. C. and Trivedi, M. M. "Video-Based Lane Estimation and Tracking for Driver Assistance: Survey, System, and Evaluation". IEEE Transactions on Intelligent Transportation Systems, 7(1):20 to 37, 2006. DOI: 10.1109/TITS.2006.869595. Survey of the classical stack and of how lane estimates are evaluated.
  • Hillel, A. B., Lerner, R., Levi, D., and Raz, G. "Recent Progress in Road and Lane Detection: A Survey". Machine Vision and Applications, 25(3):727 to 745, 2014. DOI: 10.1007/s00138-011-0404-2.

Image processing

  • Sobel, I. and Feldman, G. "A 3x3 Isotropic Gradient Operator for Image Processing". Presented at the Stanford Artificial Intelligence Project, 1968, and reproduced in Duda, R. O. and Hart, P. E. "Pattern Classification and Scene Analysis", Wiley, 1973. https://www.researchgate.net/publication/239398674. The gradient operator used by the thresholding stage.
  • Rosenfeld, A. and Pfaltz, J. L. "Sequential Operations in Digital Picture Processing". Journal of the ACM, 13(4):471 to 494, 1966. DOI: 10.1145/321356.321357. Connected component labelling.
  • Bolelli, F., Allegretti, S., Baraldi, L., and Grana, C. "Spaghetti Labeling: Directed Acyclic Graphs for Block-Based Connected Components Labeling". IEEE Transactions on Image Processing, 29:1999 to 2012, 2020. DOI: 10.1109/TIP.2019.2946979. The algorithm behind cv2.connectedComponentsWithStats.
  • Serra, J. "Image Analysis and Mathematical Morphology". Academic Press, 1982. https://shop.elsevier.com/books/image-analysis-and-mathematical-morphology/serra/978-0-12-637240-3. Opening and closing, used to clean the obstacle mask.
  • Smith, A. R. "Color Gamut Transform Pairs". ACM SIGGRAPH Computer Graphics, 12(3):12 to 19, 1978. DOI: 10.1145/965139.807361. The hue, lightness, saturation space the colour thresholds are defined in.

Road geometry standards

  • Federal Highway Administration. "Manual on Uniform Traffic Control Devices for Streets and Highways", 11th edition, 2023. https://mutcd.fhwa.dot.gov/. Part 3 for the yellow line on the left of a one-way carriageway, the white line on the right, and the broken line pattern of a 3 m mark on a 12 m cycle used by the scene generator.
  • American Association of State Highway and Transportation Officials. "A Policy on Geometric Design of Highways and Streets", 7th edition, 2018. https://store.transportation.org/item/collectiondetail/180. Lane widths of 2.7 m to 3.6 m and the minimum curve radii that set the acceptance band on curvature.

Alternatives considered

  • Duda, R. O. and Hart, P. E. "Use of the Hough Transformation to Detect Lines and Curves in Pictures". Communications of the ACM, 15(1):11 to 15, 1972. DOI: 10.1145/361237.361242.
  • Fischler, M. A. and Bolles, R. C. "Random Sample Consensus: A Paradigm for Model Fitting with Applications to Image Analysis and Automated Cartography". Communications of the ACM, 24(6):381 to 395, 1981. DOI: 10.1145/358669.358692.
  • Southall, B. and Taylor, C. J. "Stochastic Road Shape Estimation". In IEEE International Conference on Computer Vision, 2001, pp. 205 to 212. DOI: 10.1109/ICCV.2001.937519.
  • Badino, H., Franke, U., and Pfeiffer, D. "The Stixel World: A Compact Medium Level Representation of the 3D-World". In Pattern Recognition, DAGM 2009, Lecture Notes in Computer Science volume 5748, pp. 51 to 60. DOI: 10.1007/978-3-642-03798-6_6.
  • Pan, X., Shi, J., Luo, P., Wang, X., and Tang, X. "Spatial As Deep: Spatial CNN for Traffic Scene Understanding". In AAAI Conference on Artificial Intelligence, 32(1), 2018. DOI: 10.1609/aaai.v32i1.12301.
  • Neven, D., De Brabandere, B., Georgoulis, S., Proesmans, M., and Van Gool, L. "Towards End-to-End Lane Detection: an Instance Segmentation Approach". In IEEE Intelligent Vehicles Symposium, 2018, pp. 286 to 291. DOI: 10.1109/IVS.2018.8500547.
  • Tabelini, L., Berriel, R., Paixao, T. M., Badue, C., De Souza, A. F., and Oliveira-Santos, T. "Keep Your Eyes on the Lane: Real-Time Attention-Guided Lane Detection". In IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp. 294 to 302. DOI: 10.1109/CVPR46437.2021.00036.

Dependencies

Package Version Purpose Licence
numpy >= 2.0 Array storage, least squares and singular value decomposition, seeded random number generation BSD-3-Clause
opencv-python-headless >= 4.10 Calibration solver, chessboard corner detection and subpixel refinement, pose refinement, image remapping and perspective warping, colour conversion, Sobel filtering, morphology, connected components Apache-2.0 for the OpenCV library, MIT for the Python packaging
matplotlib >= 3.9 Calibration, pipeline stage, lane geometry, and obstacle figures Matplotlib licence, a BSD-compatible licence derived from the Python Software Foundation licence
pytest >= 8.3 Test runner, development only MIT
pytest-cov >= 6.0 Coverage measurement, development only MIT
ruff >= 0.8 Linter, development only MIT
mypy >= 1.13 Static type checker, development only MIT

Citations for the runtime dependencies:

  • Harris, C. R. et al. "Array Programming with NumPy". Nature, 585:357 to 362, 2020. DOI: 10.1038/s41586-020-2649-2.
  • Bradski, G. "The OpenCV Library". Dr. Dobb's Journal of Software Tools, 2000. https://opencv.org/.
  • Hunter, J. D. "Matplotlib: A 2D Graphics Environment". Computing in Science and Engineering, 9(3):90 to 95, 2007. DOI: 10.1109/MCSE.2007.55.

License

Released under the MIT license. See LICENSE.

About

Lane detection and obstacle extraction with camera calibration and a bird's-eye transform stage.

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages