Segment aerial line-transect marine mammal survey data for distance sampling and density surface modeling.
Takes survey data in the format of the North Atlantic Right Whale Consortium
(NARWC) sightings database and produces effort segments — each with a location,
an amount of effort, and counts of what was seen along it — ready for Distance
and dsm.
The input format is specified by the NARWC users' guide: Kenney, R.D. (2023) The North Atlantic Right Whale Consortium Database: A Guide for Users and Contributors, Version 8. NARWC Reference Document 2023-01. Read the handbook (PDF) · database landing page
Every code book, column definition, and effort rule in this package cites a section of it. Chapter 8 sections are numbered alphabetically by variable, so they shift between versions —
distsampcites Version 8 throughout. See docs/06-references.md for what each section is relied on for.
# install.packages("remotes")
remotes::install_github("chross22/distsamp")The bundled example is mock data.
narwc-example.csvis a synthetic two-day survey generated bydata-raw/make-fixture.R, modelled on the hypothetical survey in Figure 2 of the handbook. No real NARWC records are distributed with this package, and none have been used to develop it — every number below is invented. Right whale sighting data is not ours to redistribute; request it from the NARWC.
library(distsamp)
# Synthetic example data - see the note above
path <- system.file("extdata", "narwc-example.csv", package = "distsamp")
dat <- read_narwc(path) # standardise names and types
validate_narwc(dat) # report problems, never stop
segs <- segment_survey(dat, seg_length = 5, seed = 1)
segs
#> <distsamp_segments>
#> segments: 20
#> tracks: 7
#> total effort: 92.23 km
#> segment length: median 4.44 km, range 2.22-7.78 km
#> target length: 5 km seed: 1
#> species: FIWH, RIWH, SEWH
segments_wide(segs) # one count column per species, for dsm
segs$detections # one row per sighting, with perpendicular distanceTwo vignettes:
vignette("segmenting-narwc-data", package = "distsamp") # the full walkthrough
vignette("from-segments-to-density", package = "distsamp") # handing off to Distance and dsmread_narwc -> validate_narwc
|
make_leg_id -> flag_circling -> flag_effort -> point_to_point_effort
|
split_tracks -> track_effort -> plan_segments -> cut_segments
|
attach_circling_sightings -> segment_midpoints -> segment_sightings
Where the survey recorded ANGLEL/ANGLER declination angles, perpendicular
distances are computed too, and segs$detections comes out in the shape a
detection function wants.
segment_survey() runs all of it. Every stage is also exported, so you can
intervene between them.
Reading a file is not the same as being ready to segment one. Four steps sit in between, and their order matters in ways that are quiet when you get them wrong:
dat <- narwcr::read_narwc("extract.csv")
air <- prepare_aerial(dat)
segs <- segment_survey(air, seg_length = 10, seed = 1)prepare_aerial() builds line occupations, reconstructs the line state where
the file recorded it only at change points, classifies the platform by speed,
filters to the aerial records, and flags effort — in that order. Every step is
also exported, so you can run them by hand and intervene between them; the
wrapper exists because the order is not obvious and getting it wrong costs
records silently. Its help says which step depends on which, and why.
Two things it deliberately leaves to you, because they are assertions about a particular file rather than facts about NARWC data: mapping a declination angle out of a column the handbook does not name, and correcting an altitude recorded in feet. Both are covered in docs/08-onboarding-a-real-extract.md.
The failure this package guards against is not the one that errors. It is the one that returns a plausible number: effort that includes the ferry between two survey days, an altitude read as metres when the column was recorded in feet, a shipboard survey segmented as though it were flown.
diagnose_pipeline(air, days = "auto")
plot_survey(air, "occupations")runs the whole pipeline against one representative survey day and reports what
it finds — merged days, an implausible altitude, a platform moving at ten
knots, which criterion excluded which records, and how the recorded
along-track distance compares with the reconstructed one. days also takes a
count or specific dates.
plot_survey() is the same question as a picture: it draws the survey at
whatever stage it has reached, and a survey line drawn in three pieces — or
three drawn as one — is visible on a map and invisible in a table. "positions"
is the view that needs nothing but coordinates, so it draws even on a file none
of this has touched; add sightings = TRUE and it shows where the aircraft went
and what was seen from it, in one map.
A decade of survey on one map is a smear, so every plotting function takes
dates, years, and months — years = 2019, months = 8 is August 2019.
filter_days() is that selection on its own.
The scales are Okabe-Ito, which stays legible to colourblind readers, falling
back to viridis past eight levels. Legends are capped: past max_legend groups,
tracks and occupations recycle a small palette and drop the legend, and species
beyond the cap are gathered into one "other" entry rather than losing their
colour. All three of these were real failures on a real archive — a 130-entry
legend that crushed the map, and 28 species drawn with no colour at all.
Each of these checks exists because a real extract produced a believable wrong answer. docs/08-onboarding-a-real-extract.md is the order to bring a new dataset in, and the traps in it are quoted with the magnitudes they actually caused.
The effort criteria are the handbook's aerial ones and perp_distance() is a
declination angle from an aircraft, so a shipboard survey does not belong here —
but a NARWC extract can hold both with nothing in the data to say so.
segment_survey() checks the platform's speed and refuses to segment a vessel
track silently. Split with narwcr::classify_platform(), after
make_leg_id(): filtering before it merges two occupations of one line and
counts the ferry between them as survey effort.
Segments are cut following the approach of Becker et al. (2010): divide continuous portions of survey effort into segments of approximately equal length, assign sightings to the segment they fall in, and take habitat covariates at each segment's mid-point. Concretely, fit as many whole segments of the target length as a continuous track allows, then either give the remainder its own segment or absorb it into a randomly chosen one, depending on whether it clears half the target length.
The remainder handling is a property of Becker's segchopr implementation rather
than of the published methods, which state the target length but not what happens
to the leftover — see docs/06-references.md
if you are writing this up.
Effort criteria, sighting filters, and validation all read from the code books in
Kenney (2023), The North Atlantic Right Whale Consortium Database: A Guide for
Users and Contributors, Version 8 — LEGTYPE (8.A.21), LEGSTAGE (8.A.20),
IDREL (8.A.16), VISIBLTY (8.A.38), and the rest. narwc_codes() prints them.
A few consequences worth knowing:
VISIBLTYcarries two encodings. Since 2004 it is a distance in nautical miles; before that it was a code, folded into the same field as a negative number, where-1means clear for at least 2 nautical miles. A plainVISIBLTY >= 2test throws away every legacy record.visibility_ok()handles both.- Pilot sightings do not count.
LEGSTAGE == 6marks a sighting by someone other than an on-duty observer, which handbook 4.2 says "cannot be included in a density estimate". Excluded by default, along withLEGSTAGE == 7(photographic) andIDREL1 and 9. - Declination angles give perpendicular distance.
ANGLEL/ANGLERare angles below the horizon, taken when the sighting is abeam (8.A.2), so distance isALT / tan(angle). They replacedSTRIPin 2022 because survey altitudes had to rise for offshore wind, andSTRIP's fixed intervals shift with altitude while an angle does not. - Circling sightings do count, conditionally. A group sighted from the track and then circled has its position recorded off effort. Those are attached back to the segment they came from, following the CETAP rule that only further groups of the same species count with the original.
Segmentation makes two random choices. Pass a seed and a run repeats exactly;
your session's own random stream is never disturbed.
identical(
segment_survey(dat, seg_length = 5, seed = 1)$segments,
segment_survey(dat, seg_length = 5, seed = 1)$segments
)
#> TRUEdist_methods()"becker" (Becker's segchopr) and "kenney" (Kenney and Winn 1986) are both
selectable — and are the same formula, so they return identical values. The
default "haversine" uses the same sphere but is numerically stable at the short
ranges between consecutive survey records, where the law of cosines loses
precision.
docs/ covers how this was built:
the plan,
implementation notes,
defects found in the original code,
verification,
next steps,
references — including how to cite this in a methods
section — where fitting lives, and
onboarding a real extract.
This package was rewritten from a set of internal research scripts. Those are
not distributed — they contain collaborators' working notes and local paths.
docs/ records what was carried over, what was fixed, and why.
citation("distsamp")That returns three entries — the package, the segmentation method it implements, and the data format — because a methods section usually needs all three:
Ross, C. distsamp: Segment Aerial Line-Transect Survey Data for Distance
Sampling. R package. https://github.com/chross22/distsamp
Becker, E.A., Forney, K.A., Ferguson, M.C., Foley, D.G., Smith, R.C., Barlow, J.
and Redfern, J.V. (2010) Comparing California Current cetacean-habitat models
developed using in situ and remotely sensed sea surface temperature data. Marine
Ecology Progress Series 413:163-183. doi:10.3354/meps08696
Kenney, R.D. (2023) The North Atlantic Right Whale Consortium Database: A Guide
for Users and Contributors, Version 8. NARWC Reference Document 2023-01.
The package version comes from DESCRIPTION at install time, so it is always the
version you actually have.
Record the seed you used — without it the segmentation is not reproducible.
Suggested methods-section wording is in docs/06-references.md, along with a note on which parts of the algorithm are attributable to the published methods and which are properties of Becker's implementation.
Alphabetical. What each source is relied on for is set out in
docs/06-references.md. These are checked monthly by
CI — DOIs still resolve and still describe the paper we cite, hosted PDFs are
still there, and the NARWC handbook is still at the version we cite. Run it
yourself with Rscript tools/check-citations.R.
Becker, E.A., Forney, K.A., Ferguson, M.C., Foley, D.G., Smith, R.C., Barlow, J. and Redfern, J.V. (2010) Comparing California Current cetacean–habitat models developed using in situ and remotely sensed sea surface temperature data. Marine Ecology Progress Series 413:163–183. https://doi.org/10.3354/meps08696 — the segmentation method.
Becker, E.A., Forney, K.A., Redfern, J.V., Barlow, J., Jacox, M.G., Roberts, J.J. and Palacios, D.M. (2019) Predicting cetacean abundance and distribution in a changing climate. Diversity and Distributions 25:626–643. https://doi.org/10.1111/ddi.12867
Becker, E.A., Forney, K.A., Thayre, B.J., Debich, A.J., Campbell, G.S., Whitaker, K., Douglas, A.B., Gilles, A., Hoopes, R. and Hildebrand, J.A. (2017) Habitat-based density models for three cetacean species off southern California illustrate pronounced seasonal differences. Frontiers in Marine Science 4:121. https://doi.org/10.3389/fmars.2017.00121
Buckland, S.T., Anderson, D.R., Burnham, K.P., Laake, J.L., Borchers, D.L. and Thomas, L. (2001) Introduction to Distance Sampling: Estimating Abundance of Biological Populations. Oxford University Press, New York, NY.
CETAP (1982) A Characterization of Marine Mammals and Turtles in the Mid- and North-Atlantic Areas of the U.S. Outer Continental Shelf, Final Report. Cetacean and Turtle Assessment Program, University of Rhode Island. Bureau of Land Management, Washington, DC. — the survey protocol behind the effort defaults.
Hedley, S.L. and Buckland, S.T. (2004) Spatial models for line transect sampling. Journal of Agricultural, Biological, and Environmental Statistics 9:181–199. https://doi.org/10.1198/1085711043578
Kenney, R.D. (2002) Quality-control Issues for Data Submissions to the North Atlantic Right Whale Consortium Database. NARWC Reference Document 2002-02. University of Rhode Island, Graduate School of Oceanography, Narragansett, RI.
Kenney, R.D. (2023) The North Atlantic Right Whale Consortium Database: A Guide for Users and Contributors, Version 8. NARWC Reference Document 2023-01. University of Rhode Island, Graduate School of Oceanography, Narragansett, RI. https://www.narwc.org/sightings-database.html — the input format.
Kenney, R.D. and Scott, G.P. (1981) Calibration of the Beechcraft AT-11 forward observation bubble for population estimation purposes. Pp. III.1–III.11 in CETAP, A Characterization of Marine Mammals and Turtles in the Mid- and North-Atlantic Areas of the U.S. Outer Continental Shelf, Annual Report for 1979. Bureau of Land Management, Washington, DC.
Kenney, R.D. and Shoop, C.R. (2012) Aerial surveys for marine turtles. Pp. 264–271 in R.W. McDiarmid, M.S. Foster, C. Guyer, J.W. Gibbons and N. Chernoff, eds. Reptile Biodiversity: Standard Methods for Inventory and Monitoring. University of California Press, Berkeley, CA.
Kenney, R.D. and Winn, H.E. (1986) Cetacean high-use habitats of the northeast
United States continental shelf. Fishery Bulletin 84(2):345–357. — the
"kenney" distance formula, p. 347, and the CETAP on-effort criteria.
Miller, D.L., Burt, M.L., Rexstad, E.A. and Thomas, L. (2013) Spatial models for distance sampling data: recent developments and future directions. Methods in Ecology and Evolution 4:1001–1010. https://doi.org/10.1111/2041-210X.12105
Miller, D.L., Rexstad, E., Thomas, L., Marshall, L. and Laake, J.L. (2019) Distance sampling in R. Journal of Statistical Software 89(1):1–28. https://doi.org/10.18637/jss.v089.i01
Pebesma, E. (2018) Simple features for R: standardized support for spatial vector data. The R Journal 10(1):439–446. https://doi.org/10.32614/RJ-2018-009
R Core Team (2026) R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria. https://www.R-project.org/
Shoemake, K. (1985) Animating rotation with quaternion curves. ACM SIGGRAPH Computer Graphics 19(3):245–254. https://doi.org/10.1145/325165.325242
Sinnott, R.W. (1984) Virtues of the haversine. Sky and Telescope 68(2):159.