Skip to content

n_scientificName counts a scientific name once per deployment instead of once per group #432

Description

@northfox

Summary

summarize_observations() computes n_scientificName per deployment and then sums these values across deployments. When group_by does not include deploymentID, a scientific name observed in several deployments is therefore counted several times.

This contradicts the documentation of n_scientificName:

If scientificName is in group_by, n_scientificName is equal to 1 or 0, if scientificName = NA (unidentified animals).

and the expected output described in #367 ("this will be 1 for all if scientificName is selected in group_by").

Reproducible example

library(camtraptor)
library(dplyr, warn.conflicts = FALSE)

x <- example_dataset()
summarize_observations(x, group_by = "scientificName") %>%
  select(scientificName, n_scientificName) %>%
  filter(n_scientificName > 1)
#> # A tibble: 2 × 2
#> # Groups:   scientificName [2]
#>   scientificName     n_scientificName
#>   <chr>                         <int>
#> 1 Anas platyrhynchos                2
#> 2 Aves                              2

Expected: 1 for both, since each row represents a single scientific name. The value 2 is the number of deployments in which the name was observed. The same over-counting can occur for any grouping that combines several deployments when the same scientific name occurs in more than one of them, e.g. group_by = "locationName" when a location has more than one deployment, or group_by = c("latitude", "longitude").

The default group_by includes deploymentID, so the default output and the deprecated get_n_species() are not affected. n_events, n_observations and sum_count are additive across deployments and remain correct. The rai_* values are calculated from the aggregated counts and effort and are unaffected.

Cause

calc_obs_feature() uses n_distinct(scientificName) per deployment as formula_per_deployment and sum(n_scientificName) as formula_total. Distinct counts are not additive across deployments.

Proposed fix

Count distinct scientific names at the final output-group level (group_by columns + group_time_by) instead of summing per-deployment distinct counts, e.g. directly from the enriched observations or by retaining the distinct names until the final aggregation. NA would still be excluded, so groups with only unidentified animals keep n_scientificName = 0.

I'd be happy to open a PR with a regression test through summarize_observations() and a NEWS.md entry.

Environment

macOS 26.6.2, R 4.6.1, camtraptor 1.0.0.9000 (main @ 56e4384), camtrapdp 0.6.0, dplyr 1.2.1.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions