Summary
summarize_observations() computes n_scientificName per deployment and then sums these values across deployments. When group_by does not include deploymentID, a scientific name observed in several deployments is therefore counted several times.
This contradicts the documentation of n_scientificName:
If scientificName is in group_by, n_scientificName is equal to 1 or 0, if scientificName = NA (unidentified animals).
and the expected output described in #367 ("this will be 1 for all if scientificName is selected in group_by").
Reproducible example
library(camtraptor)
library(dplyr, warn.conflicts = FALSE)
x <- example_dataset()
summarize_observations(x, group_by = "scientificName") %>%
select(scientificName, n_scientificName) %>%
filter(n_scientificName > 1)
#> # A tibble: 2 × 2
#> # Groups: scientificName [2]
#> scientificName n_scientificName
#> <chr> <int>
#> 1 Anas platyrhynchos 2
#> 2 Aves 2
Expected: 1 for both, since each row represents a single scientific name. The value 2 is the number of deployments in which the name was observed. The same over-counting can occur for any grouping that combines several deployments when the same scientific name occurs in more than one of them, e.g. group_by = "locationName" when a location has more than one deployment, or group_by = c("latitude", "longitude").
The default group_by includes deploymentID, so the default output and the deprecated get_n_species() are not affected. n_events, n_observations and sum_count are additive across deployments and remain correct. The rai_* values are calculated from the aggregated counts and effort and are unaffected.
Cause
calc_obs_feature() uses n_distinct(scientificName) per deployment as formula_per_deployment and sum(n_scientificName) as formula_total. Distinct counts are not additive across deployments.
Proposed fix
Count distinct scientific names at the final output-group level (group_by columns + group_time_by) instead of summing per-deployment distinct counts, e.g. directly from the enriched observations or by retaining the distinct names until the final aggregation. NA would still be excluded, so groups with only unidentified animals keep n_scientificName = 0.
I'd be happy to open a PR with a regression test through summarize_observations() and a NEWS.md entry.
Environment
macOS 26.6.2, R 4.6.1, camtraptor 1.0.0.9000 (main @ 56e4384), camtrapdp 0.6.0, dplyr 1.2.1.
Summary
summarize_observations()computesn_scientificNameper deployment and then sums these values across deployments. Whengroup_bydoes not includedeploymentID, a scientific name observed in several deployments is therefore counted several times.This contradicts the documentation of
n_scientificName:and the expected output described in #367 ("this will be
1for all ifscientificNameis selected ingroup_by").Reproducible example
Expected:
1for both, since each row represents a single scientific name. The value2is the number of deployments in which the name was observed. The same over-counting can occur for any grouping that combines several deployments when the same scientific name occurs in more than one of them, e.g.group_by = "locationName"when a location has more than one deployment, orgroup_by = c("latitude", "longitude").The default
group_byincludesdeploymentID, so the default output and the deprecatedget_n_species()are not affected.n_events,n_observationsandsum_countare additive across deployments and remain correct. Therai_*values are calculated from the aggregated counts and effort and are unaffected.Cause
calc_obs_feature()usesn_distinct(scientificName)per deployment asformula_per_deploymentandsum(n_scientificName)asformula_total. Distinct counts are not additive across deployments.Proposed fix
Count distinct scientific names at the final output-group level (
group_bycolumns +group_time_by) instead of summing per-deployment distinct counts, e.g. directly from the enriched observations or by retaining the distinct names until the final aggregation.NAwould still be excluded, so groups with only unidentified animals keepn_scientificName = 0.I'd be happy to open a PR with a regression test through
summarize_observations()and aNEWS.mdentry.Environment
macOS 26.6.2, R 4.6.1, camtraptor 1.0.0.9000 (
main@ 56e4384), camtrapdp 0.6.0, dplyr 1.2.1.