Count distinct scientific names per group in summarize_observations() - #434
Conversation
…ons() `n_scientificName` was calculated as the number of distinct scientific names per deployment, summed over all deployments of a group. As distinct counts are not additive, a scientific name observed in more than one deployment was counted more than once when `group_by` did not contain `deploymentID`, e.g. `n_scientificName = 2` when grouping by `scientificName`, which contradicts the documentation. The scientific names are now retained per deployment by grouping by `scientificName` as well, and counted over all deployments of a group afterwards. Refs inbo#432
|
Thanks @northfox. I will review it this week, meanwhile I let automatic workflows/checks running on it. |
|
Thanks @damianooldoni. I've updated Also, a note on the failing The other R-CMD-check matrix jobs pass. This matches the current CRAN macOS R 4.6 binary-compression issue confirmed upstream:
A rerun may succeed once the affected CRAN binary has been refreshed. |
Mention the new contributor as well.
damianooldoni
left a comment
There was a problem hiding this comment.
Thanks @northfox for fixing the bug, adding tests and news.
I have just slightly modified the news.
Fixes #432.
n_scientificNamesummed distinct counts per deployment, so a scientific name observed in several deployments of a group was counted more than once. It is now counted once per group.summarize_observations(): retainscientificNameper deployment and count distinct names at the final group level. Other features are unchanged, as is the default output.summarize_observations()(fail onmain, pass here).NEWS.mdentry.DESCRIPTION: added myself as a contributor, as requested inn_scientificNamecounts a scientific name once per deployment instead of once per group #432. I don't have an ORCID, so I left it out; please let me know if you'd prefer otherwise.