Skip to content

Count distinct scientific names per group in summarize_observations() - #434

Merged
damianooldoni merged 3 commits into
inbo:mainfrom
northfox:fix/432-n-scientific-name
Sep 25, 2026
Merged

damianooldoni merged 3 commits into
inbo:mainfrom
northfox:fix/432-n-scientific-name

Conversation

@northfox

Copy link
Copy Markdown

Fixes #432.

n_scientificName summed distinct counts per deployment, so a scientific name observed in several deployments of a group was counted more than once. It is now counted once per group.

  • summarize_observations(): retain scientificName per deployment and count distinct names at the final group level. Other features are unchanged, as is the default output.
  • Added tests through summarize_observations() (fail on main, pass here).
  • Added a NEWS.md entry.
  • DESCRIPTION: added myself as a contributor, as requested in n_scientificName counts a scientific name once per deployment instead of once per group #432. I don't have an ORCID, so I left it out; please let me know if you'd prefer otherwise.

northfox added 2 commits September 22, 2026 08:11
…ons()

`n_scientificName` was calculated as the number of distinct scientific
names per deployment, summed over all deployments of a group. As distinct
counts are not additive, a scientific name observed in more than one
deployment was counted more than once when `group_by` did not contain
`deploymentID`, e.g. `n_scientificName = 2` when grouping by
`scientificName`, which contradicts the documentation.

The scientific names are now retained per deployment by grouping by
`scientificName` as well, and counted over all deployments of a group
afterwards.

Refs inbo#432
@damianooldoni

Copy link
Copy Markdown
Member

Thanks @northfox. I will review it this week, meanwhile I let automatic workflows/checks running on it.
Also, please add your name, surname and an ORCID, if you have one, in the DESCRIPTION. I would like to know the person behind a GitHub account if it goes about contributing to an open software project as this one. Thanks.

@northfox

Copy link
Copy Markdown
Author

Thanks @damianooldoni. I've updated DESCRIPTION with my name and surname. I don't have an ORCID, so I left that out.

Also, a note on the failing macos-latest (release) check: it appears unrelated to this PR. The job fails during dependency installation, before the package check runs, when pak cannot extract sf_1.1-3.tgz:

Cannot extract .../sf_1.1-3.tgz,
unknown archive type.

The other R-CMD-check matrix jobs pass.

This matches the current CRAN macOS R 4.6 binary-compression issue confirmed upstream:

A rerun may succeed once the affected CRAN binary has been refreshed.

Comment thread DESCRIPTION
@damianooldoni
damianooldoni self-requested a review September 25, 2026 08:14
Mention the new contributor as well.

@damianooldoni damianooldoni left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @northfox for fixing the bug, adding tests and news.
I have just slightly modified the news.

@damianooldoni damianooldoni self-assigned this Sep 25, 2026
@damianooldoni
damianooldoni merged commit 5b800c2 into inbo:main Sep 25, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

n_scientificName counts a scientific name once per deployment instead of once per group

2 participants