Skip to content

Auditor.compute_audit_result() returns is_model_fair=True when no finite loss ratios remain (or on an exact tie at the threshold), because a NaN p-value falls through to the "fair" branch #72

Description

@shaurya416

inFairness/auditor/auditor.py:58-114

    def compute_audit_result(self, loss_ratios, threshold=None, confidence=0.95):
        loss_ratios = loss_ratios[np.isfinite(loss_ratios)]

        lossratio_mean = np.mean(loss_ratios)
        lossratio_std = np.std(loss_ratios)
        N = loss_ratios.shape[0]

        z = norm.ppf(confidence)
        lower_bound = lossratio_mean - z * lossratio_std / np.sqrt(N)

        ...
        else:
            tval = (lossratio_mean - threshold) / lossratio_std
            tval *= np.sqrt(N)

            pval = 1 - norm.cdf(tval)
            is_model_fair = False if pval < (1 - confidence) else True

compute_audit_result produces the verdict for Auditor.audit() in both SenSeIAuditor (sensei_auditor.py:167) and SenSRAuditor (sensr_auditor.py:172). Both pass the output of compute_loss_ratio in unchanged.

Line 82 drops non-finite loss ratios. Nothing sets a minimum on the remaining N, and nothing records how many rows were dropped: AuditorResponse has no sample-count field. Two inputs make tval/pval at lines 98-101 NaN:

  • N falls to 0 after filtering.
  • The remaining ratios all equal threshold exactly (std == 0 and mean - threshold == 0, so 0/0).

In Python, NaN < x is always False. The ternary on line 102 therefore returns is_model_fair=True, the same value as a genuine pass. numpy emits RuntimeWarnings on the way ("Mean of empty slice", "invalid value encountered in scalar divide"), but the returned AuditorResponse looks like a normal passing result.

The zero-sample case is the worst one. compute_loss_ratio computes loss_adversarial / loss_original. If every original-sample loss is exactly 0, each ratio is +inf (adversarial loss > 0) or NaN (adversarial loss also 0). A finite ratio of +inf would mean the worst possible individual-fairness violation. Instead, all ratios are filtered out and the model is reported fair. A per-sample loss of exactly 0 is realistic: float32 cross-entropy on saturated logits produces it, and so does an exact regression fit.

Measured

Reproduced by copying the function body and the AuditorResponse dataclass verbatim into an isolated script (numpy 2.3.1 and scipy.stats.norm only; no inFairness modules imported) and calling it directly. Commit 701d6d0.

Case Input Result
Zero valid samples (inf) loss_ratios = np.full((50,), np.inf), threshold=1.25 (every original loss 0, adversarial loss > 0) lossratio_mean=nan, lossratio_std=nan, lower_bound=nan, pval=nan, is_model_fair=True (RuntimeWarnings emitted)
Zero valid samples (0/0) loss_ratios = np.full((50,), np.nan), threshold=1.25 (original and adversarial losses both 0) lossratio_mean=nan, pval=nan, is_model_fair=True
Exact tie at threshold loss_ratios = [1.25, 1.25, 1.25], threshold=1.25 lossratio_mean=1.25, lossratio_std=0.0, pval=nan, is_model_fair=True
Exact tie, unchanged inputs loss_ratios = np.ones(20), threshold=1.0 (worst-case examples identical to originals) pval=nan, is_model_fair=True
Control: must fail 200 finite samples, mean≈3.0, sd 0.2, threshold=1.25 is_model_fair=False, pval=0.0 (correct)
Control: must pass 200 finite samples, mean≈1.0, sd 0.05, threshold=1.25 is_model_fair=True, pval=1.0 (correct)
Control: zero variance, above threshold loss_ratios = [3.0, 3.0], threshold=1.25 is_model_fair=False, pval=0.0 (correct; tval=+inf)
Partial contamination 190 finite (mean≈3.0) + 10 inf, threshold=1.25 is_model_fair=False (correct verdict), computed on N=190, no warning, and no field in AuditorResponse records that 10 of 200 rows were dropped

Consequence

A caller that audits a model and gets zero usable loss ratios back receives is_model_fair=True, which looks the same as a real pass on real data. Every original-sample loss being exactly 0 is enough to cause this. The same happens when the remaining ratios fall exactly on the threshold. AuditorResponse is a plain dataclass with no valid-sample-count field, so a caller cannot tell "verified fair" apart from "the test could not run". When only some rows are dropped, the verdict is computed on fewer samples than the caller supplied, and the response does not say so.

Suggested fix

Handle invalid or degenerate input the way the rest of the codebase does: distances/distance.py:27 and distances/logistic_sensitive_subspace.py:99 raise AssertionError, and utils/plackett_luce.py:68,124 raise ValueError. Concretely:

  • When N == 0 after the finite filter, raise, or return an explicit insufficient-data response (for example is_model_fair=None).
  • Guard the tval/pval computation so a NaN p-value can never reach the else branch of the fairness ternary. For example, handle lossratio_std == 0 explicitly by comparing lossratio_mean to threshold, and treat a non-finite pval as "undetermined" rather than fair.
  • Add an n_samples (or n_dropped) field to AuditorResponse, so a caller can see a partially contaminated audit even when the verdict is correct.

tests/auditor/test_auditor.py covers only the output shape of compute_loss_ratio. tests/auditor/test_sensei_auditor.py and tests/auditor/test_sensr_auditor.py reach compute_audit_result through audit(), but only with random, well-conditioned inputs, and they assert only that the fields have the right types. A direct regression test with all-non-finite and tied inputs would cover this path.

Happy to open the PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions