Skip to content

fix: floor the eigenvalues of the residual block input - #10

Merged
Matthieu-Gallet merged 5 commits into
mainfrom
fix/rresnet-input-floor
Sep 28, 2026
Merged

Matthieu-Gallet merged 5 commits into
mainfrom
fix/rresnet-input-floor

Conversation

@Matthieu-Gallet

Copy link
Copy Markdown
Contributor

Stacked on #9.

Why

ResidualBlock computes its unit affine-invariant step norm with a Cholesky of its input. After a batch normalization, that input can be numerically singular:

  • GBWBN folds matrices outside the injectivity domain of the Bures–Wasserstein exponential. Each BW step is a "square", (I+V/2)^2, (I+c(X^{1/2}-I))^2 or (I+Z)G(I+Z), and becomes singular when an inner eigenvalue crosses 0.
    • On HDM05 with the GBWBN paper's RResNet, at epoch 27, the minimal eigenvalue is 1.8e-10 after the bias step and about 1e-20 after the θ=1/2 post-transform.
    • Every θ in {0.25, 0.5, 1} failed.
  • GAH / adaptive GAH on HDM05, paper hyperparameters: the Cholesky failed in 3 of 4 runs. The same configuration with 2 threads instead of 4 trained without failure (78.6% test accuracy vs 72.2% without BN), so this sits at the edge of float64 precision.

The reference implementations (CUAI, GBWBN) have the same structure: a Cholesky norm, with projx on the output only.

Change

The block floors its input eigenvalues at 1e-8, the lower bound of affine_invariant_projx on its output. It uses ReEig's manually differentiated EighReLu (Daleckii–Krein), which is well defined for repeated eigenvalues. On well-conditioned inputs it is the identity, in values and Jacobian.

Tests

  • New TestResidualBlockSingularInput: inputs with three eigenvalues at 1e-20, where torch.linalg.cholesky raises. Output is finite and SPD with finite gradients, for both metrics and both gradient paths. These 4 tests fail on the previous code.
  • The floor is a no-op in values and Jacobian on well-conditioned inputs.
  • Full suite: 4007 passed, 2 skipped.

🤖 Generated with Claude Code

https://claude.ai/code/session_01PQdVCDbXCd8gvf1Y4TufJR

The unit affine-invariant step computes its norm with a Cholesky of the
block input. After a batch normalization this input can be numerically
singular:

- GBWBN folds matrices outside the injectivity domain of the
  Bures-Wasserstein exponential (HDM05, paper RResNet: minimal eigenvalue
  1.8e-10 after the bias step, ~1e-20 after the theta = 1/2
  post-transform), and every theta failed;
- with GAH / adaptive GAH on HDM05 the Cholesky failed or not depending on
  the number of threads, i.e. at the edge of float64 precision.

The block now floors its input eigenvalues at 1e-8 (the lower bound of
affine_invariant_projx on its output) through ReEig's manually
differentiated EighReLu, well defined for repeated eigenvalues. It is the
identity, values and Jacobian, on well-conditioned inputs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PQdVCDbXCd8gvf1Y4TufJR
Copilot AI lite review requested due to automatic review settings September 27, 2026 11:27

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The autograd path still uses an eigenvalue backward that is undefined for repeated eigenvalues; use the repeated-eigenvalue-safe EighReLu backward.

Review effort: Lite
Findings: 1 High severity

Open (1)
What changed in this PR

This pull request hardens ResidualBlock against numerically singular SPD inputs by flooring eigenvalues before residual computation.

Changes:

  • Adds eigenvalue flooring at 1e-8.
  • Adds singular-input, gradient, and no-op regression tests.
File Summary
tests/​nn/​test_rresnet_layers.py Tests finite outputs, gradients, and identity behavior.
src/​yetanotherspdnet/​nn/​rresnet_layers.py Floors residual block inputs before geometric operations.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

self.exp_map = affine_invariant_exp
self.logm = lambda x: logm_SPD(x)[0]
self.expm = lambda x: expm_symmetric(x)[0]
self.floor = lambda x: eigh_relu(x, INPUT_EIGVAL_FLOOR)[0]
- Furo theme instead of pydata-sphinx-theme.
- Curated reference (docs/reference/*.md) with autodoc instead of the
  exhaustive sphinx-autoapi dump: one page per part of the library, each
  opening with the context of its objects and a summary table; short names,
  parameter types next to each parameter, one parameter per line in long
  signatures, attributes as fields (no duplicated objects).
- dualpath-table directive (docs/_ext/dualpath.py): pairs every snake_case
  function with its CamelCase manual-backward Function in one table.
- New guides: "How the library fits together" (levels, shapes through a
  network, conventions), "Layers and parametrizations" (BiMap, static vs
  dynamic parametrization with its equations, ReEig/ReEigBias/LogEig,
  estimation layers), "Models and their options" (constructor arguments
  grouped by role, configurations of the published experiments).
- Batch normalization and residual guides: adaptive GAH, input floor.
- SPDnet docstring: correct default batchnorm_mean_type.
- Builds with -W and no warnings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PQdVCDbXCd8gvf1Y4TufJR
Matthieu-Gallet and others added 2 commits September 28, 2026 08:59
The "Gradients and precision" guide becomes "Numerical and optimization
techniques", with the equations of every trick the library relies on:
Daleckii-Krein backward and its degenerate case, Sylvester equations inside
backward passes (whitening, Bures-Wasserstein exponential and parallel
transport), means as iterations, implicit differentiation at a fixed point
(M-estimators), parametrizations, batch normalization mechanics (running
statistics, minibatch smoothing, single-matrix batches, matrix-power
scaling, Bures-Wasserstein injectivity domain), the residual step (unit
norm, Cholesky, floor and projx, other retractions), ReEig, precision.

Seven SVG diagrams generated by docs/_diagrams/make_diagrams.py (readable in
light and dark themes), also used in the layers, batch normalization and
residual guides.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PQdVCDbXCd8gvf1Y4TufJR
docs: Furo theme, curated API reference and linking guides
Base automatically changed from feat/gbwbn-conformity to main September 28, 2026 08:27
ruff 0.16 also formats the Python blocks of the Markdown pages; the CI lint
job failed on three user guide pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PQdVCDbXCd8gvf1Y4TufJR
@Matthieu-Gallet
Matthieu-Gallet merged commit fb361d5 into main Sep 28, 2026
7 checks passed
@Matthieu-Gallet
Matthieu-Gallet deleted the fix/rresnet-input-floor branch September 28, 2026 12:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants