fix: floor the eigenvalues of the residual block input - #10
Merged
Merged
Conversation
The unit affine-invariant step computes its norm with a Cholesky of the block input. After a batch normalization this input can be numerically singular: - GBWBN folds matrices outside the injectivity domain of the Bures-Wasserstein exponential (HDM05, paper RResNet: minimal eigenvalue 1.8e-10 after the bias step, ~1e-20 after the theta = 1/2 post-transform), and every theta failed; - with GAH / adaptive GAH on HDM05 the Cholesky failed or not depending on the number of threads, i.e. at the edge of float64 precision. The block now floors its input eigenvalues at 1e-8 (the lower bound of affine_invariant_projx on its output) through ReEig's manually differentiated EighReLu, well defined for repeated eigenvalues. It is the identity, values and Jacobian, on well-conditioned inputs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PQdVCDbXCd8gvf1Y4TufJR
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The autograd path still uses an eigenvalue backward that is undefined for repeated eigenvalues; use the repeated-eigenvalue-safe EighReLu backward.
Review effort: Lite
Findings: 1
What changed in this PR
This pull request hardens ResidualBlock against numerically singular SPD inputs by flooring eigenvalues before residual computation.
Changes:
- Adds eigenvalue flooring at
1e-8. - Adds singular-input, gradient, and no-op regression tests.
| File | Summary |
|---|---|
tests/nn/test_rresnet_layers.py |
Tests finite outputs, gradients, and identity behavior. |
src/yetanotherspdnet/nn/rresnet_layers.py |
Floors residual block inputs before geometric operations. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| self.exp_map = affine_invariant_exp | ||
| self.logm = lambda x: logm_SPD(x)[0] | ||
| self.expm = lambda x: expm_symmetric(x)[0] | ||
| self.floor = lambda x: eigh_relu(x, INPUT_EIGVAL_FLOOR)[0] |
- Furo theme instead of pydata-sphinx-theme. - Curated reference (docs/reference/*.md) with autodoc instead of the exhaustive sphinx-autoapi dump: one page per part of the library, each opening with the context of its objects and a summary table; short names, parameter types next to each parameter, one parameter per line in long signatures, attributes as fields (no duplicated objects). - dualpath-table directive (docs/_ext/dualpath.py): pairs every snake_case function with its CamelCase manual-backward Function in one table. - New guides: "How the library fits together" (levels, shapes through a network, conventions), "Layers and parametrizations" (BiMap, static vs dynamic parametrization with its equations, ReEig/ReEigBias/LogEig, estimation layers), "Models and their options" (constructor arguments grouped by role, configurations of the published experiments). - Batch normalization and residual guides: adaptive GAH, input floor. - SPDnet docstring: correct default batchnorm_mean_type. - Builds with -W and no warnings. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PQdVCDbXCd8gvf1Y4TufJR
The "Gradients and precision" guide becomes "Numerical and optimization techniques", with the equations of every trick the library relies on: Daleckii-Krein backward and its degenerate case, Sylvester equations inside backward passes (whitening, Bures-Wasserstein exponential and parallel transport), means as iterations, implicit differentiation at a fixed point (M-estimators), parametrizations, batch normalization mechanics (running statistics, minibatch smoothing, single-matrix batches, matrix-power scaling, Bures-Wasserstein injectivity domain), the residual step (unit norm, Cholesky, floor and projx, other retractions), ReEig, precision. Seven SVG diagrams generated by docs/_diagrams/make_diagrams.py (readable in light and dark themes), also used in the layers, batch normalization and residual guides. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PQdVCDbXCd8gvf1Y4TufJR
docs: Furo theme, curated API reference and linking guides
ruff 0.16 also formats the Python blocks of the Markdown pages; the CI lint job failed on three user guide pages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PQdVCDbXCd8gvf1Y4TufJR
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Stacked on #9.
Why
ResidualBlockcomputes its unit affine-invariant step norm with a Cholesky of its input. After a batch normalization, that input can be numerically singular:(I+V/2)^2,(I+c(X^{1/2}-I))^2or(I+Z)G(I+Z), and becomes singular when an inner eigenvalue crosses 0.The reference implementations (CUAI, GBWBN) have the same structure: a Cholesky norm, with
projxon the output only.Change
The block floors its input eigenvalues at 1e-8, the lower bound of
affine_invariant_projxon its output. It uses ReEig's manually differentiatedEighReLu(Daleckii–Krein), which is well defined for repeated eigenvalues. On well-conditioned inputs it is the identity, in values and Jacobian.Tests
TestResidualBlockSingularInput: inputs with three eigenvalues at 1e-20, wheretorch.linalg.choleskyraises. Output is finite and SPD with finite gradients, for both metrics and both gradient paths. These 4 tests fail on the previous code.🤖 Generated with Claude Code
https://claude.ai/code/session_01PQdVCDbXCd8gvf1Y4TufJR