Skip to content

fix(nip46): make every sign-in error actionable and surface init failures - #77

Merged
oth-body merged 1 commit into
masterfrom
fix/nip46-signin-error-message
Sep 15, 2026
Merged

oth-body merged 1 commit into
masterfrom
fix/nip46-signin-error-message

Conversation

@oth-body

Copy link
Copy Markdown
Owner

Bug Description

The user's bug report quoted an unhelpful message — "web socket failure, response, no exception, no..." — that maps to the bare "no response yet" polling-timeout string in Session.CheckConnection (and the dial-failure path that discards per-URL error context). The error path during Amber sign-in either shows nothing actionable ("Generating connection..." forever) or shows a bare unhelpful string.

The actual substring "web socket failure" is not in the source. It's either a wrapped go-nostr error or a stale binary from before PR #69/70/73/74/75. Either way, every error path in nip46.go was leaking insufficient context, so the user could not diagnose.

Fixes #76

Root Cause

Three call sites produced unhelpful error messages:

  1. Session.CheckConnection returned the bare "no response yet" after a polling timeout. No mention of Amber, no mention of approval, no mention of what to check.
  2. Session.ConnectRelays discarded per-URL dial errors and only reported "could not connect to any relay (tried: ...)". A user with a stale relays.txt couldn't tell which entry was dead.
  3. tui.initQR swallowed errors from OnInitQR via return nil. Users saw the QR screen with "Generating connection..." forever — no error message ever reached the TUI.
  4. tui.checkQRConnection retried every error indefinitely, including non-retryable hard failures (all relays dropped, proxy refused, etc.). After a hard failure, the user kept seeing "Waiting for connection..." forever.

Fix

  • nip46.ErrTimeout wraps the polling timeout with Amber/approval context. Its Error() method also guards against callers passing a bare unhelpful Reason — it appends an actionable suffix.
  • nip46.RelayDialFailure carries per-URL dial error strings so users see the actual failure cause (DNS, TLS, timeout, connection refused) for each relay, not just a list of URLs.
  • nip46.IsRetryable lets callers (the TUI) distinguish transient errors (keep polling) from hard failures (stop and surface to user).
  • tui.qrInitErrorMsg is the new error message type. The TUI's initQR now surfaces errors and the View() renders them below the QR block.
  • tui.checkQRConnection stops polling on non-retryable errors via nip46.IsRetryable.
  • warnUnreachableRelays now logs the per-URL failure cause inline (e.g. wss://relay.example.com (dial tcp: lookup ...: no such host)) so users can identify the dead entry without enabling debug logging.
  • Session.connect-time errors are renamed to "relays connected but every one rejected our subscription" so the user knows the issue is subscription, not dial.

How to Verify

  1. go run . and pick "Scan QR with Amber".
  2. With a bad relays.txt (e.g. one entry pointing at 127.0.0.1:1), confirm the sign-in error names the failing URL and the per-URL cause.
  3. Without scanning Amber (so the 3-second timeout fires), confirm the timeout error mentions Amber and the connection prompt.
  4. With HTTPS_PROXY=socks5h://127.0.0.1:9050 go run ., confirm the proxy error mentions HTTPS_PROXY and 9050.
  5. Verify the regression test TestUserFacingErrorsAreActionable and TestErrTimeoutMessageIsActionable pass.

Test Plan

  • Regression test TestRelayDialFailureMessage — asserts the per-URL cause is included
  • Regression test TestErrTimeoutIsRetryable — covers all IsRetryable branches
  • Regression test TestErrTimeoutMessageIsActionable — asserts the message includes Amber/approval context, with sabotage-run verified
  • Regression test TestUserFacingErrorsAreActionable — meta-test asserting no user-facing error contains "websocket failure" or lacks an actionable noun
  • Existing tests pass (go test ./...)
  • go vet ./... clean
  • Manual verification of the four scenarios above

Risk Assessment

Low. This change is purely error-message wrapping plus plumbing for the new error type. No behavioural change to successful sign-in flows. The TUI changes are:

  • A new error-message type that's only created when an error exists.
  • IsRetryable returns false for non-timeout errors, so the previous retry-loop behaviour is preserved for transient timeouts and ONLY changed for hard failures (which previously looped forever with no message — strictly worse).

The ErrTimeout.Error() defence-in-depth guard could theoretically append unwanted context if a caller passes an exotic Reason, but the production caller (Session.CheckConnection) always passes a Reason that contains an actionable noun, so the guard never fires in practice.

…ures

The user's bug report: 'web socket failure, response, no exception, no...'
The 'no...' substring maps to the bare 'no response yet' string at
nip46.go:371 (renamed the polling timeout to ErrTimeout with a
contextful default). The 'web socket failure' substring never appears
in source — it was either a wrapper error or a stale binary.

Changes:
- Add nip46.ErrTimeout (wraps the bare 'no response yet' string with
  Amber/approval context), nip46.RelayDialFailure (carries per-URL
  dial errors instead of discarding them), and nip46.IsRetryable so
  callers can distinguish transient vs hard failures.
- ConnectRelays now returns *RelayDialFailure with per-URL error
  strings, so users see exactly which relays.txt entry is dead.
- warnUnreachableRelays takes the failure map and logs the cause
  inline (DNS, TLS, timeout) instead of just the URL.
- TUI's initQR previously swallowed OnInitQR errors via 'return nil'
  — now surfaces them as qrInitErrorMsg with the actionable message
  appended below the QR block. checkQRConnection stops polling on
  non-retryable errors via nip46.IsRetryable instead of looping
  forever on a hard failure.
- Every user-facing error in the package now mentions the failing
  URL, env var, or an actionable noun (Amber, signer, network,
  relays.txt). Regression guards in nip46_test.go assert no message
  leaks the old bare 'no response yet' or 'websocket failure'
  substrings, and that the bare Reason case is wrapped with context
  by Error() as a defence-in-depth.

Tests: +5 new, full suite green, go vet clean.
@oth-body
oth-body merged commit bd1658e into master Sep 15, 2026
7 checks passed
@oth-body
oth-body deleted the fix/nip46-signin-error-message branch September 15, 2026 21:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bug: 'web socket failure' / 'no response yet' during NIP-46 Amber scan shows no actionable info

1 participant