The official push client keeps the notify_push WebSocket healthy with an active heartbeat and bounds authentication retries. src/libsync/pushnotifications.cpp:
- PING_INTERVAL = 30s: sends a WS ping (
pingWebSocketServer), arms a 30s _pingTimedOutTimer; if no pong arrives in time (onPingTimedOut) it tears down and reconnects (setup()).
- MAX_ALLOWED_FAILED_AUTHENTICATION_ATTEMPTS = 3: after 3 failed auth attempts
tryReconnectToWebSocket returns false and stops retrying.
_reconnectTimerInterval = 20s for reconnect attempts.
Current state in nextsync-rs: push.rs handles inbound Message::Ping (responds with pong, line 621-625) but never sends its own ping and has no pong-timeout detection, so a half-dead TCP connection (like the stuck dhclient that took down hub.cloudless.club on 26-Aug) can sit silently with no notifications flowing and no reconnect. Auth retries exist (AuthRequired -> state + reconnect via BACKOFF_SECONDS [2,5,10,30,60,300]) but there is no cap on total failed attempts, so an invalid credential could retry forever.
Proposed work:
- Send a WS ping every ~30s and arm a pong-timeout timer; on timeout, treat the channel as dead, emit a disconnect and reconnect (the existing generation/backoff machinery already drops stale attempts).
- Count failed auth attempts and stop retrying after 3 (mirroring the official cap), surfacing
PushState::AuthRequired and requiring an explicit trigger (network restored / settings change / manual) to retry.
The official push client keeps the notify_push WebSocket healthy with an active heartbeat and bounds authentication retries.
src/libsync/pushnotifications.cpp:pingWebSocketServer), arms a 30s_pingTimedOutTimer; if no pong arrives in time (onPingTimedOut) it tears down and reconnects (setup()).tryReconnectToWebSocketreturns false and stops retrying._reconnectTimerInterval = 20sfor reconnect attempts.Current state in nextsync-rs:
push.rshandles inboundMessage::Ping(responds with pong, line 621-625) but never sends its own ping and has no pong-timeout detection, so a half-dead TCP connection (like the stuck dhclient that took down hub.cloudless.club on 26-Aug) can sit silently with no notifications flowing and no reconnect. Auth retries exist (AuthRequired-> state + reconnect via BACKOFF_SECONDS [2,5,10,30,60,300]) but there is no cap on total failed attempts, so an invalid credential could retry forever.Proposed work:
PushState::AuthRequiredand requiring an explicit trigger (network restored / settings change / manual) to retry.