Skip to content

Egress gateway stays on a demoted Redis replica after a failover, and every egress grant fails with READONLY #277

Description

@sisve

Summary

The ioredis client in service/src/egress-ledger.ts has no reconnectOnError. When Sentinel demotes a master to replica, Redis keeps the existing client connections open. An egress gateway that was connected to that node stays on it, and every ledger script it runs is refused:

READONLY You can't write against a read only replica. script: …

Since the socket stays up, retryStrategy never runs. Readiness keeps passing too: it calls pingEgressLedger(), and PING works on a replica. /live doesn't look at Redis. Code execution is down until someone restarts the egress gateway by hand.

queue.ts, file-server.ts and tool-call-server.ts pass reconnectOnError returning true on READONLY, so after a failover they reconnect on their first refused write.

Environment

  • Code as of main (836d001). Both problems date back to the initial public release (904438d).
  • Kubernetes, hardenedSandboxMode: true, egressGrant.ledgerRequired: true
  • Redis 7.4 in replication mode with 3 Sentinels, reached through a Service that selects the current master

Steps to reproduce

  1. Run the egress gateway against a Sentinel-managed Redis and let it connect to the master.
  2. Trigger a failover with SENTINEL FAILOVER <master-name>. Draining the master's node or rolling the Redis StatefulSet has the same effect.
  3. Start any code execution.

Expected

The first refused write makes the gateway reconnect to the new master, and executions work again on their own.

Actual

  • POST /internal/egress-grants returns 500 for every execution, and each refused write bumps the READONLY count in INFO errorstats on the demoted node.
  • LibreChat shows Error from sandbox: [Internal server error].
  • Readiness stays green, so the pod stays in the Service and is never restarted.
  • Restarting the egress-gateway deployment restores service.

Related: the same client stops reconnecting after five attempts

retryStrategy in the same function returns null after five attempts:

const retryStrategy: CommonRedisOptions['retryStrategy'] = times => {
  if (times > 5) return null;
  return 2000;
};

A null tells ioredis to stop reconnecting for good. If Redis is unreachable for longer than those five attempts (for example a failover that drops the connection rather than demoting the node under it), the gateway stays disconnected. Readiness fails then, but /live still passes and the pod keeps running. queue.ts has retried indefinitely with redisReconnectDelay since #196.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions