Summary
The ioredis client in service/src/egress-ledger.ts has no reconnectOnError. When Sentinel demotes a master to replica, Redis keeps the existing client connections open. An egress gateway that was connected to that node stays on it, and every ledger script it runs is refused:
READONLY You can't write against a read only replica. script: …
Since the socket stays up, retryStrategy never runs. Readiness keeps passing too: it calls pingEgressLedger(), and PING works on a replica. /live doesn't look at Redis. Code execution is down until someone restarts the egress gateway by hand.
queue.ts, file-server.ts and tool-call-server.ts pass reconnectOnError returning true on READONLY, so after a failover they reconnect on their first refused write.
Environment
- Code as of
main (836d001). Both problems date back to the initial public release (904438d).
- Kubernetes,
hardenedSandboxMode: true, egressGrant.ledgerRequired: true
- Redis 7.4 in replication mode with 3 Sentinels, reached through a Service that selects the current master
Steps to reproduce
- Run the egress gateway against a Sentinel-managed Redis and let it connect to the master.
- Trigger a failover with
SENTINEL FAILOVER <master-name>. Draining the master's node or rolling the Redis StatefulSet has the same effect.
- Start any code execution.
Expected
The first refused write makes the gateway reconnect to the new master, and executions work again on their own.
Actual
POST /internal/egress-grants returns 500 for every execution, and each refused write bumps the READONLY count in INFO errorstats on the demoted node.
- LibreChat shows
Error from sandbox: [Internal server error].
- Readiness stays green, so the pod stays in the Service and is never restarted.
- Restarting the egress-gateway deployment restores service.
Related: the same client stops reconnecting after five attempts
retryStrategy in the same function returns null after five attempts:
const retryStrategy: CommonRedisOptions['retryStrategy'] = times => {
if (times > 5) return null;
return 2000;
};
A null tells ioredis to stop reconnecting for good. If Redis is unreachable for longer than those five attempts (for example a failover that drops the connection rather than demoting the node under it), the gateway stays disconnected. Readiness fails then, but /live still passes and the pod keeps running. queue.ts has retried indefinitely with redisReconnectDelay since #196.
Summary
The ioredis client in
service/src/egress-ledger.tshas noreconnectOnError. When Sentinel demotes a master to replica, Redis keeps the existing client connections open. An egress gateway that was connected to that node stays on it, and every ledger script it runs is refused:Since the socket stays up,
retryStrategynever runs. Readiness keeps passing too: it callspingEgressLedger(), andPINGworks on a replica./livedoesn't look at Redis. Code execution is down until someone restarts the egress gateway by hand.queue.ts,file-server.tsandtool-call-server.tspassreconnectOnErrorreturningtrueonREADONLY, so after a failover they reconnect on their first refused write.Environment
main(836d001). Both problems date back to the initial public release (904438d).hardenedSandboxMode: true,egressGrant.ledgerRequired: trueSteps to reproduce
SENTINEL FAILOVER <master-name>. Draining the master's node or rolling the Redis StatefulSet has the same effect.Expected
The first refused write makes the gateway reconnect to the new master, and executions work again on their own.
Actual
POST /internal/egress-grantsreturns 500 for every execution, and each refused write bumps theREADONLYcount inINFO errorstatson the demoted node.Error from sandbox: [Internal server error].Related: the same client stops reconnecting after five attempts
retryStrategyin the same function returnsnullafter five attempts:A
nulltells ioredis to stop reconnecting for good. If Redis is unreachable for longer than those five attempts (for example a failover that drops the connection rather than demoting the node under it), the gateway stays disconnected. Readiness fails then, but/livestill passes and the pod keeps running.queue.tshas retried indefinitely withredisReconnectDelaysince #196.