Repository navigation
Replies: 3 comments 2 replies
|
Note I have very little experience with Traefik, so I used ChatGPT to help me write this reply. Thank you for the detailed write-up and the working Compose example. I can see why this is confusing, because the same hostname and port can appear in several places, but they describe different network paths. The short version is:
The different network pathsIt helps to separate the system into these paths:
In your setup, the browser path should be: The fact that the browser uses HTTPS on port 443 does not mean the xyOps container itself must listen on HTTPS or port 443. Traefik terminates TLS and forwards ordinary HTTP traffic to port 5522 inside the Docker network. The WebSocket used by the web UI follows the same route. Traefik supports WebSocket upgrades automatically, so xyOps should not need special WebSocket header middleware. An xySat connecting through Traefik would use: That is where these settings matter: XYOPS_satellite__config__host: xyops.server.example.com
XYOPS_satellite__config__secure: "true"
XYOPS_satellite__config__port: "443"They tell xySat how to reach the conductor. They do not change how the conductor listens, and they do not change how Traefik reaches the container. They also affect the installation URL shown in the Add Server dialog. That URL is deliberately constructed from the satellite connection configuration, because the address used by the browser is not necessarily reachable from worker servers. This is described in the Satellite Configuration and Overriding the Connect URL sections. Why the restart problem was probably unrelatedChanging: does not alter: Therefore, those satellite properties cannot directly make the xyOps HTTP listener disappear from port 5522. A Traefik A few possibilities after the restart are:
The After a restart, try: docker exec xyops-conductor-01 \
curl -sS -i http://127.0.0.1:5522/healthAnd from outside Docker: curl -k -i https://xyops.server.example.com/healthThe results narrow the problem down considerably:
Environment variables versus
|
|
Thank you very much for doing all of this additional testing and for providing the complete crash trace. This explains the behavior perfectly, and you were right that the restart problem was correlated with the satellite configuration change. You have also uncovered a real crasher bug in xyOps. I will be fixing that right away! The crash occurs in the conductor's satellite release API when the Servers page asks for the available xySat versions. Regardless of which component owns the code, an API request should never be able to crash the entire conductor, even when the configuration is incomplete or invalid. That part is definitely a bug on our side. Why it crashedSo, the root cause is the structure that was manually added to {
"secret_key": "redacted_secret_key",
"satellite": {
"config": {
"port": 443,
"secure": true
}
}
}The Internally, the file is a sparse collection of individual configuration paths. For example, when the xyOps UI changes these two settings, it writes something conceptually like this: {
"secret_key": "redacted_secret_key",
"satellite.config.port": 443,
"satellite.config.secure": true
}Each dotted key identifies one specific leaf in the main configuration tree. When the manually edited file instead contains a top-level key named The normal {
"satellite": {
"list_url": "https://api.github.com/repos/pixlcore/xysat/releases",
"base_url": "https://github.com/pixlcore/xysat/releases",
"image": "ghcr.io/pixlcore/xysat",
"version": "latest",
"enable_version_checks": true,
"cache_ttl": 3600,
"config": {
"port": 5522,
"secure": false
}
}
}After your nested override was loaded, the effective value became only: {
"satellite": {
"config": {
"port": 443,
"secure": true
}
}
}This removed That also explains why the environment variables do not produce the same failure. This environment variable: is translated into the specific path So your original observation was correct. The environment variables and the manually edited file really did behave differently in this case. The reason was not a different implementation for xySat connectivity, but the fact that the nested Why everything initially appeared to workSeveral parts of xyOps only need the two properties that remained: For example, generating the "Add Server" installation command only needs enough information to build the connection URL. Existing satellites and workflows also do not need the conductor's xySat release metadata during normal operation. The browser may also have already cached the latest xySat version information in memory. This can temporarily hide the missing When the Servers page is opened in a new session, xyOps asks the conductor for the available xySat releases so it can identify outdated agents. The conductor eventually reaches code equivalent to: self.request.json( sat.list_url, ... );At that point, That exception is not currently caught, so the entire conductor process exits. Docker restarts the process, and Traefik understandably returns The message saying that a crash log was found at startup is xyOps reporting the crash record left by the previous process. The later uncaught exception after opening Servers is the new crash that starts the loop again. This matches every part of the behavior you reported. What should be used in DockerFor Docker Compose, I recommend keeping these deployment-specific settings in the environment: environment:
XYOPS_satellite__config__host: xyops.server.example.com
XYOPS_satellite__config__port: "443"
XYOPS_satellite__config__secure: "true"Alternatively, settings can be changed through the xyOps configuration UI, which will write the appropriate sparse keys into The existing For recovery, the manually added nested The crasher is still our bugEven though the nested shape is not the intended format for At minimum, the satellite release API should verify that I will fix this crasher right away. I will also clarify the documentation around Thank you again for continuing to investigate this. Your crash trace made the problem immediately identifiable. About the endpoint-list ideaYour broader suggestion about giving xySat an explicit list of endpoints is reasonable, especially for a topology where workers should try a load balancer first and then fall back to individual conductors. The current xySat model has:
As you correctly observed, this cannot cleanly express a sequence such as: with different routing behavior for each endpoint. When the static If this area were redesigned, a list of full endpoint URLs might be clearer and more flexible than separate host, port, and secure properties. For example: {
"endpoints": [
"wss://loadbalancer.example.com:443",
"wss://server1.example.com:5523",
"wss://server2.example.com:5523"
]
}Each endpoint would then carry its own protocol and port. There are some xyOps-specific details that would need careful consideration. In particular, xySat does not assume that the first conductor in its list is primary. It can connect to a conductor, learn which conductor currently owns the primary role, and reconnect appropriately. The conductor list can also be updated automatically when the cluster membership changes. A load balancer endpoint would need to coexist with that discovery and failover behavior without accidentally leaving workers connected to a backup conductor. Still, the use case you described is valid, and I appreciate the suggestion. I would consider that a separate feature proposal from the crash itself, but it is useful feedback on where the current And finally, thank you for confirming the Thanks again! I really appreciate all your work and detail in these posts. |

Uh oh!
There was an error while loading. Please reload this page.
I'm currently testing xyops in my homelab and it looks very promising. However, I’ve encountered some issues setting it up with docker and traefik.
This is an
docker-compose.ymlexample for a single xyops container that works with traefik:The important settings are:
hostname- Otherwise the "conductor host id" (aka. satellite.config.hosts) and url shown when adding a server is the randomly generated container id.XYOPS_satellite__config__secure=trueandXYOPS_satellite__config__port=443(mentioned in the tailscale documentation).Other settings
XYOPS_mastersleads to an error as the port5522is not used externaly:Error > PeerSocket | comm | Socket Error: connect ECONNREFUSED 192.168.2.15:5522XYOPS_xysat_local=trueas the satellite is set up directly on the host afterwards.Issues and Questions
I've faced the following issues when setting up xyops the quick-start way but with the exception of using traefik and not exposing the ports:
config.satellite.port=443andconfig.satellite.secure=truein theoverride.jsonfile.Questions
Is the frontend considered a "satellite" or why changing the
satellite.confighas implications on how the frontend behaves?Why is setting
XYOPS_satellite__config__*as environment variables vs. setting them inoverride.jsoncircumventing this behaviour?Is there a "correct" way of using xyops with traefik that I'm not aware of?
I understand the reasoning behind some behaviour, as explained in this issue.
But in my opinion, the implementation for the different possible configuration methods of how the xysat connects to the server is somewhat complicated.
Anyway, thank you for this great software.
All reactions