We take the security of fleet seriously. Thank you for helping keep fleet and its users safe.
Please do not report security vulnerabilities through public GitHub issues, pull requests, or discussions.
Instead, report them privately by email to hello@elcanotek.com.
Please include as much of the following as you can, so we can triage quickly:
- A description of the vulnerability and its potential impact.
- Steps to reproduce, or a proof-of-concept.
- The affected component(s) and version / commit.
- Any suggested remediation, if you have one.
If you would like to encrypt your report, mention that in an initial email and we will arrange a secure channel.
- Acknowledgement within 3 business days of your report.
- An initial assessment (severity and likely remediation path) within 7 business days.
- Progress updates as we work on a fix, and credit in the release notes once the issue is resolved — unless you prefer to remain anonymous.
We ask that you give us a reasonable opportunity to remediate the issue before any public disclosure.
fleet ships from a rolling release train: every green push to main is tagged
vYYYY.MM.DD.N automatically, and only the newest release is supported.
Fixes land on main and reach operators on their next fleet update; there are
no maintained release branches and no backports to an older date. Please
reproduce against current main (or the newest release tag) before reporting,
and include the output of fleet version — it names the release the box is on
and how many commits past it, which is exactly what we need to tell "already
fixed" from "still broken". See docs/VERSIONING.md.
CI runs gitleaks — gitleaks dir . --redact --exit-code 1, over the whole working tree, not just the diff — and
fails the build on any new, un-ignored secret. It is the one lane that is
deliberately not gated on the docs-only classifier, because a secret can be
pasted into a markdown file.
To be precise about coverage, since "on every push" would overstate it: the scan
runs on pull requests into main and on pushes to main. Those are the
only events the CI workflow subscribes to — main is the only long-lived branch
(ADR-0062) — so a push to a personal
feature branch runs no CI at all — including no secret scan — until a PR is
opened against main. Treat pre-PR local hygiene accordingly.
If you are contributing, never commit real credentials — the generic
config/default bundle ships with no connector secrets, and all deployment
secrets live in an operator-managed 0600 env file outside the repo (see the
README).
Two static analyzers gate merges, and an auditor will want them by name. The full
design notes are docs/SCANNING.md (who checks what, and what
actually gates) and docs/CODEQL.md; the threshold decision is
ADR-0048.
CodeQL — advanced setup, four languages, security-extended.
.github/workflows/codeql.yml analyzes go, python, javascript-typescript
and actions, each with the security-extended query suite (the broader security
set, not the default one). Go builds via autobuild with
GOFLAGS=-tags=fleet_host_executor, so internal/sandbox/host.go — the
unsandboxed host executor, fenced out of the default build — is inside the
database rather than the one file the analysis cannot see. It is a reusable
workflow: ci.yml calls it as a job, so its result lands in CI gate on every
PR into main and every push to main. There is also a weekly schedule, because a
CodeQL verdict is a function of the query pack as well as the commit.
Four properties of that gate matter for an audit, and each is a deliberate limit rather than an oversight:
- The threshold is the High band, not "any finding". A finding blocks when its
rule publishes
security-severity >= 7.0(CodeQL's own High/Critical cut), or — for a rule that publishes no security-severity — when its SARIF level iserror/warning. Level is deliberately not used for rules that do publish a security-severity: nearly every CodeQL security query is@problem.severity error,go/log-injectionat severity 6.1 included, so banding on level would block on everything. Findings below the band are advisory — printed in the job log and uploaded to the Security tab, not blocking. - Accepted findings are a reviewed register, not a disabled query.
.github/codeql-accepted-findings.jsonholds accepted(rule, file)pairs, each with a mandatory written reason that must say why the finding is not exploitable there. It is per-file on purpose: aquery-filtersexclude would switch a security-severity 9.1 query off repo-wide, whereas a register entry leaves it live everywhere else. An in-source// codeql[rule-id]comment is the second waiver route. Widening the register appears in the PR diff, andscripts/check_codeql_register_test.gofails the test suite on an entry that names a missing file, lacks a reason, or has become decoupled from the workflow. - A
pull_requestrun certifies a diff, not a tree. On PR events the CodeQL action runs diff-informed: it builds the full database and evaluates every query, then reports only results inside the PR's diff. Tree-wide verdicts come only from the push and scheduled runs. Any "the scanners are green, therefore the tree is clean" claim resting on a PR run is unsound — this repo learned that the expensive way, and ADR-0048 records it. _test.gofiles are outside the Go database.autobuildbuilds packages, not tests, so every_test.gofile in the tree (625 at the time of writing) is unanalyzed. Unchanged from GitHub's default setup; bringing them in would requirebuild-mode: manual.
The gate fails closed: a missing register, a missing filter file, unparseable
SARIF, or findings present with zero rule metadata resolved all fail the job
rather than reporting clean. One shared classifier (.github/codeql-gate.jq)
feeds both the report and the gate, so the two cannot disagree, and the job log
prints three tiers — BLOCKING, ACCEPTED (by name), ADVISORY. Note that
dismissing an alert in the Security tab does not turn the check green: the
gate reads the run's own SARIF and never consults the code-scanning API.
Semgrep — all four registry packs, blocking on any finding.
.github/workflows/semgrep.yml runs p/github-actions, p/golang,
p/javascript and p/python with --error and no continue-on-error, also as a
reusable workflow inside both gates, also with a weekly schedule. The tree is at
zero unsuppressed findings; seven false positives are waived at the line with
nosemgrep: <rule-id> plus a reason, and each waiver was mutation-tested
(removing it makes the finding reappear, so a green scan means the waivers work
rather than the rules having silently stopped matching). One honest limitation:
the rule packs are fetched from the registry at scan time and cannot be vendored —
the Semgrep Rules License v1.0 forbids redistribution, and this is a public MIT
repo — so a registry-side rule addition can turn the lane red with no commit to
blame. The binary version is pinned; the rules are not.
All 53 third-party action references across the 13 workflow files are pinned to
40-hex commit SHAs with the version in a trailing comment, which is also the
form Dependabot updates. Two of those pins were subtly wrong — taken from the
annotated tag object of a mutable major tag rather than the commit it points at,
so they looked like commit pins while resolving to a moving tag — and both are now
the peeled commit, with scripts/check_action_pins_test.go asserting the shape.
Where enforcement actually lands. All of the above is wired into the one
aggregate gate job, and a gate job only blocks a merge where it is a required
status check. CI gate is required on main, and main is the only long-lived
branch, so every scanner above blocks the merge. (Until the dev integration
branch was retired in ADR-0062, its
Dev gate was red-but-not-required; that closed gap and what it interacted with
is recorded under "Known gaps" in docs/SCANNING.md.)
Fleet pulls third-party code from three ecosystems — Go modules at the repo root,
npm packages under web/, and Fedora RPMs inside the sandbox image — and relies
on several deliberate controls to keep a
compromised or fresh-and-unvetted release from reaching main:
-
Go module integrity is verified, with the defaults intact. The repo commits a complete
go.sum, and the build does not set any ofGOFLAGS,GOPROXY,GOSUMDB,GONOSUMCHECK,GONOSUMDB,GOPRIVATE, orGOINSECURE— so Go's out-of-the-box posture applies: modules are fetched through the public proxy (proxy.golang.org) and every download is checked against the committedgo.sumand the public checksum database (sum.golang.org). The checksum DB is not disabled. This is a conscious choice: there is no private module source to exempt, so weakening these (e.g.GONOSUMCHECK,GOFLAGS=-insecure, or aGOPRIVATEcarve-out) would only remove protection. Do not add such an override without a documented reason and a corresponding entry here. -
Dependency-CVE scanning. CI runs
govulncheckagainst the Go module on every PR (thegovulncheckjob in.github/workflows/ci.yml), failing the build on a known-vulnerable dependency that fleet actually calls into. -
npm dependency-CVE scanning.
npm audit --audit-level=lowruns in thewebCI job, lockfile-only (before the install, so a vulnerable lockfile fails fast) against theweb/tree — the only npm tree since thescripts/rampart-servicereference service was removed (ADR-0063) — failing the build on any severity. Like govulncheck, its verdict is a function of the clock as well as the commit: a newly published advisory can redden an unchanged tree, which is the point. There is no accepted-advisory register for this gate: the one advisory no bump could ever clear (adm-zip GHSA-vwc7-r8mq-g2x9, reachable only through the rampart service's ML runtime) was resolved by removing the tree that carried it. -
Container-image CVE scanning. CI also scans the rootless-Podman sandbox image (built from
config/default/sandbox/Containerfile) with Grype in thegrype-scanjob, a surfacegovulncheck(Go modules only) cannot see.scripts/check-grype-policy.shfails the build on a fixable CRITICAL or HIGH CVE — and only in the image's RPM packages. Be precise about that restriction, because it is a real, deliberate limit: Grype also catalogs the Pythondist-infothat Fedora RPMs ship as independent PyPI artifacts, and those records use upstream versions and advisories, so one can claim a fix exists when Fedora has already backported it or has not published an RPM update. Those language records are uploaded to SARIF but do not gate; treating them as a merge gate previously produced hand-maintained pip replacements layered over a coherent distro package set. MEDIUM and below are reported, not blocking. Findings upload to GitHub Security → Code scanning (categorygrype-sandbox-image), and a weekly scheduled scan (.github/workflows/grype-scheduled.yml) catches newly-disclosed CVEs against the existing image between PRs.Scope note:
grype-scanlives inci.yml, so the image is scanned on every non-docs-only PR intomain, on pushes tomain, and weekly. -
Release cooldown — on the two ecosystems that support it.
.github/dependabot.ymlapplies acooldownto the gomod and npm surfaces so Dependabot waits a few days (3 for patch, 7 for minor, 14 for major) before proposing a freshly published release. This blunts fast typosquat / account-takeover attacks, where a malicious version is published and then yanked once the ecosystem flags it. Cooldown applies to version updates only — Dependabot security updates are never delayed, so urgent CVE fixes still flow immediately.The
github-actionsecosystem is the exception, and it is the one where a cooldown would matter most. Dependabot supportscooldownfor gomod and npm only, so the one ecosystem whose "dependency" is the CI definition itself — agithub-actionsbump rewrites.github/workflows/*and therefore changes what CI executes — cannot be made to wait, and it is configured daily againstmain. What contains that is that such a bump meets the fullCI gatelike any other PR (see "Static analysis" above), and that every Dependabot PR takes a human: automatic merging was removed from this repository, so no dependency bump of any ecosystem or bump level reachesmainwithout someone looking at it.
The cooldown reduces the window for a fast attack but is not a guarantee:
a patient attacker who waits out the cooldown, or a compromise the ecosystem
never flags, would still slip through. The committed go.sum + checksum DB,
govulncheck and npm audit are the stronger, always-on controls; the cooldown
is defense-in-depth on top of human review, and it does not cover
github-actions at all.
State-mutating orchestrator routes (POST /tasks, POST /upload, …) accept the
shared elcano_auth session cookie. Cookie-authenticated requests are protected
against cross-site request forgery by a stateless Origin check
(CSRFMiddleware, applied globally before every route group): a mutating request
on the cookie path must carry an Origin header whose host matches the server's
(X-Forwarded-Host when behind a proxy, else Host). A missing, malformed, or
cross-origin Origin is rejected with 403 Cross-origin request blocked.
Requests authenticated with X-API-Key, X-Registration-Token,
Authorization: Bearer …, or the Next-proxy X-Orchestrator-Server-Token are
exempt — browsers do not auto-attach custom headers cross-origin, so those
paths are not CSRF-reachable. On the proxy path the browser's Origin has
already been enforced by the Next.js layer (web/src/app/lib/csrf.ts), which
does not forward the header upstream.
Two operator-facing contracts make this defense-in-depth complete:
- The auth service MUST set
SameSite=Lax(orStrict) on theelcano_authcookie. Fleet reads and deletes that cookie but does not mint it;SameSiteis the browser's first line of defense and blocks the overwhelming majority of CSRF vectors before any server check runs. The Origin check is the backstop. - Non-browser clients that authenticate with the
elcano_authcookie must setOriginexplicitly on everyPOST/PUT/DELETE/PATCH, e.g.Origin: https://fleet.example.com. Requests that omit it receive403. Clients usingX-API-Key/X-Registration-Token/Authorization: Bearerare unaffected.
A deployment's behavior comes from an external client-config bundle (a git
repo FLEET_CLIENT_CONFIG_DIR points at). Treat write access to that repo as
production access, because the bundle is effectively root-equivalent on the
box:
- Its
sandbox/Containerfileis built and run on the host. - Its MCP servers run as host-side subprocesses of the dedicated MCP broker with the selected per-account credentials placed in their environment.
So the README's "credentials never enter the sandbox" guarantee is about the
sandbox — brokered secrets do reach the bundle's own host-side MCP servers
by design. Anyone who can push to the bundle's tracked branch gains host-side
code execution under the fleet service identity and access to those secrets on
the next update. Protect the bundle repo accordingly: restrict who can push,
require signed commits / branch protection, and pin the checkout to a reviewed
commit when you can.
Connector execution runs in a separate fleet mcp-broker process, and the agent
loop's process scrubs every connector environment key once that child is verified
up. Two things are worth stating plainly about the shape of that boundary, so it
is not read as more than it is.
What it protects. After boot, the agent-loop process cannot read a bundle
connector secret, and it cannot exercise one past the gates: the credential owner
re-derives each server's tool allowlist and enabled-server set from its own copy
of the bundle, refuses a scope selection naming a server it does not have, and
refuses a call outside the scope's bound set — so a gating bug in the agent-loop
process restricts a run rather than unbounding it. See ADR-0042 and
docs/MCP-BROKER-SCOPES.md.
What it does not protect: remote-MCP OAuth tokens at rest. Per-user hosted MCP runs resolve their tokens in the child (ADR-0040), but the OAuth control plane — connect, callback, and connection CRUD — is parent-side by design, and the parent therefore installs the at-rest cipher on the chat store. Code running in the parent process can decrypt any user's stored remote-MCP tokens. This is an accepted boundary, not an oversight: browser redirects and user credential intake are host control-plane operations, they are not model-callable, and moving them behind the child would require a second authenticated control protocol without changing where model-initiated MCP calls execute.
The threat model that follows is explicit: compromise of the fleet parent process implies compromise of stored remote-MCP tokens (and of the other secrets the same store cipher protects). The broker boundary is an address-space-isolation claim about connector credentials and agent-initiated calls, not a claim that a fully compromised parent cannot reach the database or its encryption key. Full control-plane isolation is a possible future change, tracked separately rather than left ambiguous.
Recorded here because a migration is applied history, so a question parked in
one is a question nobody can close in place. Migration
internal/sched/db/migrations/022_add_task_credential_allowlist.up.sql points at
this section.
What is stored. tasks.credential_allowlist holds (server, account) pair
names only — never credential values. The values never enter the database:
they live in the process env file and are brokered host-side (internal/creds,
ADR-0003, ADR-0042). The same is true of the sibling columns that name MCP
servers or accounts: tasks.mcp_selection, the approval and remote-MCP seat
columns on the chat store, and the per-conversation MCP account selection. None
of them is a secret store.
What is unsettled. Whether an account name is itself sensitive. A name can carry tenant or customer identity ("acme-prod"), so a deployment that treats its customer list as confidential may want these columns encrypted at rest.
Current posture: not encrypted, and deliberately so. Encrypting one of these columns while its siblings stay plaintext would be theater — the same names appear in several tables, in logs, and in the agent's own system-prompt roster. The honest scope of a fix is store-wide at-rest encryption, which is the same decision the remote-MCP token boundary above already frames: a compromised parent process reaches the store and its key either way. So this is a threat-model decision for a deployment, tracked here rather than silently resolved: if account names are in scope for you, the control is disk/volume encryption plus restricted database access, not a per-column cipher.
This policy covers the code in this repository. Deployments are configured by a separate, operator-supplied client-config bundle and environment file; the security of a given deployment also depends on how those are managed (see above).