Skip to content

fix: refuse create_filter actions that apply TRASH or SPAM - #25

Merged
jgalea merged 1 commit into
jgalea:mainfrom
Kaliratech:fix/refuse-destructive-filter-labels
Sep 3, 2026
Merged

jgalea merged 1 commit into
jgalea:mainfrom
Kaliratech:fix/refuse-destructive-filter-labels

Conversation

@Kaliratech

Copy link
Copy Markdown
Contributor

Fixes #23.

The problem

create_filter passed add_label into addLabelIds with no validation of the value:

if (args.add_label) {
  action.addLabelIds = await (provider as GmailProvider).resolveLabelIds(
    [args.add_label as string], { create: args.create_label as boolean | undefined }
  );
}

Gmail's system label IDs are literally TRASH and SPAM, so a single call could install a filter that silently destroys or junks every future matching message.

A filter is not a one-off action, it is persistent server-side state. It keeps applying to mail that arrives long after the call that created it, and nothing in the client transcript shows it working. Combined with #22, where attacker-controlled email text reached the model outside the fences, an injected instruction had a durable and invisible payload available to it.

For anyone comparing against older versions: 0.2.0 had an allowlist here, but it validated which action keys could appear (addLabelIds, removeLabelIds), not which label values were acceptable, so TRASH passed through there too. This is a longstanding gap rather than a regression.

The fix

Resolve first, then refuse if the resolved IDs include a destructive system label. Checking after resolution is deliberate: resolveLabelIds maps the name "Trash" to the ID TRASH through its byName lookup, so a check on the raw argument would be bypassable by passing the name instead of the ID.

archive and mark_read are untouched, since they only ever remove INBOX and UNREAD.

If you would rather keep the capability reachable, an explicit opt-in argument or an env gate would also close the injected-call path. The goal is only that a model reading untrusted mail cannot arrive here by accident.

Tests

Two added to tests/tools/gmail-only.test.ts.

The load-bearing assertion is expect(filters.create).not.toHaveBeenCalled(), not merely that an error came back: the point is that no filter reaches Gmail on the refusal path. Both fail without the guard:

⎯⎯⎯ Failed Tests 2 ⎯⎯⎯
Tests  2 failed | 19 passed (21)

The SPAM test passes the ID form rather than a lowercase name, with a comment explaining why: the test mock passes unknown labels through untouched, so "spam" would never become "SPAM" in the test and the assertion would have exercised the mock instead of the guard.

Suite

256 passed across 22 files, excluding tests/signal-resilience.test.ts. That file passes 4/4 on its own but is timing-sensitive under parallel load; it failed the same way on unmodified main on this machine before any of my changes, so it is unrelated.

Also verified against a real Gmail account: a create_filter call with add_label: "TRASH" is refused, and list_filters output on the target account is identical before and after, so nothing is created on the refusal path.

Fixes jgalea#23.

`add_label` reached `addLabelIds` with no validation of the value, and Gmail's
system label IDs are literally `TRASH` and `SPAM`, so one call could install a
filter that silently destroys or junks every future matching message. A filter
is persistent server-side state: it keeps working long after the call that
created it, and nothing in the client transcript shows it running.

That matters more than a normal mislabel here because this server feeds
attacker-controlled email text to a model (see jgalea#22), so an injected instruction
had a durable, invisible payload available to it.

The check runs after label resolution rather than on the raw argument, so a
label *named* "Trash" cannot smuggle the system ID through: resolveLabelIds
maps that name to the TRASH id via its byName lookup.

`archive` and `mark_read` are untouched; they only ever remove INBOX and UNREAD.

Two tests added. Both fail without the guard (the filter gets created), and the
assertion that matters is `filters.create` never being called, not just the
error being returned. The SPAM test deliberately passes the ID form and says
why in a comment: the test mock passes unknown labels through untouched, so a
lowercase name would have tested the mock instead of the guard.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

create_filter passes add_label straight to addLabelIds, so a filter can apply TRASH or SPAM

2 participants