Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 33 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -9,3 +9,36 @@ DOT_TREASURY_GRAPHQL_URL=https://eco-api.dotreasury.com/graphql

# Identifiers for the SubSquare API endpoints
SUBSQUARE_IDENTITY_SERVER_URL=https://id2.statescan.io

# ---- Rate limiting ---------------------------------------------------------
# Fixed-window rate limit on the /mcp endpoint. Rules are optional: every
# variable has a default, so you can omit them entirely. Defaults match the
# SubSquare GraphQL gateway rate limit plugin.

# Toggle rate limiting on/off (default: true)
RATE_LIMIT_ENABLED=true

# Requests carrying a valid API token are limited per user (token `sub`) using
# the built-in quota plan named by the token's `quota` claim (standard /
# premium — see README). The knobs below only apply to tokenless (anonymous)
# requests, which are limited per IP.

# Max requests allowed per anonymous IP within the window (default: 100)
RATE_LIMIT_MAX=100

# Window length in seconds (default: 1)
RATE_LIMIT_WINDOW=1

# Semicolon-separated IPs exempted from rate limiting (default: none)
RATE_LIMIT_WHITELIST_IPS=

# Max number of distinct clients tracked in the in-memory store (default: 10000)
RATE_LIMIT_STORE_SIZE=10000

# HS256 secret used to verify API tokens sent as `Authorization: Bearer <jwt>`.
# DEDICATED to MCP tokens — independent from subsquare-backend's own
# user-session JWT_SECRET_KEY. It must match the key the backend uses to sign
# MCP tokens. When set, the token's `sub` (user ID) and `quota` (plan) claims
# drive rate limiting; invalid or expired tokens are rejected with 401. Leave
# empty to disable token verification (per-IP limiting only).
MCP_API_JWT_SECRET_KEY=
76 changes: 76 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,3 +37,79 @@ Available `--scope` values:
```
http://127.0.0.1:3210/mcp
```

# Rate limiting

The HTTP endpoint is protected by a fixed-window rate limit, applied before
the MCP transport handles a request. Rules mirror the SubSquare GraphQL gateway
and are configurable through environment variables (see `.env.example`); the
server must be restarted after changing them.

## Anonymous requests (per IP)

Requests without a token are treated as anonymous and limited per client IP
using the operator-tunable knobs below (defaults match the `standard` quota
plan).

| Env var | Default | Description |
| -------------------------- | ------- | --------------------------------------------- |
| `RATE_LIMIT_ENABLED` | `true` | Set to `0`/`false` to disable rate limiting |
| `RATE_LIMIT_MAX` | `100` | Max requests per anonymous IP per window |
| `RATE_LIMIT_WINDOW` | `1` | Window length in seconds |
| `RATE_LIMIT_WHITELIST_IPS` | (none) | Semicolon-separated IPs exempt from the limit |
| `RATE_LIMIT_STORE_SIZE` | `10000` | Max clients tracked in the in-memory store |
| `MCP_API_JWT_SECRET_KEY` | (unset) | HS256 secret verifying API tokens (see below) |

The client IP is derived from reverse-proxy headers when present
(`cf-connecting-ip`, then `x-real-ip`, then `x-forwarded-for`), falling back to
the TCP peer address. Put the server behind a trusted proxy (e.g. Cloudflare)
if those headers can be spoofed in your network.

## API token requests (per user)

When `MCP_API_JWT_SECRET_KEY` is set, a client can present an HS256 JWT as
`Authorization: Bearer <token>`. Valid tokens are rate limited **per user**
(`sub` claim) instead of per IP, using the quota plan named by the token's
`quota` claim. An invalid or expired token is rejected with `401 Unauthorized`.

`MCP_API_JWT_SECRET_KEY` is a **dedicated key** for MCP tokens: it is separate
from subsquare-backend's own user-session `JWT_SECRET_KEY`, and must match the
key the backend signs MCP tokens with. Tokens are issued by subsquare-backend
using the same [`jsonwebtoken`](https://www.npmjs.com/package/jsonwebtoken)
package and conventions — `jwt.sign(content, secret, { expiresIn })` (HS256),
verified here with `jwt.verify(token, secret)`.

Token payload:

```json
{
"sub": "user_9527",
"jti": "mcp_tk_abc123",
"exp": 1716283700,
"quota": "premium"
}
```

| Claim | Required | Meaning |
| ------- | -------- | ------------------------------------------------ |
| `sub` | yes | Unique user ID; the rate limit bucket key |
| `exp` | yes | Expiry timestamp (seconds); hard safety floor |
| `jti` | no | Unique token ID (reserved for future revocation) |
| `quota` | no | Rate limit plan; defaults to `standard` |

Built-in quota plans (hardcoded for now):

| Plan | Max requests per window | Window |
| ---------- | ----------------------- | ------ |
| `standard` | 100 | 1s |
| `premium` | 600 | 1s |

An unknown or missing `quota` falls back to `standard`.

## Responses

A client that exceeds its limit receives `429 Too Many Requests` with
`Retry-After` and `X-RateLimit-*` headers, plus a JSON-RPC error body
(`code -32029`) that MCP clients can surface as a retryable error. An invalid
or expired token receives `401 Unauthorized` (`code -32002`, plus
`WWW-Authenticate: Bearer`).
3 changes: 2 additions & 1 deletion package.json
Original file line number Diff line number Diff line change
Expand Up @@ -17,9 +17,10 @@
"@koa/router": "^15.7.0",
"@modelcontextprotocol/sdk": "^1.30.0",
"bignumber.js": "^11.1.5",
"dotenv": "^17.2.3",
"jsonwebtoken": "^8.5.1",
"koa": "^3.2.1",
"lodash": "^4.18.1",
"dotenv": "^17.2.3",
"lru-cache": "^11.5.2",
"p-limit": "^7.3.2",
"undici": "^6.28.0",
Expand Down
95 changes: 95 additions & 0 deletions pnpm-lock.yaml

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

43 changes: 43 additions & 0 deletions src/config/quota.js
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
// Built-in rate limit quota plans.
//
// An API token's `quota` claim (see README) selects one of these plans, which
// defines how many requests that user may send per window. Two tiers are
// hardcoded for now; make this configurable when more tiers are needed.
//
// | Plan | Max requests per window | Window |
// |------------|-------------------------|--------|
// | standard | 100 | 1s |
// | premium | 600 | 1s |
//
// Tokenless (anonymous) requests do not carry a quota claim; they use
// `anonymousPlan` below, derived from the operator-tunable rate limit knobs.

import { rateLimitConfig } from "./rateLimit.js";

export const QUOTA_PLANS = Object.freeze({
standard: Object.freeze({ maxRequests: 100, windowSeconds: 1 }),
premium: Object.freeze({ maxRequests: 600, windowSeconds: 1 }),
});

// Plan applied when a token carries no quota claim or an unknown one.
export const DEFAULT_QUOTA = "standard";

export function getQuotaPlan(quota) {
const name = Object.hasOwn(QUOTA_PLANS, quota) ? quota : DEFAULT_QUOTA;
const plan = QUOTA_PLANS[name];

return Object.freeze({
name,
maxRequests: plan.maxRequests,
windowMs: plan.windowSeconds * 1000,
});
}

// Requests without a token are treated as anonymous and limited per IP with
// the operator-tunable environment knobs (their defaults match the `standard`
// quota plan).
export const anonymousPlan = Object.freeze({
name: "anonymous",
maxRequests: rateLimitConfig.maxRequests,
windowMs: rateLimitConfig.windowMs,
});
Loading