Skip to content

Add upload_image: re-host an external image on Substack - #20

Merged
marcomoauro merged 12 commits into
mainfrom
image-upload
Aug 8, 2026
Merged

marcomoauro merged 12 commits into
mainfrom
image-upload

Conversation

@marcomoauro

Copy link
Copy Markdown
Owner

Summary

Adds an upload_image MCP tool that takes an http(s) image URL, downloads it, re-encodes it as a base64 data URI, and uploads it to POST /api/v1/image, returning a Substack-hosted URL that goes straight into image2.src in set_post_body.

This overturns a documented limit. CLAUDE.md recorded the endpoint as unreachable — "hangs in all three tested encodings, do not implement." That record was wrong. Re-measured live on 2026-08-08 from the authenticated dashboard, the endpoint answers 200 in ~300ms. The three earlier attempts failed because they sent the wrong thing, not for a header detail or a Cloudflare wall: the body is JSON {image: "data:<mime>;base64,…"}, a data URI built in the editor from canvas.toDataURL().

Proven end to end on a real draft before any code was written: upload → captionedImage node → PUT → the editor renders it through Substack's CDN (substackcdn.com/image/fetch/…).

What changed

  • SubstackApi.uploadImage({image, post_id}) — JSON data-URI POST to /api/v1/image, post_id mapped to postId and omitted when absent.
  • src/tools/upload_image.js — schema (url + optional post_id) and handler: guard → fetch → validate → encode → upload → map the response.
  • src/logger.js — elides a base64 data: URI's payload (keeps the mime prefix + omitted length). SubstackApi logs every request body at info, so without this a real upload would put hundreds of KB of base64 on one line. A post body, being prose, is still logged in full.
  • Registered in the tools registry; README section added; CLAUDE.md corrected.

Why the tool downloads the image itself

Substack server-fetches only its own S3 bucket — an external URL passed as image answers 400 "Failed to fetch image". So the download happens here, and that fetch (the one place this server contacts a caller-chosen host) is guarded:

  • http/https only.
  • SSRF block on private/loopback/link-local addresses after DNS resolution, re-checked on every redirect hop (redirects are followed manually — fetch's default follow would contact a redirect target before we ever saw its host). IPv6 hosts are de-bracketed before lookup, and IPv4-mapped IPv6 is caught in both the dotted and the compressed-hex form the URL parser emits (::ffff:a9fe:a9fe = 169.254.169.254).
  • image/* content-type, HEIC (and -sequence variants) refused early with a convert-first message.
  • A 10 MB cap — ours, not Substack's (its MAX_FILE_SIZE could not be read from the minified bundle) — enforced against a declared Content-Length before the body is read, then against the buffered length; a per-request timeout bounds an unbounded chunked response.

Test plan

  • 706 tests pass on Node 24 (.nvmrc) and Node 22 (engines floor).
  • Every guard is mutation-tested (broken on purpose, confirmed the test fails, restored).
  • Real entrypoint over stdio advertises upload_image in tools/list with additionalProperties: false, required: ['url'].
  • Full upload→render loop verified live on a real draft (then reverted).

For the reviewer

One thing is unverified: every live measurement used the browser session cookie, not SUBSTACK_SESSION_TOKEN in a header as SubstackApi sends it. Equivalent in principle, unconfirmed through SubstackApi — the first thing to check if the tool misbehaves against the real API. Accepted residual risks, documented in-code: the DNS-rebinding TOCTOU (validated then fetched separately) and exotic IPv4-in-IPv6 encodings (6to4/Teredo/NAT64).

🤖 Generated with Claude Code

marcomoauro and others added 12 commits August 8, 2026 10:10
POST /api/v1/image was recorded in CLAUDE.md as unreachable in all three
tested encodings. Re-measured live on 2026-08-08: it works. The body is a
JSON data URI, not a file or a URL, which is why multipart and
form-urlencoded failed. External URLs are rejected (Substack only
server-fetches its own S3 bucket), so the tool downloads and re-encodes.
End-to-end proven on a real draft: upload renders through Substack's CDN.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Nine TDD tasks: SubstackApi.uploadImage, the tool (fetch/encode/upload),
content-type and HEIC validation, scheme + DNS-resolved SSRF guards, size
cap, server registration, logging assertions, docs correction, and
two-Node verification.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
First of nine tasks adding an upload_image MCP tool. This adds only the
thin API method and the MSW test scaffolding (IMAGE_URL,
IMAGE_UPLOAD_RESPONSE, imageUploadHandler) it needs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
fetchImpl was called with fetch's default redirect: 'follow', but assertPublicUrl only
validated the original host - a public host that 3xx-redirects to a private address
(e.g. 169.254.169.254) bypassed the SSRF guard entirely. fetchGuarded now follows
redirects manually, validating each hop's host before contacting it.

Also: upload_image now logs args and parse failures like every other tool
(comment_on_post, get_draft, restack_item); the unreachable try/catch in
assertPublicUrl is gone now that the schema's .url() already guarantees a parseable
URL; and three comments were corrected to match what the code actually does (IPv6
embedding coverage, the memory cap's real purpose). Removed the unused captureLogs
import from the spec.
The tool's own lines never carry the data URI, but SubstackApi logs every
request body at info — so a real image upload would put hundreds of KB of
base64 on one line. The logger now elides a base64 data URI's payload,
keeping the mime prefix and the omitted length; a post body, being prose,
is still logged in full. This is what makes "the data URI is never logged"
true end to end, not just in the tool.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds the import and registry entry, plus the README section the sync test
requires, and updates the two hardcoded tool lists (server.spec, index.spec)
and the set_post_body note that claimed images could not be uploaded.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Rewrites the "cannot be uploaded / do not implement" record with what was
measured live on 2026-08-08 — the data-URI shape, the S3-only server-fetch,
the guards, the logger elision, and the header-auth caveat — and updates the
fork-survey note that called image-in-post impossible.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Closes a latent SSRF hole and two robustness gaps the whole-feature review
surfaced:
- IPv6-mapped host in the compressed hex form the URL parser emits
  (::ffff:a9fe:a9fe = 169.254.169.254) is now caught; the dotted-only regex
  missed it. IPv6 hosts are de-bracketed before the DNS lookup, which also
  fixes IPv6 URLs erroring before the guard.
- The size cap is enforced against a declared Content-Length before the body
  is read, then against the buffered length; an unbounded chunked body is
  bounded by a new per-request fetch timeout.
- HEIC -sequence variants join the reject set.

Security-critical additions are mutation-verified. Also corrects the CLAUDE.md
claim that create_draft_post is the only reader of ZodError.issues.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@marcomoauro
marcomoauro merged commit a33864d into main Aug 8, 2026
2 checks passed
@marcomoauro
marcomoauro deleted the image-upload branch August 8, 2026 12:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant