Skip to content

Performance & cost: gpt-4o-mini, gzip + caching, parallel GeoJSON#4

Open
joaquinOEF wants to merge 2 commits into
mainfrom
feat/perf-cost
Open

Performance & cost: gpt-4o-mini, gzip + caching, parallel GeoJSON#4
joaquinOEF wants to merge 2 commits into
mainfrom
feat/perf-cost

Conversation

@joaquinOEF

Copy link
Copy Markdown
Owner

Low-risk performance/cost wins.

  • gpt-4o → gpt-4o-mini for the chat summary — ~50× cheaper, still fast; the response is short prose so quality holds.
  • gzip (compression) — the ~1.9 MB of GeoJSON ships at ~200–300 KB.
  • Static cachingexpress.static maxAge: 1d so repeat visits don't re-download assets/GeoJSON (index.html stays uncached via the SPA fallback).
  • Parallel GeoJSONosmService loads the 6 risk-zone files with Promise.all instead of sequentially.
  • Embed batch 50 → 500 — far fewer OpenAI round-trips → faster startup indexing.

Deferred (intentionally)

Persisting embeddings so the server doesn't re-embed ~6k features on every boot. On Replit, a disk cache doesn't reliably survive cold starts, so doing it right needs a persistent vector store (Qdrant/pgvector) — that's an M-effort follow-up, not a low-risk one-liner. The larger batch size already cuts the re-embed time.

⚠️ Stacked on #3 (security) — merge #3 first, then this. npm run build passes (the vite.ts:39 type error is pre-existing).

🤖 Generated with Claude Code

Joaquin van Peborgh and others added 2 commits June 8, 2026 17:13
… OpenAI timeout

The /api/chat/* and /api/osm/* endpoints called OpenAI / Overpass with no auth, no rate
limit, no input validation, and no body-size cap — anyone could run up the OpenAI bill, use
us as a free proxy, or OOM the server with a huge payload.

- express-rate-limit (40/min/IP) on /api/chat and /api/osm
- zod-validate /api/chat/query (string ≤500 chars + optional bbox tuple); cap index-assets
  (≤5000), index-geojson features (≤20000), overpass query length (≤10k)
- express.json/urlencoded limit 1mb
- OpenAI client timeout 15s + maxRetries 2 (a hung call can't block the server)
- health check no longer leaks OPENAI_API_KEY presence
- .gitignore .env*

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…, bigger embed batch

- generateSummary: gpt-4o → gpt-4o-mini (~50× cheaper, fast; summary is short prose)
- compression() middleware → the ~1.9MB GeoJSON ships gzipped (~200-300KB)
- express.static maxAge 1d → repeat visits don't re-download the static assets/GeoJSON
- osmService: load the 6 risk-zone files in parallel (Promise.all) instead of sequentially
- embedding batch 50 → 500 (far fewer OpenAI round-trips → faster startup indexing)

Deferred (needs a persistent vector store to be reliable on Replit cold starts, so not a
low-risk one-liner): persisting embeddings so the server doesn't re-embed on every boot.
Mitigated here by the larger batch size. Tracked as a follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant