Skip to content

Architecture

t957095 edited this page Jun 15, 2026 · 4 revisions

ShelfWise Architecture

ShelfWise is a linear data pipeline wrapped in a FastAPI backend and a vanilla JavaScript frontend. This page explains the major components and how data flows through the system.

High-Level Data Flow

User Input (UPCs / CSV)
    ↓
Multi-Source Scraper
    ↓
Product Reasoning Agent
    ↓
Image Verification Pipeline
    ↓
Optional LLM / Foundry IQ Enrichment
    ↓
SQLite Storage
    ↓
Frontend UI + Marketplace Exports

Component Diagram

Layer Responsibility Key Files
Input Layer Accept UPCs manually or via CSV upload. frontend/index.html, frontend/app.js, main.py
Scraper Layer Query public product databases concurrently. backend/scraper.py, backend/scraper_registry.py, backend/scraper_registry.json
Reasoning Layer Consolidate conflicting source data. backend/foundry_agent.py
Image Layer Download, score, deduplicate images; marketplace-aware sizing; manual uploads. backend/image_verifier.py, backend/image_search.py
LLM Layer Optional enrichment via Azure OpenAI / GitHub Models / Ollama. backend/foundry_agent.py, backend/foundry_iq.py
Storage Layer Persist products and jobs. backend/database.py (shelfwise.db)
Output Layer Render products and export to marketplaces. frontend/app.js, backend/main.py

Backend Structure

main.py

FastAPI application entry point. Responsibilities:

  • Lifespan initialization of SQLite, httpx client, scraper, reasoning agent, and Foundry IQ service.
  • CORS configuration for local development.
  • Static file serving for the frontend at /app.
  • Background task process_upc() orchestrating cache → scrape → reason → store → knowledge graph.

scraper.py

UPCScraper.scrape_all(upc) handles concurrent lookups across:

  • Core sources with dedicated parsers (Open Food Facts, UPCItemDB, BarcodeLookup, Go-UPC, Buycott, EANdata, Lookify, UPCDatabase, Brave Search, Google Search).
  • Registry sources loaded from scraper_registry.json (271 configurable sources; default 10 queried, configurable via SHELFWISE_MAX_REGISTRY_SOURCES).

Resilience & observability features:

  • Circuit breaker pattern (threshold/recovery configurable via env vars)
  • Exponential backoff retry with error classification (no retry on 4xx except 429)
  • Rotating user-agents
  • Request coalescing
  • Per-source ScraperHealth tracking calls, success rate, and average latency
  • Demo fallback data for 3 canonical UPCs

foundry_agent.py

ProductReasoningAgent.consolidate(upc, raw_data_list) performs:

  1. Source weighting and filtering
  2. Name resolution with Jaccard deduplication
  3. Brand and category resolution
  4. Attribute merging
  5. Description generation
  6. Verified image selection
  7. Confidence scoring
  8. Citation generation
  9. Optional LLM enrichment

foundry_iq.py

Local simulation of a Foundry IQ knowledge graph:

  • ProductKnowledgeGraph: in-memory graph with BM25-style semantic search.
  • FoundryIQService: query knowledge, reason over products, export ontology, permission model, query history.
  • Catalog ingestion from SQLite.

image_verifier.py

ProductImageVerifier downloads and scores images on:

  • White / clean background
  • Quality and aspect ratio
  • Focus / edge density
  • Perceptual-hash deduplication

Optional LLM vision verification if an OpenAI-compatible endpoint is configured.

database.py

SQLite persistence layer:

  • Tables: products, jobs
  • WAL mode for concurrency
  • Indexed columns: name, brand, category, confidence, status
  • Thread-local connections and bulk upserts

cache.py

Thread-safe TTL LRU cache:

  • upc_cache: 2000 entries, 10-minute TTL
  • http_cache: cached HTTP responses

Frontend Structure

index.html

Single-page application shell:

  • Skip link and semantic header
  • UPC textarea and CSV upload
  • Submit / demo buttons
  • Live job status with progress bar
  • Product grid with sort and search
  • Portfolio analytics section
  • Export grid (11 formats including DoorDash, Uber Eats, Grubhub)
  • Reasoning trace modal, image lightbox, toast container

app.js

Frontend logic:

  • State management for products, job ID, and SSE connection
  • Event handlers for submit, demo, CSV upload, export, sort, search, clear
  • startJobStream(jobId): connects to /api/jobs/{job_id}/stream
  • renderProductCard(): accessible product cards with confidence badges and citations
  • showReasoningTrace(): modal with step-by-step trace
  • Keyboard shortcuts: Ctrl/Cmd+Enter submit, / focus search, Esc close modals

styles.css

  • Dark theme using CSS custom properties
  • Responsive grid and flexbox layouts
  • WCAG 2.1 AA focus indicators
  • prefers-reduced-motion and prefers-contrast support

sw.js

Service worker caches static assets and API GET responses for offline support.

Security & Operations

  • .env and database files are gitignored.
  • Dockerfile runs as a non-root user.
  • No secrets are required for the deterministic local mode.

Clone this wiki locally