Repository navigation
Backend Guide
This guide is for developers who want to understand, extend, or debug the ShelfWise backend.
backend/
├── main.py # FastAPI app and endpoints
├── models.py # Pydantic models
├── database.py # SQLite persistence
├── scraper.py # Async UPC scraper
├── scraper_registry.py # Dynamic registry loader
├── scraper_registry.json # 271 source definitions
├── foundry_agent.py # Product reasoning agent
├── foundry_iq.py # Foundry IQ knowledge graph simulation
├── image_verifier.py # Image scoring pipeline
├── cache.py # TTL LRU caches
├── .env.example # Environment template
└── shelfwise.db # SQLite database (generated)
cd backend
python -m uvicorn main:app --reload --host 127.0.0.1 --port 8000The --reload flag enables auto-reload during development.
- Defines all FastAPI endpoints.
- Manages application lifespan: opens SQLite, initializes scraper/agent/IQ service.
- Serves static frontend files at
/app. - Background task
process_upc()coordinates the enrichment pipeline.
The UPCScraper class is the heart of data acquisition.
scraper = UPCScraper(http_client, registry)
results = await scraper.scrape_all(upc)Core concepts:
- Sources: each source has a parser, weight, and optional parser override.
-
Concurrency: all sources are queried with
asyncio.gather. - Circuit breakers: temporarily skip failing sources.
- Retries: exponential backoff on transient errors.
- Request coalescing: duplicate in-flight requests for the same UPC are deduplicated.
The ProductReasoningAgent turns raw source data into a single consolidated product.
agent = ProductReasoningAgent(...)
product = await agent.consolidate(upc, raw_data_list)Steps:
- Filter sources below a minimum weight.
- Cluster names by Jaccard similarity and pick the winning cluster.
- Resolve brand and category by weighted vote.
- Merge attributes from all sources.
- Generate a description from the consolidated fields.
- Select the highest-scoring verified image.
- Compute confidence from source coverage and field completeness.
- Build citations showing which source contributed which fields.
ProductImageVerifier downloads each candidate image and scores it.
Scores:
-
white_background: ratio of near-white pixels -
quality: resolution and aspect ratio -
focus: edge density using Sobel filter -
dedup: perceptual hash Hamming distance
verifier = ProductImageVerifier()
scored_images = await verifier.verify_images(image_urls)SQLite layer with WAL mode enabled.
Tables:
-
products: consolidated product records with JSON columns -
jobs: batch job status
Important functions:
upsert_product(product)get_products(**filters)create_job(job_id, total)update_job(job_id, ...)
Local knowledge graph for natural-language queries.
-
ProductKnowledgeGraph: nodes (products) and edges (brand/category/attribute similarity). -
FoundryIQService: query, reason, ingest, ontology export, history. - Permission levels:
guest,user,admin.
- Add a source definition to
scraper_registry.json:
{
"name": "ExampleDB",
"url_template": "https://api.example.com/{upc}",
"weight": 0.6,
"enabled": true
}-
If the source needs custom parsing, add a parser function in
scraper.pyand reference it in the registry. -
Restart the backend and test with a known UPC.
Run backend tests from the project root:
pytest tests/ -vLint and format:
ruff check backend/
ruff format backend/- Enable FastAPI debug logs:
UVICORN_LOG_LEVEL=debug - Inspect the SQLite database directly:
sqlite3 backend/shelfwise.db - Use
/api/healthand/api/metricsto diagnose scraper health. - View reasoning traces in the frontend modal to understand consolidation decisions.
ShelfWise — AI Product Portfolio Builder · GitHub · MIT License
- Home
- Getting Started
- Use Cases
- Roadmap
- Architecture
- API Reference
- Configuration
- Backend Guide
- Frontend Guide
- Scraping & Reasoning
- Testing
- Deployment
- Changelog
Quick Start
docker-compose up --build
# open http://localhost:8000/app