Skip to content

Intelligent Ingredient Extraction for Recipe Import #90

Description

@tsueri

Problem Statement

When importing recipes from Swiss cooking sites (swissmilk.ch, fooby.ch, bettybossi.ch), the scraped ingredient lines contain descriptive noise mixed in with the ingredient name. The regex-based IngredientLineParser leaves adjectives, preparation instructions, alternative ingredients, embedded weight specifications, and non-food equipment items in the parsed name field. Analysis of three swissmilk.ch recipes revealed that 55% of ingredient lines are "hard" — containing prep notes after commas ("grüne Spargeln, in 3-4 cm langen Stücken"), German adjective-first word order ("gehackter Peterli"), alternatives joined by "oder", embedded weight specs ("Camembert Suisse à 300 g"), parenthetical clarifications, or equipment items tagged as recipeIngredient ("Runde Gratinform / Kuchenblech 30 cm"). These noisy names produce low-confidence or incorrect matches in the IngredientNormalizer, forcing users to manually clean every imported ingredient line.

Solution

A two-tier intelligent extraction pipeline that cleans ingredient names before normalization, using the project's existing IngredientLineParser for quantity/unit extraction and adding model-based name cleaning.

  • Tier 1: A fine-tuned German NER model (distilbert-base-german-cased, token classification) strips adjectives, preparation notes, parentheticals, and regional modifiers from ingredient names. Fast (~25ms/line), runs in-process on CPU, handles the majority of cases.
  • Tier 2: A local LLM via Ollama Docker sidecar handles critical edge cases requiring reasoning — alternatives ("X oder Y"), equipment detection, embedded weight specs ("à 300 g"), and word-form quantities ("fünf Löffel Zucker"). Batch-processed with the household's canonical ingredient list for disambiguation.
  • Degradation: When Ollama is unavailable, the pipeline falls through to the current IngredientNormalizer. Imports never block on Tier 2.
  • Ingredient creation: When no canonical Ingredient match is found, Tier 2 proposes a new ingredient name; the import form shows a "Neu" badge for user confirmation.

User Stories

  1. As a household member importing a recipe from swissmilk.ch, I want ingredient names like "grüne Spargeln, in 3-4 cm langen Stücken" automatically cleaned to "Spargeln", so I don't manually edit every line.
  2. As a household member, I want adjective modifiers ("gehackter Peterli" → "Peterli") and parenthetical clarifications ("Rahmjoghurt (griechische Art)" → "Rahmjoghurt") stripped automatically.
  3. As a household member, I want the system to recognize alternative ingredients ("Weisswein oder Zitronensaft") and pick the primary one, producing a clean match.
  4. As a household member, I want equipment items ("Runde Gratinform / Kuchenblech 30 cm", "Backpapier") detected and excluded from imported RecipeIngredients.
  5. As a household member, I want embedded weight specifications ("Camembert Suisse à 300 g") stripped from names, using the already-parsed amount field.
  6. As a household member, I want word-form quantities ("fünf Löffel Zucker") recognized and converted to numeric amounts with the correct unit.
  7. As a household member, when no canonical Ingredient matches a cleaned name, I want to see a suggested new ingredient name I can confirm or edit before creation.
  8. As a household member, when the LLM (Tier 2) is unavailable, I want the import to still work using current fuzzy matching as fallback.
  9. As a developer, I want a repeatable CLI tool that scrapes sitemaps from recipe sites, uses a frontier API to generate labeled training data, so the Tier 1 model can be retrained.
  10. As a developer, I want a fine-tuning script that consumes labeled data and produces a model artifact bundled into the Docker image.

Implementation Decisions

Architecture

The import endpoint routes scraped ingredient lines through a new IngredientCleanupPipeline service: parser → pattern router → Tier 1 NER → (Tier 2 LLM dispatch) → IngredientNormalizer fallback.

The existing IngredientLineParser stays — it handles all known Swiss quantity/unit formats correctly and is battle-tested. The model pipeline cleans only the name field.

Tier 1: NER model

  • Base: distilbert-base-german-cased (~260 MB), HuggingFace AutoModelForTokenClassification
  • Task: token classification with labels B-ING, I-ING, O — labeling only the core ingredient tokens
  • Loading: at app startup via create_app(), held in process memory
  • Bundled in Docker image at build time

Tier 2: Local LLM via Ollama

  • Docker sidecar in docker-compose.yml, reachable at http://ollama:11434
  • 3-4B parameter model (q4 quantized), model name configurable via env var
  • Batch processing: all hard lines sent in one POST with canonical ingredient list in context
  • Output per line: cleaned_name, corrected_quantity, corrected_unit, ingredient_id, confidence, is_equipment, suggested_name
  • Timeout: 10 seconds; on failure returns confidence=0, falls through to normalizer

Routing rules

Lines go to Tier 2 when: (a) regex detects critical patterns (oder, à, equipment denylist, word-form numbers, slash between nouns), OR (b) Tier 1 confidence is below threshold (0.7) and IngredientNormalizer also fails to match. All other lines go Tier 1 only.

Ingredient creation

When Tier 2 returns suggested_name, the ScrapedIngredientItem carries it in the import response. The form shows a "Neu" badge pre-filled with the suggestion, editable. On save, the backend creates a new Ingredient row and IngredientAlias. Measurement defaults are left null — UnitConverter handles this.

Training data pipeline

  • Script: backend/scripts/generate_training_data.py — scrapes sitemaps, calls frontier API, outputs labeled JSON pairs
  • Script: backend/scripts/fine_tune.py — trains distilbert-base-german-cased on labeled data, saves model artifact

Schema changes

ScrapedIngredientItem gains optional fields: tier1_cleaned_name, tier2_cleaned_name, corrected_quantity, corrected_unit, suggested_ingredient_name, is_equipment. No database schema changes needed — Ingredient and IngredientAlias already support the creation flow.

Phasing

Phase 1: Training data CLI + generation → Phase 2: Tier 1 NER integration → Phase 3: Tier 2 Ollama integration → Phase 4: Ingredient creation UI

Testing Decisions

Follow existing test patterns: pytest service-level unit tests, no database dependencies, mock external dependencies.

Modules to test:

  • IngredientNameCleaner: mock model, test adjective stripping, prep note removal, edge cases (empty string, single-word names)
  • IngredientCleanupPipeline: mock Tier 1 and Tier 2, test routing logic, verify Tier 2 dispatch decisions, verify fallback behavior
  • IngredientLLMResolver: mock HTTP responses, test batch construction, JSON parsing, error handling, timeout
  • POST /api/recipes/import endpoint: extend existing test_recipes.py with mocked pipeline

Prior art: test_normalizer.py (service unit tests, no DB), test_ingredient_line_parser.py (parser unit tests), test_recipes.py (endpoint integration with mocks).

Out of Scope

  • UI improvements to the import form (cursor visibility, Enter key behavior, match state clarity) — deferred to follow-up issue
  • Measurement defaults (grams_per_el, etc.) for new ingredients — left null
  • Tier 2 model management (pulling, updating, switching) — manual for now
  • Training data augmentation from user corrections — future enhancement
  • Languages other than German

Review Notes (2026-06-05)

German-only scope: The plan targets only German-language recipes. The user's source sites (swissmilk.ch, fooby.ch, bettybossi.ch) are all German. French (scout.ch), Italian (migusto.ch/it) Swiss sites are explicitly out of scope.

Tier 1 failure mode: If the bundled NER model fails to load at startup (corrupted file, incompatible version, OOM), the app fails fast — a clear error is raised and the import endpoint is unavailable until the model is fixed. This is intentional: a broken model should not silently degrade.

Training data scope (task 1.2): Target ~200 recipe URLs (150 from fooby.ch user list + 50 from swissmilk.ch and bettybossi.ch sitemaps), covering all hard patterns from the analysis (adjectives, prep commas, oder alternatives, embedded weights, equipment, parentheticals, adjective-first word order). Output: ~3000-5000 labeled pairs. Quality: frontier API output accepted, 5% manual spot-check.

Docker regeneration: After fine-tuning the Tier 1 model, the Docker image must be rebuilt to bundle the new model artifact. Task 1.4 added for this.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions