Problem Statement
When importing recipes from Swiss cooking sites (swissmilk.ch, fooby.ch, bettybossi.ch), the scraped ingredient lines contain descriptive noise mixed in with the ingredient name. The regex-based IngredientLineParser leaves adjectives, preparation instructions, alternative ingredients, embedded weight specifications, and non-food equipment items in the parsed name field. Analysis of three swissmilk.ch recipes revealed that 55% of ingredient lines are "hard" — containing prep notes after commas ("grüne Spargeln, in 3-4 cm langen Stücken"), German adjective-first word order ("gehackter Peterli"), alternatives joined by "oder", embedded weight specs ("Camembert Suisse à 300 g"), parenthetical clarifications, or equipment items tagged as recipeIngredient ("Runde Gratinform / Kuchenblech 30 cm"). These noisy names produce low-confidence or incorrect matches in the IngredientNormalizer, forcing users to manually clean every imported ingredient line.
Solution
A two-tier intelligent extraction pipeline that cleans ingredient names before normalization, using the project's existing IngredientLineParser for quantity/unit extraction and adding model-based name cleaning.
- Tier 1: A fine-tuned German NER model (distilbert-base-german-cased, token classification) strips adjectives, preparation notes, parentheticals, and regional modifiers from ingredient names. Fast (~25ms/line), runs in-process on CPU, handles the majority of cases.
- Tier 2: A local LLM via Ollama Docker sidecar handles critical edge cases requiring reasoning — alternatives ("X oder Y"), equipment detection, embedded weight specs ("à 300 g"), and word-form quantities ("fünf Löffel Zucker"). Batch-processed with the household's canonical ingredient list for disambiguation.
- Degradation: When Ollama is unavailable, the pipeline falls through to the current IngredientNormalizer. Imports never block on Tier 2.
- Ingredient creation: When no canonical Ingredient match is found, Tier 2 proposes a new ingredient name; the import form shows a "Neu" badge for user confirmation.
User Stories
- As a household member importing a recipe from swissmilk.ch, I want ingredient names like "grüne Spargeln, in 3-4 cm langen Stücken" automatically cleaned to "Spargeln", so I don't manually edit every line.
- As a household member, I want adjective modifiers ("gehackter Peterli" → "Peterli") and parenthetical clarifications ("Rahmjoghurt (griechische Art)" → "Rahmjoghurt") stripped automatically.
- As a household member, I want the system to recognize alternative ingredients ("Weisswein oder Zitronensaft") and pick the primary one, producing a clean match.
- As a household member, I want equipment items ("Runde Gratinform / Kuchenblech 30 cm", "Backpapier") detected and excluded from imported RecipeIngredients.
- As a household member, I want embedded weight specifications ("Camembert Suisse à 300 g") stripped from names, using the already-parsed amount field.
- As a household member, I want word-form quantities ("fünf Löffel Zucker") recognized and converted to numeric amounts with the correct unit.
- As a household member, when no canonical Ingredient matches a cleaned name, I want to see a suggested new ingredient name I can confirm or edit before creation.
- As a household member, when the LLM (Tier 2) is unavailable, I want the import to still work using current fuzzy matching as fallback.
- As a developer, I want a repeatable CLI tool that scrapes sitemaps from recipe sites, uses a frontier API to generate labeled training data, so the Tier 1 model can be retrained.
- As a developer, I want a fine-tuning script that consumes labeled data and produces a model artifact bundled into the Docker image.
Implementation Decisions
Architecture
The import endpoint routes scraped ingredient lines through a new IngredientCleanupPipeline service: parser → pattern router → Tier 1 NER → (Tier 2 LLM dispatch) → IngredientNormalizer fallback.
The existing IngredientLineParser stays — it handles all known Swiss quantity/unit formats correctly and is battle-tested. The model pipeline cleans only the name field.
Tier 1: NER model
- Base: distilbert-base-german-cased (~260 MB), HuggingFace AutoModelForTokenClassification
- Task: token classification with labels B-ING, I-ING, O — labeling only the core ingredient tokens
- Loading: at app startup via create_app(), held in process memory
- Bundled in Docker image at build time
Tier 2: Local LLM via Ollama
- Docker sidecar in docker-compose.yml, reachable at http://ollama:11434
- 3-4B parameter model (q4 quantized), model name configurable via env var
- Batch processing: all hard lines sent in one POST with canonical ingredient list in context
- Output per line: cleaned_name, corrected_quantity, corrected_unit, ingredient_id, confidence, is_equipment, suggested_name
- Timeout: 10 seconds; on failure returns confidence=0, falls through to normalizer
Routing rules
Lines go to Tier 2 when: (a) regex detects critical patterns (oder, à, equipment denylist, word-form numbers, slash between nouns), OR (b) Tier 1 confidence is below threshold (0.7) and IngredientNormalizer also fails to match. All other lines go Tier 1 only.
Ingredient creation
When Tier 2 returns suggested_name, the ScrapedIngredientItem carries it in the import response. The form shows a "Neu" badge pre-filled with the suggestion, editable. On save, the backend creates a new Ingredient row and IngredientAlias. Measurement defaults are left null — UnitConverter handles this.
Training data pipeline
- Script: backend/scripts/generate_training_data.py — scrapes sitemaps, calls frontier API, outputs labeled JSON pairs
- Script: backend/scripts/fine_tune.py — trains distilbert-base-german-cased on labeled data, saves model artifact
Schema changes
ScrapedIngredientItem gains optional fields: tier1_cleaned_name, tier2_cleaned_name, corrected_quantity, corrected_unit, suggested_ingredient_name, is_equipment. No database schema changes needed — Ingredient and IngredientAlias already support the creation flow.
Phasing
Phase 1: Training data CLI + generation → Phase 2: Tier 1 NER integration → Phase 3: Tier 2 Ollama integration → Phase 4: Ingredient creation UI
Testing Decisions
Follow existing test patterns: pytest service-level unit tests, no database dependencies, mock external dependencies.
Modules to test:
- IngredientNameCleaner: mock model, test adjective stripping, prep note removal, edge cases (empty string, single-word names)
- IngredientCleanupPipeline: mock Tier 1 and Tier 2, test routing logic, verify Tier 2 dispatch decisions, verify fallback behavior
- IngredientLLMResolver: mock HTTP responses, test batch construction, JSON parsing, error handling, timeout
- POST /api/recipes/import endpoint: extend existing test_recipes.py with mocked pipeline
Prior art: test_normalizer.py (service unit tests, no DB), test_ingredient_line_parser.py (parser unit tests), test_recipes.py (endpoint integration with mocks).
Out of Scope
- UI improvements to the import form (cursor visibility, Enter key behavior, match state clarity) — deferred to follow-up issue
- Measurement defaults (grams_per_el, etc.) for new ingredients — left null
- Tier 2 model management (pulling, updating, switching) — manual for now
- Training data augmentation from user corrections — future enhancement
- Languages other than German
Review Notes (2026-06-05)
German-only scope: The plan targets only German-language recipes. The user's source sites (swissmilk.ch, fooby.ch, bettybossi.ch) are all German. French (scout.ch), Italian (migusto.ch/it) Swiss sites are explicitly out of scope.
Tier 1 failure mode: If the bundled NER model fails to load at startup (corrupted file, incompatible version, OOM), the app fails fast — a clear error is raised and the import endpoint is unavailable until the model is fixed. This is intentional: a broken model should not silently degrade.
Training data scope (task 1.2): Target ~200 recipe URLs (150 from fooby.ch user list + 50 from swissmilk.ch and bettybossi.ch sitemaps), covering all hard patterns from the analysis (adjectives, prep commas, oder alternatives, embedded weights, equipment, parentheticals, adjective-first word order). Output: ~3000-5000 labeled pairs. Quality: frontier API output accepted, 5% manual spot-check.
Docker regeneration: After fine-tuning the Tier 1 model, the Docker image must be rebuilt to bundle the new model artifact. Task 1.4 added for this.
Problem Statement
When importing recipes from Swiss cooking sites (swissmilk.ch, fooby.ch, bettybossi.ch), the scraped ingredient lines contain descriptive noise mixed in with the ingredient name. The regex-based IngredientLineParser leaves adjectives, preparation instructions, alternative ingredients, embedded weight specifications, and non-food equipment items in the parsed name field. Analysis of three swissmilk.ch recipes revealed that 55% of ingredient lines are "hard" — containing prep notes after commas ("grüne Spargeln, in 3-4 cm langen Stücken"), German adjective-first word order ("gehackter Peterli"), alternatives joined by "oder", embedded weight specs ("Camembert Suisse à 300 g"), parenthetical clarifications, or equipment items tagged as recipeIngredient ("Runde Gratinform / Kuchenblech 30 cm"). These noisy names produce low-confidence or incorrect matches in the IngredientNormalizer, forcing users to manually clean every imported ingredient line.
Solution
A two-tier intelligent extraction pipeline that cleans ingredient names before normalization, using the project's existing IngredientLineParser for quantity/unit extraction and adding model-based name cleaning.
User Stories
Implementation Decisions
Architecture
The import endpoint routes scraped ingredient lines through a new IngredientCleanupPipeline service: parser → pattern router → Tier 1 NER → (Tier 2 LLM dispatch) → IngredientNormalizer fallback.
The existing IngredientLineParser stays — it handles all known Swiss quantity/unit formats correctly and is battle-tested. The model pipeline cleans only the name field.
Tier 1: NER model
Tier 2: Local LLM via Ollama
Routing rules
Lines go to Tier 2 when: (a) regex detects critical patterns (oder, à, equipment denylist, word-form numbers, slash between nouns), OR (b) Tier 1 confidence is below threshold (0.7) and IngredientNormalizer also fails to match. All other lines go Tier 1 only.
Ingredient creation
When Tier 2 returns suggested_name, the ScrapedIngredientItem carries it in the import response. The form shows a "Neu" badge pre-filled with the suggestion, editable. On save, the backend creates a new Ingredient row and IngredientAlias. Measurement defaults are left null — UnitConverter handles this.
Training data pipeline
Schema changes
ScrapedIngredientItem gains optional fields: tier1_cleaned_name, tier2_cleaned_name, corrected_quantity, corrected_unit, suggested_ingredient_name, is_equipment. No database schema changes needed — Ingredient and IngredientAlias already support the creation flow.
Phasing
Phase 1: Training data CLI + generation → Phase 2: Tier 1 NER integration → Phase 3: Tier 2 Ollama integration → Phase 4: Ingredient creation UI
Testing Decisions
Follow existing test patterns: pytest service-level unit tests, no database dependencies, mock external dependencies.
Modules to test:
Prior art: test_normalizer.py (service unit tests, no DB), test_ingredient_line_parser.py (parser unit tests), test_recipes.py (endpoint integration with mocks).
Out of Scope
Review Notes (2026-06-05)
German-only scope: The plan targets only German-language recipes. The user's source sites (swissmilk.ch, fooby.ch, bettybossi.ch) are all German. French (scout.ch), Italian (migusto.ch/it) Swiss sites are explicitly out of scope.
Tier 1 failure mode: If the bundled NER model fails to load at startup (corrupted file, incompatible version, OOM), the app fails fast — a clear error is raised and the import endpoint is unavailable until the model is fixed. This is intentional: a broken model should not silently degrade.
Training data scope (task 1.2): Target ~200 recipe URLs (150 from fooby.ch user list + 50 from swissmilk.ch and bettybossi.ch sitemaps), covering all hard patterns from the analysis (adjectives, prep commas, oder alternatives, embedded weights, equipment, parentheticals, adjective-first word order). Output: ~3000-5000 labeled pairs. Quality: frontier API output accepted, 5% manual spot-check.
Docker regeneration: After fine-tuning the Tier 1 model, the Docker image must be rebuilt to bundle the new model artifact. Task 1.4 added for this.