From 7b0496a0c1e19939b7503ee8260f430018ecd3dd Mon Sep 17 00:00:00 2001 From: nathan nelson Date: Fri, 24 Jul 2026 11:23:16 -0600 Subject: [PATCH] =?UTF-8?q?feat:=20stub-extraction=20guard=20=E2=80=94=20t?= =?UTF-8?q?he=20roster/directory=20class?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Found by the 2026-07-24 post-lift audit. A page with no prices, no table and no wall gave every existing guard nothing to arm on, so a content-bearing page could return its boilerplate and drop everything else, silently: sbir.gov/awards 10,698 visible -> 421 chars: only the "official website of the United States government" banner. Zero award data. noaa SBIR archive 6,373 visible -> 361 chars: 1 of 3+ success stories kept. apex/our-team 1,994 visible -> 161 chars: a marketing sentence; all four staff entries dropped. nomic.ai/pricing 3,043 visible -> 259 chars: marketing copy, zero prices (the price guard could not arm). This is NOT the bare yield ratio measured and rejected in v0.1.2. That design fires on neither real pricing failure (smith.ai 13%, quo 21% retention) and false-positives on nav-heavy pages. Retention is half the condition; the other half is an absolute stub-sized output, and that floor is what separates the groups. The rejection reasoning is preserved in the code comment so this cannot be "simplified" back. Calibrated by replaying every content-bearing page in a live 162-page cache: fires on 5 (3.1%), each verified a genuine omission. Silent on the LinkedIn false-positive page (10.2%), heyrosie (14.0%), smith.ai (13.2%), quo (20.7%) and eur-lex (5.2%, a legitimate terse result list, which is why the floor must not rise past ~570). Calibration note worth keeping: measure with a FRESH extract_content, never a cached markdown column. Stored rows can carry include_links=True, which inflated noaa to 614 chars and hid it as "healthy" during the first pass. Gate: ruff, mypy (16 files), 247 passed / 6 deselected. Co-Authored-By: Claude Opus 4.8 (1M context) --- CHANGELOG.md | 11 ++++++++ README.md | 11 ++++++++ src/tearsheet/content.py | 57 +++++++++++++++++++++++++++++++++++--- tests/test_quality.py | 60 ++++++++++++++++++++++++++++++++++++++-- 4 files changed, 133 insertions(+), 6 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index b7bbf25..3948d9e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,17 @@ failure. For the trust model and the evaluation harness behind those calls, see ## [Unreleased] ### Added +- **Stub-extraction guard** — the roster/directory class. A page carrying real content + (≥ 1,000 chars of visible text) whose extraction is a stub (< 500 chars) retaining under + 10% now emits a `warning:` line. Closes a gap where a page with no prices, no table, and + no wall gave every existing guard nothing to arm on: `idahoapexaccelerator.com/our-team` + returned a marketing sentence and dropped all four staff entries, `sbir.gov/awards` + returned only the government-site banner, and a NOAA SBIR archive kept 1 of 3+ stories — + all silently. Deliberately **not** the bare yield ratio rejected in v0.1.2 (it fires on + neither real pricing failure and false-positives on nav-heavy pages); pairing retention + with an absolute output floor is what separates the two groups. Calibrated by replaying + all 162 content-bearing pages in a live cache: fires on 5 (3.1%), every one a genuine + omission, and stays silent on the LinkedIn false-positive page and every good pricing page. - MIT `LICENSE`, project URLs, and a GitHub Actions CI gate (`ruff` + `mypy` + `pytest`, matrixed over Python 3.12 and 3.13). - Synthetic false-positive fixture so the probation corpus republishes no real individual's page. diff --git a/README.md b/README.md index 38edccd..c164e30 100644 --- a/README.md +++ b/README.md @@ -126,6 +126,17 @@ a research tool for reading the public web — not for evading paywalls or bot d reported and never cached, and previously poisoned cache rows are evicted, not replayed. Note `--raw` deliberately uses the plain fetch, not the browser: a rendered DOM can be *worse* (on smith.ai the consent overlay replaced the pricing table the raw fetch still carried). +- **Listings and rosters can come back as a boilerplate stub.** A page with no prices, no + table, and no wall gave the earlier guards nothing to arm on, so a directory page could + return its marketing sentence and silently drop every row. Measured 2026-07-24 across the + live cache: `idahoapexaccelerator.com/our-team` returned a 161-char blurb from a 1,994-char + roster and dropped all four staff entries; `sbir.gov/awards` returned only the "official + website of the United States government" banner; a NOAA SBIR news archive kept 1 of 3+ + stories. Since **v0.1.5** a `warning:` line fires when an extraction is a stub (< 500 chars) + of a content-bearing page (>= 1,000 chars of visible text) it retained under 10% of. + It is deliberately **not** a bare yield ratio — that fires on neither real pricing failure + and false-positives on nav-heavy pages; the absolute floor is what separates them. Replayed + over the whole cache it fires on 3.1% of pages, every one a genuine omission. - **Figures no guard can see**: prices rendered by JS into tabs/RSC payloads (dialpad), served as literal `"null"` placeholders (aircall), or drawn in images — invisible to any non-interactive fetch, tearsheet included. And a page carrying a **single** figure sits diff --git a/src/tearsheet/content.py b/src/tearsheet/content.py index f1ea675..6319e92 100644 --- a/src/tearsheet/content.py +++ b/src/tearsheet/content.py @@ -68,6 +68,40 @@ _PRICING_TITLE = re.compile(r"pricing|plans? and pricing|tarif|preise", re.I) _MONEY_MIN_TITLED = 3 +# Stub-extraction guard (the roster/directory class, 2026-07-24 audit). A content-bearing +# page whose extraction is boilerplate: idahoapexaccelerator.com/our-team returned a +# 161-char marketing sentence and dropped every consultant name; sbir.gov/awards returned +# only the "official website of the United States government" banner. Neither carries a +# price, a table, or a wall, so no other guard can arm. +# +# NOT the bare yield ratio that was measured and rejected — that fires on neither real +# pricing failure (smith.ai 13%, quo 21% retention here). Retention is only half the +# condition; the other half is an absolute stub-sized output. Calibrated by replaying +# every content-bearing page in the live cache (162 pages; fires on 5, i.e. 3.1%): +# +# page visible md ratio verdict +# sbir.gov/awards 10,698 421 3.9% FIRE (only the .gov banner) +# noaa SBIR archive 6,373 361 5.7% FIRE (1 of 3+ stories kept) +# apex/our-team 1,994 161 8.1% FIRE (marketing line only) +# nomic.ai/pricing 3,043 259 8.5% FIRE (marketing, zero prices) +# pocatello Fees 2,134 206 9.7% FIRE (already price-warned) +# eur-lex search 11,022 571 5.2% silent (real result list) +# dronedominance 16,661 1,553 9.3% silent +# linkedin FP 17,890 1,827 10.2% silent (MUST stay silent) +# smith.ai 10,961 1,452 13.2% silent +# heyrosie 16,218 2,263 14.0% silent +# quo 17,961 3,716 20.7% silent +# +# Measure with a FRESH extract_content, not a cached `markdown` column — stored rows may +# have been written with include_links=True, which inflates length by ~70% and hid the +# noaa failure during the first pass of this calibration. +# +# The markdown floor is what separates the two groups; do NOT drop it and leave a bare +# ratio, and do NOT raise it past ~570 (eur-lex, a legitimate terse result list). +_STUB_MIN_PAGE_TEXT = 1_000 +_STUB_MAX_MARKDOWN = 500 +_STUB_MAX_RATIO = 0.10 + # A collapsed table column reads as a run of identical consecutive rows # (quo: `Unlimited* / Unlimited* / Unlimited*`). One run happens naturally; two is a pattern. _COLLAPSE_RUN_LEN = 3 @@ -137,10 +171,14 @@ def _has_collapsed_columns(markdown: str) -> bool: def assess_extraction(html: bytes, extracted: ExtractedContent | None) -> ExtractionQuality: """Judge an extraction against the page it came from. - Catches the two failure classes that reached real research (2026-07-14): a consent - banner served as content, and a pricing table whose figures never survived. Deliberately - NOT a markdown/text yield ratio — measured against both real failures, a ratio fires on - neither (smith.ai yields 34%, quo 21%) and would only add false positives. + Catches the failure classes that reached real research: a consent banner served as + content and a pricing table whose figures never survived (2026-07-14), plus a + content-bearing page whose extraction is a boilerplate stub (2026-07-24). + + A BARE markdown/text yield ratio is still deliberately absent — measured against the + pricing failures it fires on neither and only adds false positives. The stub guard + pairs retention with an absolute output floor, which is what separates the real + failures from legitimately nav-heavy pages; see the calibration table above it. """ quality = ExtractionQuality() if extracted is None: @@ -164,6 +202,17 @@ def assess_extraction(html: bytes, extracted: ExtractedContent | None) -> Extrac "independently." ) + if ( + len(page_text) >= _STUB_MIN_PAGE_TEXT + and len(markdown) < _STUB_MAX_MARKDOWN + and len(markdown) / len(page_text) < _STUB_MAX_RATIO + ): + quality.warnings.append( + f"extraction is a {len(markdown)}-char stub of a page holding ~{len(page_text)} " + "chars of visible text — names, rows, or listings on this page are probably " + "missing entirely. Use raw=true, or fetch independently." + ) + on_page = set(_MONEY.findall(page_text)) title_armed = ( extracted.title is not None diff --git a/tests/test_quality.py b/tests/test_quality.py index d6c63e0..8736272 100644 --- a/tests/test_quality.py +++ b/tests/test_quality.py @@ -19,9 +19,10 @@ more than 3 near each other). A page-wide count cannot tell those apart; a cluster can, with margin on both sides (3 vs 5). -An earlier design used a markdown/visible-text yield ratio; it is kept out on +An earlier design used a BARE markdown/visible-text yield ratio; it stays out on purpose because it provably fires on neither bad page (smith.ai yields 34%, -quo 21%) while risking false positives on legitimately terse pages. +quo 21%) while risking false positives on legitimately terse pages. The +low-yield guard added 2026-07-24 is NOT that design — see TestLowYield. """ from tearsheet.content import ExtractedContent, assess_extraction, html_to_text @@ -279,3 +280,58 @@ def test_single_repeated_triple_is_tolerated(self) -> None: markdown = "Included\n\nYes\n\nYes\n\nYes\n\n" + "Real prose about the product. " * 30 quality = assess_extraction(page("Included"), extracted(markdown)) assert not any("column" in w for w in quality.warnings) + + +class TestLowYield: + """The roster/directory class (2026-07-24 audit): a content-bearing page whose + extraction is a boilerplate stub, with no price, table, or wall for any other + guard to arm on. + + Calibrated by replaying every content-bearing page in the live cache (162 pages; + the guard fires on 5, i.e. 3.1%). Retention alone is NOT the signal — nav-heavy + pages legitimately retain under 10% (dronedominance 9.3% is a HEALTHY extraction). + The discriminator is retention AND an absolute stub-sized output: + + page visible md ratio verdict + sbir.gov/awards 10,698 421 3.9% FIRE (only the .gov banner) + noaa SBIR archive 6,373 361 5.7% FIRE (1 of 3+ stories kept) + idahoapex /our-team 1,994 161 8.1% FIRE (only a marketing line) + atlas.nomic.ai/pricing 3,043 259 8.5% FIRE (marketing, zero prices) + pocatello.gov/812/Fees 2,134 206 9.7% FIRE (already price-warned) + eur-lex search 11,022 571 5.2% silent (real result list) + dronedominance 16,661 1,553 9.3% silent (healthy) + linkedin post (the FP) 17,890 1,827 10.2% silent (MUST stay silent) + smith.ai receptionists 10,961 1,452 13.2% silent + heyrosie.com 16,218 2,263 14.0% silent + quo.com/pricing 17,961 3,716 20.7% silent + + Measure with a FRESH extract_content, never a cached `markdown` column: stored rows + may carry include_links=True, which inflated noaa to 614 chars and hid it as + "healthy" during the first pass of this calibration. + """ + + def test_stub_extraction_from_a_content_bearing_page_warns(self) -> None: + """idahoapexaccelerator.com/our-team: 1,994 chars of roster -> 161-char blurb.""" + body = "Meet the team. " + "Name Title City. " * 110 + stub = "Meet the passionate team members who are committed to excellence." + quality = assess_extraction(page(body), extracted(stub)) + assert any("stub" in w for w in quality.warnings) + + def test_low_ratio_with_substantial_output_stays_silent(self) -> None: + """The noaa/dronedominance class: nav-heavy page, healthy article extracted.""" + body = "Site navigation. " * 400 + real = "A genuine article body. " * 40 # ~960 chars, well past a stub + quality = assess_extraction(page(body + real), extracted(real)) + assert not any("stub" in w for w in quality.warnings) + + def test_short_page_extracted_faithfully_stays_silent(self) -> None: + """A genuinely brief page is not a failure — high ratio, small output.""" + body = "A short notice. It says one thing and stops." + quality = assess_extraction(page(body), extracted(body)) + assert not any("stub" in w for w in quality.warnings) + + def test_tiny_page_cannot_arm_the_guard(self) -> None: + """Below the page-substance floor there is nothing to have missed.""" + body = "Hello. " * 20 # ~140 chars of visible text + quality = assess_extraction(page(body), extracted("Hello.")) + assert not any("stub" in w for w in quality.warnings)