Skip to content

simplifyHtml: front matter swamped by application-state <meta> tags; multiple <title>s concatenated #82

Description

@morisil

Summary

simplifyHtml lifts every <meta name="…" content="…"> in <head> into the front matter, filtered only by a denylist of known-noise names (isNoiseMetaName).
On single-page apps that use <meta> as a transport for application state, that turns the front matter into most of the document.
A related bug in the same metadata extraction concatenates the text of every <title> in <head> into one title.

Both were found while driving LinkedIn through umwelt, which uses the markanywhere-html pipeline (0.4.0) for its Markdown dumps.
The dumps below come from a signed-in session. Only metadata names and sizes are quoted here, no page content.

1. Front matter swamped by application-state <meta> tags

Observed

A LinkedIn company "People" page (/company/<name>/people/):

whole Markdown dump 823,354 bytes
front matter 782,266 bytes (95%) in 136 entries
page content ~41 KB

The largest entries:

key size of the YAML line
spark/hash-includes 559,942
__init 151,991
voyager-web/config/asset-manifest 42,345
storage-inventory 13,507
voyager-web/config/environment 5,808
applicationInstance 216
trusted-types 156
103 more <app>/config/environment 33–142 each

The feed page (/feed/) shows the same pattern at a smaller scale: 14 KB of 52 KB, via storage-inventory, como-t, como-err, como-pk and trusted-types.

What survives and is actually page metadata: lang, title and description (empty on this page).

The entries fall into three shapes:

  • Raw JSON in content: __init, spark/hash-includes, storage-inventory, applicationInstance, como-t, como-err, e.g. "{\"title\":\"Something went wrong\",\"retry\":\"Try again\"}".
  • Percent-encoded JSON, which is Ember's <meta name="<app>/config/environment" content="%7B…%7D"> convention: every */config/environment entry, down to "jam/config/environment": "%7B%7D" (an empty object).
  • Plain application flags and ids: isGuest, treeID, serviceVersion, i18nLocale, baseCDNUrl, bigpipeResponseTimeout, …

Why it matters

The front matter is the first thing an LLM reads, and umwelt's agents read every dump whole.
Here an agent pays for ~780 KB of tokens (about 200k) before reaching the first person on the page, and a dump of this size can exceed a context window on its own.
Every markanywhere-html consumer gets the same bloat, which is why this report is here and not in umwelt: umwelt only appends its own typed status entry after the transform.

Cause

SimplifyHtml.kt, the match("meta") rule:

match("meta") { event ->
    val name = event["name"]
    val content = event["content"]
    if (name != null && content != null && !isNoiseMetaName(name)) {
        metadata[name] = content
    }
}

isNoiseMetaName is a denylist (viewport, robots, msapplication-*, *-verification, …).
A denylist of names cannot keep up with names that are private to each site's own framework, such as spark/hash-includes, __init and como-t.
Its KDoc gives a good reason for choosing a denylist over an allowlist, which is to keep unknown but possibly useful names like description, author, og:* and article:*, so the fix below keeps it.

Proposal

Keep the name denylist, and add rules based on the value, which is what actually tells metadata apart from state:

  1. Drop a content that is a JSON object or array. Trim it; if it starts with { or [, or with their percent-encoded forms %7B / %5B (case-insensitive), it is serialized application state. No human-facing metadata value begins that way.
    This alone removes the six raw-JSON entries and all 104 */config/environment entries, %7B%7D included.
  2. Drop a value longer than a cap of about 1,024 characters.
    Real metadata is short: description and og:description are usually under 300 characters.
    This is the backstop for opaque blobs of any other shape, such as base64 or hash lists.
    It could be exposed as a parameter (maxMetaValueLength) with a default, in the style of keepAttributes, for a caller that really wants long values.
  3. (optional) Drop a name containing /. Standard and de-facto metadata names (description, og:title, twitter:card, article:published_time, citation_author) never contain one, while framework config namespaces (<app>/config/…) always do.
    Rule 1 already covers the entries seen here, so this is only a cheap extra guard.
  4. (optional) Drop an empty content: description: "" tells the reader nothing.

On the LinkedIn page, rules 1 and 2 leave about 25 small flag/id entries (isGuest, treeID, i18nLocale, …). Those are a few hundred bytes of noise instead of 780 KB, and they would be a case for the denylist if they ever matter.

Alternative considered: an allowlist (title, description, author, keywords, og:*, twitter:*, article:*, citation_*, dc.*).
It is tighter, but it is the trade-off the existing KDoc argues against, and it would silently drop legitimate site-specific metadata.

Tests to add (SimplifyHtmlTest)

  • <meta name="__init" content="{&quot;a&quot;:1}"> produces no entry.
  • <meta name="feed/config/environment" content="%7B%22x%22%3A1%7D"> and content="%7B%7D" produce no entry.
  • A content longer than the cap produces no entry; one just under it is kept.
  • description, og:title, author and article:published_time with ordinary values are still kept (a regression guard for the denylist rationale).
  • A value that merely contains { ("Save {50%} today") is kept, because only a leading {/[ counts.

2. Two <title> elements are concatenated into one title

Observed

A dump of LinkedIn's feed, right after an in-app sign-in, had:

title: Feed | LinkedInFeed | LinkedIn

The browser's own document.title at the same moment was Feed | LinkedIn.
A later dump of another page had a single, correct title, which fits the page sometimes carrying two <title> elements in <head>, for example after a client-side navigation.

Cause

titleText is one StringBuilder per collection, and the match("title") rule appends to it without ever clearing it:

val titleText = StringBuilder()          // once per collection
…
match("title") {
    children(mode = "titleText")
    afterClose {
        val trimmed = titleText.toString().trim()
        if (trimmed.isNotEmpty()) metadata["title"] = trimmed
    }
}
matchText(mode = "titleText") { titleText.append(it) }

A second <title> appends its text to the first one's, and the afterClose of the second overwrites metadata["title"] with the concatenation.

Proposal

Clear titleText when a <title> opens, or collect into a local buffer per element, and keep the first non-blank title. That matches the HTML spec, where document.title is the first <title> in the document.

Test to add

<head><title>A</title><title>B</title></head> gives title: A, not AB and not B.

Downstream

Once released, umwelt raises its markanywhere catalog pin and adds an e2e fixture page with a JSON-carrying <meta> and a doubled <title>, so the fix is pinned in umwelt as well.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions