Summary
simplifyHtml lifts every <meta name="…" content="…"> in <head> into the front matter, filtered only by a denylist of known-noise names (isNoiseMetaName).
On single-page apps that use <meta> as a transport for application state, that turns the front matter into most of the document.
A related bug in the same metadata extraction concatenates the text of every <title> in <head> into one title.
Both were found while driving LinkedIn through umwelt, which uses the markanywhere-html pipeline (0.4.0) for its Markdown dumps.
The dumps below come from a signed-in session. Only metadata names and sizes are quoted here, no page content.
1. Front matter swamped by application-state <meta> tags
Observed
A LinkedIn company "People" page (/company/<name>/people/):
|
|
| whole Markdown dump |
823,354 bytes |
| front matter |
782,266 bytes (95%) in 136 entries |
| page content |
~41 KB |
The largest entries:
| key |
size of the YAML line |
spark/hash-includes |
559,942 |
__init |
151,991 |
voyager-web/config/asset-manifest |
42,345 |
storage-inventory |
13,507 |
voyager-web/config/environment |
5,808 |
applicationInstance |
216 |
trusted-types |
156 |
103 more <app>/config/environment |
33–142 each |
The feed page (/feed/) shows the same pattern at a smaller scale: 14 KB of 52 KB, via storage-inventory, como-t, como-err, como-pk and trusted-types.
What survives and is actually page metadata: lang, title and description (empty on this page).
The entries fall into three shapes:
- Raw JSON in
content: __init, spark/hash-includes, storage-inventory, applicationInstance, como-t, como-err, e.g. "{\"title\":\"Something went wrong\",\"retry\":\"Try again\"}".
- Percent-encoded JSON, which is Ember's
<meta name="<app>/config/environment" content="%7B…%7D"> convention: every */config/environment entry, down to "jam/config/environment": "%7B%7D" (an empty object).
- Plain application flags and ids:
isGuest, treeID, serviceVersion, i18nLocale, baseCDNUrl, bigpipeResponseTimeout, …
Why it matters
The front matter is the first thing an LLM reads, and umwelt's agents read every dump whole.
Here an agent pays for ~780 KB of tokens (about 200k) before reaching the first person on the page, and a dump of this size can exceed a context window on its own.
Every markanywhere-html consumer gets the same bloat, which is why this report is here and not in umwelt: umwelt only appends its own typed status entry after the transform.
Cause
SimplifyHtml.kt, the match("meta") rule:
match("meta") { event ->
val name = event["name"]
val content = event["content"]
if (name != null && content != null && !isNoiseMetaName(name)) {
metadata[name] = content
}
}
isNoiseMetaName is a denylist (viewport, robots, msapplication-*, *-verification, …).
A denylist of names cannot keep up with names that are private to each site's own framework, such as spark/hash-includes, __init and como-t.
Its KDoc gives a good reason for choosing a denylist over an allowlist, which is to keep unknown but possibly useful names like description, author, og:* and article:*, so the fix below keeps it.
Proposal
Keep the name denylist, and add rules based on the value, which is what actually tells metadata apart from state:
- Drop a
content that is a JSON object or array. Trim it; if it starts with { or [, or with their percent-encoded forms %7B / %5B (case-insensitive), it is serialized application state. No human-facing metadata value begins that way.
This alone removes the six raw-JSON entries and all 104 */config/environment entries, %7B%7D included.
- Drop a value longer than a cap of about 1,024 characters.
Real metadata is short: description and og:description are usually under 300 characters.
This is the backstop for opaque blobs of any other shape, such as base64 or hash lists.
It could be exposed as a parameter (maxMetaValueLength) with a default, in the style of keepAttributes, for a caller that really wants long values.
- (optional) Drop a name containing
/. Standard and de-facto metadata names (description, og:title, twitter:card, article:published_time, citation_author) never contain one, while framework config namespaces (<app>/config/…) always do.
Rule 1 already covers the entries seen here, so this is only a cheap extra guard.
- (optional) Drop an empty
content: description: "" tells the reader nothing.
On the LinkedIn page, rules 1 and 2 leave about 25 small flag/id entries (isGuest, treeID, i18nLocale, …). Those are a few hundred bytes of noise instead of 780 KB, and they would be a case for the denylist if they ever matter.
Alternative considered: an allowlist (title, description, author, keywords, og:*, twitter:*, article:*, citation_*, dc.*).
It is tighter, but it is the trade-off the existing KDoc argues against, and it would silently drop legitimate site-specific metadata.
Tests to add (SimplifyHtmlTest)
<meta name="__init" content="{"a":1}"> produces no entry.
<meta name="feed/config/environment" content="%7B%22x%22%3A1%7D"> and content="%7B%7D" produce no entry.
- A
content longer than the cap produces no entry; one just under it is kept.
description, og:title, author and article:published_time with ordinary values are still kept (a regression guard for the denylist rationale).
- A value that merely contains
{ ("Save {50%} today") is kept, because only a leading {/[ counts.
2. Two <title> elements are concatenated into one title
Observed
A dump of LinkedIn's feed, right after an in-app sign-in, had:
title: Feed | LinkedInFeed | LinkedIn
The browser's own document.title at the same moment was Feed | LinkedIn.
A later dump of another page had a single, correct title, which fits the page sometimes carrying two <title> elements in <head>, for example after a client-side navigation.
Cause
titleText is one StringBuilder per collection, and the match("title") rule appends to it without ever clearing it:
val titleText = StringBuilder() // once per collection
…
match("title") {
children(mode = "titleText")
afterClose {
val trimmed = titleText.toString().trim()
if (trimmed.isNotEmpty()) metadata["title"] = trimmed
}
}
matchText(mode = "titleText") { titleText.append(it) }
A second <title> appends its text to the first one's, and the afterClose of the second overwrites metadata["title"] with the concatenation.
Proposal
Clear titleText when a <title> opens, or collect into a local buffer per element, and keep the first non-blank title. That matches the HTML spec, where document.title is the first <title> in the document.
Test to add
<head><title>A</title><title>B</title></head> gives title: A, not AB and not B.
Downstream
Once released, umwelt raises its markanywhere catalog pin and adds an e2e fixture page with a JSON-carrying <meta> and a doubled <title>, so the fix is pinned in umwelt as well.
Summary
simplifyHtmllifts every<meta name="…" content="…">in<head>into the front matter, filtered only by a denylist of known-noise names (isNoiseMetaName).On single-page apps that use
<meta>as a transport for application state, that turns the front matter into most of the document.A related bug in the same metadata extraction concatenates the text of every
<title>in<head>into one title.Both were found while driving LinkedIn through umwelt, which uses the
markanywhere-htmlpipeline (0.4.0) for its Markdown dumps.The dumps below come from a signed-in session. Only metadata names and sizes are quoted here, no page content.
1. Front matter swamped by application-state
<meta>tagsObserved
A LinkedIn company "People" page (
/company/<name>/people/):The largest entries:
spark/hash-includes__initvoyager-web/config/asset-manifeststorage-inventoryvoyager-web/config/environmentapplicationInstancetrusted-types<app>/config/environmentThe feed page (
/feed/) shows the same pattern at a smaller scale: 14 KB of 52 KB, viastorage-inventory,como-t,como-err,como-pkandtrusted-types.What survives and is actually page metadata:
lang,titleanddescription(empty on this page).The entries fall into three shapes:
content:__init,spark/hash-includes,storage-inventory,applicationInstance,como-t,como-err, e.g."{\"title\":\"Something went wrong\",\"retry\":\"Try again\"}".<meta name="<app>/config/environment" content="%7B…%7D">convention: every*/config/environmententry, down to"jam/config/environment": "%7B%7D"(an empty object).isGuest,treeID,serviceVersion,i18nLocale,baseCDNUrl,bigpipeResponseTimeout, …Why it matters
The front matter is the first thing an LLM reads, and umwelt's agents read every dump whole.
Here an agent pays for ~780 KB of tokens (about 200k) before reaching the first person on the page, and a dump of this size can exceed a context window on its own.
Every
markanywhere-htmlconsumer gets the same bloat, which is why this report is here and not in umwelt: umwelt only appends its own typedstatusentry after the transform.Cause
SimplifyHtml.kt, thematch("meta")rule:isNoiseMetaNameis a denylist (viewport,robots,msapplication-*,*-verification, …).A denylist of names cannot keep up with names that are private to each site's own framework, such as
spark/hash-includes,__initandcomo-t.Its KDoc gives a good reason for choosing a denylist over an allowlist, which is to keep unknown but possibly useful names like
description,author,og:*andarticle:*, so the fix below keeps it.Proposal
Keep the name denylist, and add rules based on the value, which is what actually tells metadata apart from state:
contentthat is a JSON object or array. Trim it; if it starts with{or[, or with their percent-encoded forms%7B/%5B(case-insensitive), it is serialized application state. No human-facing metadata value begins that way.This alone removes the six raw-JSON entries and all 104
*/config/environmententries,%7B%7Dincluded.Real metadata is short:
descriptionandog:descriptionare usually under 300 characters.This is the backstop for opaque blobs of any other shape, such as base64 or hash lists.
It could be exposed as a parameter (
maxMetaValueLength) with a default, in the style ofkeepAttributes, for a caller that really wants long values./. Standard and de-facto metadata names (description,og:title,twitter:card,article:published_time,citation_author) never contain one, while framework config namespaces (<app>/config/…) always do.Rule 1 already covers the entries seen here, so this is only a cheap extra guard.
content:description: ""tells the reader nothing.On the LinkedIn page, rules 1 and 2 leave about 25 small flag/id entries (
isGuest,treeID,i18nLocale, …). Those are a few hundred bytes of noise instead of 780 KB, and they would be a case for the denylist if they ever matter.Alternative considered: an allowlist (
title,description,author,keywords,og:*,twitter:*,article:*,citation_*,dc.*).It is tighter, but it is the trade-off the existing KDoc argues against, and it would silently drop legitimate site-specific metadata.
Tests to add (
SimplifyHtmlTest)<meta name="__init" content="{"a":1}">produces no entry.<meta name="feed/config/environment" content="%7B%22x%22%3A1%7D">andcontent="%7B%7D"produce no entry.contentlonger than the cap produces no entry; one just under it is kept.description,og:title,authorandarticle:published_timewith ordinary values are still kept (a regression guard for the denylist rationale).{("Save {50%} today") is kept, because only a leading{/[counts.2. Two
<title>elements are concatenated into one titleObserved
A dump of LinkedIn's feed, right after an in-app sign-in, had:
The browser's own
document.titleat the same moment wasFeed | LinkedIn.A later dump of another page had a single, correct title, which fits the page sometimes carrying two
<title>elements in<head>, for example after a client-side navigation.Cause
titleTextis oneStringBuilderper collection, and thematch("title")rule appends to it without ever clearing it:A second
<title>appends its text to the first one's, and theafterCloseof the second overwritesmetadata["title"]with the concatenation.Proposal
Clear
titleTextwhen a<title>opens, or collect into a local buffer per element, and keep the first non-blank title. That matches the HTML spec, wheredocument.titleis the first<title>in the document.Test to add
<head><title>A</title><title>B</title></head>givestitle: A, notABand notB.Downstream
Once released, umwelt raises its
markanywherecatalog pin and adds an e2e fixture page with a JSON-carrying<meta>and a doubled<title>, so the fix is pinned in umwelt as well.