diff --git a/CLAUDE.md b/CLAUDE.md index e7747f9..90e1790 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -9,7 +9,8 @@ Both developers and AI agents are expected to add entries as they encounter surp - **Add an entry** when you encounter something unexpected: a build quirk, a non-obvious constraint, a dependency gotcha, or any behavior that would surprise the next agent or developer. - **Add an entry** when a developer flags an anti-pattern produced by AI — describe the anti-pattern and the preferred alternative. - **Do not** add codebase overviews, directory listings, or anything discoverable by reading the source. -- Keep entries concise: one line per lesson, grouped under a heading if a theme emerges. +- Keep entries concise: the lesson and its non-obvious *why*, grouped under a heading if a theme emerges. + Name a symbol or test as a pointer, but don't restate what its source comment already says. ## Conventions @@ -411,6 +412,8 @@ If the fix is bounded (one inline construct, one or two lines of lookahead), it' `renderDumpMarkdown` (Gradle task, `RenderDumpMarkdown.kt` in jvmTest) runs the full `transformHtmlToMarkdown` pipeline over every dump to `build/renderedMarkdown/.md` — the canonical way to regenerate the per-dump golden strings (`dumps/OpenjurTest`, `dumps/HackerNewsTest`) after a pipeline change. - A `Regex` used with `matches()` in `commonMain` must be **explicitly anchored** (`^(?:a|b)$`) when it contains a top-level alternation: Kotlin/JS resolves `matches` through the leftmost `find`, so the first branch wins on a *prefix* and the whole-input check fails (`0x1F` matched `[-+]?[0-9]+|0x[0-9a-fA-F]+` as `0`; a full timestamp matched the date-only branch) — JVM backtracks across the branches and never shows it. The YAML scalar typing in `markanywhere-yaml` hit exactly this: green on `jvmTest`, red on `jsBrowserTest`. Dev builds run the JS tests for `yaml`/`parse`/`render`/`html`, so a JVM-only test run is not enough evidence for a regex change. +- A `commonMain` `Regex` applied to untrusted input must not **repeat a group** (`(?:_?[0-9])*`, `(?::[0-9])+`): `java.util.regex` matches each repetition of a group one stack frame deeper, so a few thousand repetitions throw `StackOverflowError` on the JVM (an `Error`, which nothing catches — it takes the whole pipeline down). + Repeat a character class instead (`[0-9_]*`) and check what the group expressed by hand (`YamlScalarType.kt`). - Facts of the HTML standard needed by more than one of `parse` / `render` / `html` — HTML whitespace (`isHtmlWhitespace` / `isHtmlBlank`), void and raw text elements, `Mark.classList` — live in the leaf module `markanywhere-html-spec` (depends only on `api`), never in `markanywhere-api` (which describes only the event model) and never as private copies. Only spec facts go there, not project policy (e.g. which tags render as blocks in Markdown); the parser's `isFlankWhitespace` is deliberately a different set (CommonMark "Unicode whitespace", which includes NBSP). - `MutableMap.putIfAbsent` is **JVM-only** — not in the common stdlib, so it can't be used in `commonMain`. @@ -543,11 +546,15 @@ If the fix is bounded (one inline construct, one or two lines of lookahead), it' The key is an attribute, not the mark name, on purpose: keys collide with tag names (`title` would hit `simplifyHtml`'s head-mode `match("title")`) and `og:title` would look like `` custom markup. `entry`/`item` are Markdown-native **only inside** a `frontmatter` (the Markdown renderer creates a `YamlWriter` on the mark, routes every following event into `YamlWriter.collect`, and treats the first `unmark` arriving at `depth == 0` as the wrapper's own close; outside one they fall through to the raw-tag path), so they are deliberately NOT in `MARKDOWN_NATIVE_MARK_NAMES`. Any other mark is **transparent** to the writer (children lay out as if it were absent), which is what lets `renderYaml()` accept a `frontmatter`-wrapped stream unchanged. - The writer's quoting rules reuse the parser's `yamlScalarType` (one copy of the reserved literals / int / float / timestamp rules) and are pinned against it by `YamlRoundTripTest` in `markanywhere-yaml`; `FrontMatterRoundTripTest` in `markanywhere-parse` (whose tests depend on render) covers only the fences and the wrapper — extend the yaml one when you touch either side of the codec. + The writer's quoting rules reuse the parser's `yamlScalarType` (`YamlScalarType.kt`, one copy of the reserved literals / int / float / timestamp rules) and are pinned against it by `YamlRoundTripTest` in `markanywhere-yaml`; `FrontMatterRoundTripTest` in `markanywhere-parse` (whose tests depend on render) covers only the fences and the wrapper — extend the yaml one when you touch either side of the codec. + Typing follows the front matter readers (Psych/Jekyll, PyYAML, go-yaml v2): a shape some reader types must be typed by the **parser** (writer-only quoting rewrote Jekyll's `date: 2016-01-01 12:00:00 -0500` into a quoted string), a shape none types must not be (or `type=int` stops meaning a number), and a shape they read **differently** from one another (`[a:, b]`: a mapping, an error, the string `"a:"`) stays a verbatim line — picking one reading rewrites the document for the others. + Verify a change by running the readers themselves (PyYAML `safe_load`, Ruby `Psych.unsafe_load`, `gopkg.in/yaml.v2` into `map[string]interface{}`) over generated strings, never from their docs or the spec — reasoning from the spec got each of the three rules above wrong at least once. + The one writer-only quoting is a shape a reader **refuses to load** (`=`, `0x_`, `2024-13-45`), pinned as `DIVERGENCE` in `YamlRoundTripTest`. `renderYaml()` **keeps the trailing newline** (unlike `renderMarkdown()` / `renderHtml()`, which suppress one): a `|+` block scalar at the end of the document needs it, and dropping it broke the round-trip for `z\n\n`. - Front matter **detection** requires line 2 to be a mapping key (identifier or quoted, then `:` + whitespace/EOL — `isYamlKeyLine`, a public function of `markanywhere-yaml`, NOT a private copy in `FrontMatterFilter`); a `# comment` cannot be line 2 (`---` + `# Heading` is a thematic break + heading), and neither can `http://…` or a `- item`. Consequence for the YAML writer: `YamlWriter.renderKey` quotes every key that is not identifier-shaped (`"a b"`, `"42"`, `"og:title"`), even where YAML would not require it, so whatever `simplifyHtml` emits first (a `` with a space, say) still re-detects. It also quotes a key that is a reserved literal (`"yes":`) — a YAML 1.1 reader would take a bare `yes:` as a boolean key. + Key typing is not value typing: go-yaml v2 decodes a **top-level** front matter key into a string but a nested one as `interface{}`, so `y:` must stay plain at the top level (it was once rewritten to `"y":`) and be quoted below it (`isTypedPlainKey`). The writer's plain-key rule (`isIdentifierKey`) and the detector share their character rules in `YamlKeyLine.kt`, and `YamlWriterTest` "should write every key as a line the front matter detector accepts" pins the two — widen one without the other and every document whose *first* entry has such a key silently loses its front matter on re-parse. **DIVERGENCE (no round-trip)**: a `frontmatter` whose root is a **sequence** (top-level `item`s) renders as `---\n- a\n…`, which line 2 rejects — Markdown reads `---` + `- a` as a thematic break + list, and no front matter consumer (Jekyll, Hugo, `wrapInHtmlDocument`) accepts a non-mapping — so the shape is pinned as such in `FrontMatterRoundTripTest` rather than widened into detection. - `wrapInHtmlDocument` and `ensureFrontmatterTitle` read the front matter by **depth** (a direct child `entry` is depth 1 / 2 respectively, counted from the `frontmatter` mark): only top-level scalar entries become ``/``, only a top-level `entry key=title` counts as an existing title. @@ -565,6 +572,7 @@ If the fix is bounded (one inline construct, one or two lines of lookahead), it' - For semantic event flow testing, use the overloaded `sameAs` infix on `Flow<SemanticEvent>` (defined in `markanywhere-test`) against a `semanticEvents { ... }` builder — this is event-stream comparison, unrelated to HTML. - For asserting rendered HTML output (e.g. in `markanywhere-render` tests), use `sameAsHtml` (not the generic string `sameAs`) — provides syntax highlighting in the IDE. - GFM example test naming convention: `example N - <description>` for spec-conformant tests, `example N - DIVERGENCE - <description>` when the parser intentionally diverges from GFM. + Any other test pinning an intentional divergence (from GFM, YAML or a reader's behaviour) starts with `DIVERGENCE - ` (`DIVERGENCE - should keep …`), never `should DIVERGENCE …` or a trailing `… DIVERGENCE`. Keep `DIVERGENCE` in the name even when updating expectations to match the divergent behavior. - When asserting on a boolean expression (counts, equality between two nullables, `contains`, etc.), use `assert(expr)` from `com.xemantic.kotlin.test.assert` — it is power-assert-instrumented in the convention plugin, so a failure prints the full expression with subvalues. Do NOT fall back to `kotlin.test.assertEquals` / `assertTrue` / `assertNotNull` — they discard that information. diff --git a/markanywhere-parse/src/commonTest/kotlin/FrontMatterRoundTripTest.kt b/markanywhere-parse/src/commonTest/kotlin/FrontMatterRoundTripTest.kt index 7341179..ffa85cf 100644 --- a/markanywhere-parse/src/commonTest/kotlin/FrontMatterRoundTripTest.kt +++ b/markanywhere-parse/src/commonTest/kotlin/FrontMatterRoundTripTest.kt @@ -128,7 +128,7 @@ class FrontMatterRoundTripTest { } @Test - fun `should DIVERGENCE not re-detect a root sequence front matter`() = runTest { + fun `DIVERGENCE - should not re-detect a root sequence front matter`() = runTest { // given — YAML allows a root sequence, but front matter detection // requires a mapping key on line 2 (`---` + `- a` is a thematic // break followed by a list in Markdown, and no front matter consumer diff --git a/markanywhere-parse/src/commonTest/kotlin/FrontMatterTest.kt b/markanywhere-parse/src/commonTest/kotlin/FrontMatterTest.kt index 99c5f9a..ef49706 100644 --- a/markanywhere-parse/src/commonTest/kotlin/FrontMatterTest.kt +++ b/markanywhere-parse/src/commonTest/kotlin/FrontMatterTest.kt @@ -168,7 +168,7 @@ class FrontMatterTest { } @Test - fun `should DIVERGENCE auto-close unterminated front matter at EOF`() = runTest { + fun `DIVERGENCE - should auto-close unterminated front matter at EOF`() = runTest { // given — opener but no closer. DIVERGENCE: spec-correct behavior would // require buffering the whole document to detect the missing closer // and fall back to a thematic break + paragraphs. The streaming parser diff --git a/markanywhere-parse/src/commonTest/kotlin/HtmlBlockInListItemTest.kt b/markanywhere-parse/src/commonTest/kotlin/HtmlBlockInListItemTest.kt index 67308ac..b3197d7 100644 --- a/markanywhere-parse/src/commonTest/kotlin/HtmlBlockInListItemTest.kt +++ b/markanywhere-parse/src/commonTest/kotlin/HtmlBlockInListItemTest.kt @@ -98,7 +98,7 @@ class HtmlBlockInListItemTest { } @Test - fun `type 1 pre block inside list item DIVERGENCE`() = runTest { + fun `DIVERGENCE - type 1 pre block inside list item`() = runTest { // given val textFlow = /* language=markdown */ """ - intro @@ -298,7 +298,7 @@ class HtmlBlockInListItemTest { } @Test - fun `blank line inside list-internal html block stays in raw mode DIVERGENCE`() = runTest { + fun `DIVERGENCE - blank line inside list-internal html block stays in raw mode`() = runTest { // given: a blank line inside a `<div>` block. At top level this would // transition the block to sub-parse mode and the `- nested` line would // open a Markdown list. Inside a list item we stay in raw-text mode @@ -444,7 +444,7 @@ class HtmlBlockInListItemTest { } @Test - fun `type 1 pre block with multi-line opener inside list item DIVERGENCE`() = runTest { + fun `DIVERGENCE - type 1 pre block with multi-line opener inside list item`() = runTest { // given: `<pre` on the marker line, attribute on the next line, `>` // closes the opener. Exercises the buffered-opener path in the type-1 // branch of `streamListHtmlBlockLine` — `firstLineBuffer` accumulates diff --git a/markanywhere-parse/src/commonTest/kotlin/HtmlParsingTest.kt b/markanywhere-parse/src/commonTest/kotlin/HtmlParsingTest.kt index 2550d3d..f756625 100644 --- a/markanywhere-parse/src/commonTest/kotlin/HtmlParsingTest.kt +++ b/markanywhere-parse/src/commonTest/kotlin/HtmlParsingTest.kt @@ -401,7 +401,7 @@ class HtmlParsingTest { * over spec-faithful escaping. */ @Test - fun `DIVERGENCE inline disallowed script tag and its body are dropped`() = runTest { + fun `DIVERGENCE - inline disallowed script tag and its body are dropped`() = runTest { // given val src = "before<script>alert(1)</script>after\n" @@ -419,7 +419,7 @@ class HtmlParsingTest { * including HTML-like substrings, is dropped. */ @Test - fun `DIVERGENCE inline disallowed style tag and its body are dropped`() = runTest { + fun `DIVERGENCE - inline disallowed style tag and its body are dropped`() = runTest { // given val src = "x<style>a::before{content:\"<x>\"}</style>y\n" @@ -437,7 +437,7 @@ class HtmlParsingTest { * the final `>` (`</script >`), mirroring the block-1 close rule. */ @Test - fun `DIVERGENCE inline disallowed close tag tolerates trailing whitespace`() = runTest { + fun `DIVERGENCE - inline disallowed close tag tolerates trailing whitespace`() = runTest { // given val src = "a<script>js</script\t >b\n" @@ -539,7 +539,7 @@ class HtmlParsingTest { * always close their `<script>`/`<style>`, so this edge is benign in practice. */ @Test - fun `DIVERGENCE unclosed inline disallowed opener drops only the rest of its line`() = runTest { + fun `DIVERGENCE - unclosed inline disallowed opener drops only the rest of its line`() = runTest { // given val src = "keep<script>dropped to end of line\nnext line\n" diff --git a/markanywhere-parse/src/commonTest/kotlin/gfm/Gfm_06_11_Test.kt b/markanywhere-parse/src/commonTest/kotlin/gfm/Gfm_06_11_Test.kt index 54c40b7..1088118 100644 --- a/markanywhere-parse/src/commonTest/kotlin/gfm/Gfm_06_11_Test.kt +++ b/markanywhere-parse/src/commonTest/kotlin/gfm/Gfm_06_11_Test.kt @@ -36,7 +36,7 @@ import kotlin.test.Test class Gfm_06_11_Test { @Test - fun `example 657 - DIVERGENCE disallowed inline tags are dropped not escaped`() = runTest { + fun `example 657 - DIVERGENCE - disallowed inline tags are dropped not escaped`() = runTest { // given val textFlow = buildString { +"<strong> <title> <style> <em>\n" diff --git a/markanywhere-yaml/README.md b/markanywhere-yaml/README.md index 18f70ab..eb61603 100644 --- a/markanywhere-yaml/README.md +++ b/markanywhere-yaml/README.md @@ -39,7 +39,10 @@ The vocabulary is deliberately small and the key lives in an attribute, so a key - `entry` (attribute `key`) is a mapping entry; `item` is a sequence element. A root mapping is a run of top-level `entry` marks, a root sequence a run of top-level `item` marks — there is no wrapper mark. - The kind of value follows from the children: text only is a scalar, nested `entry` marks are a mapping, nested `item` marks are a sequence. -- `type` is present only on a non-string scalar: `bool`, `int`, `float`, `null`, `timestamp` (YAML 1.2 core schema, plus the 1.1 `yes`/`no`/`on`/`off` booleans so a Jekyll / Hugo document round-trips as written). +- `type` is present only on a non-string scalar: `bool`, `int`, `float`, `null`, `timestamp`. + A plain scalar is typed exactly when some reader front matter is written for reads it as something other than a string — the YAML 1.2 core schema, plus the YAML 1.1 shapes Psych (Jekyll), PyYAML and go-yaml v2 (Hugo) still type — so a Jekyll / Hugo document round-trips as written. + Beyond the core schema that means booleans and nulls in any letter case and `y` / `n`, numbers with `_` / `,` separators or a binary / octal / upper-case base prefix (`1_000`, `1,000`, `0b101`, `0X1F`), base-60 numbers (`12:30` is an `int`), and timestamps with a one-digit month or day or a colonless offset (`2016-01-01 12:00:00 -0500`); a date only when it is in the calendar, and a number only go-yaml v2 reads only within its 64-bit range (`1e999` stays a string). + So `type="int"` / `type="float"` says a reader reads a number, not that the text is a Kotlin literal: reading the value yourself takes that reader's rules (drop the separators, resolve the prefix or the base 60). No text and no type is an empty string; a bare `key:` is `type="null"`; an empty `[]` / `{}` is `type="seq"` / `type="map"` with no children. - All marks are untagged. @@ -50,13 +53,14 @@ Block mappings and sequences (a sequence may sit at its key's own indentation), Streaming: a `key: value` line commits on its newline. A bare `key:` (or `-`) is held for one line to decide between a nested block and a null value; a block scalar is buffered until it closes — never past the enclosing construct. -DIVERGENCE (never throws, never loses content): a line outside that subset — a complex `? key`, a directive, a document marker, a multi-line flow collection or quoted scalar, a plain scalar's continuation line, a `- item` inside a mapping — is emitted **verbatim** (with its `\n`) as a text child of the container it sits in, and the writer writes it back as-is. +DIVERGENCE (never throws, never loses content): a line outside that subset — a complex `? key`, a directive, a document marker, a multi-line flow collection or quoted scalar, a plain scalar's continuation line, a `- item` inside a mapping, a shape the front matter readers read differently from one another (`[draft:, x]`, `{a:[1]}`, `|-#note`) — is emitted **verbatim** (with its `\n`) as a text child of the container it sits in, and the writer writes it back as-is. Anchors, aliases and tags are not resolved: a plain scalar starting with `&`, `*` or `!` is just a string. ## Writing `renderYaml()` emits block style, one line per scalar, two spaces per nesting level. -A string that would re-parse as another type or shape (`"true"`, `"42"`, `"- dash"`, `"a: b"`, surrounding whitespace, control characters) is double-quoted; a multi-line string becomes a literal block scalar with the chomping indicator its trailing newlines call for; a key is written plain only when identifier-shaped and not a reserved literal. +A string that would re-parse as another type or shape (`"true"`, `"y"`, `"42"`, `"12:30"`, `"- dash"`, `"a: b"`, surrounding whitespace) is double-quoted, with every character YAML does not allow in a document `\u`-escaped; a multi-line string becomes a literal block scalar with the chomping indicator its trailing newlines call for; a key is written plain only when identifier-shaped and not a typed word (`"yes"`, `"y"`, `"null"`). +DIVERGENCE: a string some reader refuses to load when plain — `=`, `<<`, a base prefix with no digit (`0x_`), a date not in the calendar (`2024-13-45`) — is quoted too, although the parser reads it back as a plain string, so the first render quotes such a value in the source. Any mark other than `entry` / `item` is transparent, so a `frontmatter` wrapper renders the same with or without it. Every line is terminated — a non-empty document ends with `\n`, which a `|+` block scalar at the end needs. diff --git a/markanywhere-yaml/src/commonMain/kotlin/YamlKeyLine.kt b/markanywhere-yaml/src/commonMain/kotlin/YamlKeyLine.kt index f953e37..89bdedc 100644 --- a/markanywhere-yaml/src/commonMain/kotlin/YamlKeyLine.kt +++ b/markanywhere-yaml/src/commonMain/kotlin/YamlKeyLine.kt @@ -44,9 +44,7 @@ public fun isYamlKeyLine(line: String): Boolean { j } while (i < line.length && line[i] == ' ') i++ - if (i >= line.length || line[i] != ':') return false - i++ - return i == line.length || line[i] == ' ' || line[i] == '\t' + return i < line.length && line.isMappingColonAt(i) } // A key [YamlWriter] may write plain: identifier-shaped, so it passes diff --git a/markanywhere-yaml/src/commonMain/kotlin/YamlParser.kt b/markanywhere-yaml/src/commonMain/kotlin/YamlParser.kt index de47a5a..001449b 100644 --- a/markanywhere-yaml/src/commonMain/kotlin/YamlParser.kt +++ b/markanywhere-yaml/src/commonMain/kotlin/YamlParser.kt @@ -35,8 +35,12 @@ import kotlinx.coroutines.flow.FlowCollector * sequence is nested `item` marks. The root is a mapping (top-level * `entry` marks) or a sequence (top-level `item` marks) — no wrapper mark. * - A non-string scalar carries `type` = `bool` / `int` / `float` / `null` / - * `timestamp` (YAML 1.2 core schema, plus the 1.1 `yes`/`no`/`on`/`off` - * booleans so a Jekyll / Hugo document round-trips as written). An empty + * `timestamp`: typed exactly when a front matter reader — the YAML 1.2 + * core schema, or the YAML 1.1 shapes of Psych (Jekyll), PyYAML and + * go-yaml v2 (Hugo) — reads it as other than a string (`y`, `yEs`, + * `1_000`, `0X1F`, `12:30`, `2024-5-1`), so a Jekyll / Hugo document + * round-trips as written. A number keeps that reader's syntax (separators, + * base prefixes, base 60) in its text. An empty * flow collection carries `type` = `seq` / `map`. An empty string is an * entry with no text and no type; a missing value (`key:`) is `type=null`. * @@ -48,7 +52,10 @@ import kotlinx.coroutines.flow.FlowCollector * * DIVERGENCE (verbatim fallback): a line the subset does not understand — * complex keys, directives, multi-line flow collections or quoted scalars, - * a plain scalar's continuation line, a `- item` inside a mapping — is + * a plain scalar's continuation line, a `- item` inside a mapping, or a + * shape the front matter readers read differently from one another (a + * flow scalar's `:` before `,` / `[` / `]` / `{` / `}`, a block scalar + * header followed by `#` with no space) — is * emitted **verbatim** (with its `\n`) as a text child of the container it * sits in, so nothing is lost and [YamlWriter] can write it back as-is. * Anchors, aliases and tags are not resolved: a plain scalar starting with @@ -56,6 +63,13 @@ import kotlinx.coroutines.flow.FlowCollector * directives are not recognised — a multi-document stream is outside the * subset. * + * DIVERGENCE (lenient plain scalars): a mapping entry's plain value + * containing a mapping indicator — `k: Note: see`, `k: ends:` — is read as + * the string after the first `: `, where YAML (Psych, PyYAML) rejects the + * line. A sequence item is not lenient: `- Note: see` is a compact mapping + * (an item holding the entry `Note`), as in YAML. This does not make such a + * value safe to write plain: [YamlWriter] still quotes it. + * * Streaming: a `key: value` line commits on its newline. A bare `key:` (or * `-`) is held for one line to decide between a nested block and a null * value. A block scalar is buffered until it closes — never past the @@ -335,17 +349,15 @@ public class YamlParser( val (key, end) = scanQuoted(content) ?: return null var i = end while (i < content.length && content[i] == ' ') i++ - if (i >= content.length || content[i] != ':') return null - if (i + 1 < content.length && content[i + 1] != ' ' && content[i + 1] != '\t') return null + if (i >= content.length || !content.isMappingColonAt(i)) return null return key to content.substring(i + 1) } if (first in KEY_FORBIDDEN_START) return null if (content == "?" || content.startsWith("? ")) return null var i = 0 while (i < content.length) { - val c = content[i] - if (c == ':' && (i + 1 == content.length || content[i + 1] == ' ' || content[i + 1] == '\t')) break - if (c == '#' && i > 0 && content[i - 1] == ' ') return null + if (content.isMappingColonAt(i)) break + if (content.isCommentStartAt(i)) return null i++ } if (i == content.length) return null @@ -388,7 +400,10 @@ public class YamlParser( } i++ } - if (!isBlankOrComment(s, i)) return null + // PyYAML wants whitespace before a comment here, unlike after a + // quoted scalar or a flow collection (`|-#x` is refused) + val rest = skipSpaces(s, i) + if (rest < s.length && !s.isCommentStartAt(rest)) return null return Value.Block(folded, chomp, digit) } @@ -471,8 +486,8 @@ public class YamlParser( var i = start while (i < s.length) { val c = s[i] - if (c == ':' && (i + 1 >= s.length || s[i + 1] == ' ' || s[i + 1] == ',' || s[i + 1] == '}')) break - if (c == ',' || c == '}' || c == ']' || c == '[' || c == '{') return null + if (s.isMappingColonAt(i)) break + if (c in FLOW_INDICATORS || s.isFlowColonAt(i) || s.isCommentStartAt(i)) return null i++ } val key = s.substring(start, i).trim() @@ -487,8 +502,7 @@ public class YamlParser( while (i < s.length) { val c = s[i] if (c == ',' || c == ']' || c == '}') break - if (c == ':' && (i + 1 >= s.length || s[i + 1] == ' ')) return null - if (c == '#' && i > start && s[i - 1] == ' ') return null + if (s.isMappingColonAt(i) || s.isFlowColonAt(i) || s.isCommentStartAt(i)) return null i++ } val text = s.substring(start, i).trim() @@ -522,6 +536,16 @@ private const val ITEM = "item" // supported subset (`-` is handled by the sequence check first). private const val KEY_FORBIDDEN_START = "[]{}&*!|>%@`," +private const val FLOW_INDICATORS = ",[]{}" + +// A `:` followed by a flow indicator inside a flow collection, where the +// readers disagree: PyYAML ends the plain scalar there (`[a:, b]` holds the +// mapping `{a: null}`), Psych refuses the line and go-yaml v2 reads the +// colon as content (`"a:"`). No reading is right for all three, so the +// line is left outside the subset and kept as written. +private fun String.isFlowColonAt(i: Int): Boolean = + this[i] == ':' && i + 1 < length && this[i + 1] in FLOW_INDICATORS + private fun leadingSpaces(s: String): Int { var i = 0 while (i < s.length && s[i] == ' ') i++ @@ -537,18 +561,22 @@ private fun skipSpaces(s: String, from: Int): Int { private fun isSequenceEntry(content: String): Boolean = content == "-" || content.startsWith("- ") || content.startsWith("-\t") -// True when only whitespace or a `#` comment follows index [from]. +// True when only whitespace or a `#` comment follows index [from], the end +// of a quoted scalar or a flow collection. Not [isCommentStartAt]: that is +// the rule inside a plain scalar, while after a closed token the front +// matter readers (PyYAML, Psych, go-yaml v2 — all libyaml's scanner) start a +// comment at any `#`, even with no space (`"x"#b`). DIVERGENCE: YAML 1.2 +// §6.6 wants whitespace first, and a strict reader (npm `yaml`, +// snakeyaml-engine) refuses such a line; the writer never produces one. private fun isBlankOrComment(s: String, from: Int): Boolean { val i = skipSpaces(s, from) - return i >= s.length || (s[i] == '#' && (i == from || s[i - 1] == ' ' || s[i - 1] == '\t')) + return i >= s.length || s[i] == '#' } // A ` #` (hash preceded by whitespace) starts a trailing comment. private fun stripTrailingComment(value: String): String { for (i in 1 until value.length) { - if (value[i] == '#' && (value[i - 1] == ' ' || value[i - 1] == '\t')) { - return value.substring(0, i) - } + if (value.isCommentStartAt(i)) return value.substring(0, i) } return value } @@ -631,39 +659,3 @@ private fun StringBuilder.appendCodePointCompat(code: Int) { append(((v shr 10) + 0xD800).toChar()) append(((v and 0x3FF) + 0xDC00).toChar()) } - -private val YAML_NULLS = setOf("null", "Null", "NULL", "~") - -// YAML 1.2 core schema booleans plus the 1.1 forms still common in front -// matter (Jekyll's and Hugo's YAML readers accept them). -private val YAML_BOOLS = setOf( - "true", "True", "TRUE", "false", "False", "FALSE", - "yes", "Yes", "YES", "no", "No", "NO", - "on", "On", "ON", "off", "Off", "OFF" -) - -// Each pattern is explicitly anchored: Kotlin/JS resolves `matches` through -// the leftmost match, so an unanchored alternation lets the first branch win -// on a prefix (`0x1F` matched `[0-9]+` as `0`, a full timestamp matched its -// date-only branch) and the whole-input check then fails — JVM backtracks -// across the branches and never showed it. -private val YAML_INT = Regex("""^(?:[-+]?[0-9]+|0o[0-7]+|0x[0-9a-fA-F]+)$""") - -private val YAML_FLOAT = Regex( - """^(?:[-+]?(\.[0-9]+|[0-9]+(\.[0-9]*)?)([eE][-+]?[0-9]+)?|[-+]?\.(inf|Inf|INF)|\.(nan|NaN|NAN))$""" -) - -private val YAML_TIMESTAMP = Regex( - """^(?:[0-9]{4}-[0-9]{2}-[0-9]{2}""" + - """|[0-9]{4}-[0-9]{1,2}-[0-9]{1,2}([Tt]|[ \t]+)[0-9]{1,2}:[0-9]{2}:[0-9]{2}(\.[0-9]*)?([ \t]*(Z|[-+][0-9]{1,2}(:[0-9]{2})?))?)$""" -) - -// The `type` a plain scalar resolves to, or null for a string. -internal fun yamlScalarType(text: String): String? = when { - text in YAML_NULLS -> "null" - text in YAML_BOOLS -> "bool" - YAML_INT.matches(text) -> "int" - YAML_FLOAT.matches(text) -> "float" - YAML_TIMESTAMP.matches(text) -> "timestamp" - else -> null -} diff --git a/markanywhere-yaml/src/commonMain/kotlin/YamlScalarType.kt b/markanywhere-yaml/src/commonMain/kotlin/YamlScalarType.kt new file mode 100644 index 0000000..87327f8 --- /dev/null +++ b/markanywhere-yaml/src/commonMain/kotlin/YamlScalarType.kt @@ -0,0 +1,357 @@ +/* + * Copyright 2026 Kazimierz Pogoda / Xemantic + * + * Licensed under the Apache License, Version 2.0 (the "License"); + * you may not use this file except in compliance with the License. + * You may obtain a copy of the License at + * + * https://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, software + * distributed under the License is distributed on an "AS IS" BASIS, + * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + * See the License for the specific language governing permissions and + * limitations under the License. + */ + +package com.xemantic.markanywhere.yaml + +// The one copy of the plain-scalar typing rules: [YamlParser] resolves a +// plain scalar's `type` with them, [YamlWriter] quotes a string they would +// type. A plain scalar is typed exactly when a reader front matter is +// written for reads it as something other than a string — the YAML 1.2 +// core schema, and the YAML 1.1 shapes of Psych (Jekyll), PyYAML and go-yaml +// v2 (Hugo). The number rules are therefore transcribed per reader, each +// less the shapes that reader then refuses to construct, rather than merged +// into one pattern: every reader has its own separator rules (Psych takes +// `,` only before a digit, PyYAML takes `_` anywhere after the first digit, +// go-yaml v2 deletes every `_` first), and a merged pattern would type +// shapes none of them does (`1,`, `1_,2`). So `type=int` / `type=float` means +// some reader reads a number — reading it yourself needs that reader's +// rules (drop `_` and `,`; base 60, binary and octal prefixes). +// Covered beyond the core schema: +// - booleans and nulls in any letter case (`yEs`, `nULL` — Psych ignores +// case), and the one-letter booleans `y` / `n` (go-yaml v2); +// - the special floats in any letter case (`.Nan`, `+.InF`); +// - integers with separators, binary, upper-case and signed prefixes +// (`1_000`, `1,000`, `0b101`, `0X1F`, `+0x1F`, `+_1`); +// - floats with separators (`1_000.5`, `1_e5`), and a leading-zero decimal +// that is not octal (`08`, go-yaml v2); +// - base-60 numbers (`12:30` an int, `190:20:30.15` a float); +// - timestamps with a one-digit month / day (`2024-5-1`) or an offset +// without a colon (`2016-01-01 12:00:00 -0500`, Jekyll's documented +// format) — a date only when it is in the calendar, a date and time only +// within the ranges Psych's `Time` accepts (it normalises `2023-02-31`). +// A shape go-yaml v2 alone reads as a number is typed only within the +// range it parses (`1e999`, `0X` and 17 hex digits are strings to it). +// Typing such a value keeps a source document as written: the writer writes +// a typed scalar bare, so only a *string* of one of these shapes is quoted. + +private val YAML_NULLS = setOf("null", "~") + +private val YAML_BOOLS = setOf("true", "false", "yes", "no", "on", "off", "y", "n") + +private val YAML_SPECIAL_FLOATS = setOf(".inf", "+.inf", "-.inf", ".nan") + +// the longest word above, so a long string is never lowercased +private val YAML_WORD_MAX_LENGTH = + (YAML_NULLS + YAML_BOOLS + YAML_SPECIAL_FLOATS).maxOf { it.length } + +// Each pattern is explicitly anchored: Kotlin/JS resolves `matches` through +// the leftmost match, so an unanchored alternation lets the first branch win +// on a prefix (`0x1F` matched `[0-9]+` as `0`, a full timestamp matched its +// date-only branch) and the whole-input check then fails — JVM backtracks +// across the branches and never showed it. +// No pattern repeats a group (`(?:_?[0-9])*`), not even a bounded number of +// times: the JVM engine matches each repetition of a group one stack frame +// deeper, so a long digit run overflows the stack. A digit run is a +// character class instead, and the rule the group expressed is checked by +// hand ([hasSeparatorsBeforeDigits], [hasUnderscoresBetweenDigits], +// [isBase60Tail]). +private fun anchored(pattern: String) = Regex("^(?:$pattern)$") + +// PyYAML's resolver, less a base prefix with no digit (`0x_`), which it +// resolves and then fails to construct. The digits after a base prefix are +// written `_*[01][01_]*`, not `[01_]*[01][01_]*`: the same strings, but the +// latter can split a digit run many ways and backtracks quadratically over +// a long near-miss. +private val PYYAML_INT = anchored( + "[-+]?0b_*[01][01_]*|[-+]?0[0-7_]+|[-+]?(?:0|[1-9][0-9_]*)" + + "|[-+]?0x_*[0-9a-fA-F][0-9a-fA-F_]*" +) + +private val PYYAML_FLOAT = anchored( + """[-+]?[0-9][0-9_]*\.[0-9_]*(?:[eE][-+][0-9]+)?|\.[0-9][0-9_]*(?:[eE][-+][0-9]+)?""" +) + +// The base-60 int and float: PyYAML's `[1-9][0-9_]*(?::[0-5]?[0-9])+` and +// Psych's `[0-9][0-9_]*(?::[0-5]?[0-9]){1,2}` (so a leading `0` only with at +// most two segments), and for the float `[0-9][0-9_]*(?::[0-5]?[0-9])+` +// with a fraction — Psych's float takes the same with at most two segments, +// a subset. The group is written as a run it is the tail of, which +// [isBase60Tail] then checks; the first group is the leading digit. +private val BASE60_INT = anchored("[-+]?([0-9])[0-9_]*(:[0-9:]*)") + +private val BASE60_FLOAT = anchored("""[-+]?[0-9][0-9_]*(:[0-9:]*)\.[0-9_]*""") + +// Psych's scalar scanner, with its default (legacy) integers that allow `,` +private val PSYCH_INT = anchored( + "[-+]?0b[_,]*[01][01_,]*|[-+]?0[0-7_,]+|[-+]?0" + + "|[-+]?0x[_,]*[0-9a-fA-F][0-9a-fA-F_,]*" +) + +// … and its decimal, `[1-9](?:[0-9]|[,_][0-9])*`: a separator only before a +// digit, which [hasSeparatorsBeforeDigits] checks +private val PSYCH_DECIMAL_INT = anchored("[-+]?[1-9][0-9,_]*") + +private val PSYCH_FLOAT = anchored("""[-+]?(?:[0-9][0-9_,]*\.[0-9]*|\.[0-9]+)(?:[eE][-+][0-9]+)?""") + +// go-yaml v2 deletes every `_` from a text starting with a sign or a digit, +// then tries Go's `ParseInt` with base 0 (which reads a `0x` / `0o` / `0b` +// prefix in either case, and a bare leading `0` as octal — so `08` fails +// here and falls through to the float), a `0b` prefix of its own (which +// lets a sign follow an unsigned prefix: `0b-1`, not `-0b-1`) and a YAML 1.2 +// float +private val GO_YAML_INT = anchored( + "[-+]?(?:0[0-7]*|[1-9][0-9]*|0[xX][0-9a-fA-F]+|0[oO][0-7]+|0[bB][01]+)|0b[-+][01]+" +) + +private val GO_YAML_FLOAT = anchored("""[-+]?(?:\.[0-9]+|[0-9]+(?:\.[0-9]*)?)(?:[eE][-+]?[0-9]+)?""") + +// … but hands a text starting with `.` to Go's `ParseFloat` as is, which +// takes an `_` only between two digits +// ([hasUnderscoresBetweenDigits] checks that) +private val GO_YAML_DOT_FLOAT = anchored("""\.[0-9][0-9_]*(?:[eE][-+]?[0-9][0-9_]*)?""") + +// The union of the Psych and PyYAML timestamp shapes (go-yaml v2 decodes a +// timestamp into a string); the groups are the date, the time and the +// offset digits, so [isTimestampInRange] can check their values and +// [isPyYamlTimestamp] PyYAML's narrower shape. +private val YAML_TIMESTAMP = anchored( + """(-?[0-9]{4})-([0-9]{1,2})-([0-9]{1,2})""" + + """(?:(?:[Tt]|[ \t]+)([0-9]{1,2}):([0-9]{2}):([0-9]{2})(?:\.[0-9]*)?""" + + """(?:[ \t]*(?:Z|[-+]([0-9]{1,2}(?::?[0-9]{2})?)))?)?""" +) + +// Every character a numeric or timestamp shape can hold, so a string with +// any other one is a string without running a pattern: digits, signs, +// separators, base prefixes and hex digits, exponents, and a timestamp's +// `T` / `Z` and spaces. +private fun Char.canBeNumeric(): Boolean = + this in '0'..'9' || this in "+-._,:" || this in 'a'..'f' || this in 'A'..'F' || + this in "xXoO tTZ\t" + +// The `type` a plain scalar resolves to, or null for a string. +internal fun yamlScalarType(text: String): String? { + if (text.isEmpty()) return null + if (text.length <= YAML_WORD_MAX_LENGTH) { + val word = text.lowercase() + if (word in YAML_NULLS) return "null" + if (word in YAML_BOOLS) return "bool" + if (word in YAML_SPECIAL_FLOATS) return "float" + } + // every numeric shape starts with a sign, a digit or a dot, so the + // patterns only run on a string that can match them + val first = text[0] + if (first != '-' && first != '+' && first != '.' && first !in '0'..'9') return null + if (!text.all { it.canBeNumeric() }) return null + // a date is the only shape with a `-` after four digits, and a time + // follows one, so a timestamp never runs the number patterns + if (text.isTimestampShaped()) { + return if (timestampVerdict(text) == TYPED) "timestamp" else null + } + // … and base 60 is the only number with a `:` + if (':' in text) { + val int = BASE60_INT.matchEntire(text) + if (int != null && isBase60Int(int.groupValues)) return "int" + val float = BASE60_FLOAT.matchEntire(text) + return if (float != null && isBase60Tail(float.groupValues[1])) "float" else null + } + if (PYYAML_INT.matches(text) || PSYCH_INT.matches(text)) return "int" + if (PSYCH_DECIMAL_INT.matches(text) && text.hasSeparatorsBeforeDigits()) return "int" + if (PYYAML_FLOAT.matches(text) || PSYCH_FLOAT.matches(text)) return "float" + if (first == '.') { + if (GO_YAML_DOT_FLOAT.matches(text) && text.hasUnderscoresBetweenDigits() && + text.replace("_", "").toDouble().isFinite() + ) { + return "float" + } + } else { + val plain = if ('_' in text) text.replace("_", "") else text + if (GO_YAML_INT.matches(plain) && plain.fitsGoYamlInt()) return "int" + if (GO_YAML_FLOAT.matches(plain) && plain.toDouble().isFinite()) return "float" + } + return null +} + +// Whether go-yaml v2 parses this integer (`_` already deleted): Go's +// `ParseInt` takes a sign and up to 64 bits signed, `ParseUint` no sign and +// up to 64 bits unsigned; its own `0b` prefix hands the rest after it, +// sign included, to the same two. +private fun String.fitsGoYamlInt(): Boolean { + val binarySigned = startsWith("0b-") || startsWith("0b+") + val number = if (binarySigned) substring(2) else this + val sign = number[0].takeIf { it == '-' || it == '+' } + val unsigned = if (sign != null) number.substring(1) else number + val (radix, digits) = when { + binarySigned -> 2 to unsigned + unsigned.length > 1 && unsigned[0] == '0' -> when (unsigned[1]) { + 'x', 'X' -> 16 to unsigned.substring(2) + 'o', 'O' -> 8 to unsigned.substring(2) + 'b', 'B' -> 2 to unsigned.substring(2) + else -> 8 to unsigned.substring(1) + } + else -> 10 to unsigned + } + val magnitude = digits.toULongOrNull(radix) ?: return false + return when (sign) { + '-' -> magnitude <= Long.MIN_VALUE.toULong() + '+' -> magnitude <= Long.MAX_VALUE.toULong() + else -> true + } +} + +// Whether the text starts like a [YAML_TIMESTAMP]: four digits, after an +// optional `-`, then a `-`. +private fun String.isTimestampShaped(): Boolean { + val start = if (startsWith('-')) 1 else 0 + return length > start + 4 && this[start + 4] == '-' && + (start until start + 4).all { this[it] in '0'..'9' } +} + +// Whether every `,` / `_` is followed by a digit. +private fun String.hasSeparatorsBeforeDigits(): Boolean = indices.all { i -> + this[i] != ',' && this[i] != '_' || i + 1 < length && this[i + 1] in '0'..'9' +} + +// Whether every `_` sits between two digits. +private fun String.hasUnderscoresBetweenDigits(): Boolean = indices.all { i -> + this[i] != '_' || i > 0 && this[i - 1] in '0'..'9' && i + 1 < length && this[i + 1] in '0'..'9' +} + +// Whether [tail] is one or more `:` segments of `[0-5]?[0-9]`. +private fun isBase60Tail(tail: String): Boolean = + tail.split(':').drop(1).all { segment -> + segment.length == 1 || segment.length == 2 && segment[0] in '0'..'5' + } + +// Whether a [BASE60_INT] match is PyYAML's (no leading `0`, any number of +// segments) or Psych's (at most two). +private fun isBase60Int(groups: List<String>): Boolean { + val tail = groups[2] + return isBase60Tail(tail) && (groups[1] != "0" || tail.count { it == ':' } <= 2) +} + +private enum class TimestampVerdict { NONE, TYPED, REFUSED } + +// For a timestamp-shaped text: [TimestampVerdict.TYPED] when some reader +// types it, [TimestampVerdict.REFUSED] when PyYAML resolves its shape and +// then refuses the value, [TimestampVerdict.NONE] for a plain string. +private fun timestampVerdict(text: String): TimestampVerdict { + val groups = YAML_TIMESTAMP.matchEntire(text)?.groupValues ?: return NONE + return when { + isTimestampInRange(groups) -> TYPED + isPyYamlTimestamp(groups) -> REFUSED + else -> NONE + } +} + +// Whether a [YAML_TIMESTAMP] match is also PyYAML's own shape: no sign on the +// year, a two-digit month and day for a date alone, and a colon in an +// offset with minutes. +private fun isPyYamlTimestamp(groups: List<String>): Boolean { + if (groups[1].startsWith('-')) return false + if (groups[4].isEmpty()) return groups[2].length == 2 && groups[3].length == 2 + val offset = groups[7] + return offset.length <= 2 || ':' in offset +} + +// Psych checks a date against the calendar, but normalises a date and time +// (`2023-02-31T10:00:00` is March 3rd), refusing only the values its `Time` +// cannot take — hour 24 only as `24:00:00`, an offset under a day; a +// date-only shape with a sign is not a timestamp at all. +private fun isTimestampInRange(groups: List<String>): Boolean { + val year = groups[1].toInt() + val month = groups[2].toInt() + val day = groups[3].toInt() + if (month !in 1..12) return false + if (groups[4].isNotEmpty()) { + val hour = groups[4].toInt() + val minute = groups[5].toInt() + val second = groups[6].toInt() + val offset = groups[7] + return day in 1..31 && minute <= 59 && second <= 60 && + (hour <= 23 || hour == 24 && minute == 0 && second == 0) && + (offset.isEmpty() || offsetMinutes(offset) < 24 * 60) + } + // `-0000` is a signed year too, though its value is not negative + if (groups[1].startsWith('-')) return false + return day in 1..daysInMonth(year, month) +} + +// The offset as Psych's `parse_time` splits it: at the colon when there is +// one, else up to two digits of hours and the rest as minutes (`+530` is +// 53 hours, `+070` is 7). Minutes past 59 carry into the hours (`+05:99`). +private fun offsetMinutes(offset: String): Int { + val colon = offset.indexOf(':') + val split = if (colon >= 0) colon else minOf(2, offset.length) + val hours = offset.substring(0, split).toInt() + val minutes = offset.substring(if (colon >= 0) split + 1 else split) + return hours * 60 + (if (minutes.isEmpty()) 0 else minutes.toInt()) +} + +private fun daysInMonth(year: Int, month: Int): Int = when (month) { + 2 -> if (year % 4 == 0 && (year % 100 != 0 || year % 400 == 0)) 29 else 28 + 4, 6, 9, 11 -> 30 + else -> 31 +} + +// Shapes some reader resolves and then refuses to load, so the whole +// document fails: PyYAML's value / merge tags (`=`, `<<`), which its +// SafeLoader cannot construct; a base prefix with no digit (`0x_`, PyYAML +// and Psych); an exponent with no mantissa digit (`.e+4`, Psych); and, in +// [isUnsafePlainScalar], a PyYAML timestamp shape ([isPyYamlTimestamp]) out +// of range, which +// PyYAML refuses (`2024-13-45`, `2024-01-01 24:30:00`) — the parser does +// not type it either, or it would not get there. There is no type to report +// for them, so the parser keeps them strings — DIVERGENCE: the writer still +// quotes them, which rewrites such a plain source value on its first render. +private val YAML_REJECTED = anchored("""=|<<|[-+]?0[bx][_,]+|[-+]?\.[eE][-+][0-9]+""") + +// Whether a plain scalar would not read back as this string in every +// reader: the parser types it, or some reader refuses it. +internal fun isUnsafePlainScalar(text: String): Boolean { + // typed or refused, decided by one match of the timestamp pattern + if (text.isTimestampShaped()) return timestampVerdict(text) != NONE + if (yamlScalarType(text) != null) return true + // every other refused shape starts with one of these, so the pattern + // only runs on a string that can match it + val first = text.firstOrNull() ?: return false + return (first == '=' || first == '<' || first == '0' || first == '+' || first == '-' || first == '.') && + YAML_REJECTED.matches(text) +} + +// Whether a plain identifier-shaped mapping key would be read as something +// other than a string: Psych types the boolean and null words in any letter +// case (PyYAML only some cases of them), and go-yaml v2 its one-letter +// booleans `y` / `n` too — except in a [topLevel] key, which it decodes into +// a front matter's `map[string]interface{}`, while a nested mapping gets +// `interface{}` keys (`params:\n y: 1` has the key `true`). +// The parser never types a key, so this is quoting the writer alone does. +internal fun isTypedPlainKey(key: String, topLevel: Boolean): Boolean { + if (key.length > YAML_WORD_MAX_LENGTH) return false + val word = key.lowercase() + return word in YAML_NULLS || word in YAML_BOOLS && (!topLevel || word != "y" && word != "n") +} + +// The two indicators that can end a plain scalar in block context, shared by +// [YamlParser] (where it splits a key and strips a comment) and [YamlWriter] +// (which quotes a string holding either), so the two cannot drift apart: +// a `:` followed by whitespace or the end is a mapping colon, a `#` after +// whitespace starts a comment. Anywhere else both are plain content. + +internal fun String.isMappingColonAt(i: Int): Boolean = + this[i] == ':' && (i + 1 == length || this[i + 1] == ' ' || this[i + 1] == '\t') + +internal fun String.isCommentStartAt(i: Int): Boolean = + this[i] == '#' && i > 0 && (this[i - 1] == ' ' || this[i - 1] == '\t') diff --git a/markanywhere-yaml/src/commonMain/kotlin/YamlWriter.kt b/markanywhere-yaml/src/commonMain/kotlin/YamlWriter.kt index 682d35a..e9fada5 100644 --- a/markanywhere-yaml/src/commonMain/kotlin/YamlWriter.kt +++ b/markanywhere-yaml/src/commonMain/kotlin/YamlWriter.kt @@ -35,8 +35,10 @@ import com.xemantic.markanywhere.SemanticEvent * - a scalar with a `type` (`bool`, `int`, `float`, `null`, `timestamp`) is * written bare; a string is written plain unless a plain scalar would * re-parse as something else — reserved literals, numbers, timestamps, - * a leading indicator, `: `, ` #`, surrounding whitespace, a control - * character — in which case it is double-quoted with the YAML escapes; + * including the YAML 1.1 shapes the parser types (`12:30`, `1_000`, + * `yEs`), a leading indicator, `: ` or a trailing `:`, ` #`, surrounding + * whitespace, a line break or tab, a character YAML requires escaped — + * in which case it is double-quoted with the YAML escapes; * - a multi-line string becomes a literal block scalar (`|`, `|-`, `|+` * according to its trailing newlines) unless its first line starts with * whitespace or it has no content, which fall back to a quoted scalar; @@ -44,9 +46,9 @@ import com.xemantic.markanywhere.SemanticEvent * text is a bare `key:`; `type=seq` / `type=map` with no children are * `[]` / `{}`; * - a key is written plain only when identifier-shaped (letters, digits, - * `_`, `-`, `.`, starting with a letter or `_`) and not a reserved - * literal — a YAML 1.1 reader takes a bare `yes:` as a boolean key — and - * double-quoted otherwise, even where YAML would not require it; + * `_`, `-`, `.`, starting with a letter or `_`) and not a key a reader + * would type — Psych takes a bare `yes:` or `nULL:` as a boolean / null + * key, go-yaml v2 a `y:` below the top level — and double-quoted otherwise, even where YAML would not require it; * - a `text` child of a container (the parser's verbatim fallback for a * line outside its YAML subset) is written as-is, on its own line(s). * @@ -70,7 +72,9 @@ public class YamlWriter( // the column this node's own line starts at val indent: Int, // the column this node's children start at - val childIndent: Int + val childIndent: Int, + // no entry or item encloses it — a key of the root mapping + val topLevel: Boolean ) { var state: State = UNDECIDED val scalar = StringBuilder() @@ -78,7 +82,7 @@ public class YamlWriter( } private val stack = ArrayDeque<Node>().apply { - addLast(Node(OTHER, key = null, type = null, indent = 0, childIndent = 0)) + addLast(Node(OTHER, key = null, type = null, indent = 0, childIndent = 0, topLevel = true)) } // True right after `-` was written for an item whose first child @@ -109,7 +113,8 @@ public class YamlWriter( } val indent = parent.childIndent val childIndent = if (kind == OTHER) indent else indent + 2 - stack.addLast(Node(kind, event.attributes["key"], event.attributes["type"], indent, childIndent)) + val topLevel = parent.topLevel && !parent.isValue + stack.addLast(Node(kind, event.attributes["key"], event.attributes["type"], indent, childIndent, topLevel)) } private fun text(text: String) { @@ -151,14 +156,14 @@ public class YamlWriter( atLineStart = false afterDash = true } else { - out(linePrefix(node) + renderKey(node.key) + ":\n") + out(linePrefix(node) + renderKey(node) + ":\n") atLineStart = true } if (pendingVerbatim != null) writeVerbatim(pendingVerbatim) } private fun writeLine(node: Node, value: String) { - val head = if (node.kind == ITEM) "-" else renderKey(node.key) + ":" + val head = if (node.kind == ITEM) "-" else renderKey(node) + ":" out(linePrefix(node) + head + (if (value.isEmpty()) "" else " $value") + "\n") atLineStart = true } @@ -200,7 +205,11 @@ public class YamlWriter( private fun canBlock(text: String): Boolean { val first = text[0] if (first == ' ' || first == '\t' || first == '\n') return false - for (c in text) if (c != '\n' && c != '\t' && c < ' ') return false + for (i in text.indices) { + val c = text[i] + // a carriage return would be read as a line break + if (c == '\r' || text.isYamlUnprintableAt(i)) return false + } return text.trimEnd('\n').isNotEmpty() } @@ -224,14 +233,14 @@ public class YamlWriter( return sb.toString() } - // A key is written plain only when it is identifier-shaped and not a - // reserved literal: the Markdown parser's front matter detection requires + // A key is written plain only when it is identifier-shaped and would not + // be typed (`yes`, `nULL`, a nested `y`): the Markdown parser's front matter detection requires // the first line to pass `isYamlKeyLine` (such a key, or a quoted one), // so quoting everything else keeps whatever entry comes first // re-detectable. The character rules are shared with that check. - private fun renderKey(key: String?): String { - val k = key ?: "" - return if (isIdentifierKey(k) && yamlScalarType(k) == null) k else quoted(k) + private fun renderKey(node: Node): String { + val k = node.key ?: "" + return if (isIdentifierKey(k) && !isTypedPlainKey(k, node.topLevel)) k else quoted(k) } } @@ -239,32 +248,59 @@ public class YamlWriter( // double-quoted output (see YAML 1.2 §6.4 / §6.6). private const val YAML_INDICATORS = "-?:,[]{}#&*!|>'\"%@`" -// A character a plain scalar cannot carry: C0 / DEL / NEL and the Unicode -// line and paragraph separators (YAML 1.2 §5.1 printable characters). -private val Char.isYamlControl: Boolean - get() = this < ' ' || this == '\u007f' || this == '\u0085' || this == '\u2028' || this == '\u2029' +// Whether the character at [i] may only appear `\u`-escaped: it is outside +// YAML's printable set (YAML 1.2 §5.1 — C0 controls other than tab, line +// feed and carriage return, DEL, the C1 controls, U+FFFE / U+FFFF, a +// surrogate not part of a pair), which Psych and PyYAML refuse anywhere in a +// document; a byte order mark, which must not appear inside a document +// (§5.2); or a line break for a YAML 1.1 reader (NEL, the Unicode line and +// paragraph separators). A lone surrogate has no valid YAML representation +// at all — raw it is unprintable, and Psych and go-yaml refuse its `\u` +// escape — so [quoted] writes it as U+FFFD (DIVERGENCE: lossy). Tab, line feed +// and carriage return are printable — each writer path decides where it can +// keep them. +private fun String.isYamlUnprintableAt(i: Int): Boolean { + val c = this[i] + return when { + c.isHighSurrogate() -> i + 1 == length || !this[i + 1].isLowSurrogate() + c.isLowSurrogate() -> i == 0 || !this[i - 1].isHighSurrogate() + else -> (c < ' ' && c != '\t' && c != '\n' && c != '\r') || + c in '\u007f'..'\u009f' || c == '\u2028' || c == '\u2029' || + c == '\ufeff' || c == '\ufffe' || c == '\uffff' + } +} -// A plain scalar the parser would not read back as the same string: one it -// would type (`yamlScalarType`), or whose shape is an indicator, a comment, -// a mapping colon, surrounding whitespace or a control character. +// A plain scalar a reader would not read back as the same string: one the +// parser would type or a reader refuses (`isUnsafePlainScalar`), or whose +// shape is an indicator, a comment (`#` after whitespace), a mapping colon +// (`:` before whitespace or at the end), surrounding whitespace, a line +// break or tab, or an unprintable character. A `:` or `#` anywhere else +// is plain content (`https://x/y`, `C#`), as is an inner `"` or `\`. private fun needsQuoting(s: String): Boolean { if (s.isEmpty()) return true if (s.first().isWhitespace() || s.last().isWhitespace()) return true if (s[0] in YAML_INDICATORS) return true - for (c in s) if (c == ':' || c == '#' || c == '"' || c == '\\' || c.isYamlControl) return true - return yamlScalarType(s) != null + for (i in s.indices) { + val c = s[i] + if (c == '\n' || c == '\r' || c == '\t' || s.isYamlUnprintableAt(i)) return true + if (s.isMappingColonAt(i) || s.isCommentStartAt(i)) return true + } + return isUnsafePlainScalar(s) } private fun quoted(s: String): String = buildString { +'"' - for (c in s) when { - c == '\\' -> +"\\\\" - c == '"' -> +"\\\"" - c == '\n' -> +"\\n" - c == '\r' -> +"\\r" - c == '\t' -> +"\\t" - c.isYamlControl -> +("\\u" + c.code.toString(16).padStart(4, '0')) - else -> +c + for (i in s.indices) when (val c = s[i]) { + '\\' -> +"\\\\" + '"' -> +"\\\"" + '\n' -> +"\\n" + '\r' -> +"\\r" + '\t' -> +"\\t" + else -> when { + !s.isYamlUnprintableAt(i) -> +c + c.isSurrogate() -> +'\uFFFD' + else -> +("\\u" + c.code.toString(16).padStart(4, '0')) + } } +'"' } diff --git a/markanywhere-yaml/src/commonTest/kotlin/YamlParserTest.kt b/markanywhere-yaml/src/commonTest/kotlin/YamlParserTest.kt index 7b46280..ef3d56a 100644 --- a/markanywhere-yaml/src/commonTest/kotlin/YamlParserTest.kt +++ b/markanywhere-yaml/src/commonTest/kotlin/YamlParserTest.kt @@ -23,6 +23,7 @@ import com.xemantic.markanywhere.test.sameAs import kotlinx.coroutines.flow.asFlow import kotlinx.coroutines.flow.flow import kotlinx.coroutines.test.runTest +import com.xemantic.kotlin.test.assert import kotlin.test.Test /** @@ -301,6 +302,350 @@ class YamlParserTest { } } + @Test + fun `should type the plain scalars a YAML 1_1 reader types`() = runTest { + // given — YAML 1.2 reads these as strings, but Psych (Jekyll), PyYAML + // or go-yaml v2 type them, and front matter is written for those + // readers: Jekyll's documented date format carries a colonless offset + val textFlow = """ + date: 2016-01-01 12:00:00 -0500 + short: 2024-5-1 + answer: y + shout: N + mixed: yEs + nil: nULL + big: 1_000 + grouped: 1,000 + bin: 0b101 + signed: +0x1F + money: 1_000.5 + nan: .NaN + time: 12:30 + clock: 12:30:45:10 + angle: 190:20:30.15 + """.trimIndent().chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "date", "type" to "timestamp") { +"2016-01-01 12:00:00 -0500" } + "entry"("key" to "short", "type" to "timestamp") { +"2024-5-1" } + "entry"("key" to "answer", "type" to "bool") { +"y" } + "entry"("key" to "shout", "type" to "bool") { +"N" } + "entry"("key" to "mixed", "type" to "bool") { +"yEs" } + "entry"("key" to "nil", "type" to "null") { +"nULL" } + "entry"("key" to "big", "type" to "int") { +"1_000" } + "entry"("key" to "grouped", "type" to "int") { +"1,000" } + "entry"("key" to "bin", "type" to "int") { +"0b101" } + "entry"("key" to "signed", "type" to "int") { +"+0x1F" } + "entry"("key" to "money", "type" to "float") { +"1_000.5" } + "entry"("key" to "nan", "type" to "float") { +".NaN" } + "entry"("key" to "time", "type" to "int") { +"12:30" } + "entry"("key" to "clock", "type" to "int") { +"12:30:45:10" } + "entry"("key" to "angle", "type" to "float") { +"190:20:30.15" } + } + } + + @Test + fun `should type a timestamp offset the way Psych splits and bounds it`() = runTest { + // given — Psych reads up to two digits as the offset hours and the + // rest as its minutes, and only bounds the whole offset under a day: + // minutes past 59 carry into the hours, and a three-digit offset is + // two hour digits and one minute digit + val textFlow = """ + a: 2024-01-01 10:00:00 +05:99 + b: 2024-01-01 10:00:00 +0599 + c: 2024-01-01 10:00:00 +19:70 + d: 2024-01-01 10:00:00 +070 + e: 2024-01-01 10:00:00 -2359 + """.trimIndent().chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "a", "type" to "timestamp") { +"2024-01-01 10:00:00 +05:99" } + "entry"("key" to "b", "type" to "timestamp") { +"2024-01-01 10:00:00 +0599" } + "entry"("key" to "c", "type" to "timestamp") { +"2024-01-01 10:00:00 +19:70" } + "entry"("key" to "d", "type" to "timestamp") { +"2024-01-01 10:00:00 +070" } + "entry"("key" to "e", "type" to "timestamp") { +"2024-01-01 10:00:00 -2359" } + } + } + + @Test + fun `should not type a timestamp whose offset ends in a colon`() { + // given — a `:` ending a plain scalar is a mapping indicator, so + // Psych refuses the line: the shape can never be a timestamp + val values = listOf("2024-01-01 12:00:00 +5:", "2024-01-01 12:00:00 -05:") + + // when + val types = values.map { yamlScalarType(it) } + + // then + assert(types == listOf(null, null)) + } + + @Test + fun `should type the number shapes go-yaml v2 reads once underscores are removed`() = runTest { + // given — go-yaml v2 (Hugo) deletes every `_` before it parses a + // number, so an underscore next to a sign, an exponent or a base + // prefix still makes one; it also takes an upper-case base prefix + val textFlow = """ + a: 1_e5 + b: 1e5_ + c: 1_E5 + d: +_1 + e: +_1. + f: 0_x1 + g: 0X1F + """.trimIndent().chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "a", "type" to "float") { +"1_e5" } + "entry"("key" to "b", "type" to "float") { +"1e5_" } + "entry"("key" to "c", "type" to "float") { +"1_E5" } + "entry"("key" to "d", "type" to "int") { +"+_1" } + "entry"("key" to "e", "type" to "float") { +"+_1." } + "entry"("key" to "f", "type" to "int") { +"0_x1" } + "entry"("key" to "g", "type" to "int") { +"0X1F" } + } + } + + @Test + fun `should not type a shape no reader types`() = runTest { + // given — a trailing or doubled separator, a dot before an + // underscore, an exponent with no mantissa digit, a date that is not + // in the calendar, a date-only shape with a sign, a minute Psych's + // Time refuses, a sign after a signed binary prefix, an offset + // Psych splits into 53 hours or bounds at a day and a date-only + // shape with a signed zero year: PyYAML, Psych and go-yaml v2 all + // read strings + val textFlow = """ + a: 1, + b: 1,,2 + c: ._0 + d: .e5 + e: 2024-2-30 + f: -2024-01-01 + g: 2024-01-01T12:60:00 + h: -0b-1 + i: -0b+1 + j: 2024-01-01 10:00:00 +530 + k: 2024-01-01 10:00:00 +2400 + l: -0000-01-01 + """.trimIndent().chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "a") { +"1," } + "entry"("key" to "b") { +"1,,2" } + "entry"("key" to "c") { +"._0" } + "entry"("key" to "d") { +".e5" } + "entry"("key" to "e") { +"2024-2-30" } + "entry"("key" to "f") { +"-2024-01-01" } + "entry"("key" to "g") { +"2024-01-01T12:60:00" } + "entry"("key" to "h") { +"-0b-1" } + "entry"("key" to "i") { +"-0b+1" } + "entry"("key" to "j") { +"2024-01-01 10:00:00 +530" } + "entry"("key" to "k") { +"2024-01-01 10:00:00 +2400" } + "entry"("key" to "l") { +"-0000-01-01" } + } + } + + @Test + fun `should not type a number go-yaml v2 alone reads and then overflows`() = runTest { + // given — PyYAML and Psych need a `.` and a signed exponent for a + // float and take only lower-case `0x` / `0b` prefixes, so go-yaml v2 + // alone reads these shapes as numbers, and it falls back to a string + // when the value overflows its int64 / uint64 / float64 + val textFlow = """ + a: 1e999 + b: .5e999 + c: 0XFFFFFFFFFFFFFFFFF + d: 0o7777777777777777777777777 + e: +0XFFFFFFFFFFFFFFFF + f: -0X8000000000000001 + """.trimIndent().chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "a") { +"1e999" } + "entry"("key" to "b") { +".5e999" } + "entry"("key" to "c") { +"0XFFFFFFFFFFFFFFFFF" } + "entry"("key" to "d") { +"0o7777777777777777777777777" } + "entry"("key" to "e") { +"+0XFFFFFFFFFFFFFFFF" } + "entry"("key" to "f") { +"-0X8000000000000001" } + } + } + + @Test + fun `should type a number go-yaml v2 alone reads within its range`() = runTest { + // given — the largest uint64 and the smallest int64 still fit, a + // decimal too large for an integer is read as a float, and an + // underflowing float is zero rather than an error + val textFlow = """ + a: 0XFFFFFFFFFFFFFFFF + b: -0X8000000000000000 + c: 0b-1000000000000000000000000000000000000000000000000000000000000000 + d: -_99999999999999999999 + e: 1e-999 + """.trimIndent().chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "a", "type" to "int") { +"0XFFFFFFFFFFFFFFFF" } + "entry"("key" to "b", "type" to "int") { +"-0X8000000000000000" } + "entry"("key" to "c", "type" to "int") { + +"0b-1000000000000000000000000000000000000000000000000000000000000000" + } + "entry"("key" to "d", "type" to "float") { +"-_99999999999999999999" } + "entry"("key" to "e", "type" to "float") { +"1e-999" } + } + } + + @Test + fun `DIVERGENCE - should not type a plain scalar a reader refuses to load`() = runTest { + // given — PyYAML resolves `=` and `<<` to its value / merge tags and + // then refuses to construct them; it refuses a date not in the + // calendar or a time / offset out of range (which Psych reads as a + // string), and PyYAML / Psych refuse a base prefix with no digit or + // an exponent with no mantissa: none of them has a type to report + val textFlow = """ + a: = + b: << + c: 2024-13-45 + d: 0x_ + e: .e+4 + f: 2024-01-01 24:30:00 + g: 2024-01-01 10:00:00 +23:99 + """.trimIndent().chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "a") { +"=" } + "entry"("key" to "b") { +"<<" } + "entry"("key" to "c") { +"2024-13-45" } + "entry"("key" to "d") { +"0x_" } + "entry"("key" to "e") { +".e+4" } + "entry"("key" to "f") { +"2024-01-01 24:30:00" } + "entry"("key" to "g") { +"2024-01-01 10:00:00 +23:99" } + } + } + + @Test + fun `should resolve a long near-miss of a base prefixed number`() { + // given — a base prefix, a long digit run and one character outside + // the base but still numeric-shaped, so the number patterns do run: + // a pattern that lets the digits match two ways backtracks + // quadratically, which at this length is hours instead of + // milliseconds — a regression hangs the test rather than failing an + // assertion, since a wall-clock bound would be flaky across targets + val values = listOf("0x", "+0x").map { it + "1".repeat(500_000) + "." } + + listOf("0b", "-0b").map { it + "1".repeat(500_000) + "2" } + + for (value in values) { + // when + val type = yamlScalarType(value) + val unsafe = isUnsafePlainScalar(value) + + // then + assert(type == null) + assert(!unsafe) + } + } + + @Test + fun `should type a long run of digits without exhausting the stack`() { + // given — a repeated group is matched recursively by the JVM regex + // engine, one frame per repetition, so a pattern that repeats one + // over a digit run overflows the stack on a long enough value + val digits = "1".repeat(100_000) + val values = mapOf( + digits to "int", + "1,".repeat(50_000) + "1" to "int", + ".$digits" to "float", + ".1" + "_1".repeat(50_000) to "float", + "1" + ":1".repeat(50_000) to "int", + "1" + ":1".repeat(50_000) + ".5" to "float", + "${digits}g" to null, + "1,".repeat(50_000) + "g" to null, + ".${digits}g" to null, + ".1" + "_1".repeat(50_000) + "g" to null, + "1" + ":1".repeat(50_000) + "g" to null, + ) + + // when + val types = values.keys.map { yamlScalarType(it) } + + // then + assert(types == values.values.toList()) + } + + @Test + fun `should type a leading-zero decimal go-yaml v2 reads as a float`() = runTest { + // given — PyYAML and Psych read these as strings; go-yaml v2 takes + // the leading zero as an octal prefix, fails on the 8 or 9 and falls + // back to a float + val textFlow = """ + a: 08 + b: 019 + c: 0_9 + d: -08 + e: 07 + """.trimIndent().chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "a", "type" to "float") { +"08" } + "entry"("key" to "b", "type" to "float") { +"019" } + "entry"("key" to "c", "type" to "float") { +"0_9" } + "entry"("key" to "d", "type" to "float") { +"-08" } + "entry"("key" to "e", "type" to "int") { +"07" } + } + } + + @Test + fun `should not type a float shape without a digit`() = runTest { + // given — no reader types a lone dot or a sign and a dot + val textFlow = """ + dot: . + signed: +. + under: ._ + """.trimIndent().chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "dot") { +"." } + "entry"("key" to "signed") { +"+." } + "entry"("key" to "under") { +"._" } + } + } + @Test fun `should keep a quoted scalar a string`() = runTest { // given — quoting suppresses the type resolution @@ -636,10 +981,148 @@ class YamlParserTest { } } + @Test + fun `should read a colon before a tab as a mapping indicator in a flow mapping`() = runTest { + // given + val textFlow = "geo: {lat:\t1.5}\n".chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "geo") { + "entry"("key" to "lat", "type" to "float") { +"1.5" } + } + } + } + + @Test + fun `DIVERGENCE - should read a hash right after a closed token as a comment`() = runTest { + // given — YAML 1.2 §6.6 wants whitespace before a comment, but the + // front matter readers (PyYAML, Psych, go-yaml v2 — all libyaml's + // scanner) start one at any `#` after a quoted scalar or a flow + // collection, and read these lines as the values below; a strict + // reader (npm `yaml`, snakeyaml-engine) refuses them instead + val textFlow = "a: \"x\"#b\nc: 'y'#d\ntags: [e]#f\ngeo: {g: 1}#h\n".chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "a") { +"x" } + "entry"("key" to "c") { +"y" } + "entry"("key" to "tags") { + "item" { +"e" } + } + "entry"("key" to "geo") { + "entry"("key" to "g", "type" to "int") { +"1" } + } + } + } + // --- verbatim fallback (DIVERGENCE) ----------------------------------- @Test - fun `should DIVERGENCE keep an unrecognised line verbatim`() = runTest { + fun `should read a hash after a tab as a comment in a flow sequence`() = runTest { + // given — the comment leaves the sequence unterminated, which every + // reader refuses, so the line is outside the subset + val textFlow = "title: x\ntags: [a\t#b]\n".chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "title") { +"x" } + +"tags: [a\t#b]\n" + } + } + + @Test + fun `should read a hash after a tab as a comment in a key line`() = runTest { + // given — `#` after any whitespace starts a comment, so the colon + // behind it is not a mapping indicator and the line is not an entry + val textFlow = "title: x\na\t#b: c\n".chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "title") { +"x" } + +"a\t#b: c\n" + } + } + + @Test + fun `should read a hash after a space as a comment in a flow mapping key`() = runTest { + // given — the comment leaves the mapping unterminated, which every + // reader refuses, so the line is outside the subset + val textFlow = "title: x\ngeo: {a #b: c}\n".chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "title") { +"x" } + +"geo: {a #b: c}\n" + } + } + + @Test + fun `DIVERGENCE - should keep a flow scalar with a colon before a flow indicator verbatim`() = runTest { + // given — the readers disagree on a plain flow scalar whose `:` is + // followed by `,` / `[` / `]` / `{` / `}`: PyYAML reads a mapping + // (`[{draft: null}, x]`), Psych refuses the line, go-yaml v2 reads + // the colon as content (`"draft:"`), so no single reading is right and + // the line is kept as written; a colon followed by anything else is + // content for all three + val textFlow = """ + title: x + tags: [draft:, x] + one: [a:] + geo: {a:[1, 2]} + nested: {a:{b: 1}} + empty: {a:, b: 1} + last: {a:} + time: [a:b] + """.trimIndent().chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "title") { +"x" } + +"tags: [draft:, x]\none: [a:]\ngeo: {a:[1, 2]}\nnested: {a:{b: 1}}\nempty: {a:, b: 1}\nlast: {a:}\n" + "entry"("key" to "time") { + "item" { +"a:b" } + } + } + } + + @Test + fun `should keep a block scalar header with a hash right after it verbatim`() = runTest { + // given — Psych and go-yaml v2 read `|-#x` as a header and a comment, + // but PyYAML refuses it, so the line is outside the subset; with + // whitespace before the `#` every reader takes the comment + val textFlow = "a: |-#x\nb: | #y\n z\n".chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + +"a: |-#x\n" + "entry"("key" to "b") { +"z\n" } + } + } + + @Test + fun `DIVERGENCE - should keep an unrecognised line verbatim`() = runTest { // given — a complex key, a multi-line flow sequence, an unterminated // quote, a `key:value` without a space and a tab-indented line are // outside the subset; each line survives verbatim (with its newline) @@ -667,7 +1150,7 @@ class YamlParserTest { } @Test - fun `should DIVERGENCE keep a sequence line inside a mapping verbatim`() = runTest { + fun `DIVERGENCE - should keep a sequence line inside a mapping verbatim`() = runTest { // given — a `- item` at the mapping's own indentation is a YAML error val textFlow = """ title: x @@ -685,7 +1168,7 @@ class YamlParserTest { } @Test - fun `should DIVERGENCE parse anchors aliases and tags as plain strings`() = runTest { + fun `DIVERGENCE - should parse anchors aliases and tags as plain strings`() = runTest { // given val textFlow = """ base: &b value @@ -704,6 +1187,24 @@ class YamlParserTest { } } + @Test + fun `DIVERGENCE - should accept a mapping indicator inside a plain value`() = runTest { + // given — YAML readers (Psych, PyYAML) reject both lines + val textFlow = """ + note: Note: see + end: ends: + """.trimIndent().chunkedRandomly().asFlow() + + // when + val parsed = textFlow.parseYaml() + + // then + parsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "note") { +"Note: see" } + "entry"("key" to "end") { +"ends:" } + } + } + @Test fun `should be ready for a new document after finish`() = runTest { // given — the first document ends with a pending `key:` that finish diff --git a/markanywhere-yaml/src/commonTest/kotlin/YamlRoundTripTest.kt b/markanywhere-yaml/src/commonTest/kotlin/YamlRoundTripTest.kt index da884ba..c38a12d 100644 --- a/markanywhere-yaml/src/commonTest/kotlin/YamlRoundTripTest.kt +++ b/markanywhere-yaml/src/commonTest/kotlin/YamlRoundTripTest.kt @@ -64,6 +64,369 @@ class YamlRoundTripTest { } } + @Test + fun `should write a colon or hash that is not an indicator plain`() = runTest { + // given — `:` not followed by a space and `#` not after a space are + // plain-scalar content for every YAML reader (issue #80) + val values = listOf( + "https://xemantic.com/contact", "a:b", "C#", "a#b", "say \"hi\"", "C:\\path", + ) + for (value in values) { + val source = "k: $value\n" + + // when + val rendered = flowOf(source).parseYaml().renderYaml() + val reparsed = flowOf(rendered).parseYaml() + + // then + rendered sameAs source + reparsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "k") { +value } + } + } + } + + @Test + fun `should write a colon or hash that is not an indicator plain in a sequence item`() = runTest { + // given — an item's content is first checked for a compact mapping + // (`- key: value`), so a `:` or `#` there takes a different path + // through the parser than an entry value + val source = "k:\n - https://xemantic.com/contact\n - a:b\n - C#\n - a#b\n" + + // when + val rendered = flowOf(source).parseYaml().renderYaml() + val reparsed = flowOf(rendered).parseYaml() + + // then + rendered sameAs source + reparsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "k") { + "item" { +"https://xemantic.com/contact" } + "item" { +"a:b" } + "item" { +"C#" } + "item" { +"a#b" } + } + } + } + + @Test + fun `should keep quoting a colon or hash that is an indicator and YAML 1_1 sexagesimals`() = runTest { + // given — a mapping colon, a comment, a leading indicator, and base-60 + // numbers a YAML 1.1 reader (Psych, PyYAML) would type as int / float + val values = listOf( + "Note: see", "ends:", "a #b", "12:30", "+1:30", "190:20:30.15", "#tag", ": x", + ) + for (value in values) { + val events = semanticEvents { + "entry"("key" to "k") { +value } + } + + // when + val rendered = events.renderYaml() + val reparsed = flowOf(rendered).parseYaml() + + // then + rendered sameAs "k: \"$value\"\n" + reparsed.mergeAdjacentText() sameAs events + } + } + + @Test + fun `should quote a string a YAML 1_1 reader would type`() = runTest { + // given — plain scalars YAML 1.2 reads as strings, but Psych (Jekyll) + // or PyYAML type as a number, time, boolean or null, or refuse + val values = listOf( + "2024-05-01T10:00:00+0100", "2024-05-01 10:00:00 +0100", "2024-5-1", + "1_000", "1,000", "0b101", "+0x1F", "1_000.5", "1,000.5", "1.5_0", ".e+4", + "yEs", "oFF", "nULL", "y", "Y", "n", "N", ".Nan", ".iNf", "+.InF", "-.INf", "=", "<<", + ) + for (value in values) { + val events = semanticEvents { + "entry"("key" to "k") { +value } + } + + // when + val rendered = events.renderYaml() + val reparsed = flowOf(rendered).parseYaml() + + // then + rendered sameAs "k: \"$value\"\n" + reparsed.mergeAdjacentText() sameAs events + } + } + + @Test + fun `should quote a key a YAML 1_1 reader would type`() = runTest { + // given — Psych reads booleans and nulls in any letter case + val events = semanticEvents { + "entry"("key" to "yEs") { +"v" } + "entry"("key" to "nULL") { +"v" } + } + + // when + val rendered = events.renderYaml() + val reparsed = flowOf(rendered).parseYaml() + + // then + rendered sameAs "\"yEs\": v\n\"nULL\": v\n" + reparsed.mergeAdjacentText() sameAs events + } + + @Test + fun `should not quote a one-letter boolean key`() = runTest { + // given — go-yaml v2 reads a plain `y` / `n` value as a boolean, but + // decodes a front matter key into a string, and neither Psych nor + // PyYAML types them at all + val source = "x: 1\ny: 2\nn: 3\nY: 4\nN: 5\n" + + // when + val rendered = flowOf(source).parseYaml().renderYaml() + + // then + rendered sameAs source + } + + @Test + fun `should quote a one-letter boolean key below the top level`() = runTest { + // given — go-yaml v2 decodes only the top-level keys into strings; a + // nested mapping, in an entry or in an item, is decoded with + // interface{} keys, so a plain `y:` there becomes the key `true` + val events = semanticEvents { + "entry"("key" to "params") { + "entry"("key" to "y") { +"1" } + "entry"("key" to "N") { +"2" } + } + "entry"("key" to "list") { + "item" { + "entry"("key" to "n") { +"3" } + } + } + "entry"("key" to "y") { +"4" } + } + + // when + val rendered = events.renderYaml() + val reparsed = flowOf(rendered).parseYaml() + + // then + rendered sameAs """ + params: + "y": "1" + "N": "2" + list: + - "n": "3" + y: "4" + """.trimIndent() + "\n" + reparsed.mergeAdjacentText() sameAs events + } + + @Test + fun `should keep YAML 1_1 typed front matter values as written`() = runTest { + // given — Jekyll's documented date format, a short date, go-yaml v2 + // booleans and Psych numbers: the parser types them, so the writer + // writes them back bare instead of quoting them into strings + val source = """ + date: 2016-01-01 12:00:00 -0500 + short: 2024-5-1 + answer: y + mixed: yEs + big: 1_000 + time: 12:30 + offset: 2024-01-01 10:00:00 +05:99 + """.trimIndent() + "\n" + + // when + val rendered = flowOf(source).parseYaml().renderYaml() + + // then + rendered sameAs source + } + + @Test + fun `should keep a line the readers disagree on as written`() = runTest { + // given — PyYAML, Psych and go-yaml v2 read a colon before a flow + // indicator three different ways, and PyYAML alone refuses a hash + // right after a block scalar header, so these lines stay verbatim + val source = """ + tags: [draft:, x] + geo: {a:[1, 2]} + text: |-#x + """.trimIndent() + "\n" + + // when + val rendered = flowOf(source).parseYaml().renderYaml() + + // then + rendered sameAs source + } + + @Test + fun `should quote a string go-yaml v2 reads as a number once underscores are removed`() = runTest { + // given + val values = listOf("1_e5", "1e5_", "1_E5", "+_1", "+_1.", "0_x1", "0X1F") + for (value in values) { + val events = semanticEvents { + "entry"("key" to "k") { +value } + } + + // when + val rendered = events.renderYaml() + val reparsed = flowOf(rendered).parseYaml() + + // then + rendered sameAs "k: \"$value\"\n" + reparsed.mergeAdjacentText() sameAs events + } + } + + @Test + fun `should not quote a shape no reader types`() = runTest { + // given + val values = listOf( + "1,", "1,,2", "._0", ".e5", "2024-2-30", "2024-5-32", "2024-01-01 10:00:00 +530", + "1e999", ".5e999", "0XFFFFFFFFFFFFFFFFF", "0o7777777777777777777777777", + ) + for (value in values) { + val events = semanticEvents { + "entry"("key" to "k") { +value } + } + + // when + val rendered = events.renderYaml() + val reparsed = flowOf(rendered).parseYaml() + + // then + rendered sameAs "k: $value\n" + reparsed.mergeAdjacentText() sameAs events + } + } + + @Test + fun `DIVERGENCE - should quote a plain scalar a reader refuses to load`() = runTest { + // given — the parser leaves these strings (no reader has a type for + // them), but PyYAML or Psych refuse the whole document when they are + // plain, so the writer quotes them and the first render rewrites the + // source; PyYAML refuses a timestamp shape out of range + val source = """ + a: = + b: << + c: 2024-13-45 + d: 2024-01-01 24:30:00 + e: 0x_ + f: .e+4 + g: 2024-01-01 10:00:00 +23:99 + """.trimIndent() + "\n" + + // when + val rendered = flowOf(source).parseYaml().renderYaml() + + // then + rendered sameAs """ + a: "=" + b: "<<" + c: "2024-13-45" + d: "2024-01-01 24:30:00" + e: "0x_" + f: ".e+4" + g: "2024-01-01 10:00:00 +23:99" + """.trimIndent() + "\n" + } + + @Test + fun `should not quote a float shape without a digit`() = runTest { + // given + val values = listOf(".", "+.", "._") + for (value in values) { + val events = semanticEvents { + "entry"("key" to "k") { +value } + } + + // when + val rendered = events.renderYaml() + val reparsed = flowOf(rendered).parseYaml() + + // then + rendered sameAs "k: $value\n" + reparsed.mergeAdjacentText() sameAs events + } + } + + @Test + fun `should escape every character YAML does not allow in a document`() = runTest { + // given — C1 controls (mojibake like a Windows-1252 apostrophe read + // as Latin-1) and the U+FFFE / U+FFFF non-characters are outside + // YAML's printable set (YAML 1.2 §5.1): Psych and PyYAML refuse + // the whole document when they appear raw, even in a block scalar; + // a byte order mark must not appear inside a document (§5.2) + val values = listOf( + "It\u0092s" to "\"It\\u0092s\"", + "a\u0080b" to "\"a\\u0080b\"", + "\u009f" to "\"\\u009f\"", + "x\uFFFEy" to "\"x\\ufffey\"", + "x\uFFFF" to "\"x\\uffff\"", + "first\u0092\nsecond" to "\"first\\u0092\\nsecond\"", + "\uFEFFTitle" to "\"\\ufeffTitle\"", + "first\n\uFEFFsecond" to "\"first\\n\\ufeffsecond\"", + "smile 😀" to "smile 😀", + ) + for ((value, expected) in values) { + val events = semanticEvents { + "entry"("key" to "k") { +value } + } + + // when + val rendered = events.renderYaml() + val reparsed = flowOf(rendered).parseYaml() + + // then + rendered sameAs "k: $expected\n" + reparsed.mergeAdjacentText() sameAs events + } + } + + @Test + fun `DIVERGENCE - should write a lone surrogate as the replacement character`() = runTest { + // given — a lone surrogate (a JS string cut mid-pair) has no YAML + // representation: raw it is outside the printable set, and Psych and + // go-yaml refuse its `\u` escape as an invalid code point, so either + // loses the whole document; U+FFFD keeps it loadable, but the + // surrogate does not survive the round-trip + val values = listOf( + "lone \uD800 high" to "lone \uFFFD high", + "lone \uDC00 low" to "lone \uFFFD low", + "\uDC00\uD800" to "\uFFFD\uFFFD", + ) + for ((value, replaced) in values) { + val events = semanticEvents { + "entry"("key" to "k") { +value } + } + + // when + val rendered = events.renderYaml() + val reparsed = flowOf(rendered).parseYaml() + + // then + rendered sameAs "k: \"$replaced\"\n" + reparsed.mergeAdjacentText() sameAs semanticEvents { + "entry"("key" to "k") { +replaced } + } + } + } + + @Test + fun `DIVERGENCE - should write a lone surrogate in a key as the replacement character`() = runTest { + // given + val events = semanticEvents { + "entry"("key" to "a\uD800") { +"v" } + } + + // when + val rendered = events.renderYaml() + + // then + rendered sameAs "\"a\uFFFD\": v\n" + } + @Test fun `should round-trip keys that need quoting`() = runTest { // given diff --git a/markanywhere-yaml/src/commonTest/kotlin/YamlWriterTest.kt b/markanywhere-yaml/src/commonTest/kotlin/YamlWriterTest.kt index 5cd9634..0e482d4 100644 --- a/markanywhere-yaml/src/commonTest/kotlin/YamlWriterTest.kt +++ b/markanywhere-yaml/src/commonTest/kotlin/YamlWriterTest.kt @@ -241,7 +241,7 @@ class YamlWriterTest { dash: "- not an item" padded: " spaced " tabbed: "a\tb" - url: "http://x#y" + url: http://x#y page.section: News """.trimIndent() + "\n" } @@ -317,7 +317,7 @@ class YamlWriterTest { // after its dash; the line must still be terminated val flow = semanticEvents { "item" { "x" { } } - "item" { +"y" } + "item" { +"next" } } // when @@ -326,7 +326,7 @@ class YamlWriterTest { // then yaml sameAs """ - - - y + - next """.trimIndent() + "\n" }