Skip to content

Split a caller-supplied TokenType::Eof off before the grammar runs - #450

Merged
tobilg merged 2 commits into
tobilg:mainfrom
geoHeil:fix/eof-token-is-end-of-input
Sep 17, 2026
Merged

tobilg merged 2 commits into
tobilg:mainfrom
geoHeil:fix/eof-token-is-end-of-input

Conversation

@geoHeil

@geoHeil geoHeil commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Follow-up to #447. A caller-supplied TokenType::Eof must not change the parse; this makes
that hold by construction.

The invariant

Tokenizer never emits the variant, but Parser::new and friends are public, so a caller can
hand the parser a stream carrying one. It means end of input, so a terminated stream must
parse identically to an unterminated one.
#447 added explicit_eof_token_tests asserting
exactly that, and pinned the one case where it did not hold.

0.11.0 fixed that case by adding && !self.check(TokenType::Eof) to the scan it surfaced in.
That scan was not special: there are 109 places in parser.rs that compare an index
against self.tokens.len() to ask whether the input has run out, and to every one of them a
terminator left in the stream is a token.

The fix

The terminator is split off in the constructors. The tokens before it are the input, and
the grammar never sees it:

fn split_parser_terminator(mut tokens: Vec<ParserToken>) -> (Vec<ParserToken>, Option<TokenType>) {
    let Some(at) = tokens.iter().position(|t| t.token_type == TokenType::Eof) else {
        return (tokens, None);
    };
    let followed_by = tokens.get(at + 1).map(|t| t.token_type);
    tokens.truncate(at);
    (tokens, followed_by)
}

Two streams that differ only by a terminator become the same stream, so the invariant is
structural rather than a property each of those 109 readers has to remember. It also covers
peek, whose span every parse error is built from — without that, two streams could produce
the same error at different positions.

Tokens after a terminator are a stream built wrongly rather than input. A constructor cannot
fail, so the offending token's type is recorded and refused by ensure_complexity_guards, the
one hook all three public entry points — parse, parse_statement,
parse_standalone_data_type — already run, exactly once, before any grammar.

What it fixes

Sweeping all 38 fixture files, 11,311 distinct inputs, parsed with and without an appended
terminator and compared:

revision statement-level differences
0.11.0 as it ships 98
recognising the terminator in is_at_end only 13
this 0

That first number corrects something in this PR's own history, so it is worth being explicit
about. My earlier sweeps used four fixture files — 994 inputs — and reported 1, which is
where the BEGIN framing came from. Over all 38 files it is 98, spanning ALTER TABLE … MODIFY COLUMN … (eleven variants), the BEGIN family, CALL a.b.c(x, y), CREATE STAGE …,
1 div and more. The defect was considerably more widespread on main than either of us had
measured; the small corpus undersampled it.

BEGIN is still the clearest single case: with a terminator it stopped being a Transaction
and became an opaque Command { this: "BEGIN " } — a different AST node, so anything matching
on the AST saw something else.

The middle row is the first revision of this PR, and it is why the mechanism moved. Answering
the question in one reader left the others asking a different one:

what asked the wrong question input effect
parse_standalone_data_type's !is_at_end() trailing check INT, Eof, SELECT, 2 returned Ok(Int), dropping the trailing statement
advance_text testing the stream length SET, Eof, x, =, 1 read the terminator as the variable name → SetStatement with an empty identifier
is_last_expression_token testing the next index SELECT 1 is + 10 more trailing keyword read as an alias unterminated, error terminated; SELECT * FROM t LIMIT 10% went from parsing to failing
peek returning the zero-span terminator 9 of the 13 identical error message, different position

All of them are fixed by the terminator not being there.

Empty and terminator-first streams

Normalizing a stream that begins with a terminator leaves no tokens at all, and the first
revision of this change did not allow for that: parse_error built its span from peek, which
asserts at least one token, so all three entry points panicked with Token list should not be empty where the base returned an error. Found by @tobilg's probes. Three things close it:

  • parse_error takes its span the way end_of_input_error beside it always has —
    tokens.get(current), falling back to last(), then unwrap_or_default().
  • The offending token's span is carried alongside its type, so a placement error points at
    that token. Better than the first revision managed even on non-empty streams, which built the
    span from peek on the truncated remainder: Eof → SELECT → 1 reported line 0, column 0
    and now reports line 1, column 7.
  • parse_statement and parse_standalone_data_type answer end of input on an empty stream
    instead of dispatching on a token that is not there. parse needs no guard: no statements is
    the right answer for no tokens.
stream entry point base after
…tokens… Eof …more… all three Err("Unexpected token: Eof") Err("Unexpected token after end of input: …"), at the offending token
Eof alone parse Err("Unexpected token: Eof") Ok([]) — like parse_sql(""), parse_sql(" ") and an empty token vector
Eof alone parse_statement, parse_standalone_data_type Err Err("Unexpected end of input")
Vec::new() parse_statement, parse_standalone_data_type panic Err("Unexpected end of input")

That last row is not a regression from this branch: an empty token vector has always been
constructible through Parser::new, and those two entry points panicked on it before any of
this. Normalization added a second door to the same assumption by making a terminator-only
stream empty, so the guard closes both.

The comparisons stay

All eight, including the one 0.11.0 added, and the checks in is_at_end, advance and
advance_text. They are unreachable now that no constructor admits an embedded terminator,
and kept for the reason @tobilg gave for keeping the original five: removing a comparison
against this variant is what the first revision of #447 got wrong.

Tests

  • parser.rs, explicit_eof_token_tests:
    test_an_explicit_eof_token_does_not_change_the_parse (36 statements),
    test_tokens_after_an_eof_token_are_refused_by_every_entry_point (all three entry points
    plus the SET case), test_a_trailing_keyword_is_read_the_same_with_and_without_a_terminator
    (all 13 from the sweep), test_a_stream_of_only_an_eof_token_parses_as_empty,
    test_a_raw_unset_clause_stops_at_an_explicit_eof,
    test_unset_property_with_an_explicit_eof_is_not_a_raw_clause.
    test_a_stream_beginning_with_an_eof_token_errors_rather_than_panicking covers leading,
    duplicate and terminator-only streams across all three entry points and all three public
    constructors — new, with_config, with_source share the normalization — including an
    assertion that the placement error carries the offending token's position.
  • tests/data_type_api.rs:
    parse_standalone_data_type_treats_an_eof_token_as_end_of_input — the token-stream entry
    point, which parse_data_type cannot reach through a string — and
    parse_standalone_data_type_errors_on_an_empty_or_eof_first_stream. 11 → 13 tests.

Every case was checked against the revision it fails on, rather than assumed.

Verification

Against 47342ba. cargo fmt --all -- --check clean.

make test-rust-verify rerun after the adjustment — exit 0, every step green: lib
1236 passed / 0 failed / 2 ignored, generic identity 977/977, dialect identity 4086/4086,
transpilation 6058/6058 (4 known failures), transpile generic 154/154, parser 32/32,
pretty-print 23/23, custom dialect 276/276 + 347/347, ClickHouse parser 9417 parsed /
0 failed, ClickHouse coverage 100% on every group, FFI 75 passed.

Also cargo test --test error_handling 62 passed and cargo test --test data_type_api 13
passed.


🤖 Generated with Claude Code

@geoHeil

geoHeil commented Sep 16, 2026

Copy link
Copy Markdown
Contributor Author

make test-rust-verify has finished: exit 0, every step green, so the two steps the
description left open are confirmed.

step result
ClickHouse coverage (release) 100% on every group
FFI 75 passed

Full run, against 058cf91: lib 1235 passed / 0 failed / 2 ignored, generic identity
977/977, dialect identity 4086/4086, transpilation 6058/6058 (4 known failures), transpile
generic 154/154, parser 32/32, pretty-print 23/23, custom dialect 276/276 + 347/347,
ClickHouse parser 9417 parsed / 0 failed, ClickHouse coverage 100%, FFI 75 passed.
cargo fmt --all -- --check clean.


Two things I looked at while chasing this and am reporting rather than pushing, since I
found no live defect and neither is worth a large speculative diff:

peek / peek_text at end of input. #447 flagged these as the same shape of hazard as
advance was — they return the last token rather than reporting the end, so a
peek_text().eq_ignore_ascii_case("…") past the end compares against a token that was
already consumed, which misparses rather than hangs. I instrumented peek to count
past-end fallbacks and swept truncated prefixes of the fixture corpus looking for prefixes
that parsed successfully while touching a stale token. The only ones that came back were
bare CREATE, which is the deliberate Command passthrough rather than a misparse. So it
stays a latent hazard with no instance I can point at, and making 270 call sites fallible
on that basis did not seem like a good trade. Happy to do it if you would rather have the
type-level guarantee.

parse_sql("CREATE") returns Ok(Command { this: "CREATE " }). A bare CREATE is
accepted, with a trailing space in the command text. Whether that should be a parse error
depends on how much the Command passthrough is meant to swallow, which is your call
rather than mine — flagging it because it turned up in the same sweep.

skip() is unchanged from #447: 769 sites that do nothing at end of input, and the
termination oracle still finds no loop that spins on one.

@geoHeil
geoHeil force-pushed the fix/eof-token-is-end-of-input branch from a61e6db to f92d2f1 Compare September 16, 2026 10:59
@geoHeil

geoHeil commented Sep 16, 2026

Copy link
Copy Markdown
Contributor Author

Rebased onto 47342ba — the branch had been cut from 058cf91, before the CI fix — and
CI is green on the rebased head: run 35087914548.

job
quality pass, 2m35s
python-fast-check pass, 2m47s
rust-test pass, 41m6s
go-sdk pass, 14m34s
sdk-build pass, 25m44s

No failed or cancelled checks. Worth noting for the record that rust-test covers five
steps make test-rust-verify does not, so these are the ones the description's local table
could not speak for, and all five pass: deep-nesting regressions, the standalone example
check, the benches check, the WASM unit tests and FFI tests plus the FFI release build, and
both feature-gate checks. The deep-nesting one was the one I actually wanted to see, since
is_at_end is about as hot as a function in this parser gets.

@tobilg

tobilg commented Sep 17, 2026

Copy link
Copy Markdown
Owner

Thanks for the follow-up and the thorough verification. The shared is_at_end() change fixes the BEGIN AST discrepancy, and I confirmed that CI is green. Locally, the 1,235 unit tests, 20 deep-nesting tests, and 11 data-type API tests also pass.

I’d recommend a small follow-up before merging. Additional token-stream probes identified one regression and two remaining gaps in the intended EOF guarantees:

  1. Standalone type parsing can silently ignore tokens after EOF.
    Parser::parse_standalone_data_type() rejects INT → Eof → SELECT → 2 on the base branch, but returns Ok(Int) on this PR. Its end-of-input check now stops at EOF, while the new trailing-token validation only runs in parse(). Could EOF placement validation also cover this entry point?

  2. A parser can consume EOF before the new guard runs.
    SET → Eof → x → = → 1 still succeeds, producing a SetStatement with an empty variable name. advance_text() checks physical token exhaustion, so it consumes the marker and the parser reaches the end before the guard can detect it. This behavior predates the PR, but means the new guard does not cover every malformed stream. Validating EOF placement before grammar dispatch, and making consuming helpers respect logical EOF, would close this gap.

  3. One existing fixture still changes behavior when EOF is appended.
    SELECT 1 is parses successfully with is as an alias, but fails with an appended EOF. This also predates the PR: is_last_expression_token() recognizes physical exhaustion but not an EOF token in lookahead. My sweep of 1,160 distinct fixture inputs found two statement-level differences before this change and one afterward—BEGIN is fixed, while this case remains.

Could you please add these cases to the existing EOF and data-type test files, covering both parse() and parse_statement() where applicable, and rerun make test-rust-verify after the adjustments?

These look addressable within the current approach; I don’t think they require a new parser API or a wholesale rewrite of peek().

@geoHeil
geoHeil force-pushed the fix/eof-token-is-end-of-input branch from f92d2f1 to 5d51206 Compare September 17, 2026 09:20
@geoHeil geoHeil changed the title Recognise a caller-supplied TokenType::Eof in is_at_end Split a caller-supplied TokenType::Eof off before the grammar runs Sep 17, 2026
@geoHeil

geoHeil commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

Thanks — all three are fixed, and the first one was mine. Rather than fix them where they
surfaced, I moved the mechanism, because the three of them are the same bug three times.

What the three had in common

Your probes and my own follow-up sweep say the same thing: making is_at_end report the end
at the terminator answered the question in one reader, while every other reader went on
asking a different one. There are 109 places in parser.rs that compare an index against
self.tokens.len() to ask whether the input has run out, and to all of them a terminator left
in the stream is a token.

So the terminator is now split off in the constructors — the tokens before it are the
input, and the grammar never sees it. Two streams differing only by a terminator become the
same stream
, which makes the invariant structural rather than something each reader has to
remember. That also covers a fourth reader neither of us probed directly: peek, whose span
every parse error is built from.

Tokens after a terminator are a stream built wrongly rather than input. A constructor cannot
fail, so the offending token's type is recorded and refused by ensure_complexity_guards —
the one hook all three public entry points already run, exactly once, before any grammar.
That is the "validate EOF placement before grammar dispatch" you suggested, and it is what
covers parse_standalone_data_type without a second copy of the check.

Your three points

1. parse_standalone_data_type silently ignoring tokens after EOF. Confirmed, and it was
a regression I introduced: its trailing-token check is !self.is_at_end(), which my change
made true at the terminator, while the new validation lived in parse alone. INT, Eof,
SELECT, 2 now refuses again, and INT, Eof still parses as INT.

2. A parser consuming EOF before the guard runs. Confirmed: advance_text asked whether
the stream had run out, so SET, Eof, x, =, 1 read the terminator as the variable
name and produced a SetStatement with an empty identifier. It is refused now, from both
parse and parse_statement. advance and advance_text also ask is_at_end rather than
the stream length, so a terminator could not be consumed as a word even if one reached them.

3. SELECT 1 is. Confirmed, and there were more. is_last_expression_token asks whether
the next index is past the stream to decide whether a trailing keyword is an alias. Sweeping
all 38 fixture files rather than the four I used the first time — 11,311 distinct inputs —
found 13 differences under the previous revision, not one:

revision statement-level differences (11,311 inputs)
0.11.0 as it ships 98
previous revision of this PR 13
this revision 0

The 98 corrects my own earlier figure: I had swept four fixture files (994 inputs) and reported
1, which is where the BEGIN-only framing came from. Over all 38 files the base differs on 98,
including eleven ALTER TABLE … MODIFY COLUMN … variants, the BEGIN family, CALL a.b.c(x, y), CREATE STAGE … and 1 div. The small corpus undersampled it, and your probes finding
things it missed is what prompted me to widen it.

Nine of the thirteen were the same error message at a different position, because peek
returned the zero-span terminator instead of the last real token. The other four were real,
and one changed from parsing to failing: SELECT * FROM t LIMIT 10%. All thirteen are in the
new test alongside SELECT 1 is.

Tests

In the two files you named, covering parse, parse_statement and
parse_standalone_data_type:

  • parser.rs, explicit_eof_token_tests:
    test_tokens_after_an_eof_token_are_refused_by_every_entry_point (all three entry points,
    plus the SET case) and
    test_a_trailing_keyword_is_read_the_same_with_and_without_a_terminator (all thirteen).
    The previous revision's parse-only test is gone, subsumed by the first of those.
  • tests/data_type_api.rs:
    parse_standalone_data_type_treats_an_eof_token_as_end_of_input — the token-stream entry
    point, which parse_data_type cannot reach through a string. 11 tests there → 12.

Each case was checked against the previous revision to confirm it actually fails there, rather
than assumed.

Verification

Against 47342ba. cargo fmt --all -- --check clean; cargo clippy -p polyglot-sql --lib --tests adds no warnings this branch is responsible for.

make test-rust-verify rerun after the adjustments — exit 0, every step green:

step result
Lib unit tests 1236 passed, 0 failed, 2 ignored
Generic identity 977/977
Dialect identity 4086/4086
Transpilation 6058/6058 (4 known failures)
Transpile generic 154/154
Parser 32/32
Pretty-print (release) 23/23
Custom dialect 276/276 identity, 347/347 transpilation
ClickHouse parser (release) 9417 parsed, 0 failed (9474 files, 57 skipped)
ClickHouse coverage (release) 100% on every group
FFI 75 passed

Also cargo test --test error_handling 62 passed, and cargo test --test data_type_api 12
passed (was 11).

is_at_end, advance and advance_text keep their checks even though no constructor now
admits an embedded terminator. Unreachable, and kept for the same reason you asked me to keep
the original five comparisons.

@tobilg

tobilg commented Sep 17, 2026

Copy link
Copy Markdown
Owner

Thanks for the update and the expanded regression coverage. I rechecked commit 5d512064 and confirmed that all three earlier findings are addressed. Normalizing the token stream in the constructors also makes the EOF behavior more consistent across grammar paths.

I’d recommend one focused follow-up before merging: EOF-first streams can now panic instead of returning a parse error.

I reproduced these cases:

Token stream Entry point Result
Eof → SELECT → 1 parse(), parse_statement(), parse_standalone_data_type() Panic
Eof → Eof All three entry points Panic
Eof parse_statement(), parse_standalone_data_type() Panic
Eof parse() Empty result, as intended

The panic is Token list should not be empty, and reproduces through Parser::new, with_config, and with_source. The base revision returns errors for the cases that now panic.

The cause appears to be the interaction between normalization and error construction: truncating at a leading EOF leaves an empty vector, while parse_error() calls peek(), which assumes at least one token remains. EOF-only statement/type parsing reaches the same assumption.

Could we please:

  • Make error construction safe for an empty token vector.
  • Consider retaining the offending post-EOF token’s span alongside its type, so placement errors point to that token.
  • Extend the existing EOF and data-type tests with leading EOF, duplicate EOF, and EOF-only cases across the three entry points and public constructors.
  • Rerun make test-rust-verify after the adjustment.

For context, all 1,330 locally executed tests passed, with two additional tests ignored, and formatting passed. An additional 44,154 appended-EOF comparisons also passed; the missing coverage is specifically the empty-prefix cases above.

The constructor-normalization approach still looks sound. This should be addressable without changing the overall design.

@geoHeil

geoHeil commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

Thanks — reproduced all four rows, and it was mine. Fixed, along with one adjacent panic that
turned out not to be.

The cause

Exactly as you diagnosed. Normalizing a stream that begins with a terminator truncates to an
empty vector, and parse_error built its span from peek, which asserts at least one
token. end_of_input_error next to it had always allowed for an empty stream; parse_error
had never needed to, because before normalization no constructor could produce one from a
non-empty input.

Three changes

Error construction is safe for an empty vector. parse_error now takes its span the same
way end_of_input_error does — tokens.get(current), falling back to last(), then
unwrap_or_default().

The offending token's span is carried alongside its type, as you suggested, so a placement
error points at the token rather than at whatever survived normalization. It is a strict
improvement on what the previous revision did even when the stream was not empty, because
that built the span from peek on the truncated remainder:

Eof → SELECT → 1     before: line 0, column 0
                      now:    line 1, column 7   ← the SELECT

parse_statement and parse_standalone_data_type answer end of input on an empty stream
rather than dispatching on a token that is not there. parse needs no such guard — no
statements is the right answer for no tokens, and that row was already correct.

One of these was not my regression

Worth separating, because it changes what the fix is worth. Parser::new(Vec::new()) followed
by parse_statement or parse_standalone_data_type panicked with the same message on the
base revision
, before this branch existed — an empty token vector has always been
constructible. What normalization did was make that reachable through a second door, by turning
a stream that is only a terminator into an empty one. Both doors are closed now.

token stream entry point base previous revision now
Eof → SELECT → 1 all three error panic error, pointing at SELECT
Eof → Eof all three error panic error, naming the second terminator
Eof parse error Ok([]) Ok([]) — as intended
Eof parse_statement, parse_standalone_data_type error panic Unexpected end of input
Vec::new() parse Ok([]) Ok([]) Ok([])
Vec::new() parse_statement, parse_standalone_data_type panic panic Unexpected end of input

Reproduced through Parser::new, with_config and with_source, as you found — they share
the normalization, so all three are covered by the tests rather than only the first.

Tests

  • parser.rs, explicit_eof_token_tests:
    test_a_stream_beginning_with_an_eof_token_errors_rather_than_panicking — leading EOF,
    duplicate EOF and EOF-only, across all three entry points and all three public constructors,
    including an assertion that the placement error carries the offending token's position.
  • tests/data_type_api.rs:
    parse_standalone_data_type_errors_on_an_empty_or_eof_first_stream — the same shapes for
    that entry point, plus Vec::new(). 12 tests → 13.

Verification

Against 47342ba. cargo fmt --all -- --check clean.

make test-rust-verify rerun after the adjustment — exit 0, every step green: lib
1236 passed / 0 failed / 2 ignored, generic identity 977/977, dialect identity 4086/4086,
transpilation 6058/6058 (4 known failures), transpile generic 154/154, parser 32/32,
pretty-print 23/23, custom dialect 276/276 + 347/347, ClickHouse parser 9417 parsed /
0 failed, ClickHouse coverage 100% on every group, FFI 75 passed.

Also cargo test --test error_handling 62 passed and cargo test --test data_type_api 13
passed.

The appended-terminator sweep is unchanged at 0 of 11,311 — these fixes touch error
construction and the empty case, not the normalization itself, so your 44,154 appended-EOF
comparisons should be unaffected too.

`Tokenizer` never emits the variant, but `Parser::new` and friends are public, so a
caller can hand the parser a stream carrying one. It means end of input, and such a
stream must parse identically to one without it.

Recognising it in `is_at_end` answered that in one reader while over a hundred others
went on asking whether the index had reached `self.tokens.len()`. So the terminator
is split off in the constructors instead: the tokens before it are the input, and the
grammar never sees it. Two streams that differ only by a terminator become the same
stream, which makes the invariant structural rather than a property every reader has
to remember -- including `peek`, whose span every parse error is built from.

Over all 38 fixture files, 11,311 distinct inputs parsed with and without an appended
terminator: 98 differences on the base revision, 0 here. `BEGIN` is the clearest of
them -- with a terminator it stopped being a `Transaction` and became an opaque
`Command("BEGIN ")`, a different AST node.

Tokens *after* a terminator are a stream built wrongly rather than input. A
constructor cannot fail, so the offending token's type and span are recorded and
refused by `ensure_complexity_guards`, the one hook all three public entry points
already run, exactly once, before any grammar. The span is kept so the error points
at that token rather than at what is left of the stream after normalizing it.

A stream that *begins* with a terminator normalizes to no tokens at all, and two
things assumed otherwise:

- `parse_error` built its span from `peek`, which asserts a token exists, so every
  entry point panicked with `Token list should not be empty` where the base returned
  an error. It now takes its span the way `end_of_input_error` beside it always has.
- `parse_statement` and `parse_standalone_data_type` owe a value an empty stream
  cannot provide, and dispatched into the grammar regardless. They answer end of
  input instead. This one predates the branch -- `Parser::new(Vec::new())` panicked
  identically before it -- and normalization only added a second door to it.

`parse` needs no such guard: no statements is the right answer for no tokens, and a
terminator-only stream now parses like `parse_sql("")` rather than erroring.

`is_at_end`, `advance` and `advance_text` keep their checks. They are unreachable now
that no constructor admits an embedded terminator, and kept for the reason the five
original comparisons were kept: this is the safe side to err on.
@geoHeil
geoHeil force-pushed the fix/eof-token-is-end-of-input branch from 5d51206 to d3e5ce4 Compare September 17, 2026 10:36
Guard the shared primary-expression path when explicit EOF normalization leaves an empty token stream, so public fragment helpers return no expression instead of panicking.

Cover all five affected helpers and all three constructors in the existing EOF tests, including repeated calls, malformed EOF placement, and valid-input controls.

Validation: make test-rust-verify; focused EOF, data-type API, and error-handling tests; cargo fmt --all -- --check; 44,154 corpus comparisons with no differences or panics.
@tobilg
tobilg merged commit ce62df4 into tobilg:main Sep 17, 2026
19 checks passed
@tobilg

tobilg commented Sep 17, 2026

Copy link
Copy Markdown
Owner

Thanks, merged with a few updates

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants