Skip to content

feat: add Vertica dialect - #467

Closed
ArjixWasTaken wants to merge 1 commit into
tobilg:mainfrom
msensis-com:feat/vertica-dialect
Closed

ArjixWasTaken wants to merge 1 commit into
tobilg:mainfrom
msensis-com:feat/vertica-dialect

Conversation

@ArjixWasTaken

@ArjixWasTaken ArjixWasTaken commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Add Vertica dialect

Adds vertica as a supported dialect across the Rust crate, FFI, Python, WASM, the TypeScript SDK and the playground. Behavior follows the vertica-sqlglot-dialect reference, and the fixture expectations are taken from its tests.

This covers core query SQL only. Vertica-only DDL and query clauses are listed under "Not included" below.

Vertica semantics

  • Types: every Vertica integer is 64-bit and every float is a double.

    • INT/INTEGER/SMALLINT/TINYINT/INT8 become BIGINT. REAL/FLOAT/FLOAT8 become DOUBLE PRECISION.
    • Adds LONG VARCHAR[(n)] and LONG VARBINARY[(n)]. BINARY VARYING(n) becomes VARBINARY(n), and TIMETZ is kept.
    • Interval types accept seconds precision: INTERVAL SECOND(3), INTERVAL DAY TO SECOND(5).
  • Operators: MINUS is read as EXCEPT. Adds // (integer division), postfix ! and prefix !! (factorial), and prefix @ (absolute value). |/ and ||/ become SQRT/CBRT.

  • Functions:

    • NVL, NVL2, DECODE and ZEROIFNULL stay as written when the output is Vertica.
    • TIMESTAMPDIFF becomes DATEDIFF. TIMESTAMPADD and DATEDIFF units are upper-cased.
    • SYSDATE becomes GETDATE(). DAYOFWEEK_ISO is supported.
    • LISTAGG(x USING PARAMETERS separator = …, max_length = …, on_overflow = …) is parsed and generated.
  • Converting from other dialects into Vertica:

    From To
    IF/IFF CASE
    DATEADD TIMESTAMPADD
    CHARINDEX INSTR
    APPROX_COUNT_DISTINCT APPROXIMATE_COUNT_DISTINCT
    GROUP_CONCAT/STRING_AGG LISTAGG
    TRY_CAST/SAFE_CAST CAST
    COUNT_IF SUM(CASE …)
    QUALIFY and semi/anti joins subqueries / EXISTS
  • Converting Vertica into other dialects:

    • GETDATE()/SYSDATE are start-of-statement times, so for PostgreSQL they become CAST(STATEMENT_TIMESTAMP() AS TIMESTAMP). GETUTCDATE() gets its UTC equivalent.
    • LISTAGG … WITHIN GROUP becomes STRING_AGG(x, sep ORDER BY …) (Postgres family, T-SQL, BigQuery) or GROUP_CONCAT (MySQL family, SQLite). The default ',' separator is written out explicitly.
    • ZEROIFNULL becomes COALESCE(x, 0).
    • TIMESTAMPADD becomes each target's date-add form, or interval arithmetic for Postgres.
    • APPROXIMATE_COUNT_DISTINCT becomes the portable approximate-distinct node.
  • NULL sort order: where Vertica puts NULLs depends on the column type (NULLS AUTO). So no default is assumed when reading Vertica, and when writing Vertica the source's NULL ordering is always spelled out.

Changes to shared code

  • ListAggFunc has a new optional max_length field. It's skipped from JSON when unset. Other dialects report it as unsupported instead of silently dropping it.
  • New tokenizer setting double_slash_int_div, off by default and enabled only for Vertica.
  • The parser has a Vertica type-aliasing hook and Vertica-only operator parsing.
  • Vertica is added to the existing target lists for native NVL2/DECODE, INSTR generation, IF→CASE, the 2-argument DATEDIFF rewrite and GETDATE/SYSDATE handling.
  • Vertica-specific source rewrites live in the new dialects/normalization/vertica.rs.
  • For set-operation output types, Vertica uses the Standard rules with a NUMERIC precision cap of 1024.

Registration and docs

  • Adds the dialect-vertica feature to both Cargo.toml files. The dialect is registered in the FFI, Python, WASM, SDK and playground lists.
  • The dialect-count assertions in the FFI and WASM tests go from 34 to 35.
  • Adds a transpile,dialect-vertica feature-gate check to the Makefile.
  • Updates the READMEs, docs/set-operation-types.md and CHANGELOG.md (Unreleased).

Tests

  • New tests/custom_fixtures/vertica/ fixtures: identity.json, types.json and transpilation.json.

  • Results:

    Suite Result
    Vertica identity fixtures 58/58
    Vertica transpilation fixtures 63/63
    Library unit tests 1280 passed
    SQLGlot identity 977/977
    SQLGlot dialect identity 4086/4086
    SQLGlot transpilation 6048/6048
    DataFusion custom fixtures unchanged, all passing
  • FFI, WASM and Python binding tests pass. cargo fmt --check and scripts/check_project_consistency.py pass.

Not included

  • Vertica-only DDL and query clauses: projections, SEGMENTED BY/KSAFE, TIMESERIES, MATCH, INTERPOLATE, and LIMIT n OVER (…).
  • The legacy INTERVALYM keyword and INTERVAL(p) '…' literals. The latter is still parsed incorrectly without an error.
  • Mapping SERIAL and JSON types to Vertica.

Existing issues found but not fixed here

  • MySQL TIMESTAMPADD(DAY, n, ts) → Postgres drops n.
  • DuckDB can't parse its own // operator.
  • Snowflake LISTAGG → Postgres isn't converted to STRING_AGG.

🤖 Generated with Claude Code

@ArjixWasTaken

Copy link
Copy Markdown
Contributor Author

This was based on the vertica plugin for sqlglot, and it was compared against it.

@tobilg

tobilg commented Sep 22, 2026

Copy link
Copy Markdown
Owner

Thanks for the PR! I'd like to ask you to at least fix the known issues, I will run other verifications on my side...

@tobilg

tobilg commented Sep 22, 2026

Copy link
Copy Markdown
Owner

My recommendation is NO-GO for the current revision / request changes.

I reviewed commit 2102318b6bcdeaa9bcba6d58dd677b75b527cc5c. The dialect registration and binding updates are consistent, and local validation passed. Targeted checks did identify several cases where accepted SQL changes meaning during transpilation, including with TranspileOptions::strict().

The main correctness concerns are:

  1. Aggregate filters are lost when converting approximate distinct counts to Vertica.
    APPROX_COUNT_DISTINCT(x) FILTER (WHERE keep) becomes APPROXIMATE_COUNT_DISTINCT(x), including rows the source excludes. The transformation copies the argument into a generic function but drops the aggregate modifiers. Please preserve these modifiers or reject unsupported combinations.
    Relevant transformation

  2. LISTAGG length and overflow options are silently discarded for some targets.
    For example, Vertica → PostgreSQL converts:

    LISTAGG(x USING PARAMETERS max_length=3, on_overflow='TRUNCATE')

    into:

    STRING_AGG(x, ',')

    The explicit size limit and truncation behavior disappear. MySQL conversion has the same problem. Please check these options before replacing the aggregate node, preserving their semantics where possible and reporting unsupported conversions otherwise.
    Relevant normalization

  3. TRY_CAST and SAFE_CAST become ordinary, potentially throwing casts.
    Snowflake TRY_CAST(value AS INTEGER) becomes Vertica CAST(value AS BIGINT). Invalid string values therefore raise an error instead of producing NULL. Vertica’s ::! operator may cover some cases, but its documented limitations around constants need consideration. Where equivalent behavior cannot be expressed, strict mode should reject the conversion.
    Relevant transformation · Vertica cast documentation

  4. The PostgreSQL DATEDIFF conversion does not preserve Vertica’s boundary-counting behavior.
    From 2026-01-01 23:59:00 to 2026-01-02 00:01:00, Vertica’s day difference is 1. The generated elapsed-seconds calculation produces 0. This originates in existing shared lowering, but the new Vertica fixture now accepts that behavior. Please add boundary-crossing regressions and correct the source-specific conversion.
    New fixture · Vertica DATEDIFF documentation

  5. The null-ordering exemption also affects contexts with defined defaults.
    Vertica WITHIN GROUP (ORDER BY y DESC) defaults to NULLS FIRST. Converting an ordered LISTAGG to DuckDB without making that explicit changed the result from A,B to B,A for rows ('A', NULL) and ('B', 1). Please distinguish top-level ordering from aggregate and analytic ordering.
    Relevant normalization · Vertica ordering documentation

Additional reproducible issues:

  • Sized binary types: CAST(x AS LONG VARBINARY(1000)) remains unchanged when targeting PostgreSQL because the type is stored as an opaque custom spelling. It needs a structured representation and target-specific handling. Parser location
  • Parameter expressions: LISTAGG(x USING PARAMETERS max_length=1024*2) fails at *, although Vertica permits integer expressions here. The parameter parser currently consumes only a primary expression. Parser location
  • Nested factorials: SELECT !! !! 3 generates SELECT 3!!, which Polyglot’s Vertica parser then rejects. Parenthesizing the nested operand should preserve round-tripping. Generator location
  • Acknowledged interval limitation: INTERVAL(3) '1.2345 SECOND' becomes INTERVAL (3) + INTERVAL '1.2345 SECOND', including in strict mode. Supporting this syntax can be deferred, but an explicit rejection would provide a safe boundary.

I also compared the reference Vertica plugin 0.2.7, using its supported SQLGlot 30.13.0 dependency. It rejects the unsupported PostgreSQL LISTAGG parameters, while sharing the safe-cast and date-difference issues described above. Reference compatibility is useful, but these cases also need verification against dialect semantics.

Local validation passed: 1,280 Rust unit tests, 744 custom dialect fixture cases—including all 121 Vertica cases—formatting, project consistency, and a minimal Vertica/transpile feature build. These checks did not include the full make test-rust-verify or execution against a live Vertica server.

Before merging, I suggest addressing the findings above, adding regressions to the existing test infrastructure, and completing make test-rust-verify with green CI.

@tobilg

tobilg commented Sep 22, 2026

Copy link
Copy Markdown
Owner

Following a broader comparison with the official Vertica 26.2 SQL reference, my recommendation remains NO-GO / request changes for commit 2102318b6bcdeaa9bcba6d58dd677b75b527cc5c.

This supplements the earlier review. The additional checks found both missing dialect coverage and further cases where accepted SQL changes meaning, including during native Vertica-to-Vertica transpilation.

The additional correctness findings are:

  1. FOR UPDATE is silently removed, including in strict mode.

    -- Vertica input
    SELECT id FROM t FOR UPDATE
    
    -- Vertica output
    SELECT id FROM t

    FOR UPDATE OF t is also removed. Vertica supports these clauses, so removing them changes locking behavior. The dialect currently sets locking_reads_supported: false.

    Please preserve the documented locking forms and handle unsupported variants separately.
    Dialect configuration · Vertica SELECT reference

  2. Array indexing is not adjusted for PostgreSQL or DuckDB.

    -- Vertica input: returns 10
    SELECT (ARRAY[10, 20])[0]
    
    -- Generated DuckDB SQL: returns NULL
    SELECT ([10, 20])[0]

    Vertica arrays are zero-based. I executed the generated DuckDB expression and confirmed the different result. Column access such as a[0] has the same omission. Vertica’s exclusive upper bound for array slices also needs explicit consideration.

    Please preserve indexing and slicing semantics across targets, including nested arrays and non-literal indices.
    Vertica ARRAY reference

  3. Collection casts are misparsed during native round-tripping.

    -- Input
    SELECT ARRAY['1', '2']::ARRAY[INT]
    
    -- Output
    SELECT CAST(ARRAY['1', '2'] AS ARRAY)[INT]

    ::SET[INT] similarly becomes CAST(... AS SET)[INT]. The element type is interpreted as a subscript. A ROW(...) column declaration is also regenerated as STRUCT(...), rather than Vertica’s documented ROW syntax.

    Please represent collection types, element types, bounds, and size limits structurally and preserve their native syntax.
    ARRAY reference · SET reference · ROW reference

  4. Binary literals change value or type across dialects.

    Vertica’s B'101100' represents binary data containing byte 0x2c. Polyglot emits the same spelling for DuckDB, where it produces the VARCHAR value b101100. PostgreSQL interprets B'...' and X'...' as bit-string constants.

    Please preserve the binary value and type through target-specific generation.
    Vertica binary literals · PostgreSQL bit-string literals

  5. COPY parser arguments lose their required call syntax.

    -- Input
    COPY t FROM '/tmp/data.json'
    PARSER FJSONPARSER(flatten_maps=true)
    
    -- Output
    COPY t FROM '/tmp/data.json'
    PARSER FJSONPARSER flatten_maps = TRUE

    This conversion succeeds in strict mode, but the documented parser invocation uses parentheses around its arguments.

    Please preserve the parser call and its parameters as a structured expression.
    Vertica COPY reference

The broader coverage audit also identified these gaps:

Area Observed behavior
Native safe casts value::!INT fails to parse.
Collection declarations Typed arrays, bounded arrays, arrays with binary-size limits, and SET[INT] fail to parse.
Collection constructors Basic arrays and named ROW fields work; SET[...] and outer ROW field aliases fail.
Function parameters APPROXIMATE_PERCENTILE(... USING PARAMETERS ...) and parameterized EXPLODE fail. Parameter handling is currently specialized to LISTAGG.
Collection expansion Documented EXPLODE(a) OVER() is rejected as unsupported in Vertica strict mode; OVER(PARTITION BEST) also fails.
Null ordering NULLS AUTO fails in aggregate and analytic ordering.
Historical queries AT EPOCH LATEST, numbered epochs, and AT TIME query prefixes fail.
Specialized query clauses Partitioned LIMIT, TIMESERIES, MATCH, and INTERPOLATE fail. These are already identified as exclusions in the PR.
Physical table definitions Tested segmentation and column-encoding clauses fail.
Projection and flex-table statements Tested statements survive as raw text, without structured AST coverage, and pass unchanged to foreign targets in strict mode.
Loading and export Tested COPY FROM LOCAL and EXPORT TO PARQUET forms fail.
Hints A native SELECT /*+LABEL(...)*/ loses its hint.
Function translation Some native function names pass through without foreign-target translation, such as NULLIFZERO to PostgreSQL.

These distinctions matter for the support contract: accepting a statement, preserving its text, exposing its structure for analysis, and translating it correctly are separate capabilities.

I compared 53 reference-driven examples across native Vertica, PostgreSQL, and DuckDB generation: 318 Polyglot checks, covering default and strict modes, and 159 reference-plugin comparisons. The reference plugin handles many of the missing forms above, including collection types, array-index adjustment, historical queries, parameterized functions, native locking, and specialized query clauses. It also has gaps, so its output still needs verification against the SQL reference.

The earlier correctness findings remain applicable, including aggregate-filter loss, discarded LISTAGG options, safe-cast substitution, date-difference and null-ordering changes, and interval corruption.

Before merging, I suggest:

  • Correcting the silent semantic changes identified in both reviews.
  • Adding an explicit coverage matrix for native parsing, structured AST support, generation, and foreign-target translation.
  • Adding regressions for the documented forms above to the existing test infrastructure.
  • Running make test-rust-verify after the fixes and completing CI successfully.

The additional checks used actual Polyglot and reference-plugin execution, with selected generated expressions executed in DuckDB.

@tobilg

tobilg commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

This PR is superceeded by #468

@tobilg tobilg closed this Sep 22, 2026
tobilg added a commit that referenced this pull request Sep 24, 2026
* feat: add Vertica dialect

* Fix Vertica semantics and dialect coverage (#467)

* Preserve bare Vertica KSAFE through JSON (#467)

* Verify Vertica bindings and document coverage (#467)

* Fixes

* Cleanup

* Cleanup

---------

Co-authored-by: ArjixWasTaken <53124886+ArjixWasTaken@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants