Repository navigation
Dialect
Dialect defines the formatting rules used when parsing CSV data.
A dialect controls the field separator, quote character, and line-ending convention.
pub const Dialect = struct {
separator: u8 = ',',
quote: u8 = '"',
line_ending: LineEnding = .lf,
};| Field | Type | Default | Description |
|---|---|---|---|
separator |
u8 |
, |
Character separating fields |
quote |
u8 |
" |
Character used to delimit quoted fields |
line_ending |
LineEnding |
.lf |
Line-ending convention |
Creating a Dialect without specifying any fields uses the standard
defaults:
const dialect = Dialect{};This is equivalent to:
const dialect = Dialect{
.separator = ',',
.quote = '"',
.line_ending = .lf,
};Any field can be overridden when creating a dialect.
For example, a semicolon-separated format:
const dialect = Dialect{
.separator = ';',
};A tab-separated format:
const dialect = Dialect{
.separator = '\t',
};Multiple options can be configured together:
const dialect = Dialect{
.separator = ';',
.quote = '\'',
.line_ending = .crlf,
};pub fn isSeparator(self: *Dialect, c: u8) boolReturns true when c is the configured field separator.
var dialect = Dialect{};
dialect.isSeparator(','); // true
dialect.isSeparator(';'); // falsepub fn isQuote(self: *Dialect, c: u8) boolReturns true when c is the configured quote character.
var dialect = Dialect{};
dialect.isQuote('"'); // true
dialect.isQuote('\''); // falsepub fn isLineEnding(
self: *Dialect,
input: []const u8,
pos: usize,
) boolChecks whether a configured line ending starts at pos in input.
The bytes checked depend on the configured LineEnding:
| Line ending | Sequence |
|---|---|
.lf |
\n |
.cr |
\r |
.crlf |
\r\n |
For .crlf, both bytes must be present for the method to return true.
Zarko provides common dialect presets in dialects.zig.
See Built-in Dialects for:
exceltsvsemicolon