Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -191,6 +191,13 @@ if result.is_valid: # no Error-level diagnostics
frontmatter cannot be read — the `frontmatter-invalid-yaml` rule, which `validate_document`,
given a parsed document, cannot report.

The public API is what the `legaldown` and `legaldown.validator` packages export (their
`__all__`). Changes to it are listed in the notes of each
[GitHub release](https://github.com/ForLegalAI/legaldown-validator/releases); before 1.0 a
minor release may change it, a patch release does not. Other modules are internal and may
change in any release; constants still only there are to be made public
([#34](https://github.com/ForLegalAI/legaldown-validator/issues/34)).

### Working with the result

Validating a document builds the indices the checks need — section numbers, resolved definitions,
Expand All @@ -206,6 +213,19 @@ so a renderer or a UI can reuse the work instead of re-deriving it:
| `sections`, `section_lookup` | Numbered section index; resolves `{{ref:}}` targets. Numbers count from the shallowest heading level, and a level a heading skips counts as 1 (`#`, `###`, `##` → 1, 1.1.1, 1.2), so no two sections share a number except alternatives and what they contain (§15.8) |
| `definition_lookup`, `party_lookup`, `side_lookup`, `attachment_lookup` | Resolved display text |
| `inline_dates`, `inline_money`, `inline_durations`, `inline_fields`, `inline_placeholders` | Field-spec values found in the body |
| `is_template` | Whether the document is a template (§15.1): it declares `questions`, carries a condition, or holds a `{{choose:}}` |
| `placed_markers` | The markers in body text that apply (§5.7, §15.3), in document order: `PlacedMarker(section, block, fragment, offset, source, identifier, condition, field, item, include_only, line)` — in fragment `fragment` of `block_fragments(block)`, at `offset`, which is the block's `field` (`text`, or `suffix` after a lifted `{{ref:}}`/`{{term:}}`); `item` is the list item it marks, counted in pre-order over all the list's items, nested and empty ones included, as `list_fragments` counts them; `identifier` is `""` where it does not apply (an include-only paragraph, §12.2). Identifiers and conditions are as written: check `is_valid` before relying on them |

A renderer builds from these decisions rather than re-deriving them, with the helpers the
validator reads the document with:

| Helper | What it gives |
|---|---|
| `lex(text)` → `Lexed` | The directives in inline text (§11.4), and a `view` of it with comments and code spans blanked; `is_escaped(text, offset)` |
| `block_fragments(block)`, `list_fragments(block)`, `list_items(block)` | The texts of a block that hold directives and markers, in the order `PlacedMarker.fragment` counts them (`Fragment(text, anchor)`); the same for a list, with the items each is in (`ListFragment(text, anchor, items)`), numbered in pre-order: an item before the items nested in it; a list's items, as `ListItem`s |
| `is_template(document)`, `is_drafting_note(block)` | The template decision without validating (§15.1); whether a quote block is a drafting note (§15.6) |
| `legaldown.validator`: `parse_condition` → `Condition`, `condition_problem`, `exclusive`, `Presence`, `ALWAYS` | Conditions (§15.3, §15.4): parse one, tell why one is invalid, tell whether two units (each the set of conditions it appears under, `Presence`) can never appear together, given the document's `questions` |
| `legaldown.validator`: `is_valid_iso_date`, `is_valid_money_amount`, `is_positive_numeric`, `IDENTIFIER_RE`, `KNOWN_CURRENCIES` | Value checks (§3.10, §10) |

### Reading and editing the document model

Expand Down
25 changes: 24 additions & 1 deletion src/legaldown/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -33,16 +33,23 @@
DELIMITER_PAIRS,
DefinitionAnchor,
DefinitionRef,
Fragment,
ListFragment,
block_fragments,
collect_definitions,
definition_lookup,
find_definition_anchors,
id_term,
list_fragments,
)
from .directives import (
DIRECTIVE_PARAMS,
KNOWN_DIRECTIVES,
Directive,
Lexed,
is_escaped,
iter_directives,
lex,
)
from .models import (
BLOCK_DEFAULTS,
Expand All @@ -62,6 +69,7 @@
document_to_dict,
empty_document,
item_text,
list_items,
metadata_from_dict,
party_from_dict,
section_from_dict,
Expand All @@ -80,8 +88,11 @@
AttachmentDefinitionsImporter,
DefinitionsImporter,
Diagnostic,
PlacedMarker,
SectionIndexEntry,
ValidationResult,
is_drafting_note,
is_template,
slugify_identifier,
validate_document,
)
Expand Down Expand Up @@ -126,8 +137,15 @@
"DELIMITER_PAIRS",
"find_definition_anchors",
"DefinitionAnchor",
# Directives (§11)
# Directives (§11): the lexer, and where text holding them is
"iter_directives",
"lex",
"Lexed",
"is_escaped",
"block_fragments",
"list_fragments",
"Fragment",
"ListFragment",
"Directive",
"DIRECTIVE_PARAMS",
"KNOWN_DIRECTIVES",
Expand All @@ -138,13 +156,18 @@
"ValidationResult",
"SectionIndexEntry",
"Diagnostic",
"PlacedMarker",
# What validate_document decides, on its own (§15.1, §15.6)
"is_template",
"is_drafting_note",
# Document model
"Document",
"Metadata",
"Section",
"Block",
"ListItem",
"item_text",
"list_items",
"Side",
"Party",
"Representative",
Expand Down
34 changes: 27 additions & 7 deletions src/legaldown/definitions.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@

from collections.abc import Callable, Iterator
from dataclasses import dataclass
from typing import NamedTuple

from .directives import Directive, Lexed, blank_fenced_code, lex, mask_directives
from .markdown import FENCE_OPEN_RE, is_indented_code
Expand Down Expand Up @@ -177,7 +178,25 @@ class DefinitionRef:
offset: int = 0


def block_fragments(block: Block) -> list[tuple[str, bool]]:
class Fragment(NamedTuple):
"""A text of a block that may hold directives and markers
(``block_fragments``): the *text*, and whether its end is an anchor
position (§5.7)."""

text: str
anchor: bool


class ListFragment(NamedTuple):
"""A ``Fragment`` of a list (``list_fragments``), with the *items* it is
in: each item's number, outermost first."""

text: str
anchor: bool
items: tuple[int, ...]


def block_fragments(block: Block) -> list[Fragment]:
"""The free-text fragments of *block* that may contain inline directives,
each with whether a ``{#id}`` at its very end is in an anchor position.

Expand All @@ -190,17 +209,18 @@ def block_fragments(block: Block) -> list[tuple[str, bool]]:
code span or comment left open — runs into the next, and code and raw
HTML in it hold none (§11.4).
"""
return [(text, position) for text, position, _items in _fragments(block)]
return [Fragment(text, position) for text, position, _items in _fragments(block)]


def list_fragments(block: Block) -> list[tuple[str, bool, tuple[int, ...]]]:
def list_fragments(block: Block) -> list[ListFragment]:
"""The fragments of a list's items, as ``block_fragments`` lists them:
each with whether its end is an anchor position — the end of an item's
first paragraph, the block it opens with (§5.7) — and the items it is
in, outermost first, each item numbered in document order among all the
list's items, nested ones included (§15.3). A list in a block quote is
no unit: its fragments are in the items the quote is in."""
return _fragments(block)
in, outermost first. Items are numbered in pre-order among all the
list's items: an item before the items nested in it, empty items
included (§15.3). A list in a block quote is no unit: its fragments are
in the items the quote is in."""
return [ListFragment(*fragment) for fragment in _fragments(block)]


def quote_ranges(block: Block) -> list[tuple[str, range]]:
Expand Down
7 changes: 4 additions & 3 deletions src/legaldown/directives.py
Original file line number Diff line number Diff line change
Expand Up @@ -315,9 +315,10 @@ def lex(text: str) -> Lexed:
"""Lex *text* for directives, outside literal regions (§11.4).

*text* is inline text, such as a paragraph's: fenced code is not looked
for. Text that can hold one (a code block's, a block quote's) is passed
through ``blank_fenced_code`` first, as block structure precedes inline
structure. It is read once, left to right, taking whichever
for. Text that can hold it (a code block's, a block quote's) has it
blanked first, as block structure precedes inline structure
(``blank_fenced_code``; ``block_fragments`` gives such text so). It is
read once, left to right, taking whichever
of a directive, a comment, or a code span opens first. A directive is
lexed from the source as written and consumes its own text, so a quoted
value may hold backticks or ``<!--`` without opening anything. Any other
Expand Down
16 changes: 14 additions & 2 deletions src/legaldown/validator/__init__.py
Original file line number Diff line number Diff line change
@@ -1,7 +1,8 @@
"""legaldown.validator — Document validation and structural analysis."""
from __future__ import annotations

from .core import AttachmentDefinitionsImporter, DefinitionsImporter, validate_document
from .conditions import ALWAYS, Condition, Presence, condition_problem, exclusive, parse_condition
from .core import AttachmentDefinitionsImporter, DefinitionsImporter, is_template, validate_document
from .helpers import (
format_section_number,
is_positive_numeric,
Expand All @@ -18,17 +19,28 @@
VALID_DURATION_UNITS,
VALID_PLACEHOLDER_TYPES,
)
from .result import Diagnostic, SectionIndexEntry, ValidationResult
from .result import Diagnostic, PlacedMarker, SectionIndexEntry, ValidationResult
from .templates import is_drafting_note

__all__ = [
# Core
"validate_document",
"is_template",
"is_drafting_note",
"DefinitionsImporter",
"AttachmentDefinitionsImporter",
# Result types
"ValidationResult",
"SectionIndexEntry",
"Diagnostic",
"PlacedMarker",
# Conditions (§15.3, §15.4): a condition's presence, a set of them
"parse_condition",
"Condition",
"condition_problem",
"exclusive",
"Presence",
"ALWAYS",
# Helpers
"slugify_identifier",
"format_section_number",
Expand Down
8 changes: 5 additions & 3 deletions src/legaldown/validator/conditions.py
Original file line number Diff line number Diff line change
Expand Up @@ -50,8 +50,8 @@ def parse_condition(text: str) -> Condition | None:


def condition_problem(text: str, questions: Any) -> str | None:
"""Why *text* is not a valid condition for *questions* (condition-invalid,
§15.3), or None."""
"""Why *text* is not a valid condition for *questions* (the document's
``metadata.questions``; condition-invalid, §15.3), or None."""
condition = parse_condition(text)
if condition is None:
return "must be '[!]question' or '[!]question:value', in the identifier format"
Expand Down Expand Up @@ -90,7 +90,9 @@ def satisfiable(presence: Iterable[Condition], questions: Any) -> bool:


def exclusive(first: Presence, second: Presence, questions: Any) -> bool:
"""True if two units can never appear together (§15.4)."""
"""True if two units, appearing under the conditions *first* and
*second* (their ``Presence``), can never appear together for any answers
to the document's *questions* (``document.metadata.questions``, §15.4)."""
return not satisfiable(first | second, questions)


Expand Down
69 changes: 59 additions & 10 deletions src/legaldown/validator/core.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@
mask_directives,
)
from ..markdown import HTML_COMMENT_RE, INLINE_HTML_RE, is_comment_only
from ..models import Amends, Document
from ..models import LIST_KINDS, Amends, Document
from ..positions import Locator
from ..specification import SPEC_VERSION, parse_version
from .conditions import ALWAYS, Presence, always_covered, condition_problem, exclusive, parse_condition, satisfiable
Expand All @@ -47,7 +47,7 @@
VALID_DURATION_UNITS,
VALID_PLACEHOLDER_TYPES,
)
from .result import Line, SectionIndexEntry, ValidationResult
from .result import Line, PlacedMarker, SectionIndexEntry, ValidationResult
from .templates import (
BRACE_STRAY,
DECISION_QUESTION_TYPES,
Expand Down Expand Up @@ -525,15 +525,21 @@ def _check_never_true(
)


def is_template(
document: Document,
markers: list[FoundMarker] | None = None,
lex_fragment: Callable[[str], Lexed] = lex,
) -> bool:
def is_template(document: Document) -> bool:
"""True if *document* is a template (§15.1): it declares ``questions``,
carries a condition — on a section, an attachment, or a placed marker —
or contains a ``{{choose:}}``, wherever it is. *markers* are
``find_markers(document, lex_fragment)`` when the caller has them."""
or contains a ``{{choose:}}``, wherever it is. ``validate_document``
reports it too (``ValidationResult.is_template``)."""
return _is_template(document, None, lex)


def _is_template(
document: Document,
markers: list[FoundMarker] | None,
lex_fragment: Callable[[str], Lexed],
) -> bool:
"""``is_template``; *markers* are ``find_markers(document,
lex_fragment)`` when the caller has them."""
from ..definitions import text_fragments # see the import note in validate_document

meta = document.metadata
Expand Down Expand Up @@ -649,7 +655,7 @@ def named(text: str, name: str) -> list[Directive]:
# Markers are found first: only a template gives a preamble paragraph's
# condition its place (§5.7).
markers = find_markers(document, lex_fragment)
template = is_template(document, markers, lex_fragment)
template = _is_template(document, markers, lex_fragment)
units = Units(document, markers, questions, template=template)
body_directives = {
directive.name
Expand Down Expand Up @@ -1547,6 +1553,49 @@ def definition_line(ref: Any) -> int | None:
line=definition_line(ref),
)

# What a renderer builds from: the template decision and the markers
# that apply, as every check above read them.
result.is_template = template
result.placed_markers = _placed_markers(document, markers, template, marker_line)

if document.filename:
result.diagnostics = [replace(d, file=document.filename) for d in result.diagnostics]
return result


def _placed_markers(
document: Document,
markers: list[FoundMarker],
template: bool,
line: Callable[[FoundMarker], int | None],
) -> list[PlacedMarker]:
"""The markers of *markers* that apply (``PlacedMarker``)."""
from ..definitions import list_fragments # see the import note in validate_document

placed = []
lists: dict[tuple[int | None, int], list[tuple[str, bool, tuple[int, ...]]]] = {} # each list walked once
for found in markers:
if found.marker is None or not found.placed(template):
continue
blocks = document.preamble if found.section is None else document.sections[found.section].blocks
block = blocks[found.block]
items: tuple[int, ...] = ()
if block.kind in LIST_KINDS:
key = (found.section, found.block)
if key not in lists:
lists[key] = list_fragments(block)
items = lists[key][found.fragment][2]
placed.append(PlacedMarker(
section=found.section,
block=found.block,
fragment=found.fragment,
offset=found.offset,
source=found.source,
identifier="" if found.include_only else found.marker.identifier,
condition=found.marker.condition,
field="suffix" if block.kind in ("ref", "term") else "text",
item=items[-1] if items else None,
include_only=found.include_only,
line=line(found),
))
return placed
Loading
Loading