Skip to content

fix(deps): update rust crate oxidize-pdf to v5 - #54

Open
renovate[bot] wants to merge 1 commit into
masterfrom
renovate/oxidize-pdf-5.x
Open

renovate[bot] wants to merge 1 commit into
masterfrom
renovate/oxidize-pdf-5.x

Conversation

@renovate

@renovate renovate Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

This PR contains the following updates:

Package Type Update Change
oxidize-pdf dependencies major 4.0.05.0.0

Release Notes

bzsanti/oxidizePdf (oxidize-pdf)

v5.1.3

Compare Source

Changed
  • Lenient PDF loading now uses optimized binary pattern search for xref
    recovery.
    This substantially reduces startup time for documents that need
    supplementary object-header scanning.
  • Public capability claims are auditable and internally consistent (#​603).
    The PDF/A and signature claims now link to versioned evidence and explicit
    limits; the adoption-monitoring contract defines privacy, deterministic
    decision and append-only audit requirements for oxidize-stats.
Fixed
  • Visible signature text fits both dimensions without losing content (#​606).
    Logos are centered behind the text without reserving a column. Shared
    SignatureAppearance::layout preflight wraps text using Helvetica metrics,
    adjusts the font size between 6 and 12 points, and returns an explicit error
    if the complete text cannot fit instead of clipping or omitting lines.
  • Text extraction preserves word boundaries across narrow font changes and
    Form XObject boundaries
    (#​602). This prevents differential word fusions
    without weakening the committed T3 baseline.
Added
  • Custom visible incremental-signature appearances (#​596). Callers can
    provide signer text, signing date, additional text, and a bounded RGB image
    watermark through SignatureAppearance. The generated font, image, and
    appearance stream are included in the signed incremental revision for both
    new and existing signature widgets.

v5.1.2

Compare Source

oxidize-pdf v5.1.2

See CHANGELOG.md for details.

Installation

Add to your Cargo.toml:

[dependencies]
oxidize-pdf = "5.1.2"

What's Changed

Full Changelog: bzsanti/oxidizePdf@v5.1.1...v5.1.2

v5.1.1

Compare Source

Fixed
  • Figure text with custom font differences is no longer emitted as reliable
    text when the PDF provides no /ToUnicode mapping
    (#​593, #​594).
    Consumers can explicitly retain that fallback text for forensic extraction.

v5.1.0

Compare Source

Added
  • Optional URI extraction from interactive link annotations (#​584, #​591).
    TextExtractor::with_link_annotation_extraction(true) appends safe /URI
    action targets in page annotation order without following or executing them.
Fixed
  • Text extraction resolves indirect font encodings and Adobe Glyph List
    differences
    (#​572), applies Type 3 /FontMatrix scaling to glyph widths
    (#​573), and preserves hyphens in numeric and punctuation-bearing identifiers
    across line wraps (#​574, #​589).
  • Standalone CR, CRLF, and Unicode line separators are normalized
    consistently
    (#​575), while TJ kerning-space detection scales with the
    active font size (#​588).
  • Signature preparation tolerates unreferenced in-use xref entries at byte
    offset zero
    while still rejecting policy references to such entries (#​585).
  • Page annotation arrays accept direct annotation dictionaries (#​590).

v5.0.1

Compare Source

Fixed
  • Layout-preserving plaintext extraction restores document reading quality
    (#​564, #​570). PlainTextExtractor::preserve_layout() now uses the complete
    text engine and its scale-relative XY-Cut reading order, retaining
    /ActualText, artifact filtering, font metrics, and error propagation. On
    the pinned OmniDocBench protocol, global text similarity improves from
    48.26% in v5.0.0 to 60.01%, above the 55% acceptance threshold, while native
    reading-order edit distance remains within its 0.25 limit at 0.22639.
Changed
  • OmniDocBench quality measurements are reproducible (#​565, #​568). The
    versioned gate pins dataset, evaluator, source, extraction configuration, and
    scored-page population provenance, validates materialized Git LFS objects,
    supports split-page PDFs, and seals prediction and summary hashes.

v5.0.0

Compare Source

Added
  • Text extraction now exposes the active PDF rendering mode (#​477, #​562).
    Every TextFragment reports its Tr mode, including invisible OCR text and
    /ActualText replacements. Graphics-state restoration preserves the mode,
    malformed operands are handled without integer truncation, and layout
    reconstruction never fuses fragments across rendering-mode boundaries.
Changed
  • Existing-PDF operations now use one policy-driven API (#​560, #​561).
    Merge, split, extraction, reordering, batch processing, signing, and related
    workflows share explicit preservation, validation, and permission policies.
    This major release retires ambiguous legacy entry points; see
    docs/migration/v5-existing-pdf-operations.md for migration guidance.
  • TextFragment is now non-exhaustive (#​477, #​562). External callers must
    construct synthetic fragments with TextFragment::new and then set any
    non-default public fields, allowing future extraction metadata to be added
    without another source-breaking struct-field change.

v4.9.0

Compare Source

Added
  • Provider-neutral incremental PDF signing (#​540, #​558). New two-phase
    preparation and finalization APIs support caller-produced CMS signatures,
    visible and invisible fields, existing-field selection, successive
    signatures, xref tables and streams, and DocMDP and FieldMDP enforcement
    while preserving the source PDF as an exact byte prefix.
  • Lossless incremental FreeText annotation editing (#​534, #​553). A typed
    editor can enumerate, add, update, and remove FreeText annotations without
    rebuilding unrelated document content.
  • Lossless incremental Ink annotation editing (#​536, #​554). Typed APIs can
    enumerate and atomically mutate ink strokes, appearance properties, and
    annotation metadata while preserving prior PDF bytes.
  • Lossless incremental geometric annotation editing (#​537, #​555). Typed
    editors support line, square, circle, polygon, and polyline annotations,
    including geometry, color, opacity, width, dash patterns, and line endings.
  • Atomic page-tree mutation batches (#​538, #​556). New planning and mutation
    APIs can reorder, insert, duplicate, and remove pages in one validated,
    lossless incremental revision.
  • Document-semantic preservation for structural operations (#​539, #​557).
    Merge, split, extraction, and page mutation APIs preserve or safely reconcile
    outlines, named destinations, page labels, AcroForm state, metadata, and
    associated document structures.
Changed
  • Obsolete CLI and API release artifacts were retired (#​552). The
    oxidize-pdf library is now the sole maintained and published artifact.

v4.8.0

Compare Source

Fixed
  • Semantic redaction no longer presents visual masking as irreversible
    removal
    (#​541). Reports explicitly identify recoverable masking risks, and
    the security-grade API now removes exact direct-page ASCII Tj operands and
    complete literal TJ arrays backed by verified non-symbolic Standard-14
    fonts. Both Tj and TJ replacements preserve the original text advance,
    including AFM glyph widths, numeric adjustments, character spacing, and word
    spacing. The API rebuilds the file without prior revisions or document-level
    auxiliary data and verifies output page streams before reporting
    irreversible success.
    It correlates each match with its declared bounding box, audits retained page
    resources and metadata during forensic reparse, and enforces input, page,
    entity, decoded-content, and operation budgets. It fails closed for
    annotations, XObjects, shadings, inline images, marked content, custom or
    ambiguous font encodings, partial or ambiguous matches, and malformed
    streams.
Added
  • Revision-aware semantic PDF comparison (#​543). A bounded comparison API
    reports visual, textual, structural, metadata, security, and serialization
    differences, attributes changes to incremental revisions, and normalizes
    irrelevant serialization and timestamp noise.
  • Lossless tagged-PDF validation and incremental editing (#​544). Public APIs
    inspect structure trees and PDF/UA findings, then apply bounded edits for
    attributes, MCID associations, ParentTree repair, element creation, and
    reparenting while preserving unrelated bytes and enforcing DocMDP policy.
  • Public outline and bookmark reading (#​548). PdfDocument can now parse
    bounded bookmark hierarchies, styles, open state, direct and named
    destinations, GoTo actions, and all standard destination view modes into
    stable zero-based page indexes.
  • Lossless incremental OCR text layers (#​542). Positioned, invisible
    Unicode text can now be appended to existing pages without rebuilding the
    document graph or changing the source-byte prefix. The editor exposes a
    policy-aware dry run, isolated streams and font resources, language,
    confidence, source-region and reading-order metadata, duplicate-layer
    detection, deterministic output, xref-table and xref-stream support, and
    validated atomic publication through PdfOcrConverter. Encrypted inputs and
    every DocMDP certification level fail closed because OCR changes page
    content.
  • Lossless incremental page reordering (#​531). The new
    reorder_pdf_pages_lossless API preserves the source bytes, indirect page
    identities, inherited page attributes, and unrelated document objects while
    atomically applying an exact page permutation. Encrypted inputs fail closed.
  • DocMDP enforcement for incremental structural edits (#​532). Lossless
    page reordering now permits ordinary approval signatures while parsing and
    enforcing certification transforms, rejecting every certified, malformed,
    ambiguous, or unsupported policy that cannot authorize the structural edit.

v4.7.0

Compare Source

Added
  • Incremental editing for highlight annotations (#​525, #​527). The new
    IncrementalHighlightEditor API can enumerate, add, update, and remove
    /Highlight annotations while preserving the original PDF bytes and
    applying validated changes as incremental revisions.
Fixed
  • Standard-14 fonts without explicit /Widths used inaccurate fallback
    advances during text extraction
    (#​523, #​524). Encoding-aware AFM metrics
    now provide per-glyph widths for simple Standard-14 fonts, preventing
    spurious spaces when text is split across consecutive showing operators.
  • CMS signature verification could accept certificates without establishing
    trust in the configured anchors
    (#​526, #​528). Trust validation now fails
    closed while preserving the existing public verification API.
  • Floating-point residue could keep mathematically contiguous text fragments
    separate
    (#​521, #​522). Same-line merging now tolerates rounding noise
    without treating visible overlaps as adjacent or inserting a spurious space.

v4.6.0

Compare Source

Added
  • Type 3 font glyph resolution for downstream renderers (#​509). Parser
    consumers can resolve bounded CharProc content streams together with the
    glyph resources and metrics needed to render Type 3 fonts safely.
  • Resolved parser font resources (#​513). The parser now exposes resolved
    Type0, CID, simple, and Symbol font data through public, renderer-oriented
    font resource types while retaining safe fallbacks for malformed PDFs.
  • Bounded in-memory image extraction (#​516). New visitor and collection
    APIs expose encoded image data without temporary files, with configurable
    limits for image count, per-image and total encoded bytes, and decoded
    pixels. Existing file-based extraction APIs remain compatible.
Fixed
  • Flat-path word-gap threshold compared a Tm-scaled pen delta against an
    unscaled font-size threshold
    (#​510). flat_space_gap_threshold derives
    its threshold from the font's real space-glyph advance at the nominal Tf
    font size, but the gap it's compared against is measured in user space,
    already scaled by Tm/CTM. PDF generators that draw at Tf 1 with the
    real point size baked into Tm instead (a common technique) hit a
    threshold far smaller, relative to the gap, than intended: an ordinary
    sub-point positioning residue between Tj runs of the same token (e.g. a
    hyphenated phone number split across separate runs) could cross it and
    insert a spurious mid-token space. The threshold is now scaled by the same
    Tm/CTM x-factor already applied to page-space widths elsewhere in
    extraction, together with the horizontal text scaling selected by Tz.
  • Indirect /DecodeParms references bypassed stream predictors (#​514).
    Filter decoding now resolves indirect parameter dictionaries before
    applying PNG and TIFF predictors, including parameters nested in filter
    arrays, while rejecting cycles and invalid references safely.

v4.5.1

Compare Source

Fixed
  • Text extraction ignored structure-element /ActualText replacements
    (#​506). The extractor now resolves each marked-content MCID through the
    page's /StructParents entry and the document /ParentTree, preserving
    inline-over-structure precedence in both flat and layout-preserving modes.
    Resolution is lazy and bounded, malformed structure trees fall back safely
    to visual text, and tagged Form XObjects use their own structural context.

v4.5.0

Compare Source

Added
  • Incremental editing for standard text-note annotations (#​493). The new
    typed IncrementalTextNoteEditor API can enumerate, add, move, edit, and
    remove /Text annotations while preserving the original PDF bytes and
    applying each validated batch as one incremental revision.
Fixed
  • AES-256 revision 5 encryption dictionaries now emit /Perms (#​492),
    restoring interoperability with external readers while retaining strict
    validation of the encrypted permissions block.
  • Flat extraction could fuse adjacent cells in label/value grids (#​495).
    Separate text objects that return backward on the same baseline now receive
    neutral whitespace when the geometry is ambiguous, while genuine long wraps
    still become newlines and legitimate repositioned overlays remain intact.
  • Composite (Type0/CID) font text extraction never consulted the CIDFont's
    /W//DW glyph widths
    (#​496). Every glyph was assumed to advance by a
    flat 0.5 * font_size, regardless of its real width. When a glyph's true
    advance diverged enough from that flat estimate relative to its neighbors
    (e.g. a genuinely wide glyph in the font), the extractor's pen-tracking
    drifted out of sync with where the next glyph was actually drawn, crossing
    the space-insertion threshold and corrupting the extracted text with a
    spurious space in the middle of a single token. /W//DW (ISO 32000-1
    §9.7.4.3) are now parsed into a CID-indexed width table and consulted for
    the composite-font width calculation on both the flat and
    preserve_layout extraction paths, falling back to the previous heuristic
    when neither is present.
  • Tagged ActualText extraction lost the marked-content association
    (#​498). Generated ActualText sequences now carry an MCID connected to the
    structure tree, so tagged extraction returns the replacement text instead
    of dropping it.
  • Real Type0 widths could hide narrow implicit word spaces (#​500). When a
    composite font has no discoverable U+0020 mapping, flat extraction now uses
    a bounded, font-derived CID-width signal that preserves narrow word gaps
    without splitting URLs, identifiers, kerned runs, or positioned overlays.

v4.4.0

Compare Source

Added
  • ActualText marked-content support (#​63, #​490). Page generation can now
    attach replacement text to marked-content sequences, allowing assistive
    technology and text extractors to consume an accessible textual alternative
    for visually rendered content.
Fixed
  • preserve_layout's global Y-sort could interleave unrelated content
    regions, corrupting hyphen-wrapped words
    (#​482). sort_and_merge_fragments
    sorts every fragment on a page by Y-coordinate with no notion of separate
    content regions; when an unrelated fragment (e.g. a digital-signature
    annotation's appearance text) happened to sit at a Y-coordinate between the
    two halves of a hyphen-wrapped word elsewhere on the page, the sort spliced
    it in between them, and the hyphen got joined to the wrong fragment —
    corrupting both the wrapped word and the unrelated text at once. The initial
    mitigation fused hyphen-wrapped continuations before sorting. Positional
    sorting is now also scoped to structural emission regions: geometric flow
    restarts separate untagged regions, while MCID ownership separates tagged
    regions. Independent body, overlay, annotation, and appearance flows can no
    longer be interleaved merely because their Y ranges overlap.

  • merge_hyphenated had no effect on the flat (default) extraction path
    (#​486). It was only wired into reconstruct_text_from_fragments
    (preserve_layout: true) and merge_into_paragraphs
    (reconstruct_paragraphs: true); every text-showing operator on the flat
    path (Tj, TJ, ', ") independently decided a '\n' separator via
    the shared append_bounded helper without ever checking for a trailing
    hyphen, so a hyphenated word or number wrapping across two lines extracted
    with a raw newline in place of the hyphen — e.g. a phone number split as
    "...3016-" / "0900" came out "...3016-\n0900" instead of
    "...30160900". append_bounded now pops the trailing - and fuses the
    next run with no separator when the caller requested '\n' and
    merge_hyphenated is enabled (the default); the actual separator applied
    is threaded back to callers so reading-order line grouping (#​448) treats a
    fused run as a continuation rather than opening a new line.

  • Literal carriage returns in decoded text strings leaked into extracted
    plain text
    (#​476). TextExtractor::with_carriage_return_handling now lets
    callers remove standalone CR bytes, replace them with a space, or preserve
    them while normalizing CRLF without adding a field to the public
    extraction-options struct. The default removes standalone CR so producer
    noise does not split words or URLs; CRLF is treated as one line ending under
    every policy. NormalizeLineEnding preserves a standalone CR because it is
    not equivalent to LF (#​481).

v4.3.0

Compare Source

Fixed
  • Three of the four legal /DescendantFonts spellings decoded CID text to
    mojibake
    (#​469). extract_font_info read the entry only when it was a
    direct array whose element is an indirect reference. ISO 32000-1 puts no
    reference requirement on either the array or its element (Table 121, §7.3.6,
    §7.3.7 — where the spec wants a reference it says so, as Table 117 does for
    the CIDFont's FontDescriptor), and producers write the element inline:
    ReportLab's UnicodeCIDFont does. For those files descendant_font stayed
    empty, decode_text_with_font skipped its Type0 branch, and the already
    correct cid_encoding (e.g. UniJIS-UCS2-H → UTF-16BE) went unused — the
    text fell through to byte-wise decoding. The same defect class as #​463, fixed
    here for /DescendantFonts: the value may now be a direct array or a
    reference to one, and the CIDFont element a dictionary or a reference.

  • Tc, Tw and Ts were parsed and stored but never applied to extraction
    (#​456). The character-spacing (Tc) and word-spacing (Tw) parameters are
    now folded into the pen advance and Ts (text rise) into the glyph baseline
    (ISO 32000-1 §9.4.4): Tc is added once per glyph, Tw once per single-byte
    space (code 32, §9.3.3), and Ts offsets the fragment's y-origin. Because the
    advance feeds the flat path's space/newline heuristics, documents that set a
    non-zero Tc/Tw now get correct separators; the " operator, which sets
    both before showing a line, now takes effect. Tr (render mode, including the
    invisible Tr 3 OCR-layer case) is unchanged — exposing it on the public
    TextFragment is a breaking change deferred to the next major.

  • Two TJ operators drawn side by side on one line were glued together
    (#​458). The Tj arm turns a forward pen jump wider than the threshold into a
    space, but the TJ arm used the same delta only to decide newlines, so
    adjacent multi-column table cells extracted as CellOneCellTwo and a list
    bullet welded onto its item text as vlarge. A boundary space now fires on
    the first glyph of a TJ array when the pen jumped forward past
    0.7 em (calibrated on the Tc/Tw-corrected advance from #​456; a leading kern
    no longer masks the jump). On the t3-stress corpus this cut word fusions
    0.0032 → 0.0014 and reading-order misplacement 0.2766 → 0.2486. An external
    corpus of 421 real-world PDFs (contributor @​oshtivi) measured regressed
    detections dropping 177 → 125 with this fix, the largest single-fix drop of
    the cycle.

Added
  • Opt-in reading-order reorder for the flat text path (#​448), via
    TextExtractor::with_reading_order(true). Off by default (the flat path stays
    byte-identical); when on, the flat .text line groups are permuted into
    reading order — left column before right, top block before bottom — using a
    scale-relative XY-cut whose gap thresholds are multiples of the region's
    median glyph size, so column detection is font-size-relative rather than an
    absolute point gap. The text inside each group is untouched. On the
    t3-stress corpus this cut reading-order misplacement 0.2486 → 0.2255 at
    identical coverage. It reorders only groups the newline heuristic already
    separated (columns drawn row-interleaved fall in one group), and orders
    /Rotate ≠ 0 pages in unrotated page space.

Configuration

📅 Schedule: (UTC)

  • Branch creation
    • At any time (no schedule defined)
  • Automerge
    • At any time (no schedule defined)

🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.

Rebasing: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox.

🔕 Ignore: Close this PR and you won't be reminded about this update again.


  • If you want to rebase/retry this PR, check this box

This PR was generated by Mend Renovate. View the repository job log.

@renovate
renovate Bot force-pushed the renovate/oxidize-pdf-5.x branch 2 times, most recently from 0b87a2d to 66d0f8d Compare September 11, 2026 17:04
@coderabbitai

coderabbitai Bot commented Sep 11, 2026

Copy link
Copy Markdown

Important

Review skipped

Bot user detected.

To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Advanced

Run ID: af00e929-87d8-4270-97ac-a40490cf04d9

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@renovate
renovate Bot force-pushed the renovate/oxidize-pdf-5.x branch 2 times, most recently from bc6eb5c to 125a392 Compare September 16, 2026 17:03
@renovate
renovate Bot force-pushed the renovate/oxidize-pdf-5.x branch from 125a392 to 5ee397e Compare September 20, 2026 23:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants