fix(deps): update rust crate oxidize-pdf to v5 - #54
Open
renovate[bot] wants to merge 1 commit into
Open
renovate[bot] wants to merge 1 commit into
renovate[bot] wants to merge 1 commit into
Conversation
renovate
Bot
force-pushed
the
renovate/oxidize-pdf-5.x
branch
2 times, most recently
from
September 11, 2026 17:04
0b87a2d to
66d0f8d
Compare
|
Important Review skippedBot user detected. To trigger a single review, invoke the ⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
renovate
Bot
force-pushed
the
renovate/oxidize-pdf-5.x
branch
2 times, most recently
from
September 16, 2026 17:03
bc6eb5c to
125a392
Compare
renovate
Bot
force-pushed
the
renovate/oxidize-pdf-5.x
branch
from
September 20, 2026 23:03
125a392 to
5ee397e
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR contains the following updates:
4.0.0→5.0.0Release Notes
bzsanti/oxidizePdf (oxidize-pdf)
v5.1.3Compare Source
Changed
recovery. This substantially reduces startup time for documents that need
supplementary object-header scanning.
The PDF/A and signature claims now link to versioned evidence and explicit
limits; the adoption-monitoring contract defines privacy, deterministic
decision and append-only audit requirements for
oxidize-stats.Fixed
Logos are centered behind the text without reserving a column. Shared
SignatureAppearance::layoutpreflight wraps text using Helvetica metrics,adjusts the font size between 6 and 12 points, and returns an explicit error
if the complete text cannot fit instead of clipping or omitting lines.
Form XObject boundaries (#602). This prevents differential word fusions
without weakening the committed T3 baseline.
Added
provide signer text, signing date, additional text, and a bounded RGB image
watermark through
SignatureAppearance. The generated font, image, andappearance stream are included in the signed incremental revision for both
new and existing signature widgets.
v5.1.2Compare Source
oxidize-pdf v5.1.2
See CHANGELOG.md for details.
Installation
Add to your
Cargo.toml:What's Changed
Full Changelog: bzsanti/oxidizePdf@v5.1.1...v5.1.2
v5.1.1Compare Source
Fixed
text when the PDF provides no
/ToUnicodemapping (#593, #594).Consumers can explicitly retain that fallback text for forensic extraction.
v5.1.0Compare Source
Added
TextExtractor::with_link_annotation_extraction(true)appends safe/URIaction targets in page annotation order without following or executing them.
Fixed
differences (#572), applies Type 3
/FontMatrixscaling to glyph widths(#573), and preserves hyphens in numeric and punctuation-bearing identifiers
across line wraps (#574, #589).
consistently (#575), while
TJkerning-space detection scales with theactive font size (#588).
offset zero while still rejecting policy references to such entries (#585).
v5.0.1Compare Source
Fixed
(#564, #570).
PlainTextExtractor::preserve_layout()now uses the completetext engine and its scale-relative XY-Cut reading order, retaining
/ActualText, artifact filtering, font metrics, and error propagation. Onthe pinned OmniDocBench protocol, global text similarity improves from
48.26% in v5.0.0 to 60.01%, above the 55% acceptance threshold, while native
reading-order edit distance remains within its 0.25 limit at 0.22639.
Changed
versioned gate pins dataset, evaluator, source, extraction configuration, and
scored-page population provenance, validates materialized Git LFS objects,
supports split-page PDFs, and seals prediction and summary hashes.
v5.0.0Compare Source
Added
Every
TextFragmentreports itsTrmode, including invisible OCR text and/ActualTextreplacements. Graphics-state restoration preserves the mode,malformed operands are handled without integer truncation, and layout
reconstruction never fuses fragments across rendering-mode boundaries.
Changed
Merge, split, extraction, reordering, batch processing, signing, and related
workflows share explicit preservation, validation, and permission policies.
This major release retires ambiguous legacy entry points; see
docs/migration/v5-existing-pdf-operations.mdfor migration guidance.TextFragmentis now non-exhaustive (#477, #562). External callers mustconstruct synthetic fragments with
TextFragment::newand then set anynon-default public fields, allowing future extraction metadata to be added
without another source-breaking struct-field change.
v4.9.0Compare Source
Added
preparation and finalization APIs support caller-produced CMS signatures,
visible and invisible fields, existing-field selection, successive
signatures, xref tables and streams, and DocMDP and FieldMDP enforcement
while preserving the source PDF as an exact byte prefix.
editor can enumerate, add, update, and remove FreeText annotations without
rebuilding unrelated document content.
enumerate and atomically mutate ink strokes, appearance properties, and
annotation metadata while preserving prior PDF bytes.
editors support line, square, circle, polygon, and polyline annotations,
including geometry, color, opacity, width, dash patterns, and line endings.
APIs can reorder, insert, duplicate, and remove pages in one validated,
lossless incremental revision.
Merge, split, extraction, and page mutation APIs preserve or safely reconcile
outlines, named destinations, page labels, AcroForm state, metadata, and
associated document structures.
Changed
oxidize-pdflibrary is now the sole maintained and published artifact.v4.8.0Compare Source
Fixed
removal (#541). Reports explicitly identify recoverable masking risks, and
the security-grade API now removes exact direct-page ASCII
Tjoperands andcomplete literal
TJarrays backed by verified non-symbolic Standard-14fonts. Both
TjandTJreplacements preserve the original text advance,including AFM glyph widths, numeric adjustments, character spacing, and word
spacing. The API rebuilds the file without prior revisions or document-level
auxiliary data and verifies output page streams before reporting
irreversible success.
It correlates each match with its declared bounding box, audits retained page
resources and metadata during forensic reparse, and enforces input, page,
entity, decoded-content, and operation budgets. It fails closed for
annotations, XObjects, shadings, inline images, marked content, custom or
ambiguous font encodings, partial or ambiguous matches, and malformed
streams.
Added
reports visual, textual, structural, metadata, security, and serialization
differences, attributes changes to incremental revisions, and normalizes
irrelevant serialization and timestamp noise.
inspect structure trees and PDF/UA findings, then apply bounded edits for
attributes, MCID associations, ParentTree repair, element creation, and
reparenting while preserving unrelated bytes and enforcing DocMDP policy.
PdfDocumentcan now parsebounded bookmark hierarchies, styles, open state, direct and named
destinations, GoTo actions, and all standard destination view modes into
stable zero-based page indexes.
Unicode text can now be appended to existing pages without rebuilding the
document graph or changing the source-byte prefix. The editor exposes a
policy-aware dry run, isolated streams and font resources, language,
confidence, source-region and reading-order metadata, duplicate-layer
detection, deterministic output, xref-table and xref-stream support, and
validated atomic publication through
PdfOcrConverter. Encrypted inputs andevery DocMDP certification level fail closed because OCR changes page
content.
reorder_pdf_pages_losslessAPI preserves the source bytes, indirect pageidentities, inherited page attributes, and unrelated document objects while
atomically applying an exact page permutation. Encrypted inputs fail closed.
page reordering now permits ordinary approval signatures while parsing and
enforcing certification transforms, rejecting every certified, malformed,
ambiguous, or unsupported policy that cannot authorize the structural edit.
v4.7.0Compare Source
Added
IncrementalHighlightEditorAPI can enumerate, add, update, and remove/Highlightannotations while preserving the original PDF bytes andapplying validated changes as incremental revisions.
Fixed
/Widthsused inaccurate fallbackadvances during text extraction (#523, #524). Encoding-aware AFM metrics
now provide per-glyph widths for simple Standard-14 fonts, preventing
spurious spaces when text is split across consecutive showing operators.
trust in the configured anchors (#526, #528). Trust validation now fails
closed while preserving the existing public verification API.
separate (#521, #522). Same-line merging now tolerates rounding noise
without treating visible overlaps as adjacent or inserting a spurious space.
v4.6.0Compare Source
Added
consumers can resolve bounded CharProc content streams together with the
glyph resources and metrics needed to render Type 3 fonts safely.
Type0, CID, simple, and Symbol font data through public, renderer-oriented
font resource types while retaining safe fallbacks for malformed PDFs.
APIs expose encoded image data without temporary files, with configurable
limits for image count, per-image and total encoded bytes, and decoded
pixels. Existing file-based extraction APIs remain compatible.
Fixed
Tm-scaled pen delta against anunscaled font-size threshold (#510).
flat_space_gap_thresholdderivesits threshold from the font's real space-glyph advance at the nominal
Tffont size, but the gap it's compared against is measured in user space,
already scaled by
Tm/CTM. PDF generators that draw atTf 1with thereal point size baked into
Tminstead (a common technique) hit athreshold far smaller, relative to the gap, than intended: an ordinary
sub-point positioning residue between
Tjruns of the same token (e.g. ahyphenated phone number split across separate runs) could cross it and
insert a spurious mid-token space. The threshold is now scaled by the same
Tm/CTM x-factor already applied to page-space widths elsewhere inextraction, together with the horizontal text scaling selected by
Tz./DecodeParmsreferences bypassed stream predictors (#514).Filter decoding now resolves indirect parameter dictionaries before
applying PNG and TIFF predictors, including parameters nested in filter
arrays, while rejecting cycles and invalid references safely.
v4.5.1Compare Source
Fixed
/ActualTextreplacements(#506). The extractor now resolves each marked-content MCID through the
page's
/StructParentsentry and the document/ParentTree, preservinginline-over-structure precedence in both flat and layout-preserving modes.
Resolution is lazy and bounded, malformed structure trees fall back safely
to visual text, and tagged Form XObjects use their own structural context.
v4.5.0Compare Source
Added
typed
IncrementalTextNoteEditorAPI can enumerate, add, move, edit, andremove
/Textannotations while preserving the original PDF bytes andapplying each validated batch as one incremental revision.
Fixed
/Perms(#492),restoring interoperability with external readers while retaining strict
validation of the encrypted permissions block.
Separate text objects that return backward on the same baseline now receive
neutral whitespace when the geometry is ambiguous, while genuine long wraps
still become newlines and legitimate repositioned overlays remain intact.
/W//DWglyph widths (#496). Every glyph was assumed to advance by aflat
0.5 * font_size, regardless of its real width. When a glyph's trueadvance diverged enough from that flat estimate relative to its neighbors
(e.g. a genuinely wide glyph in the font), the extractor's pen-tracking
drifted out of sync with where the next glyph was actually drawn, crossing
the space-insertion threshold and corrupting the extracted text with a
spurious space in the middle of a single token.
/W//DW(ISO 32000-1§9.7.4.3) are now parsed into a CID-indexed width table and consulted for
the composite-font width calculation on both the flat and
preserve_layoutextraction paths, falling back to the previous heuristicwhen neither is present.
ActualTextextraction lost the marked-content association(#498). Generated
ActualTextsequences now carry an MCID connected to thestructure tree, so tagged extraction returns the replacement text instead
of dropping it.
composite font has no discoverable U+0020 mapping, flat extraction now uses
a bounded, font-derived CID-width signal that preserves narrow word gaps
without splitting URLs, identifiers, kerned runs, or positioned overlays.
v4.4.0Compare Source
Added
ActualTextmarked-content support (#63, #490). Page generation can nowattach replacement text to marked-content sequences, allowing assistive
technology and text extractors to consume an accessible textual alternative
for visually rendered content.
Fixed
preserve_layout's global Y-sort could interleave unrelated contentregions, corrupting hyphen-wrapped words (#482).
sort_and_merge_fragmentssorts every fragment on a page by Y-coordinate with no notion of separate
content regions; when an unrelated fragment (e.g. a digital-signature
annotation's appearance text) happened to sit at a Y-coordinate between the
two halves of a hyphen-wrapped word elsewhere on the page, the sort spliced
it in between them, and the hyphen got joined to the wrong fragment —
corrupting both the wrapped word and the unrelated text at once. The initial
mitigation fused hyphen-wrapped continuations before sorting. Positional
sorting is now also scoped to structural emission regions: geometric flow
restarts separate untagged regions, while MCID ownership separates tagged
regions. Independent body, overlay, annotation, and appearance flows can no
longer be interleaved merely because their Y ranges overlap.
merge_hyphenatedhad no effect on the flat (default) extraction path(#486). It was only wired into
reconstruct_text_from_fragments(
preserve_layout: true) andmerge_into_paragraphs(
reconstruct_paragraphs: true); every text-showing operator on the flatpath (
Tj,TJ,',") independently decided a'\n'separator viathe shared
append_boundedhelper without ever checking for a trailinghyphen, so a hyphenated word or number wrapping across two lines extracted
with a raw newline in place of the hyphen — e.g. a phone number split as
"...3016-"/"0900"came out"...3016-\n0900"instead of"...30160900".append_boundednow pops the trailing-and fuses thenext run with no separator when the caller requested
'\n'andmerge_hyphenatedis enabled (the default); the actual separator appliedis threaded back to callers so reading-order line grouping (#448) treats a
fused run as a continuation rather than opening a new line.
Literal carriage returns in decoded text strings leaked into extracted
plain text (#476).
TextExtractor::with_carriage_return_handlingnow letscallers remove standalone CR bytes, replace them with a space, or preserve
them while normalizing CRLF without adding a field to the public
extraction-options struct. The default removes standalone CR so producer
noise does not split words or URLs; CRLF is treated as one line ending under
every policy.
NormalizeLineEndingpreserves a standalone CR because it isnot equivalent to LF (#481).
v4.3.0Compare Source
Fixed
Three of the four legal
/DescendantFontsspellings decoded CID text tomojibake (#469).
extract_font_inforead the entry only when it was adirect array whose element is an indirect reference. ISO 32000-1 puts no
reference requirement on either the array or its element (Table 121, §7.3.6,
§7.3.7 — where the spec wants a reference it says so, as Table 117 does for
the CIDFont's FontDescriptor), and producers write the element inline:
ReportLab's
UnicodeCIDFontdoes. For those filesdescendant_fontstayedempty,
decode_text_with_fontskipped its Type0 branch, and the alreadycorrect
cid_encoding(e.g.UniJIS-UCS2-H→ UTF-16BE) went unused — thetext fell through to byte-wise decoding. The same defect class as #463, fixed
here for
/DescendantFonts: the value may now be a direct array or areference to one, and the CIDFont element a dictionary or a reference.
Tc,TwandTswere parsed and stored but never applied to extraction(#456). The character-spacing (
Tc) and word-spacing (Tw) parameters arenow folded into the pen advance and
Ts(text rise) into the glyph baseline(ISO 32000-1 §9.4.4):
Tcis added once per glyph,Twonce per single-bytespace (code 32, §9.3.3), and
Tsoffsets the fragment's y-origin. Because theadvance feeds the flat path's space/newline heuristics, documents that set a
non-zero
Tc/Twnow get correct separators; the"operator, which setsboth before showing a line, now takes effect.
Tr(render mode, including theinvisible
Tr 3OCR-layer case) is unchanged — exposing it on the publicTextFragmentis a breaking change deferred to the next major.Two
TJoperators drawn side by side on one line were glued together(#458). The
Tjarm turns a forward pen jump wider than the threshold into aspace, but the
TJarm used the same delta only to decide newlines, soadjacent multi-column table cells extracted as
CellOneCellTwoand a listbullet welded onto its item text as
vlarge. A boundary space now fires onthe first glyph of a
TJarray when the pen jumped forward past0.7 em(calibrated on the Tc/Tw-corrected advance from #456; a leading kernno longer masks the jump). On the
t3-stresscorpus this cut word fusions0.0032 → 0.0014 and reading-order misplacement 0.2766 → 0.2486. An external
corpus of 421 real-world PDFs (contributor @oshtivi) measured regressed
detections dropping 177 → 125 with this fix, the largest single-fix drop of
the cycle.
Added
TextExtractor::with_reading_order(true). Off by default (the flat path staysbyte-identical); when on, the flat
.textline groups are permuted intoreading order — left column before right, top block before bottom — using a
scale-relative XY-cut whose gap thresholds are multiples of the region's
median glyph size, so column detection is font-size-relative rather than an
absolute point gap. The text inside each group is untouched. On the
t3-stresscorpus this cut reading-order misplacement 0.2486 → 0.2255 atidentical coverage. It reorders only groups the newline heuristic already
separated (columns drawn row-interleaved fall in one group), and orders
/Rotate ≠ 0pages in unrotated page space.Configuration
📅 Schedule: (UTC)
🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.
♻ Rebasing: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox.
🔕 Ignore: Close this PR and you won't be reminded about this update again.
This PR was generated by Mend Renovate. View the repository job log.