Phase 5: large-file / streaming processing - #5
Merged
Merged
Conversation
Large_File_Processing/: constant-memory FundsXML processing. - make_large_sample.py: streaming WRITER -> big XSD-valid file - stream_aggregate.py (lxml iterparse) + StreamAggregate.java (StAX, native, no JAXB/DOM): identical totals at flat memory (verified ~16 MiB / ~2 MiB for 20k positions, size-independent) - split.py: split into independently XSD-valid chunks - delta_diff.py: INITIAL-vs-DELTA position diff (added/removed/changed), exit 1 on differences All verified locally. CI gains a step that generates a 30k-position file, stream-aggregates (Python+Java, totals asserted), splits + XSD-validates a chunk, and runs an identical-input delta (exit 0). Top-level index updated. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phase 5 of the enterprise FundsXML showcase. Based on PR #4's branch so the diff is Phase 5 only.
Added —
Large_File_Processing/Constant-memory FundsXML processing (verified: ~16 MiB RSS Python / ~2 MiB heap Java for 20k positions, size-independent):
make_large_sample.py— streaming writer → big XSD-valid filestream_aggregate.py(lxmliterparse, clears parsed siblings) +StreamAggregate.java(StAX pull parser, native Java, no JAXB/DOM) — identical totals at flat memorysplit.py— split into independently XSD-valid chunksdelta_diff.py— INITIAL-vs-DELTA position diff (added/removed/changed; exit 1 on differences)All verified locally. CI gains a step: generate a 30k-position file → stream-aggregate (Python+Java, totals asserted) → split + XSD-validate a chunk → identical-input delta (exit 0). README includes an ETL (Camel/NiFi) integration note. Top-level index → ✅.
Roadmap remaining: Phase 6 data binding / JSON (final phase).
🤖 Generated with Claude Code