From 7542d31738f5c2a4167289e06bb33af41f890aed Mon Sep 17 00:00:00 2001 From: Oluwatobi Ogundimu Date: Fri, 31 Jul 2026 18:05:13 +0100 Subject: [PATCH] SC-3: delta-transfer algorithm (rolling checksum + block matching) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Rolling weak checksum + MD5 strong checksum, block signature generation, and delta generation/application with weak-checksum collisions disambiguated by strong checksum. Signature bundles block size with checksums so generation/application can't drift out of sync. Self-review found a dead assignment and, more importantly, that the collision-disambiguation path had no test exercising it — added a hand-constructed genuine Adler-32 collision to prove it actually works in both directions. 58 tests, including a 10-case round-trip suite. Clean on Windows and Linux. --- README.md | 58 ++++++--- internal/cli/root.go | 6 +- internal/sync/blocks.go | 29 +++++ internal/sync/blocks_test.go | 42 +++++++ internal/sync/checksum.go | 68 +++++++++++ internal/sync/checksum_test.go | 68 +++++++++++ internal/sync/delta.go | 170 +++++++++++++++++++++++++++ internal/sync/delta_test.go | 200 ++++++++++++++++++++++++++++++++ internal/sync/filter.go | 26 ++--- internal/sync/filter_test.go | 6 +- internal/sync/roundtrip_test.go | 94 +++++++++++++++ internal/sync/signature.go | 47 ++++++++ internal/sync/signature_test.go | 41 +++++++ internal/sync/uidgid_windows.go | 2 +- internal/sync/walk.go | 14 +-- 15 files changed, 830 insertions(+), 41 deletions(-) create mode 100644 internal/sync/blocks.go create mode 100644 internal/sync/blocks_test.go create mode 100644 internal/sync/checksum.go create mode 100644 internal/sync/checksum_test.go create mode 100644 internal/sync/delta.go create mode 100644 internal/sync/delta_test.go create mode 100644 internal/sync/roundtrip_test.go create mode 100644 internal/sync/signature.go create mode 100644 internal/sync/signature_test.go diff --git a/README.md b/README.md index 5984d2b..3322d03 100644 --- a/README.md +++ b/README.md @@ -4,11 +4,13 @@ An rsync-inspired file synchronization tool written in Go. ## Status -CLI parsing, file enumeration, and filter-rule matching are implemented; -data transfer is not. `internal/sync` builds a sorted file list -(`sync.Walk`) and can filter it (`sync.FilterEntries`), but nothing calls -either yet — the CLI only echoes parsed flags — and `internal/transport` is -still empty. +CLI parsing, file enumeration, filter-rule matching, and the delta-transfer +algorithm are implemented; nothing is wired together into an actual sync +yet. `internal/sync` can list a source tree (`sync.Walk`), filter it +(`sync.FilterEntries`), and compute/apply binary deltas between two +versions of a file (`sync.GenerateDelta`/`sync.ApplyDelta`) - but the CLI +only echoes parsed flags, and `internal/transport` is still empty, so none +of this runs end to end yet. ## Build @@ -47,7 +49,7 @@ argument is always the destination. | `--exclude-from FILE` | | read exclude patterns from FILE, one per line (repeatable) | | `--include-from FILE` | | read include patterns from FILE, one per line (repeatable) | -All five filter-related flags share one ordered rule list — their relative +All five filter-related flags share one ordered rule list - their relative order on the command line is preserved, matching rsync's first-match-wins semantics. See [Filter Rules](#filter-rules) below. @@ -65,7 +67,7 @@ target). Symlinks are captured via `Lstat`, never followed. | off | on | directories listed, not descended into | | on | any | full recursion | -On Windows, `UID`/`GID` are always `0` — there's no POSIX ownership concept +On Windows, `UID`/`GID` are always `0` - there's no POSIX ownership concept to read, so `0` means "unavailable," not a real value. ## Filter Rules @@ -79,21 +81,49 @@ matches. Pattern syntax: `*` matches within one path segment, `**` crosses segment boundaries, `?` matches one character. A trailing `/` makes a pattern match directories only. `--filter` also accepts `merge FILE` to inline another -rule file at that point in the list (one level deep — a merge file that +rule file at that point in the list (one level deep - a merge file that itself tries to merge another file is an error, not silently ignored). -A pattern anchors to the transfer root — matched once against the full -path, not tried at every depth — if it has a leading `/`, contains any +A pattern anchors to the transfer root - matched once against the full +path, not tried at every depth - if it has a leading `/`, contains any other `/`, or contains `**`. Only a pattern with none of those (a bare filename like `*.log`) matches at any depth, against the final path component only. This matches real rsync's actual anchoring rule. +## Delta-Transfer Algorithm + +`internal/sync` implements rsync's signature-based delta algorithm for +transferring a changed file without resending the parts that didn't +change: + +1. **Signature** (`sync.GenerateSignature`) - the receiver splits its copy + of the file into fixed-size blocks and computes two checksums per + block: a fast rolling checksum and an MD5 strong checksum. +2. **Delta** (`sync.GenerateDelta`) - the sender slides a window over its + new copy of the file one byte at a time, using the rolling checksum to + cheaply test every offset (not just block boundaries) for a match + against the receiver's signature; a weak-checksum hit is confirmed + against the strong checksum before being trusted, since two different + blocks can share a weak checksum by chance. The result is an ordered + list of operations: copy block N from the old file, or write these + literal bytes. +3. **Reconstruction** (`sync.ApplyDelta`) - the receiver replays that + operation list against its old copy to reproduce the sender's file + exactly. + +The block size is currently a fixed constant (`sync.DefaultBlockSize`). +Real rsync scales it dynamically based on file size; fixed-size blocks are +a deliberate simplification here, not a limitation of the algorithm +itself. + ## Architecture -- `cmd/grsync` — CLI entrypoint. -- `internal/cli` — flag/argument parsing (built on cobra). -- `internal/sync` — file-list generation and filter matching today; comparison/delta logic later. -- `internal/transport` — (placeholder) data movement, local and remote. +- `cmd/grsync` - CLI entrypoint. +- `internal/cli` - flag/argument parsing (built on cobra). +- `internal/sync` - file-list generation, filter matching, and the + delta-transfer algorithm today; wiring these together into an actual + sync comes later. +- `internal/transport` - (placeholder) data movement, local and remote. Goal: full feature parity with upstream rsync, including protocol/format interoperability where specified (e.g. batch mode's file format). diff --git a/internal/cli/root.go b/internal/cli/root.go index 47b23c0..64726f7 100644 --- a/internal/cli/root.go +++ b/internal/cli/root.go @@ -1,6 +1,6 @@ // Package cli defines the grsync command-line interface: argument parsing, // flags, and the command tree. It does not perform any sync or transport -// logic itself — it only collects options and hands them off (see the +// logic itself - it only collects options and hands them off (see the // options struct printed in Run below, which will later be passed to // internal/sync). package cli @@ -32,7 +32,7 @@ const ( // collects them the same way: Type records which flag produced the rule, // and relative order across *all* of them is preserved in the order the // user supplied them. For the two "-from" kinds, Pattern is a file path, -// not a filter pattern — internal/sync reads and expands it. +// not a filter pattern - internal/sync reads and expands it. type FilterRule struct { Type FilterRuleType Pattern string @@ -55,7 +55,7 @@ type options struct { // filterRuleFlag implements pflag.Value. Each of --exclude/--include/ // --filter/--exclude-from/--include-from gets its own instance, fixed to a -// single FilterRuleType, but all of them share the same backing slice — so +// single FilterRuleType, but all of them share the same backing slice - so // pflag's normal "call Set once per occurrence" behavior naturally builds // one ordered rule list regardless of which flag name was used at each // position. diff --git a/internal/sync/blocks.go b/internal/sync/blocks.go new file mode 100644 index 0000000..07df94e --- /dev/null +++ b/internal/sync/blocks.go @@ -0,0 +1,29 @@ +package sync + +// DefaultBlockSize is the fixed block size used to split files for the +// delta-transfer algorithm. Real rsync computes this dynamically per file +// (roughly proportional to the square root of the file size, within +// tunable bounds); using one fixed size here is a deliberate +// simplification for this ticket - dynamic block sizing is a future +// refinement, not something the algorithm itself depends on. +const DefaultBlockSize = 700 + +// splitBlocks splits data into fixed-size blocks of blockSize bytes each. +// The final block is shorter than blockSize whenever len(data) isn't an +// exact multiple of it; it's still included, never dropped or padded out +// to a full block. Returned slices share data's backing array rather than +// being copied. +func splitBlocks(data []byte, blockSize int) [][]byte { + if blockSize <= 0 { + return nil + } + var blocks [][]byte + for start := 0; start < len(data); start += blockSize { + end := start + blockSize + if end > len(data) { + end = len(data) + } + blocks = append(blocks, data[start:end]) + } + return blocks +} diff --git a/internal/sync/blocks_test.go b/internal/sync/blocks_test.go new file mode 100644 index 0000000..2383a23 --- /dev/null +++ b/internal/sync/blocks_test.go @@ -0,0 +1,42 @@ +package sync + +import "testing" + +func TestSplitBlocks(t *testing.T) { + tests := []struct { + name string + data []byte + blockSize int + wantLens []int + }{ + {"empty", []byte{}, 4, nil}, + {"exact multiple", []byte("aaaabbbbcccc"), 4, []int{4, 4, 4}}, + {"partial final block", []byte("aaaabbbbcc"), 4, []int{4, 4, 2}}, + {"smaller than one block", []byte("ab"), 4, []int{2}}, + {"single byte block size", []byte("abc"), 1, []int{1, 1, 1}}, + } + + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + blocks := splitBlocks(tt.data, tt.blockSize) + if len(blocks) != len(tt.wantLens) { + t.Fatalf("got %d blocks, want %d: %v", len(blocks), len(tt.wantLens), blocks) + } + for i, wantLen := range tt.wantLens { + if len(blocks[i]) != wantLen { + t.Errorf("block %d: len = %d, want %d", i, len(blocks[i]), wantLen) + } + } + // Reassembling every block must exactly reproduce the input - + // this is the property that actually matters (no bytes lost, + // duplicated, or reordered), not just the length list. + var got []byte + for _, b := range blocks { + got = append(got, b...) + } + if string(got) != string(tt.data) { + t.Errorf("reassembled blocks = %q, want %q", got, tt.data) + } + }) + } +} diff --git a/internal/sync/checksum.go b/internal/sync/checksum.go new file mode 100644 index 0000000..473d56e --- /dev/null +++ b/internal/sync/checksum.go @@ -0,0 +1,68 @@ +package sync + +import "crypto/md5" + +// strongChecksum is the collision-resistant checksum used to confirm a +// weak-checksum match. MD5 is not cryptographically safe against a +// deliberate adversary, but that's not what it's used for here - it's +// only there to catch the rare case where two different blocks happen to +// share a weak checksum (see delta generation), which stdlib md5 is more +// than sufficient for. +func strongChecksum(block []byte) [md5.Size]byte { + return md5.Sum(block) +} + +// rollingChecksumModulus is 65536 (2^16) - a power of two, not a prime. +// Real Adler-32 uses the largest prime below 65536 (65521) instead; this +// is rsync's own simpler variant, chosen deliberately because a +// power-of-two modulus makes unsigned-integer wraparound during roll() +// mathematically safe (see the comment there) without extra bounds +// handling. It rolls in O(1), which is the only property that actually +// matters here - this is not meant to be byte-compatible with stdlib +// hash/adler32. +const rollingChecksumModulus = 1 << 16 + +// weakChecksum is a rolling checksum over a fixed-size window: two 16-bit +// accumulators (a: sum of the window's bytes, b: a position-weighted sum) +// combined into one 32-bit value via sum(). +type weakChecksum struct { + a, b uint32 + length uint32 // window size; constant across every roll() call +} + +// newWeakChecksum computes the checksum for window from scratch in O(len(window)). +func newWeakChecksum(window []byte) weakChecksum { + var a, b uint32 + n := uint32(len(window)) + for i, c := range window { + a += uint32(c) + b += (n - uint32(i)) * uint32(c) + } + return weakChecksum{ + a: a % rollingChecksumModulus, + b: b % rollingChecksumModulus, + length: n, + } +} + +// sum returns the combined checksum value. +func (w weakChecksum) sum() uint32 { + return w.a + w.b*rollingChecksumModulus +} + +// roll advances the window by exactly one byte: out is the byte leaving +// at the window's start, in is the byte entering at its end. This is O(1) +// regardless of window size - the entire point of a rolling checksum, +// versus calling newWeakChecksum on the shifted window from scratch. +// +// The subtractions below can underflow as uint32 arithmetic (e.g. if +// w.a < out). That's fine, not a bug: Go's unsigned integers wrap modulo +// 2^32, and since rollingChecksumModulus (2^16) evenly divides 2^32, the +// wrapped value still reduces to the mathematically correct result mod +// 2^16 after the final "% rollingChecksumModulus". A prime modulus (real +// Adler-32's 65521) would not have this property. +func (w weakChecksum) roll(out, in byte) weakChecksum { + a := (w.a - uint32(out) + uint32(in)) % rollingChecksumModulus + b := (w.b - w.length*uint32(out) + a) % rollingChecksumModulus + return weakChecksum{a: a, b: b, length: w.length} +} diff --git a/internal/sync/checksum_test.go b/internal/sync/checksum_test.go new file mode 100644 index 0000000..b9816b2 --- /dev/null +++ b/internal/sync/checksum_test.go @@ -0,0 +1,68 @@ +package sync + +import ( + "crypto/md5" + "testing" +) + +func TestStrongChecksum(t *testing.T) { + block := []byte("some block contents") + got := strongChecksum(block) + want := md5.Sum(block) + if got != want { + t.Errorf("strongChecksum(%q) = %x, want %x", block, got, want) + } + + if strongChecksum([]byte("a")) == strongChecksum([]byte("b")) { + t.Errorf("different blocks produced the same strong checksum") + } +} + +// TestWeakChecksum_RollMatchesFromScratch is the load-bearing test for the +// entire rolling-checksum design: it proves roll() produces exactly the +// same result as recomputing from scratch, at every single offset across +// a test string, not just a couple of spot checks. If this property ever +// broke, the delta algorithm would silently miss real block matches (or +// worse, produce false ones that only got caught by the strong checksum, +// masking the bug) - so this needs direct proof, not just "it compiles". +func TestWeakChecksum_RollMatchesFromScratch(t *testing.T) { + data := []byte("the quick brown fox jumps over the lazy dog, then jumps back again") + const windowSize = 8 + + if len(data) <= windowSize { + t.Fatalf("test data too short: need more than %d bytes", windowSize) + } + + current := newWeakChecksum(data[:windowSize]) + if want := newWeakChecksum(data[0:windowSize]); current.sum() != want.sum() { + t.Fatalf("offset 0: sum = %d, want %d", current.sum(), want.sum()) + } + + for offset := 1; offset+windowSize <= len(data); offset++ { + out := data[offset-1] + in := data[offset+windowSize-1] + current = current.roll(out, in) + + want := newWeakChecksum(data[offset : offset+windowSize]) + if current.sum() != want.sum() { + t.Fatalf("offset %d: rolled sum = %d, want %d (from scratch) - a=%d/%d b=%d/%d", + offset, current.sum(), want.sum(), current.a, want.a, current.b, want.b) + } + } +} + +func TestWeakChecksum_IdenticalWindowsMatch(t *testing.T) { + a := newWeakChecksum([]byte("abcdefgh")) + b := newWeakChecksum([]byte("abcdefgh")) + if a.sum() != b.sum() { + t.Errorf("identical windows produced different sums: %d vs %d", a.sum(), b.sum()) + } +} + +func TestWeakChecksum_DifferentWindowsUsuallyDiffer(t *testing.T) { + a := newWeakChecksum([]byte("abcdefgh")) + b := newWeakChecksum([]byte("hgfedcba")) + if a.sum() == b.sum() { + t.Errorf("reversed window produced the same sum (%d) as the original - weak checksum isn't discriminating position", a.sum()) + } +} diff --git a/internal/sync/delta.go b/internal/sync/delta.go new file mode 100644 index 0000000..2a83af2 --- /dev/null +++ b/internal/sync/delta.go @@ -0,0 +1,170 @@ +package sync + +import "fmt" + +// DeltaOp is one operation in a delta stream, produced by GenerateDelta +// and consumed by ApplyDelta. It's a sealed interface - isDeltaOp is +// unexported, so CopyOp and DataOp are its only implementations; callers +// type-switch on the concrete type. +type DeltaOp interface { + isDeltaOp() +} + +// CopyOp copies block BlockIndex - an index into the Blocks of the +// Signature that GenerateDelta was given - from the receiver's old file, +// unchanged. +type CopyOp struct { + BlockIndex int +} + +func (CopyOp) isDeltaOp() {} + +// DataOp writes Bytes literally: data present in the new file that didn't +// match any block in the old file's signature. +type DataOp struct { + Bytes []byte +} + +func (DataOp) isDeltaOp() {} + +// GenerateDelta compares newData against sig - a signature of some old +// data the receiver already has - and produces an ordered delta that, +// applied to that old data via ApplyDelta, reconstructs newData. +// +// It slides a blockSize-byte window across newData one byte at a time, +// maintaining the rolling weak checksum incrementally (weakChecksum.roll, +// O(1) per byte) rather than recomputing it from scratch at every +// position - recomputing would silently make this an O(n*blockSize) +// scan, defeating the entire reason a rolling checksum exists. +func GenerateDelta(sig Signature, newData []byte) []DeltaOp { + blockSize := sig.BlockSize + if blockSize <= 0 { + blockSize = DefaultBlockSize + } + + // weak checksum -> indices of every block sharing it. A weak checksum + // collision (two different blocks that happen to produce the same + // 32-bit weak sum) is expected to happen occasionally by chance; the + // slice of candidates lets the strong-checksum check below disambiguate + // rather than assuming the first weak match is correct. + weakIndex := make(map[uint32][]int, len(sig.Blocks)) + for i, b := range sig.Blocks { + weakIndex[b.Weak] = append(weakIndex[b.Weak], i) + } + + var ops []DeltaOp + var pending []byte // literal bytes seen so far that haven't matched a block yet + + flushPending := func() { + if len(pending) == 0 { + return + } + ops = append(ops, DataOp{Bytes: pending}) + // Reset to nil (not just len 0) so the next append starts a fresh + // backing array instead of potentially growing into - and + // corrupting - the slice just handed to the DataOp above. + pending = nil + } + + // tryMatch checks whether the blockSize-byte window at newData[at:] - + // whose already-computed rolling checksum is weak - matches a + // signature block. Confirms via strong checksum before accepting. + tryMatch := func(at int, weak weakChecksum) (blockIndex int, ok bool) { + candidates, found := weakIndex[weak.sum()] + if !found { + return 0, false + } + strong := strongChecksum(newData[at : at+blockSize]) + for _, idx := range candidates { + if sig.Blocks[idx].Strong == strong { + return idx, true + } + } + return 0, false + } + + n := len(newData) + pos := 0 + for pos < n { + if pos+blockSize > n { + // Fewer than blockSize bytes remain: no full window left to + // possibly match, so the rest of the file is literal data. + // (No need to advance pos before this break - nothing reads + // it again once the loop exits.) + pending = append(pending, newData[pos:]...) + break + } + + // Freshly computed only here - once per match/skip-ahead, not per + // byte - then advanced with roll() for every subsequent byte the + // inner loop steps through without a match. + weak := newWeakChecksum(newData[pos : pos+blockSize]) + for { + if idx, ok := tryMatch(pos, weak); ok { + flushPending() + ops = append(ops, CopyOp{BlockIndex: idx}) + pos += blockSize // skip past the whole matched block, not just one byte + break + } + + pending = append(pending, newData[pos]) + pos++ + if pos+blockSize > n { + break // not enough bytes left for a full window anymore + } + weak = weak.roll(newData[pos-1], newData[pos+blockSize-1]) + } + } + + flushPending() + return ops +} + +// ApplyDelta reconstructs a file from oldData and an ordered []DeltaOp +// produced by GenerateDelta(sig, ...) against that same oldData: each +// CopyOp copies its referenced block out of oldData, and each DataOp +// writes its literal bytes, in order. +// +// sig must be the exact Signature GenerateDelta was called with - a +// CopyOp only carries a block index, not byte offsets, so BlockSize is +// the only way to recover which bytes of oldData that index refers to. +// Passing a different Signature (or a hand-built one with a mismatched +// BlockSize) would silently reconstruct the wrong bytes; that's exactly +// the class of bug bundling BlockSize into Signature (see signature.go) +// is meant to make harder, by giving ApplyDelta and GenerateDelta the +// same single source of truth for it instead of two independent +// parameters that could drift apart. +// +// The result is returned as a []byte rather than written to an io.Writer: +// every function in this package so far (Walk aside) operates on +// in-memory byte slices - there's no streaming I/O anywhere yet for this +// to plug into - so an io.Writer parameter would just be unused +// flexibility at this stage. That can change if/when a streaming +// transport is introduced later. +func ApplyDelta(oldData []byte, ops []DeltaOp, sig Signature) ([]byte, error) { + blockSize := sig.BlockSize + if blockSize <= 0 { + return nil, fmt.Errorf("invalid signature block size %d", blockSize) + } + + var out []byte + for i, op := range ops { + switch o := op.(type) { + case CopyOp: + start := o.BlockIndex * blockSize + if o.BlockIndex < 0 || start >= len(oldData) { + return nil, fmt.Errorf("op %d: CopyOp block index %d is out of range for a %d-byte old file", i, o.BlockIndex, len(oldData)) + } + end := start + blockSize + if end > len(oldData) { + end = len(oldData) // the final block may be shorter than blockSize + } + out = append(out, oldData[start:end]...) + case DataOp: + out = append(out, o.Bytes...) + default: + return nil, fmt.Errorf("op %d: unknown DeltaOp type %T", i, op) + } + } + return out, nil +} diff --git a/internal/sync/delta_test.go b/internal/sync/delta_test.go new file mode 100644 index 0000000..f993cdb --- /dev/null +++ b/internal/sync/delta_test.go @@ -0,0 +1,200 @@ +package sync + +import ( + "strings" + "testing" +) + +func countOps(ops []DeltaOp) (copies, data int) { + for _, op := range ops { + switch op.(type) { + case CopyOp: + copies++ + case DataOp: + data++ + } + } + return copies, data +} + +func TestApplyDelta_OutOfRangeBlockIndexErrors(t *testing.T) { + old := []byte("AAAABBBB") // 2 blocks of 4 + sig := GenerateSignatureWithBlockSize(old, 4) + + _, err := ApplyDelta(old, []DeltaOp{CopyOp{BlockIndex: 5}}, sig) + if err == nil { + t.Fatalf("ApplyDelta with an out-of-range CopyOp index returned nil error, want an error") + } + + _, err = ApplyDelta(old, []DeltaOp{CopyOp{BlockIndex: -1}}, sig) + if err == nil { + t.Fatalf("ApplyDelta with a negative CopyOp index returned nil error, want an error") + } +} + +func TestApplyDelta_InvalidBlockSizeErrors(t *testing.T) { + _, err := ApplyDelta([]byte("data"), nil, Signature{BlockSize: 0}) + if err == nil { + t.Fatalf("ApplyDelta with a zero block size returned nil error, want an error") + } +} + +// TestGenerateDelta_WeakChecksumCollisionDisambiguatedByStrongChecksum +// exercises the actual weak-checksum-collision path, rather than assuming +// the strong-checksum disambiguation code is correct because it looks +// right. block1, block2, and block3 below are three deliberately +// constructed, genuinely different 3-byte sequences that all produce the +// identical weak checksum (a=60, b=100 - hand-verified against +// newWeakChecksum's exact weighting convention, no modular wraparound +// needed): with weight (n-i) for byte i in a window of n=3, all three +// satisfy sum=60 and 3x+2y+z=100 simultaneously by construction. +// +// If GenerateDelta ever regressed to trusting a weak-checksum match +// without confirming it via strong checksum, this test would silently +// start producing corrupted output (a CopyOp pointing at the wrong +// block) - exactly the class of bug a weak checksum alone can't catch, +// which is the entire reason a strong checksum exists. +func TestGenerateDelta_WeakChecksumCollisionDisambiguatedByStrongChecksum(t *testing.T) { + block1 := []byte{10, 20, 30} + block2 := []byte{15, 10, 35} + block3 := []byte{12, 16, 32} // present in neither signature block + + if newWeakChecksum(block1).sum() != newWeakChecksum(block2).sum() { + t.Fatalf("test setup invalid: block1 and block2 don't actually collide") + } + if newWeakChecksum(block1).sum() != newWeakChecksum(block3).sum() { + t.Fatalf("test setup invalid: block1 and block3 don't actually collide") + } + + old := append(append([]byte{}, block1...), block2...) // 6 bytes: block index 0 = block1, index 1 = block2 + sig := GenerateSignatureWithBlockSize(old, 3) + if sig.Blocks[0].Weak != sig.Blocks[1].Weak { + t.Fatalf("test setup invalid: signature blocks don't share a weak checksum") + } + if sig.Blocks[0].Strong == sig.Blocks[1].Strong { + t.Fatalf("test setup invalid: signature blocks unexpectedly share a strong checksum too") + } + + t.Run("matches block2, not block1, despite the weak collision", func(t *testing.T) { + ops := GenerateDelta(sig, block2) + if len(ops) != 1 { + t.Fatalf("got %d ops, want 1: %+v", len(ops), ops) + } + cp, ok := ops[0].(CopyOp) + if !ok || cp.BlockIndex != 1 { + t.Errorf("op = %+v, want CopyOp{BlockIndex: 1}", ops[0]) + } + }) + + t.Run("matches block1, not block2, despite the weak collision", func(t *testing.T) { + ops := GenerateDelta(sig, block1) + if len(ops) != 1 { + t.Fatalf("got %d ops, want 1: %+v", len(ops), ops) + } + cp, ok := ops[0].(CopyOp) + if !ok || cp.BlockIndex != 0 { + t.Errorf("op = %+v, want CopyOp{BlockIndex: 0}", ops[0]) + } + }) + + t.Run("rejects a weak match when no candidate's strong checksum matches", func(t *testing.T) { + ops := GenerateDelta(sig, block3) + if len(ops) != 1 { + t.Fatalf("got %d ops, want 1: %+v", len(ops), ops) + } + d, ok := ops[0].(DataOp) + if !ok || string(d.Bytes) != string(block3) { + t.Errorf("op = %+v, want DataOp{Bytes: %v} - a weak-only match must not produce a CopyOp", ops[0], block3) + } + }) +} + +func TestGenerateDelta_IdenticalFileIsAllCopies(t *testing.T) { + data := []byte("AAAABBBBCCCCDDDD") // 16 bytes, 4 blocks of 4 + sig := GenerateSignatureWithBlockSize(data, 4) + + ops := GenerateDelta(sig, data) + + copies, dataOps := countOps(ops) + if dataOps != 0 { + t.Errorf("got %d DataOps for an identical file, want 0: %+v", dataOps, ops) + } + if copies != 4 { + t.Errorf("got %d CopyOps, want 4 (one per block): %+v", copies, ops) + } + // Order matters too: block 0 first, then 1, 2, 3. + for i, op := range ops { + cp, ok := op.(CopyOp) + if !ok || cp.BlockIndex != i { + t.Errorf("op %d = %+v, want CopyOp{BlockIndex: %d}", i, op, i) + } + } +} + +func TestGenerateDelta_CompletelyDifferentFileIsAllData(t *testing.T) { + old := []byte("AAAABBBBCCCCDDDD") + sig := GenerateSignatureWithBlockSize(old, 4) + + newData := []byte("wxyz1234!@#$%^&*") // shares no blocks with old + ops := GenerateDelta(sig, newData) + + copies, dataOps := countOps(ops) + if copies != 0 { + t.Errorf("got %d CopyOps for a completely different file, want 0: %+v", copies, ops) + } + if dataOps == 0 { + t.Fatalf("got 0 DataOps, want at least 1") + } + + // Reassembling every DataOp's bytes must reproduce newData exactly, + // since nothing was copied. + var got []byte + for _, op := range ops { + got = append(got, op.(DataOp).Bytes...) + } + if string(got) != string(newData) { + t.Errorf("reassembled data = %q, want %q", got, newData) + } +} + +func TestGenerateDelta_SingleByteChangeInMiddle(t *testing.T) { + old := []byte(strings.Repeat("0123456789", 5)) // 50 bytes, block size 10 -> 5 blocks + sig := GenerateSignatureWithBlockSize(old, 10) + + newData := []byte(strings.Repeat("0123456789", 5)) + newData[25] = 'X' // change one byte inside block index 2 (bytes 20-29) + + ops := GenerateDelta(sig, newData) + + copies, dataOps := countOps(ops) + if copies == 0 { + t.Errorf("got 0 CopyOps for a mostly-unchanged file, want several") + } + if dataOps == 0 { + t.Fatalf("got 0 DataOps despite a changed byte, want at least 1") + } + + // The whole point of the algorithm: a one-byte change should cost a + // small, bounded amount of literal data, not force the whole file (or + // even the whole surrounding block) to be retransmitted as data. + var totalDataBytes int + for _, op := range ops { + if d, ok := op.(DataOp); ok { + totalDataBytes += len(d.Bytes) + } + } + if totalDataBytes > 20 { + t.Errorf("total literal bytes = %d, want a small bounded amount for a single changed byte", totalDataBytes) + } + + // Reconstructing from ops must still reproduce newData exactly - the + // strongest check that "mostly CopyOps, one small DataOp" is actually + // correct, not just small. + reconstructed, err := ApplyDelta(old, ops, sig) + if err != nil { + t.Fatalf("ApplyDelta returned error: %v", err) + } + if string(reconstructed) != string(newData) { + t.Errorf("reconstructed = %q, want %q", reconstructed, newData) + } +} diff --git a/internal/sync/filter.go b/internal/sync/filter.go index 30e2415..408514b 100644 --- a/internal/sync/filter.go +++ b/internal/sync/filter.go @@ -21,7 +21,7 @@ const ( // RuleExclude is a direct --exclude pattern. RuleExclude RuleKind = "exclude" // RuleFilter is a raw --filter rule line, e.g. "+ *.txt", "- .git/", - // or "merge FILE" — see parseFilterLine for the subset of rsync's + // or "merge FILE" - see parseFilterLine for the subset of rsync's // filter-rule syntax this supports. RuleFilter RuleKind = "filter" // RuleExcludeFrom has a Pattern that is a file path, not a filter @@ -57,13 +57,13 @@ const ( // // Anchored matches real rsync's actual rule: a pattern anchors to the // transfer root if it has a leading "/", contains any other "/", or -// contains "**" — only a pattern with none of those (a bare filename, e.g. +// contains "**" - only a pattern with none of those (a bare filename, e.g. // "*.log") matches at any depth, against the final path component only. type Rule struct { Action Action Pattern string Anchored bool - // DirOnly means this rule only ever matches directories — set from a + // DirOnly means this rule only ever matches directories - set from a // trailing "/" on the original pattern, stripped by CompileRules just // like the anchor marker. DirOnly bool @@ -129,8 +129,8 @@ func matchSegments(patternSegs, pathSegs []string) bool { // action. Shared by direct --include/--exclude rules, each line read from // an --exclude-from/--include-from file, and --filter rule lines. // -// A pattern containing any empty "/"-separated segment — a bare "/" or "" -// overall, or an internal "//" typo like "a//b" — is rejected rather than +// A pattern containing any empty "/"-separated segment - a bare "/" or "" +// overall, or an internal "//" typo like "a//b" - is rejected rather than // silently compiled: no real FileEntry.Path segment is ever empty (Walk() // never produces one), so a Rule requiring an empty segment could never // match anything. Compiling it anyway would leave the user with a filter @@ -152,7 +152,7 @@ func compilePattern(action Action, pattern string) (Rule, error) { } // A leading "/" always anchors. So does any *other* "/" still present - // in the pattern, or a "**" anywhere in it — matching real rsync's + // in the pattern, or a "**" anywhere in it - matching real rsync's // rule, not just the "leading slash only" simplification this started // as. Only a genuinely slash-free, "**"-free pattern (a bare filename) // matches at any depth. @@ -190,7 +190,7 @@ func readPatternFile(path string) (patterns []string, err error) { } // parsedFilterLine is the result of parsing one line of --filter RULE -// syntax — whether it came directly from a --filter flag or from a line +// syntax - whether it came directly from a --filter flag or from a line // inside a merge file. type parsedFilterLine struct { isMerge bool @@ -203,7 +203,7 @@ type parsedFilterLine struct { // equivalents "include PATTERN" / "exclude PATTERN"), plus "merge FILE". // Real rsync's filter language is much larger (modifiers like "-C", // "dir-merge" with per-directory semantics, "!", exclude-if-present rules, -// and more) — none of that is implemented here; anything outside this +// and more) - none of that is implemented here; anything outside this // subset is a hard parse error rather than a silent no-op. func parseFilterLine(line string) (parsedFilterLine, error) { switch { @@ -229,7 +229,7 @@ func parseFilterLine(line string) (parsedFilterLine, error) { } // expandMergeFile reads path and parses each non-comment, non-blank line -// with parseFilterLine — the same syntax --filter itself accepts. Nested +// with parseFilterLine - the same syntax --filter itself accepts. Nested // merge (a merge file whose own lines include another "merge OTHER" // directive) is deliberately unsupported: rather than silently recursing // (risking an infinite loop on a self-referential merge file) or silently @@ -256,9 +256,9 @@ func expandMergeFile(path string) ([]Rule, error) { return rules, nil } -// CompileRules converts raw rules — --include/--exclude patterns, +// CompileRules converts raw rules - --include/--exclude patterns, // --filter rule lines (including "merge FILE"), and --exclude-from/ -// --include-from file references — into a single ordered, ready-to-match +// --include-from file references - into a single ordered, ready-to-match // Rule list. From-file and merged rules are expanded in place at the // position their flag occurred, so e.g. a --exclude-from sandwiched // between two direct --exclude flags on the command line stays sandwiched @@ -319,7 +319,7 @@ func CompileRules(raw []RawRule) ([]Rule, error) { // Included evaluates rules against a single FileEntry using rsync's // first-match-wins semantics: rules are tried in order, and the action of // the first one that matches decides the outcome. If no rule matches, the -// entry is included by default — matching rsync's own default behavior of +// entry is included by default - matching rsync's own default behavior of // transferring anything not explicitly excluded. // // entry.Path is never empty and never represents the transfer root itself: @@ -337,7 +337,7 @@ func Included(rules []Rule, entry FileEntry) bool { // FilterEntries returns the subset of entries that rules includes, in // their original order. It is a post-pass over an already-collected -// Walk() result rather than a predicate threaded into Walk() itself — see +// Walk() result rather than a predicate threaded into Walk() itself - see // the design note on this tradeoff where FilterEntries is introduced in // the accompanying documentation/commit. func FilterEntries(entries []FileEntry, rules []Rule) []FileEntry { diff --git a/internal/sync/filter_test.go b/internal/sync/filter_test.go index 9fd4ef2..787f4dd 100644 --- a/internal/sync/filter_test.go +++ b/internal/sync/filter_test.go @@ -71,7 +71,7 @@ func TestCompileRules_Anchoring(t *testing.T) { func TestCompileRules_InternalSlashAnchorsWithoutLeadingSlash(t *testing.T) { // Real rsync's actual rule: a pattern anchors to the root if it has a - // leading "/", OR contains any other "/", OR contains "**" — not just + // leading "/", OR contains any other "/", OR contains "**" - not just // on an explicit leading "/". "src/main.go" (no leading slash, but an // internal one) must behave the same as "/src/main.go" here. rules := mustCompile(t, []RawRule{ @@ -165,7 +165,7 @@ func TestIncluded_FirstMatchWinsOrderMatters(t *testing.T) { // Same two rules, opposite order: the first one to match should win in // both cases, so swapping the order must flip the outcome. If it - // didn't, evaluation wouldn't actually be "first match wins" — it'd be + // didn't, evaluation wouldn't actually be "first match wins" - it'd be // "last match wins" or "most specific wins" or something else. includeFirst := mustCompile(t, []RawRule{ {Kind: RuleInclude, Pattern: "keep.log"}, @@ -223,7 +223,7 @@ func TestCompileRules_ExcludeFromPreservesPosition(t *testing.T) { // Position matters here: the two patterns read from the file must land // between "first.txt" and "last.txt", not get appended after - // "last.txt" — that would silently reorder rules relative to what the + // "last.txt" - that would silently reorder rules relative to what the // user typed on the command line, breaking first-match-wins semantics. want := []struct { action Action diff --git a/internal/sync/roundtrip_test.go b/internal/sync/roundtrip_test.go new file mode 100644 index 0000000..0b48f81 --- /dev/null +++ b/internal/sync/roundtrip_test.go @@ -0,0 +1,94 @@ +package sync + +import ( + "strings" + "testing" +) + +// TestRoundTrip runs the full receiver/sender/receiver cycle for each +// case: generate a Signature from the old file, generate a delta from +// (old, new) against that signature, apply the delta to the old file, and +// confirm the result is byte-for-byte identical to the new file. This is +// the property the whole algorithm exists to guarantee - none of the +// individual-step tests elsewhere in this package substitute for actually +// proving the full cycle reproduces the target file exactly. +func TestRoundTrip(t *testing.T) { + const blockSize = 8 + + base := strings.Repeat("0123456789abcdef", 4) // 64 bytes, 8 blocks + + tests := []struct { + name string + old []byte + new []byte + }{ + { + name: "identical files", + old: []byte(base), + new: []byte(base), + }, + { + name: "appended bytes", + old: []byte(base), + new: []byte(base + "EXTRA-DATA-AT-THE-END"), + }, + { + name: "prepended bytes", + old: []byte(base), + new: []byte("EXTRA-DATA-AT-THE-START" + base), + }, + { + name: "single byte change in the middle", + old: []byte(base), + new: func() []byte { + b := []byte(base) + b[32] = 'X' + return b + }(), + }, + { + name: "completely different files", + old: []byte(base), + new: []byte(strings.Repeat("!@#$%^&*()_+-=[]{}", 4)), + }, + { + name: "old file empty", + old: []byte{}, + new: []byte(base), + }, + { + name: "new file empty", + old: []byte(base), + new: []byte{}, + }, + { + name: "both files empty", + old: []byte{}, + new: []byte{}, + }, + { + name: "old file smaller than one block", + old: []byte("ab"), + new: []byte(base), + }, + { + name: "new file smaller than one block", + old: []byte(base), + new: []byte("xy"), + }, + } + + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + sig := GenerateSignatureWithBlockSize(tt.old, blockSize) + ops := GenerateDelta(sig, tt.new) + got, err := ApplyDelta(tt.old, ops, sig) + if err != nil { + t.Fatalf("ApplyDelta returned error: %v", err) + } + if string(got) != string(tt.new) { + t.Errorf("reconstructed = %q, want %q", got, tt.new) + } + }) + } +} diff --git a/internal/sync/signature.go b/internal/sync/signature.go new file mode 100644 index 0000000..88eb7b3 --- /dev/null +++ b/internal/sync/signature.go @@ -0,0 +1,47 @@ +package sync + +import "crypto/md5" + +// BlockSignature is the pair of checksums computed for one block, sent +// from receiver to sender so the sender can find matching blocks in its +// own copy of the file without transferring the block's bytes. +type BlockSignature struct { + Weak uint32 + Strong [md5.Size]byte +} + +// Signature is an ordered list of per-block checksums for a file, plus +// the block size used to produce them. BlockSize travels with the +// checksums (rather than being a separate parameter callers must keep in +// sync) because delta generation and reconstruction both need to know +// exactly how the blocks were cut - a mismatched block size would make +// every checksum meaningless. A block's position in Blocks is its index, +// which is how a CopyOp (see delta.go) refers back to it. +type Signature struct { + BlockSize int + Blocks []BlockSignature +} + +// GenerateSignature splits data into DefaultBlockSize blocks and computes +// both checksums for each one, in order. +func GenerateSignature(data []byte) Signature { + return GenerateSignatureWithBlockSize(data, DefaultBlockSize) +} + +// GenerateSignatureWithBlockSize is GenerateSignature with an explicit +// block size - mainly so tests can exercise multi-block behavior with +// small fixtures instead of needing kilobyte-sized test data. +func GenerateSignatureWithBlockSize(data []byte, blockSize int) Signature { + blocks := splitBlocks(data, blockSize) + sig := Signature{ + BlockSize: blockSize, + Blocks: make([]BlockSignature, len(blocks)), + } + for i, block := range blocks { + sig.Blocks[i] = BlockSignature{ + Weak: newWeakChecksum(block).sum(), + Strong: strongChecksum(block), + } + } + return sig +} diff --git a/internal/sync/signature_test.go b/internal/sync/signature_test.go new file mode 100644 index 0000000..1181217 --- /dev/null +++ b/internal/sync/signature_test.go @@ -0,0 +1,41 @@ +package sync + +import "testing" + +func TestGenerateSignature_BlockCountAndChecksums(t *testing.T) { + data := []byte("AAAABBBBCC") // 10 bytes, block size 4 -> blocks "AAAA","BBBB","CC" + sig := GenerateSignatureWithBlockSize(data, 4) + + if sig.BlockSize != 4 { + t.Errorf("BlockSize = %d, want 4", sig.BlockSize) + } + if len(sig.Blocks) != 3 { + t.Fatalf("got %d blocks, want 3: %+v", len(sig.Blocks), sig.Blocks) + } + + wantBlocks := [][]byte{[]byte("AAAA"), []byte("BBBB"), []byte("CC")} + for i, wantBlock := range wantBlocks { + wantWeak := newWeakChecksum(wantBlock).sum() + wantStrong := strongChecksum(wantBlock) + if sig.Blocks[i].Weak != wantWeak { + t.Errorf("block %d: Weak = %d, want %d", i, sig.Blocks[i].Weak, wantWeak) + } + if sig.Blocks[i].Strong != wantStrong { + t.Errorf("block %d: Strong = %x, want %x", i, sig.Blocks[i].Strong, wantStrong) + } + } +} + +func TestGenerateSignature_EmptyData(t *testing.T) { + sig := GenerateSignatureWithBlockSize(nil, DefaultBlockSize) + if len(sig.Blocks) != 0 { + t.Errorf("got %d blocks for empty input, want 0", len(sig.Blocks)) + } +} + +func TestGenerateSignature_DefaultBlockSize(t *testing.T) { + sig := GenerateSignature([]byte("some data")) + if sig.BlockSize != DefaultBlockSize { + t.Errorf("BlockSize = %d, want DefaultBlockSize (%d)", sig.BlockSize, DefaultBlockSize) + } +} diff --git a/internal/sync/uidgid_windows.go b/internal/sync/uidgid_windows.go index afa54c7..4a5158c 100644 --- a/internal/sync/uidgid_windows.go +++ b/internal/sync/uidgid_windows.go @@ -7,7 +7,7 @@ import "os" // lookupUIDGID always reports unavailable on Windows: os.FileInfo.Sys() // there returns *syscall.Win32FileAttributeData, which carries no POSIX // uid/gid concept at all (Windows uses SIDs and ACLs instead). Callers must -// treat ok == false as "not populated", not as "owned by uid/gid 0" — +// treat ok == false as "not populated", not as "owned by uid/gid 0" - // zero is not a meaningful default here. func lookupUIDGID(_ os.FileInfo) (uid, gid uint32, ok bool) { return 0, 0, false diff --git a/internal/sync/walk.go b/internal/sync/walk.go index 5574b79..a30de4c 100644 --- a/internal/sync/walk.go +++ b/internal/sync/walk.go @@ -1,5 +1,5 @@ // Package sync builds the file list ("flist") that grsync compares between -// source and destination before any data transfer happens — the same +// source and destination before any data transfer happens - the same // planning phase upstream rsync performs before its delta algorithm runs. package sync @@ -13,7 +13,7 @@ import ( // FileEntry describes a single file, directory, or symlink discovered under // a source root. Path is always relative to that root, using "/" as the -// separator regardless of host OS — rsync's wire protocol and file lists +// separator regardless of host OS - rsync's wire protocol and file lists // are always "/"-separated, and grsync targets protocol-level // interoperability, so paths are normalized at collection time rather than // left OS-native and converted later. @@ -26,7 +26,7 @@ type FileEntry struct { GID uint32 // OwnershipAvailable reports whether UID/GID were actually populated. // A real uid/gid of 0 (root) is a valid value, so callers must check - // this rather than treating a zero UID/GID as "unavailable" — see + // this rather than treating a zero UID/GID as "unavailable" - see // uidgid_windows.go, where it is always false. OwnershipAvailable bool LinkTarget string @@ -37,13 +37,13 @@ type FileEntry struct { // -r/--recursive and -d/--dirs flags: // // - Recursive=false, Dirs=false (rsync's default with neither flag): -// directories are skipped entirely — not listed, not descended into. +// directories are skipped entirely - not listed, not descended into. // Only regular files/symlinks directly under root are collected. // - Recursive=false, Dirs=true (-d): directories are listed (so they can -// be created on the receiving end) but their contents are not — Walk +// be created on the receiving end) but their contents are not - Walk // does not descend into them. // - Recursive=true (-r): full recursion into every subdirectory, -// regardless of Dirs — this matches rsync, where -r makes -d redundant. +// regardless of Dirs - this matches rsync, where -r makes -d redundant. type WalkOptions struct { Recursive bool Dirs bool @@ -99,7 +99,7 @@ func Walk(root string, opts WalkOptions) ([]FileEntry, error) { // uidgid_windows.go): on Windows it always reports unavailable, // leaving UID/GID at their zero value. OwnershipAvailable carries // that ok flag through so callers can't mistake the zero value for - // a real uid/gid of 0 (root) — see uidgid_windows.go for why. + // a real uid/gid of 0 (root) - see uidgid_windows.go for why. entry.UID, entry.GID, entry.OwnershipAvailable = lookupUIDGID(info) // info.Mode()&fs.ModeSymlink is only ever set by Lstat (Stat