Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
58 changes: 44 additions & 14 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,11 +4,13 @@ An rsync-inspired file synchronization tool written in Go.

## Status

CLI parsing, file enumeration, and filter-rule matching are implemented;
data transfer is not. `internal/sync` builds a sorted file list
(`sync.Walk`) and can filter it (`sync.FilterEntries`), but nothing calls
either yet — the CLI only echoes parsed flags — and `internal/transport` is
still empty.
CLI parsing, file enumeration, filter-rule matching, and the delta-transfer
algorithm are implemented; nothing is wired together into an actual sync
yet. `internal/sync` can list a source tree (`sync.Walk`), filter it
(`sync.FilterEntries`), and compute/apply binary deltas between two
versions of a file (`sync.GenerateDelta`/`sync.ApplyDelta`) - but the CLI
only echoes parsed flags, and `internal/transport` is still empty, so none
of this runs end to end yet.

## Build

Expand Down Expand Up @@ -47,7 +49,7 @@ argument is always the destination.
| `--exclude-from FILE` | | read exclude patterns from FILE, one per line (repeatable) |
| `--include-from FILE` | | read include patterns from FILE, one per line (repeatable) |

All five filter-related flags share one ordered rule list — their relative
All five filter-related flags share one ordered rule list - their relative
order on the command line is preserved, matching rsync's first-match-wins
semantics. See [Filter Rules](#filter-rules) below.

Expand All @@ -65,7 +67,7 @@ target). Symlinks are captured via `Lstat`, never followed.
| off | on | directories listed, not descended into |
| on | any | full recursion |

On Windows, `UID`/`GID` are always `0` — there's no POSIX ownership concept
On Windows, `UID`/`GID` are always `0` - there's no POSIX ownership concept
to read, so `0` means "unavailable," not a real value.

## Filter Rules
Expand All @@ -79,21 +81,49 @@ matches.
Pattern syntax: `*` matches within one path segment, `**` crosses segment
boundaries, `?` matches one character. A trailing `/` makes a pattern match
directories only. `--filter` also accepts `merge FILE` to inline another
rule file at that point in the list (one level deep — a merge file that
rule file at that point in the list (one level deep - a merge file that
itself tries to merge another file is an error, not silently ignored).

A pattern anchors to the transfer root — matched once against the full
path, not tried at every depth — if it has a leading `/`, contains any
A pattern anchors to the transfer root - matched once against the full
path, not tried at every depth - if it has a leading `/`, contains any
other `/`, or contains `**`. Only a pattern with none of those (a bare
filename like `*.log`) matches at any depth, against the final path
component only. This matches real rsync's actual anchoring rule.

## Delta-Transfer Algorithm

`internal/sync` implements rsync's signature-based delta algorithm for
transferring a changed file without resending the parts that didn't
change:

1. **Signature** (`sync.GenerateSignature`) - the receiver splits its copy
of the file into fixed-size blocks and computes two checksums per
block: a fast rolling checksum and an MD5 strong checksum.
2. **Delta** (`sync.GenerateDelta`) - the sender slides a window over its
new copy of the file one byte at a time, using the rolling checksum to
cheaply test every offset (not just block boundaries) for a match
against the receiver's signature; a weak-checksum hit is confirmed
against the strong checksum before being trusted, since two different
blocks can share a weak checksum by chance. The result is an ordered
list of operations: copy block N from the old file, or write these
literal bytes.
3. **Reconstruction** (`sync.ApplyDelta`) - the receiver replays that
operation list against its old copy to reproduce the sender's file
exactly.

The block size is currently a fixed constant (`sync.DefaultBlockSize`).
Real rsync scales it dynamically based on file size; fixed-size blocks are
a deliberate simplification here, not a limitation of the algorithm
itself.

## Architecture

- `cmd/grsync` — CLI entrypoint.
- `internal/cli` — flag/argument parsing (built on cobra).
- `internal/sync` — file-list generation and filter matching today; comparison/delta logic later.
- `internal/transport` — (placeholder) data movement, local and remote.
- `cmd/grsync` - CLI entrypoint.
- `internal/cli` - flag/argument parsing (built on cobra).
- `internal/sync` - file-list generation, filter matching, and the
delta-transfer algorithm today; wiring these together into an actual
sync comes later.
- `internal/transport` - (placeholder) data movement, local and remote.

Goal: full feature parity with upstream rsync, including protocol/format
interoperability where specified (e.g. batch mode's file format).
Expand Down
6 changes: 3 additions & 3 deletions internal/cli/root.go
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
// Package cli defines the grsync command-line interface: argument parsing,
// flags, and the command tree. It does not perform any sync or transport
// logic itself — it only collects options and hands them off (see the
// logic itself - it only collects options and hands them off (see the
// options struct printed in Run below, which will later be passed to
// internal/sync).
package cli
Expand Down Expand Up @@ -32,7 +32,7 @@ const (
// collects them the same way: Type records which flag produced the rule,
// and relative order across *all* of them is preserved in the order the
// user supplied them. For the two "-from" kinds, Pattern is a file path,
// not a filter pattern — internal/sync reads and expands it.
// not a filter pattern - internal/sync reads and expands it.
type FilterRule struct {
Type FilterRuleType
Pattern string
Expand All @@ -55,7 +55,7 @@ type options struct {

// filterRuleFlag implements pflag.Value. Each of --exclude/--include/
// --filter/--exclude-from/--include-from gets its own instance, fixed to a
// single FilterRuleType, but all of them share the same backing slice — so
// single FilterRuleType, but all of them share the same backing slice - so
// pflag's normal "call Set once per occurrence" behavior naturally builds
// one ordered rule list regardless of which flag name was used at each
// position.
Expand Down
29 changes: 29 additions & 0 deletions internal/sync/blocks.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
package sync

// DefaultBlockSize is the fixed block size used to split files for the
// delta-transfer algorithm. Real rsync computes this dynamically per file
// (roughly proportional to the square root of the file size, within
// tunable bounds); using one fixed size here is a deliberate
// simplification for this ticket - dynamic block sizing is a future
// refinement, not something the algorithm itself depends on.
const DefaultBlockSize = 700

// splitBlocks splits data into fixed-size blocks of blockSize bytes each.
// The final block is shorter than blockSize whenever len(data) isn't an
// exact multiple of it; it's still included, never dropped or padded out
// to a full block. Returned slices share data's backing array rather than
// being copied.
func splitBlocks(data []byte, blockSize int) [][]byte {
if blockSize <= 0 {
return nil
}
var blocks [][]byte
for start := 0; start < len(data); start += blockSize {
end := start + blockSize
if end > len(data) {
end = len(data)
}
blocks = append(blocks, data[start:end])
}
return blocks
}
42 changes: 42 additions & 0 deletions internal/sync/blocks_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
package sync

import "testing"

func TestSplitBlocks(t *testing.T) {
tests := []struct {
name string
data []byte
blockSize int
wantLens []int
}{
{"empty", []byte{}, 4, nil},
{"exact multiple", []byte("aaaabbbbcccc"), 4, []int{4, 4, 4}},
{"partial final block", []byte("aaaabbbbcc"), 4, []int{4, 4, 2}},
{"smaller than one block", []byte("ab"), 4, []int{2}},
{"single byte block size", []byte("abc"), 1, []int{1, 1, 1}},
}

for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
blocks := splitBlocks(tt.data, tt.blockSize)
if len(blocks) != len(tt.wantLens) {
t.Fatalf("got %d blocks, want %d: %v", len(blocks), len(tt.wantLens), blocks)
}
for i, wantLen := range tt.wantLens {
if len(blocks[i]) != wantLen {
t.Errorf("block %d: len = %d, want %d", i, len(blocks[i]), wantLen)
}
}
// Reassembling every block must exactly reproduce the input -
// this is the property that actually matters (no bytes lost,
// duplicated, or reordered), not just the length list.
var got []byte
for _, b := range blocks {
got = append(got, b...)
}
if string(got) != string(tt.data) {
t.Errorf("reassembled blocks = %q, want %q", got, tt.data)
}
})
}
}
68 changes: 68 additions & 0 deletions internal/sync/checksum.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
package sync

import "crypto/md5"

// strongChecksum is the collision-resistant checksum used to confirm a
// weak-checksum match. MD5 is not cryptographically safe against a
// deliberate adversary, but that's not what it's used for here - it's
// only there to catch the rare case where two different blocks happen to
// share a weak checksum (see delta generation), which stdlib md5 is more
// than sufficient for.
func strongChecksum(block []byte) [md5.Size]byte {
return md5.Sum(block)
}

// rollingChecksumModulus is 65536 (2^16) - a power of two, not a prime.
// Real Adler-32 uses the largest prime below 65536 (65521) instead; this
// is rsync's own simpler variant, chosen deliberately because a
// power-of-two modulus makes unsigned-integer wraparound during roll()
// mathematically safe (see the comment there) without extra bounds
// handling. It rolls in O(1), which is the only property that actually
// matters here - this is not meant to be byte-compatible with stdlib
// hash/adler32.
const rollingChecksumModulus = 1 << 16

// weakChecksum is a rolling checksum over a fixed-size window: two 16-bit
// accumulators (a: sum of the window's bytes, b: a position-weighted sum)
// combined into one 32-bit value via sum().
type weakChecksum struct {
a, b uint32
length uint32 // window size; constant across every roll() call
}

// newWeakChecksum computes the checksum for window from scratch in O(len(window)).
func newWeakChecksum(window []byte) weakChecksum {
var a, b uint32
n := uint32(len(window))
for i, c := range window {
a += uint32(c)
b += (n - uint32(i)) * uint32(c)
}
return weakChecksum{
a: a % rollingChecksumModulus,
b: b % rollingChecksumModulus,
length: n,
}
}

// sum returns the combined checksum value.
func (w weakChecksum) sum() uint32 {
return w.a + w.b*rollingChecksumModulus
}

// roll advances the window by exactly one byte: out is the byte leaving
// at the window's start, in is the byte entering at its end. This is O(1)
// regardless of window size - the entire point of a rolling checksum,
// versus calling newWeakChecksum on the shifted window from scratch.
//
// The subtractions below can underflow as uint32 arithmetic (e.g. if
// w.a < out). That's fine, not a bug: Go's unsigned integers wrap modulo
// 2^32, and since rollingChecksumModulus (2^16) evenly divides 2^32, the
// wrapped value still reduces to the mathematically correct result mod
// 2^16 after the final "% rollingChecksumModulus". A prime modulus (real
// Adler-32's 65521) would not have this property.
func (w weakChecksum) roll(out, in byte) weakChecksum {
a := (w.a - uint32(out) + uint32(in)) % rollingChecksumModulus
b := (w.b - w.length*uint32(out) + a) % rollingChecksumModulus
return weakChecksum{a: a, b: b, length: w.length}
}
68 changes: 68 additions & 0 deletions internal/sync/checksum_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
package sync

import (
"crypto/md5"
"testing"
)

func TestStrongChecksum(t *testing.T) {
block := []byte("some block contents")
got := strongChecksum(block)
want := md5.Sum(block)
if got != want {
t.Errorf("strongChecksum(%q) = %x, want %x", block, got, want)
}

if strongChecksum([]byte("a")) == strongChecksum([]byte("b")) {
t.Errorf("different blocks produced the same strong checksum")
}
}

// TestWeakChecksum_RollMatchesFromScratch is the load-bearing test for the
// entire rolling-checksum design: it proves roll() produces exactly the
// same result as recomputing from scratch, at every single offset across
// a test string, not just a couple of spot checks. If this property ever
// broke, the delta algorithm would silently miss real block matches (or
// worse, produce false ones that only got caught by the strong checksum,
// masking the bug) - so this needs direct proof, not just "it compiles".
func TestWeakChecksum_RollMatchesFromScratch(t *testing.T) {
data := []byte("the quick brown fox jumps over the lazy dog, then jumps back again")
const windowSize = 8

if len(data) <= windowSize {
t.Fatalf("test data too short: need more than %d bytes", windowSize)
}

current := newWeakChecksum(data[:windowSize])
if want := newWeakChecksum(data[0:windowSize]); current.sum() != want.sum() {
t.Fatalf("offset 0: sum = %d, want %d", current.sum(), want.sum())
}

for offset := 1; offset+windowSize <= len(data); offset++ {
out := data[offset-1]
in := data[offset+windowSize-1]
current = current.roll(out, in)

want := newWeakChecksum(data[offset : offset+windowSize])
if current.sum() != want.sum() {
t.Fatalf("offset %d: rolled sum = %d, want %d (from scratch) - a=%d/%d b=%d/%d",
offset, current.sum(), want.sum(), current.a, want.a, current.b, want.b)
}
}
}

func TestWeakChecksum_IdenticalWindowsMatch(t *testing.T) {
a := newWeakChecksum([]byte("abcdefgh"))
b := newWeakChecksum([]byte("abcdefgh"))
if a.sum() != b.sum() {
t.Errorf("identical windows produced different sums: %d vs %d", a.sum(), b.sum())
}
}

func TestWeakChecksum_DifferentWindowsUsuallyDiffer(t *testing.T) {
a := newWeakChecksum([]byte("abcdefgh"))
b := newWeakChecksum([]byte("hgfedcba"))
if a.sum() == b.sum() {
t.Errorf("reversed window produced the same sum (%d) as the original - weak checksum isn't discriminating position", a.sum())
}
}
Loading