convergeKV is an open-source (MIT), distributed, AP key-value database written in Go that keeps accepting reads and writes during network partitions and node failures, and merges divergent replicas automatically using causal δ-CRDTs.
It is for backend developers and distributed-systems learners who need an always-writable store for data that tolerates eventual consistency: session stores, user profiles, shopping carts, feature flags, presence, and IoT/device state. Every node runs the same code and is an equal peer: no leader, no coordinator, no quorum, no single point of failure. It runs on Linux and macOS (Go 1.26+) and ships as a Docker image.
Trade-off in one sentence: convergeKV stays online during failures and lets replicas briefly disagree, using a conflict-free replicated data type (CRDT) that guarantees every replica converges to the same state once they can talk again (strong eventual consistency).
- Key features
- How convergeKV compares to Cassandra, Riak, etcd, and Redis
- System requirements
- How to install
- Quickstart: run a cluster
- API usage (gRPC)
- Configuration
- Data model and conflict resolution
- Architecture: how convergeKV works
- Operations: scaling, restarts, monitoring
- Performance benchmarks
- When to use convergeKV (and when not to)
- Troubleshooting
- FAQ
- Documentation
- Project status and license
- Always available (AP in CAP terms). Reads and writes continue during network partitions and node failures, as long as one of a key's 3 replicas is reachable.
- Low-latency writes, no quorum. A write is acknowledged after one
replica durably stores it (one
fsync); replication to the other two replicas is asynchronous. - Automatic conflict resolution with CRDTs. Concurrent writes merge deterministically: each JSON field is tracked independently, and concurrent edits to the same field resolve by last-writer-wins on a Hybrid Logical Clock.
- Self-healing replication. A background Merkle-tree anti-entropy process repairs any replica that fell behind. No manual repair command.
- Leaderless, symmetric nodes. Every node is identical. Membership uses SWIM
gossip (
hashicorp/memberlist); data placement uses rendezvous (HRW) hashing computed locally by every node. - Durable storage. Each write lands in an embedded LSM engine (Pebble) before it is acknowledged.
- JSON document values.
Get,Put(replace),Patch(partial update), andDeleteover a small, typed gRPC API. - Safe deletes. Tombstones carry causal context, so a late write cannot resurrect deleted data; garbage collection reclaims them automatically.
- Observable. Prometheus metrics and Go pprof on an admin port; optional Grafana dashboard stack.
- Container-ready. Tiny static binary and a minimal Docker image on GHCR.
| convergeKV | Apache Cassandra | Riak KV | etcd | Redis Cluster (OSS) | |
|---|---|---|---|---|---|
| CAP choice | AP | AP (tunable) | AP (tunable) | CP | Primary-based, async replicas |
| Write acknowledged after | 1 durable replica, always | Tunable (ONE … ALL) |
Tunable (W) |
Raft majority | Primary (in memory) |
| Leader / coordinator node | None | None | None | Raft leader | Primary per slot |
| Conflict resolution | Causal δ-CRDT: per-field LWW + add-wins removal | Per-cell LWW timestamps | Siblings or CRDT data types | None needed (linearizable) | Failover can drop writes |
| Replica repair | Continuous Merkle anti-entropy, automatic | Read repair, hinted handoff, nodetool repair |
Active anti-entropy | Raft log | Resync from primary |
| Value model | JSON document, per-field merge | Wide-column rows | Opaque blobs or CRDTs | Opaque bytes | Rich data structures |
| Query features | Key lookup only | CQL, secondary indexes | Key lookup, 2i | Key/range, watch | Rich commands |
| Language | Go | Java | Erlang | Go | C |
| License | MIT | Apache 2.0 | Apache 2.0 | Apache 2.0 | AGPLv3 / RSALv2 / SSPLv1 |
| Best for | Always-writable app state, learning CRDTs | Large-scale wide-column data | Always-writable KV at scale | Config, coordination, locks | Caching, low-latency data structures |
convergeKV is deliberately small: it is the smallest design in this table that still gives you leaderless writes, CRDT merging, and automatic repair, and it is readable end to end.
- OS: Linux or macOS (Docker image is
linux/amd64). - Go: 1.26+ (only for building from source).
- Docker + Docker Compose: for the containerized cluster.
- Disk: local disk with working
fsyncper node (Pebble data directory). - Network: per node, 3 TCP ports (client gRPC, node gRPC, admin) plus 1 TCP+UDP gossip port.
- Cluster size: 1 node works; 3+ nodes gives full 3-way replication.
Docker image (GitHub Container Registry):
docker pull ghcr.io/janthoxo/convergekv/convergekv:latestGo install (builds the kvnode binary into $GOBIN):
go install github.com/janthoXO/convergeKV/cmd/kvnode@latestFrom source:
git clone https://github.com/janthoXO/convergeKV.git
cd convergeKV
make build # produces ./dist/kvnodeTagged builds and the Docker image tarball are on the Releases page.
The repo includes a Docker Compose setup that starts a seed node plus scalable peers.
make docker-up N=3 # 3-node cluster (1 seed + 2 peers)
make docker-up N=5 # scale to any size
make docker-down # tear down and remove volumesThe seed node's client API is published on localhost:7000.
Optional monitoring stack (Prometheus + Grafana):
docker compose --profile monitoring up
# Grafana: http://localhost:3000 (anonymous admin)
# Prometheus: http://localhost:9090docker-compose.prod.yml uses the published GHCR image with hardened defaults
(read-only root FS, dropped capabilities, resource limits, health checks, no
host-exposed ports):
docker compose -f docker-compose.prod.yml up -d --scale node=3# Node 1: bootstrap a brand-new cluster (no seeds)
CONVERGEKV_DATA_DIR=./data/n1 \
CONVERGEKV_CLIENT_ADDR=:7000 \
CONVERGEKV_NODE_ADDR=:7001 \
CONVERGEKV_GOSSIP_ADDR=:7946 \
./dist/kvnode
# Node 2: join via node 1's gossip address (distinct ports on the same host)
CONVERGEKV_DATA_DIR=./data/n2 \
CONVERGEKV_CLIENT_ADDR=:7100 \
CONVERGEKV_NODE_ADDR=:7101 \
CONVERGEKV_GOSSIP_ADDR=:7947 \
CONVERGEKV_ADMIN_ADDR=:7102 \
CONVERGEKV_SEEDS=127.0.0.1:7946 \
./dist/kvnodeEmpty SEEDS starts a new cluster; non-empty SEEDS joins an existing one.
convergeKV exposes the gRPC service convergekv.KV (definitions in
pkg/proto/kv.proto):
| Method | Request | Notes |
|---|---|---|
Put |
{ key, value } |
Replace. value is the bytes of a non-empty JSON object; any field the document currently has that value omits is removed. |
Patch |
{ key, value, delete_fields } |
Partial update. Sets the fields in value and removes those named in delete_fields; fields you don't mention are kept. value may be empty when only deleting. |
Get |
{ key } |
Returns { found, value, context_hash }. |
Delete |
{ key } |
Removes the whole key (leaves an internal tombstone, reclaimed automatically). |
Replace is add-wins, not authoritative.
Put(andPatch's deletes) only remove fields the chosen owner has already observed. A field added concurrently on another node that this owner hasn't seen yet survives the merge. There is no hard whole-document replace under concurrency.
Connect to any node. A node that isn't an owner of the key forwards the request internally (at most one extra hop).
Encoding note:
valueis a protobufbytesfield. Native gRPC clients (Go, Node.js, …) pass raw JSON bytes. JSON-based gRPC tools (grpcurl, Bruno, Postman) need a base64 string.
Example with grpcurl (the document
{"name":"Alice"} base64-encodes to eyJuYW1lIjoiQWxpY2UifQ==):
# Put
grpcurl -plaintext -import-path pkg/proto -proto kv.proto \
-d '{"key":"user:1","value":"eyJuYW1lIjoiQWxpY2UifQ=="}' \
127.0.0.1:7000 convergekv.KV/Put
# Get -> { "found": true, "value": "eyJuYW1lIjoiQWxpY2UifQ==", ... }
grpcurl -plaintext -import-path pkg/proto -proto kv.proto \
-d '{"key":"user:1"}' \
127.0.0.1:7000 convergekv.KV/Get
# Patch: set "tier", remove "name" ({"tier":"gold"} → eyJ0aWVyIjoiZ29sZCJ9)
grpcurl -plaintext -import-path pkg/proto -proto kv.proto \
-d '{"key":"user:1","value":"eyJ0aWVyIjoiZ29sZCJ9","delete_fields":["name"]}' \
127.0.0.1:7000 convergekv.KV/Patch
# Delete
grpcurl -plaintext -import-path pkg/proto -proto kv.proto \
-d '{"key":"user:1"}' \
127.0.0.1:7000 convergekv.KV/DeleteA ready-made Bruno collection lives in
docs/bruno/.
Configured only through environment variables (prefix CONVERGEKV_). No
command-line flags. Bad config fails startup loudly.
| Variable | Default | Description |
|---|---|---|
CONVERGEKV_DATA_DIR |
data |
Directory for the node's identity and on-disk data. |
CONVERGEKV_CLIENT_ADDR |
:7000 |
Listen address for the client gRPC API. |
CONVERGEKV_NODE_ADDR |
:7001 |
Listen address for node-to-node gRPC. |
CONVERGEKV_GOSSIP_ADDR |
:7946 |
Bind address for cluster membership gossip. |
CONVERGEKV_ADMIN_ADDR |
:7002 |
Prometheus metrics + pprof. Empty disables it. |
CONVERGEKV_ADVERTISE_ADDR |
(derived) | Address other nodes use to reach this one. |
CONVERGEKV_SEEDS |
(empty) | Comma-separated gossip addresses to join. Empty = bootstrap a new cluster. |
CONVERGEKV_PARTITIONS |
256 |
Cluster-wide shard count. Power of two, ≤ 1024. Fixed at cluster birth. |
CONVERGEKV_CRASH_GRACE_PERIOD |
10m |
How long a dead node keeps its data slot before successors take over. |
CONVERGEKV_ANTI_ENTROPY_INTERVAL |
45s |
How often replicas reconcile via Merkle comparison. |
CONVERGEKV_REPLICATION_MAX_AGE |
20s |
Max age a queued replication update may reach before it's dropped to the anti-entropy backstop. Must be < ANTI_ENTROPY_INTERVAL / 2. |
CONVERGEKV_LOG_LEVEL |
info |
debug, info, warn, or error. Logs are JSON. |
PARTITIONSis chosen once when the cluster is created and cannot change. A node that tries to join with a different value is rejected.
- A value is a JSON object, e.g.
{"name":"Alice","tier":"gold"}. - Each top-level field is tracked independently, so concurrent edits to different fields merge without conflict.
Putreplaces the document (drops fields the new object omits);Patchsets the fields you provide and deletes the ones you list. Two clients patching different fields of the same key both survive the merge.- Two concurrent writes to the same field resolve by last-writer-wins on a Hybrid Logical Clock timestamp, deterministically, so every replica picks the same winner.
- Field values are opaque: strings, numbers, arrays, and nested objects are stored verbatim and replaced whole (no deep merge within a field).
- Field removal (
Put's implicit drops,Patch'sdelete_fields) is add-wins: only fields the owner has already observed are removed, so a concurrent add it hasn't seen survives. - A
Putvalue must be a non-empty JSON object (useDeleteto remove a whole key); aPatchmay omitvaluewhen it only deletes fields.
Clients connect to any node; nodes are symmetric peers that gossip about membership and replicate data to each other.
- Placement (rendezvous hashing). Each key hashes to one of
Ppartitions. Each partition is owned by 3 nodes chosen by a shared ranking function (HRW) that every node computes identically. No central placement service. - Writes (no quorum). Any node routes the request to one owner, the applier. The applier stores the change durably and acknowledges immediately, without waiting for the other two owners.
- Replication (asynchronous). The applier forwards the change (a CRDT delta) to the other owners in the background, fire-and-forget.
- Self-healing (Merkle anti-entropy). Owners of a partition periodically compare Merkle-tree fingerprints and exchange whatever is missing. This single mechanism guarantees convergence and makes best-effort replication safe.
- Membership (SWIM gossip). Nodes discover each other and detect failures via gossip. Placement is recomputed automatically as nodes come and go.
The write path: client acknowledged after one durable local write, replication afterward.
- Scale out: start more nodes pointing at existing seeds. Data rebalances automatically (new owners bootstrap their partitions from current owners).
- Scale in / planned removal: stop a node gracefully; it hands its partitions to successors before leaving, preserving the replication factor.
- Node restart: a quick restart resumes its data with no re-transfer. A node
gone longer than
CRASH_GRACE_PERIODrejoins as a fresh member and re-syncs. - Monitoring: scrape
CONVERGEKV_ADMIN_ADDR(/metrics). Useful series: replication backlog/drops, anti-entropy repairs, data transfer, membership size. Go pprof on the same address.
Indicative latencies from the in-repo benchmark (5-node cluster, 16 partitions, RF=3, 1 KB JSON documents, single sequential client; Apple Silicon, synced writes):
| Operation | p50 | p99 |
|---|---|---|
| Put (durable, single owner) | ~12 ms | ~21 ms |
| Get (local read from an owner) | ~59 µs | ~205 µs |
| Convergence (write byte-equal on all 3 owners) | ~30 ms | ~42 ms |
Put latency is dominated by the per-write fsync that makes "acknowledged" mean
"durable." Reads that hit an owner never touch the network. See
docs/BENCHMARKS.md to reproduce.
Good fit:
- High write availability matters more than reading your own write instantly.
- Data tolerates eventual consistency: profiles, sessions, carts, preferences, presence, device state.
- You want operational simplicity: no leader election, no quorum tuning, no manual conflict handling.
- You want to learn or teach CRDTs, anti-entropy, and leaderless replication from a readable, tested codebase.
Not a fit:
- You need linearizable / read-after-write consistency or multi-key transactions (use etcd, a SQL database, or a CP store).
- You need uniqueness constraints, secondary indexes, or range/SQL queries.
- A field's value needs merging of concurrent edits (e.g. counters, lists) rather than last-writer-wins.
InvalidArgumenton Put: value not a non-empty JSON object. Send an object like{"a":1}; useDeleteto remove a key.- grpcurl/Bruno/Postman rejects
value:bytesfield needs base64 in JSON tools (echo -n '{"a":1}' | base64). - Node refuses to join cluster: its
CONVERGEKV_PARTITIONSdiffers from the cluster's. Use the same value everywhere. - Startup fails with config error:
PARTITIONSmust be a power of two ≤ 1024, andREPLICATION_MAX_AGEmust be <ANTI_ENTROPY_INTERVAL / 2. Unavailableerror: no owner of that key reachable right now. Retry.FailedPreconditionerror: node's membership view was stale. Retry (possibly on another node).- Read returns old value right after write: expected. Replication is asynchronous; replicas converge within milliseconds normally, at most one anti-entropy interval after a failure.
- Docker build fails in mainland China: see Building behind the Great Firewall.
convergeKV is an open-source, leaderless, distributed key-value store written in Go. It stores JSON documents, replicates each key to 3 nodes, stays writable during network partitions, and uses causal δ-CRDTs so that all replicas converge to the same state automatically.
convergeKV provides strong eventual consistency, not linearizability. Any two replicas that have received the same set of updates are in the same state, regardless of delivery order or duplication. A read may briefly return a stale value.
Each top-level JSON field is a separate CRDT register. Concurrent writes to different fields both survive. Concurrent writes to the same field resolve by last-writer-wins using Hybrid Logical Clock timestamps, with a deterministic tie-break, so every replica picks the same winner.
Both sides keep accepting reads and writes for every key that has at least one reachable owner. When the partition heals, Merkle-tree anti-entropy finds the differing keys and merges them; no update is lost to a conflict.
No. A write is acknowledged once one owner persists it with an fsync.
Replication to the other two owners is asynchronous, and anti-entropy repairs
anything the replication missed.
All three are leaderless AP stores inspired by Amazon Dynamo. Cassandra and Riak offer tunable quorums and far more features at larger scale. convergeKV has no quorums at all, merges JSON documents per field with a causal δ-CRDT, relies on a single automatic repair mechanism (Merkle anti-entropy), and is a small Go codebase meant to be read and understood end to end.
Yes. A delete keeps the document's causal context as a tombstone, so an older write arriving late is recognized as already seen and discarded. Tombstones are garbage-collected only after two clean anti-entropy rounds certify every owner has converged.
It is an educational/research implementation, exercised by property, fuzz, chaos, integration, and Docker end-to-end tests. Evaluate it carefully before storing critical data.
MIT. See LICENSE.
README_DEV.md: architecture and contributor guide.docs/concepts/: 11-chapter deep dive into the design (CRDTs, HLC, placement, gossip, storage, request paths, anti-entropy, transfer, garbage collection, lifecycle).docs/BENCHMARKS.md: benchmark methodology and results.llms.txt: machine-readable project summary for AI tools.
convergeKV is an educational/research implementation of a causal δ-CRDT
key-value store, released under the MIT License. Contributions
welcome; start with README_DEV.md.