saved
Git at any scale
Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.
gist
Vicent Martí explains how Cursor’s Continuity scales Git hosting: repositories sit on fast local disks, but a write-ahead log in S3 is the source of truth. He starts with why hosting gets harder as traffic and repository counts grow. Continuity acknowledges a push only after it reaches the log, and compare-and-swap updates keep those writes in order. Every read checks the log first, any server can rebuild a replica from it, and only the primary compacts. A busy monorepo can run on hundreds of replicas, while small agent-created repositories get one each, or none while idle.
ideas
- Packfiles govern the server architecture. Git clients exchange packfiles, and local Git operations perform linked logical and physical walks through them, so storage designs must respect their access pattern.
- Distributing the wrong layer compounds latency. Object stores turn graph traversal into sequential network round trips, while network filesystems turn random packfile access into remote I/O.
- GitHub’s Spokes buys consistency with a fixed replication cost. Three-phase commit keeps every replica current, but more replicas increase tail latency and every cold repository still needs a quorum.
- Move durable truth outside the serving replica. Continuity records each push in an S3-backed write-ahead log, then treats local NVMe repositories as materialized caches that can be rebuilt anywhere.
- Make unreliable acceleration safe. Gossip speeds replication, conditional S3 reads prove freshness, and WAL compaction lets replicas download prepared packs instead of repeating CPU-heavy repacks.
quotes
“With three-phase commit, the floor is always too high, and the ceiling too low.”
“We treat repositories like a warm cache on disk, but the source of truth is always the write-ahead log in S3.”
“We never acknowledge a push until it has been fully persisted to the WAL.”