I created 150 branches on the same repository at once, each from its own git update-ref process, on a repository using Git's reftable backend. Fifty-six of them succeeded. The other ninety-four printed this and exited non-zero:
fatal: update_ref failed for ref 'refs/heads/par-17': cannot lock references
The same 150 processes against a repository using the old files backend all succeeded, every time I tried it. That is the part of this story nobody mentions when they talk about reftable being faster.
Git 2.55.0 (released 29 June 2026) ships the reftable ref storage format that the 2.51 release notes describe as "matured enough" that Git 3.0 will make it the default for new repositories. Instead of one small file per branch or tag under .git/refs/, reftable stores the whole ref database as a small number of sorted, binary-searchable table files under .git/reftable/. I wanted to know what that actually costs and saves, not just what the release notes claim.
Setting this up
Ubuntu's packaged Git is 2.43.0, which predates reftable's finished state, so I built 2.55.0 from the source tarball on kernel.org (./configure && make, about 53 seconds on four cores). I used that single binary for every comparison below and only changed --ref-format, so nothing here conflates a ref-backend difference with a Git-version difference. All numbers are from a Docker container on one machine; treat the absolute milliseconds as this-machine numbers and the ratios as the interesting part.
Writing refs
I built a commit, then fed both backends the same batch of update refs/heads/branch-NNNNNN <sha> lines through git update-ref --stdin, run against a freshly initialised bare repository each time.
| refs | files backend | reftable backend |
|---|---|---|
| 10,000 | 436–650ms (3 runs) | 39–42ms (3 runs) |
| 50,000 | 2.1s–12.4s (3 runs) | 199–213ms (3 runs) |
The reftable numbers are boringly consistent. The files-backend numbers at 50,000 refs are not: three runs on the same freshly-initialised repository gave 2.1, 5.3 and 12.4 seconds. Creating tens of thousands of loose files in one directory appears to have genuinely unpredictable tail latency on top of being slower on average, and reftable simply doesn't have that failure mode because it isn't writing one file per ref.
Disk use tells the same story more starkly. At 10,000 refs, .git/refs/ held 40MB across 10,000 files (mostly filesystem block overhead for 41-byte files); .git/reftable/ held one 266KB table plus a 43-byte list, 272KB total. At 50,000 refs it was 198MB against 1.4MB.
The number that argues against the headline
If bulk writes are 10–60x faster, I expected reads to follow. They didn't, not by nearly that much. git for-each-ref over the 10,000-ref repositories: 180–183ms (files) against 132–134ms (reftable), about 1.4x. At 50,000 refs: 897–943ms against 636–714ms, about 1.3x.
I also tried to reproduce the specific claim in the 2.51 release notes, that reftable makes git fetch 22x faster and git push 18x faster on a 10,000-ref repository. I cloned with --mirror, then again with --no-local to force the real ref-advertisement code path instead of a local-transport shortcut, and separately timed git ls-remote. Best case, at 50,000 refs, ls-remote went from 342–376ms to 73–189ms: real, but a 2–4x gain, not 22x. At 10,000 refs the gap nearly disappeared, 174–178ms against 119–169ms. I could not get anywhere near the official multiplier on one machine over file://. My best guess is that the release notes' benchmark measures a network transport where fixed per-round-trip cost is much smaller relative to ref-advertisement cost than it is here, and I'd treat the 22x/18x figures as measured on a specific rig rather than a number that travels.
One more read-side check: reftable's other selling point is "atomic reference transactions." I sent both backends a four-update batch through update-ref --stdin where one update's expected old value didn't match reality. Both backends rejected the entire batch and applied none of it, identically. update-ref --stdin has been atomic-or-nothing in the files backend for years, so this isn't a difference reftable introduces for anyone already using batched updates.
Where reftable refuses
The concurrency result at the top of this post is the one I'd actually plan around. I ran it several times to be sure the first result wasn't a fluke, at three different levels of concurrency, always the same repository freshly initialised, always distinct target refs so there was no real conflict for Git to detect:
| concurrent writers | files backend | reftable backend (default) |
|---|---|---|
| 50 | 50/50 succeed, 37–43ms | 50/50 succeed, 74–108ms |
| 100 | not retested (see below) | 56/100 succeed, both trials |
| 150 | 150/150 succeed, 106–119ms | 54–70/150 succeed, across 5 trials |
Under 50 concurrent writers reftable is just slower, which is expected: every write has to lock a single tables.list file and append to a shared stack, where the files backend locks one file per ref and lets unrelated branches proceed independently. Somewhere between 50 and 100 writers that stops being "slower" and starts being "fails." At 100 and 150 concurrent writers I saw failure rates between 30% and 63% across seven separate trials, never zero, never total, always with the same error: cannot lock references.
cat .git/reftable/tables.list shows why. Under normal sequential load it holds one to three entries; Git's geometric compaction (reftable.geometricFactor, default 2) keeps folding small tables back together after almost every write. That compaction step needs the same lock every writer needs, so a burst of genuinely simultaneous writers queues up for one file, and Git's default patience for that queue is short.
The workaround, and its price
The relevant setting is reftable.lockTimeout, documented as "Value 0 means not to retry at all; -1 means to try indefinitely. Default is 100." I set it to 5000 and reran the 150-writer test three times: 150/150 succeeded every time, zero errors. It also took 1.05–1.15 seconds, against the files backend's 106–119ms for the identical job. The workaround removes the failures. It does not remove the underlying serialisation; it just makes everyone queue politely instead of a third of them giving up.
If your workload is "one process pushes to one repository," this never bites. If it's "CI creates a branch per job and several jobs finish at once," or "several people push new branches within the same second," it's worth checking your Git version's default reftable.lockTimeout before switching a busy repository to reftable, not after.
Migrating an existing repository
git refs migrate --ref-format=reftable converts a repository in place. On the 10,000-ref files repository from above it took 168ms and every ref came through intact. It refused outright on a repository with a second worktree attached:
error: migrating repositories with worktrees is not supported yet
More subtly, migration rewrites .git/HEAD to the literal text ref: refs/heads/.invalid, even when the real branch (refs/heads/master, pointing at a real commit) still resolves correctly through git symbolic-ref HEAD, git rev-parse HEAD, git branch and a fresh git clone. This is documented and deliberate (Git's Documentation/technical/reftable.adoc): the dummy file exists so that tools which only check "does this directory look like a Git repository" still work, without leaking a real branch name into a file format that predates reftable's own bookkeeping. It does mean that anything reading .git/HEAD directly as a text file, rather than asking Git, will get .invalid back after a migration.
What I got wrong on the way
My first attempt at the atomic-transaction test used an all-zero SHA as the "expected old value" for a ref that didn't exist yet. That's the correct sentinel for "this ref must not already exist," not a conflict, so both backends happily applied the whole batch and I nearly wrote down "no atomicity difference, and also no rejection behaviour at all," which would have been wrong on the second half. I only caught it by pre-creating the ref and deliberately asking for the wrong existing value.
I also nearly missed the concurrency result entirely. My very first 150-writer run against reftable had only two failures out of 150, which looked like a rounding error I could ignore. Running it five more times over the next few minutes turned up 30 to 63 failures each time. Whatever was quiet on that first run (a warm page cache, low load, plain luck) wasn't representative, and one clean run is not evidence of anything when you're testing a queue.
Run it yourself
This needs Git built with reftable support; check with git init --ref-format=reftable /tmp/x. Ubuntu 24.04's packaged 2.43.0 doesn't have it, so building from source is the reliable route:
curl -LO https://mirrors.edge.kernel.org/pub/software/scm/git/git-2.55.0.tar.gz
tar xzf git-2.55.0.tar.gz && cd git-2.55.0
make configure && ./configure --prefix=/opt/git-2.55.0
make -j"$(nproc)" && sudo make install
export PATH=/opt/git-2.55.0/bin:$PATH
Then the concurrency test that produced the headline failure:
git init --bare --ref-format=reftable /tmp/rt-test
sha=$(git -C /tmp/rt-test hash-object -t commit --stdin -w <<< "irrelevant") # or copy a real commit's objects in
for i in $(seq 1 150); do
git -C /tmp/rt-test update-ref "refs/heads/par-$i" "$sha" 2>>/tmp/errs.log &
done
wait
grep -c "cannot lock references" /tmp/errs.log
Compare against git init --bare --ref-format=files /tmp/files-test with the same loop, and against setting git -C /tmp/rt-test config reftable.lockTimeout 5000 before rerunning.
If you're running a repository that's mostly one writer at a time, reftable's bulk-write speed and disk footprint are a straightforward win and git refs migrate --ref-format=reftable is cheap enough to just try. If several processes create refs on the same repository around the same moment, run your own concurrency test at the scale you actually see before migrating, and check what reftable.lockTimeout is set to either way.
Top comments (14)
The 40ms number tracks with the format design — reftable batches lookups into one mmap'd table instead of the lstat-per-ref walk loose refs need. One migration note for anyone converting a server-side bare repo: the ref storage flips atomically but every clone talks to it over the wire afterwards, so coordinate the window. Did you also benchmark lookup-heavy ops like
for-each-refandbranch -a, or writes only?Good question, and yes, I did test for-each-ref specifically since I expected the batching benefit to show up there too. It's a much smaller gap though: about 1.3 to 1.4x faster with reftable, not the 10 to 60x you see on bulk writes (180 to 183ms vs 132 to 134ms at 10k refs, 897 to 943ms vs 636 to 714ms at 50k). Reading is a genuinely different access pattern from writing thousands of loose files one lstat at a time, so the two numbers shouldn't really be expected to track each other. I didn't run branch -a separately since it's doing basically the same full-ref enumeration as for-each-ref, so I'd expect it to land in the same range, but that's an assumption on my part, not something I measured, worth being upfront about.
On the migration note, fair point and I didn't test that. Everything I ran against git refs migrate was on a quiet repo with no concurrent traffic, so I can't say what happens if a clone is mid-flight over the wire while the in-place conversion runs. Worth separating two things though: the ref-format flip is local storage only, it doesn't touch the wire protocol itself, so a client mid-clone shouldn't see anything different in the negotiation. Whether there's a window where a read could hit a half-migrated state on disk is a real question I just didn't test cold. Good catch, that's worth someone actually reproducing before trusting either way.
What I find interesting here is that the benchmark is really showing different system behaviors, not just different speeds. A backend can be the better choice for one workload and a worse fit for another, even when both are solving the same problem. That makes the real question less about “which backend is faster?” and more about “what kind of workload is this architecture designed to handle?”
The 10,000-ref write result and the concurrent-writer result are almost two different stories from the same system. That’s exactly why I like seeing benchmarks under different conditions rather than treating one impressive number as the conclusion.
Yeah, that framing nails it honestly. It's basically the same backend having two totally different personalities depending on what you throw at it. Bulk write = one writer just cranking through a single mmap'd table in one go, which is exactly the party trick reftable was built to do. Throw 150 writers at it at once though and suddenly they're all fighting over the same tables.list lock (plus compaction tagging along for the ride), and that party trick turns into a pileup. Same code, wildly different day depending on who shows up.
Funny enough the old files backend is the exact opposite personality: terrible at the bulk write (syscall-per-ref, ouch), but doesn't even blink under 150 concurrent writers because everyone's got their own file to lock. So neither one is "the good backend," they're just built for different crowds.
Kind of the whole reason I don't trust a benchmark that only shows me its best angle lol. One number is basically a highlight reel, you gotta see it get stress tested from a few different directions before you know if it actually matches your situation or not
Exactly. That’s why I see benchmarks more as architecture evidence than a speed contest. A number only becomes useful when you know what produced it: workload, concurrency, contention, access pattern, and the trade-offs behind the result. The interesting part is that the “slow” backend can actually be the better choice once the workload changes.
So I usually look at benchmarks less as “who won?” and more as “under which conditions does each design make sense?” That gives you something you can actually use when making an architecture decision.
Yeah exactly, and honestly that mindset shift is the whole reason benchmarks are worth reading past the headline number in the first place. "Who won" makes for a punchier title, not gonna lie, mine included lol, but "under what conditions does each one actually make sense" is the version that's still useful to you six months from now when you're staring at your own repo trying to decide.
Kind of funny that the reftable post ended up being a decent case study for your own point without me planning it that way. Wasn't trying to write "architecture evidence over speed contest: an article," it just kept happening every time I tested another angle.
the concurrency result is probably the most interesting part here. the bulk write numbers make reftable look like an obvious upgrade, but 150 concurrent writers turning into %30–63 failures changes the picture quite a bit.
Agreed, and it's worth being blunt about how stark that reversal is: at 150 concurrent writers the files backend is the one that comes out clean, 150/150 succeed in 106 to 119ms, while reftable fails 30 to 63% of the time across those trials with the default lockTimeout. The bulk-write numbers are testing something reftable is actually built for, one writer doing a lot of work. The concurrency test is testing the opposite case, many writers touching the same tables.list lock plus the geometric compaction that rides along with it, and that's exactly where the single-file-lock design costs you.
The workaround makes it worse in a different way, not better. Bumping reftable.lockTimeout to 5000 gets you 150/150 successes with zero errors, but the job that took the files backend ~110ms now takes 1.05 to 1.15 seconds, about 10x slower. So you don't get to keep the win, you're trading failures for latency, and the underlying serialization is still there, just hidden behind everyone waiting patiently instead of a third of them bailing out.
Practically that means the answer to "should I migrate" depends entirely on whether your workload looks like CI or a bot doing bursty parallel pushes into the same repo, versus normal human-scale usage where this never shows up. Worth someone checking whether a repo's actual write concurrency profile looks like the 50-writer case, which is a wash, or the 100 to 150 case, which isn't, before assuming the bulk-write headline generalizes.
Really thorough writeup, the concurrency reversal is the part that'll stick with people. One thing I didn't see addressed: you mention geometricFactor defaulting to 2 and folding tables back together after almost every write, which is what feeds the tables.list contention. Did you try raising geometricFactor instead of (or alongside) bumping lockTimeout, to see if fewer, bigger compactions collide less often under the same burst, or does the bigger merge itself become the new bottleneck at that concurrency? Feels like a cheaper knob to test before accepting the ~10x latency tax from lockTimeout.
Didn't test it, good call though, that's on me for not trying the obvious knob before reaching for the blunt one.
My guess going in, bumping geometricFactor should mean fewer compactions but each one now folds a bigger stack of tables together, so you're trading "the lock gets grabbed constantly for small merges" for "the lock gets grabbed less often but held longer per merge." Whether that nets out better at 150 writers kind of depends on whether the merge itself scales linearly with table count or gets disproportionately slower, and I genuinely don't know which without running it. There's also a second-order effect worth watching, fewer compactions means tables.list sits longer with more uncompacted entries in between merges, and every read has to walk that whole list, so you might be trading write-side contention for read-side slowdown depending on how read-heavy the workload is.
Gonna actually run it instead of guessing though, 100 and 150 writers, a couple of geometricFactor values, same error tracking as before. Appreciate the push, this is exactly the kind of thing that should've been in the post instead of jumping straight to lockTimeout.
The concurrency result is the one that will bite people, and it has a very specific shape in CI. Fan-out builds write a ref per job: refs/ci/, a tag per build, Gerrit-style change refs, a bot pushing a result ref from every matrix cell. Those are all distinct refs with no real conflict, which is exactly the case the files backend handled by locking one file per ref and letting unrelated writers proceed. That per-ref independence was quietly load-bearing for a whole class of automation, and nothing in the release notes tells you that you were depending on it.
Worth pairing with your own atomicity finding, because together they suggest the mitigation. "Cannot lock references" is contention on tables.list, not a rejected update, and update-ref --stdin is atomic-or-nothing in both backends. So a bounded retry with jitter around the batch is safe: a retried batch cannot half-apply, and you are waiting out a queue rather than resolving a conflict. That is a much better answer than reverting the backend, though it does mean every tool in your pipeline that shells out to update-ref now needs a retry policy it probably does not have.
On the 22x you could not reproduce: worth trying with cold page cache between runs. Reftable's read advantage is mostly fewer files and fewer metadata syscalls, and once ten thousand loose ref files are sitting in page cache the files backend stops paying for most of that. A warm benchmark would compress exactly the gap you saw compress. Your 1.3 to 1.4x might be the honest steady-state number and theirs the cold-start one.
The other place your disk figures matter more than they look: network filesystems. 198MB across 50,000 tiny files on NFS or SMB, where every one of those is a round trip, is a different universe from local ext4. Anyone with repos on network-mounted storage should expect reftable to win by far more than your local numbers suggest, and to feel the concurrency limit sooner.
Okay this comment's doing a lot of work, taking it point by point.
The CI fan-out thing, yeah, that's a really good catch and it's not in the post at all. Per-ref locking quietly being load-bearing for exactly the workload that never conflicts is the kind of thing nobody notices until it's gone, "nothing in the release notes tells you" is the perfect way to put it.
The retry idea actually checks out against something I already tested. The atomicity check in the post showed both backends reject a whole batch cleanly when one expected value is wrong, nothing half-applies, in either backend. So yeah, "cannot lock references" really is just contention on tables.list, not Git telling you something's wrong with your update. A bounded retry with jitter should be safe by that logic. You're right that it just moves the pain though, now it's "which of my ten CI tools that shell out to update-ref have zero retry logic" instead of "reftable is broken."
Cold vs warm cache for the 22x gap, didn't test that and it's a genuinely good guess. The only place cache temperature comes up in the post at all is one throwaway line about the concurrency test's first suspiciously-clean run possibly being a warm cache fluke, nothing controlled for the read/fetch numbers specifically. Everything ran in Docker on overlay storage too, not bare ext4, so there's a decent chance the whole rig was staying warmer than either backend would in the wild. Dropping caches between runs is a cheap enough test I should've just done.
Network filesystem point, also untested, also makes sense on both ends, reftable should win bigger there since it's trading thousands of round trips for one, and the lock contention should bite sooner too since every lock acquisition now costs network latency instead of a local syscall. Never touched NFS or SMB in this one at all.
Basically three follow-up tests worth of comment. Might actually run all three before I let myself write another one of these.
Three good tests, and I'd run them in that order — cache temperature first since it's the cheapest to falsify and would otherwise cast doubt on the other two. If NFS/SMB confirms the round-trip theory, that's probably the more interesting result to publish; the CI fan-out point already convinced me the per-ref locking win is real-world load-bearing, not just a benchmark artifact. Curious what you find.
file:/tmp/opencode/outbound/c1.txt