You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
solver: ReleaseUnreferenced can hold the cache.db write lock for extended periods under sustained load #7175
In the bbolt-backed cache metadata store, cacheManager.ReleaseUnreferenced can perform a large amount of metadata deletion under sustained build load. The current controller invokes it through two paths:
After an automatic GC run has actually reclaimed bytes. Each solve schedules a throttled GC attempt; only when that GC reports size > 0 does it schedule throttledReleaseUnreferenced. The first throttle.After invocation runs immediately, and subsequent invocations are limited to at most once every five minutes.
After an explicit Prune that produced records. This is a deferred cleanup attempt, rather than a guarantee that a complete release pass succeeds.
The relevant release path is:
cacheManager.ReleaseUnreferenced walks the cache-key metadata and its result references, using read transactions to inspect the metadata.
For each result that is no longer present in the result store, it calls Store.Release.
In the bbolt backend, each Store.Release runs its own db.Update write transaction. emptyBranchWithParents recursively removes related entries from the _result, _links, _byresult, and _backlinks buckets within that transaction.
Therefore, one release pass does not hold one write transaction for its entire traversal. Instead, it can issue many write transactions, and each transaction can hold bbolt's single-writer lock for a significant duration when the orphan metadata subgraph is large.
Why it matters
In our high-concurrency build workload, the orphan metadata population can grow substantially between release passes. We observe that large release passes create write-side tail-latency pressure on the same cache database:
AddResult, AddLink, and other metadata writes must wait behind the active bbolt write transaction.
A sequence of release transactions can keep write contention high for the duration of the pass, especially when orphan production is faster than the release pass can drain it.
Read-side operations such as Load use bbolt read transactions and do not directly queue behind the single writer in the same way. They may still see indirect latency changes from overall database and filesystem contention.
The exact duration depends on the number and shape of orphan records, the recursive cleanup work, storage latency, and concurrent database activity. In our large-build scenarios, release work has been sufficiently expensive to affect build efficiency and write-side tail latency.
The current five-minute cadence is hard-coded in control/control.go. It is not configurable for deployments with materially different build rates. Under some sustained workloads, a new orphan population can accumulate faster than a single release pass can remove it. Conversely, running a pass during a busy period can create avoidable write contention when the same work could have been performed during a known low-traffic window.
Related work
This is related to, but separate from, several cache metadata and database maintenance changes:
cache: fix cache leak #4353 (cache: fix cache leak) cleans up compression-variant leases when snapshot loading fails at initialization.
opt-in metadata database compaction #7138 (opt-in metadata database compaction) proposes physical bbolt compaction for several metadata databases. It addresses file-space reclamation after logical deletion; it does not decide when ReleaseUnreferenced should run.
The issue here is the scheduling and operational control of logical unreferenced-metadata release. Physical compaction and transaction granularity are separate concerns.
Proposed direction
There are two complementary directions, which can be implemented separately:
Consider making the automatic release cadence configurable, while retaining the current five-minute default for compatibility. The configuration shape and whether a fixed interval is preferable to a workload-based policy need maintainer feedback. This configuration is intentionally not included in buildkitd: add on-demand release of unreferenced cache metadata #7177.
The manual operation should serialize with automatic and post-prune release passes, honor cancellation while waiting or between storage operations, and not imply that the database has been physically compacted. It should also report storage errors instead of treating every attempted deletion as successful.
Reclaiming snapshotter data, blobs, or other data outside the cache metadata store.
Reproduction and measurements
We are interested in validating the relationship between orphan metadata, release duration, and write-side latency with a reproducible workload:
Run buildkitd with the bbolt-backed cache metadata store and the oci worker.
Drive sustained parallel builds that create and later orphan cache metadata, for example with multiple concurrent cache-exporting builds.
Instrument the bbolt backend to record the duration and number of db.Update transactions issued by ReleaseUnreferenced and Store.Release.
Record cache metadata bucket statistics through bbolt's DB.Stats() and bucket statistics, together with p95/p99 latency for concurrent metadata writes.
The expected observation is not that the entire release pass holds one global transaction, but that larger orphan populations produce more or longer write transactions and increased contention for concurrent metadata writers.
Environment
BuildKit: master, base commit 99bd9de47d29269020476c3eea5898f6038fa0a1.
Runtime: multi-replica buildkitd behind a consistent-hash router.
Worker: runc + overlayfs.
Workload: sustained parallel builds with a multi-GiB cache.db.
The exact fleet measurements are omitted here, but the qualitative behavior has been observed across the relevant high-throughput workloads.
What happens
In the bbolt-backed cache metadata store,
cacheManager.ReleaseUnreferencedcan perform a large amount of metadata deletion under sustained build load. The current controller invokes it through two paths:size > 0does it schedulethrottledReleaseUnreferenced. The firstthrottle.Afterinvocation runs immediately, and subsequent invocations are limited to at most once every five minutes.Prunethat produced records. This is a deferred cleanup attempt, rather than a guarantee that a complete release pass succeeds.The relevant release path is:
cacheManager.ReleaseUnreferencedwalks the cache-key metadata and its result references, using read transactions to inspect the metadata.Store.Release.Store.Releaseruns its owndb.Updatewrite transaction.emptyBranchWithParentsrecursively removes related entries from the_result,_links,_byresult, and_backlinksbuckets within that transaction.Therefore, one release pass does not hold one write transaction for its entire traversal. Instead, it can issue many write transactions, and each transaction can hold bbolt's single-writer lock for a significant duration when the orphan metadata subgraph is large.
Why it matters
In our high-concurrency build workload, the orphan metadata population can grow substantially between release passes. We observe that large release passes create write-side tail-latency pressure on the same cache database:
AddResult,AddLink, and other metadata writes must wait behind the active bbolt write transaction.Loaduse bbolt read transactions and do not directly queue behind the single writer in the same way. They may still see indirect latency changes from overall database and filesystem contention.The exact duration depends on the number and shape of orphan records, the recursive cleanup work, storage latency, and concurrent database activity. In our large-build scenarios, release work has been sufficiently expensive to affect build efficiency and write-side tail latency.
The current five-minute cadence is hard-coded in
control/control.go. It is not configurable for deployments with materially different build rates. Under some sustained workloads, a new orphan population can accumulate faster than a single release pass can remove it. Conversely, running a pass during a busy period can create avoidable write contention when the same work could have been performed during a known low-traffic window.Related work
This is related to, but separate from, several cache metadata and database maintenance changes:
cache: fix cache leak) cleans up compression-variant leases when snapshot loading fails at initialization.bboltcachestorage: only delete link after releasing result) fixesemptyBranchWithParentsrelease semantics when a sub-result is still referenced.cache/metadata: remove obsolete indexes when replacing values) removes obsolete metadata index entries when an indexed value is replaced.opt-in metadata database compaction) proposes physical bbolt compaction for several metadata databases. It addresses file-space reclamation after logical deletion; it does not decide whenReleaseUnreferencedshould run.The issue here is the scheduling and operational control of logical unreferenced-metadata release. Physical compaction and transaction granularity are separate concerns.
Proposed direction
There are two complementary directions, which can be implemented separately:
POST /debug/cache/release-unreferenced, so an operator can run the pass at a known-quiet time. A first implementation is available in buildkitd: add on-demand release of unreferenced cache metadata #7177.The manual operation should serialize with automatic and post-prune release passes, honor cancellation while waiting or between storage operations, and not imply that the database has been physically compacted. It should also report storage errors instead of treating every attempted deletion as successful.
Explicitly out of scope for this issue and #7177:
emptyBranchWithParentsalgorithm.Reproduction and measurements
We are interested in validating the relationship between orphan metadata, release duration, and write-side latency with a reproducible workload:
ociworker.db.Updatetransactions issued byReleaseUnreferencedandStore.Release.DB.Stats()and bucket statistics, together with p95/p99 latency for concurrent metadata writes.The expected observation is not that the entire release pass holds one global transaction, but that larger orphan populations produce more or longer write transactions and increased contention for concurrent metadata writers.
Environment
99bd9de47d29269020476c3eea5898f6038fa0a1.cache.db.The exact fleet measurements are omitted here, but the qualitative behavior has been observed across the relevant high-throughput workloads.
What we'd like from maintainers
buildctl debugsubcommand, or both.