Skip to content

solver: ReleaseUnreferenced can hold the cache.db write lock for extended periods under sustained load #7175

Description

@Yeaury

What happens

In the bbolt-backed cache metadata store, cacheManager.ReleaseUnreferenced can perform a large amount of metadata deletion under sustained build load. The current controller invokes it through two paths:

  • After an automatic GC run has actually reclaimed bytes. Each solve schedules a throttled GC attempt; only when that GC reports size > 0 does it schedule throttledReleaseUnreferenced. The first throttle.After invocation runs immediately, and subsequent invocations are limited to at most once every five minutes.
  • After an explicit Prune that produced records. This is a deferred cleanup attempt, rather than a guarantee that a complete release pass succeeds.

The relevant release path is:

  1. cacheManager.ReleaseUnreferenced walks the cache-key metadata and its result references, using read transactions to inspect the metadata.
  2. For each result that is no longer present in the result store, it calls Store.Release.
  3. In the bbolt backend, each Store.Release runs its own db.Update write transaction. emptyBranchWithParents recursively removes related entries from the _result, _links, _byresult, and _backlinks buckets within that transaction.

Therefore, one release pass does not hold one write transaction for its entire traversal. Instead, it can issue many write transactions, and each transaction can hold bbolt's single-writer lock for a significant duration when the orphan metadata subgraph is large.

Why it matters

In our high-concurrency build workload, the orphan metadata population can grow substantially between release passes. We observe that large release passes create write-side tail-latency pressure on the same cache database:

  • AddResult, AddLink, and other metadata writes must wait behind the active bbolt write transaction.
  • A sequence of release transactions can keep write contention high for the duration of the pass, especially when orphan production is faster than the release pass can drain it.
  • Read-side operations such as Load use bbolt read transactions and do not directly queue behind the single writer in the same way. They may still see indirect latency changes from overall database and filesystem contention.

The exact duration depends on the number and shape of orphan records, the recursive cleanup work, storage latency, and concurrent database activity. In our large-build scenarios, release work has been sufficiently expensive to affect build efficiency and write-side tail latency.

The current five-minute cadence is hard-coded in control/control.go. It is not configurable for deployments with materially different build rates. Under some sustained workloads, a new orphan population can accumulate faster than a single release pass can remove it. Conversely, running a pass during a busy period can create avoidable write contention when the same work could have been performed during a known low-traffic window.

Related work

This is related to, but separate from, several cache metadata and database maintenance changes:

The issue here is the scheduling and operational control of logical unreferenced-metadata release. Physical compaction and transaction granularity are separate concerns.

Proposed direction

There are two complementary directions, which can be implemented separately:

  1. Provide an operator-triggered release operation, for example POST /debug/cache/release-unreferenced, so an operator can run the pass at a known-quiet time. A first implementation is available in buildkitd: add on-demand release of unreferenced cache metadata #7177.
  2. Consider making the automatic release cadence configurable, while retaining the current five-minute default for compatibility. The configuration shape and whether a fixed interval is preferable to a workload-based policy need maintainer feedback. This configuration is intentionally not included in buildkitd: add on-demand release of unreferenced cache metadata #7177.

The manual operation should serialize with automatic and post-prune release passes, honor cancellation while waiting or between storage operations, and not imply that the database has been physically compacted. It should also report storage errors instead of treating every attempted deletion as successful.

Explicitly out of scope for this issue and #7177:

  • Splitting the recursive cleanup into smaller bbolt write transactions.
  • Changing the recursive emptyBranchWithParents algorithm.
  • Coupling release scheduling to opt-in metadata database compaction #7138's physical compaction scheduler.
  • Reclaiming snapshotter data, blobs, or other data outside the cache metadata store.

Reproduction and measurements

We are interested in validating the relationship between orphan metadata, release duration, and write-side latency with a reproducible workload:

  • Run buildkitd with the bbolt-backed cache metadata store and the oci worker.
  • Drive sustained parallel builds that create and later orphan cache metadata, for example with multiple concurrent cache-exporting builds.
  • Instrument the bbolt backend to record the duration and number of db.Update transactions issued by ReleaseUnreferenced and Store.Release.
  • Record cache metadata bucket statistics through bbolt's DB.Stats() and bucket statistics, together with p95/p99 latency for concurrent metadata writes.

The expected observation is not that the entire release pass holds one global transaction, but that larger orphan populations produce more or longer write transactions and increased contention for concurrent metadata writers.

Environment

  • BuildKit: master, base commit 99bd9de47d29269020476c3eea5898f6038fa0a1.
  • Runtime: multi-replica buildkitd behind a consistent-hash router.
  • Worker: runc + overlayfs.
  • Workload: sustained parallel builds with a multi-GiB cache.db.

The exact fleet measurements are omitted here, but the qualitative behavior has been observed across the relevant high-throughput workloads.

What we'd like from maintainers

  • Review the direction of buildkitd: add on-demand release of unreferenced cache metadata #7177 as an operator-triggered logical metadata release operation.
  • Confirm whether a configurable automatic release interval is desirable, and if so, whether it belongs in the same change or a follow-up.
  • Advise whether the preferred operator surface is an HTTP debug endpoint, a buildctl debug subcommand, or both.
  • If a workload-based trigger or another scheduling policy is preferred, we are happy to rework the follow-up accordingly.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions