Found while fixing #27/#29 against 7766d52, and measured on a Samsung 990 PRO / ext4.
A seal costs six fsyncs, and the manifest it commits is rewritten in full every time.
The six, all on the one SegmentStore.seal path:
| # |
site |
what |
| 1 |
promoteLiveFile, segstore.go:1156 |
fsync the sealed data file |
| 2 |
promoteLiveFile, segstore.go:1167 |
fsync the directory after the rename |
| 3 |
writeIndexFile, segstore.go:1751 |
fsync the index file |
| 4 |
writeIndexFile, segstore.go:1762 |
fsync the directory again |
| 5 |
writeManifest, segstore.go:732 |
fsync the manifest |
| 6 |
writeManifest, segstore.go:742 |
fsync the directory a third time |
Measured
Sealing a one-record tail repeatedly, reporting cost per seal and manifest size as the segment count rises:
segments µs/seal manifest bytes
250 34792 46974
500 34123 93974
750 35122 140974
1000 36343 187976
1250 36368 235226
1500 36155 282476
1750 36172 329726
2000 37124 376976
strace -c over 200 of those seals: 1207 fsyncs (six per seal) and 603 renames (three per seal). At the ~5.5 ms per fsync this disk gives (see #32), six fsyncs is ~33 ms — which is the ~36 ms measured. The fsync count is the whole cost of a seal.
Two things follow, and it is worth keeping them apart:
- The fsync count dominates today. ~36 ms per seal, flat in segment count.
- The full-manifest rewrite is real but secondary — for now. The manifest grows ~188 bytes per segment (47 KB at 250 segments, 377 KB at 2000), and rewriting it costs only ~2 ms more across that whole range. It is O(N) per commit, so it is O(N²) over a run that seals N times, and it overtakes the fsyncs eventually — but it is not what is expensive at these sizes.
ImportSegmentFile commits a manifest per adopted segment (segstore.go:1561), so a sync of N segments pays the same O(N²).
Asks, in the order they pay
- Fewer barriers per seal. The three directory fsyncs are for renames into the same directory; one at the end orders them all. The data and index files can be fsynced before a single directory sync rather than each taking its own pair. Six could plausibly be two without weakening the crash contract — the invariant that matters is that a file is durable before the manifest names it, not that each gets its own barrier.
- Do not rewrite the whole manifest per commit. An append-only record of segment additions and removals, checkpointed occasionally, makes a commit O(1) in the segment count and takes the O(N²) out of both a long run of seals and a sync.
Both are independent of #30's tiered merge, which reduces how many segments exist; this is about what each seal costs regardless. #32 is the special case where the seal has nothing to seal at all.
Found while fixing #27/#29 against
7766d52, and measured on a Samsung 990 PRO / ext4.A seal costs six fsyncs, and the manifest it commits is rewritten in full every time.
The six, all on the one
SegmentStore.sealpath:promoteLiveFile,segstore.go:1156promoteLiveFile,segstore.go:1167writeIndexFile,segstore.go:1751writeIndexFile,segstore.go:1762writeManifest,segstore.go:732writeManifest,segstore.go:742Measured
Sealing a one-record tail repeatedly, reporting cost per seal and manifest size as the segment count rises:
strace -cover 200 of those seals: 1207 fsyncs (six per seal) and 603 renames (three per seal). At the ~5.5 ms per fsync this disk gives (see #32), six fsyncs is ~33 ms — which is the ~36 ms measured. The fsync count is the whole cost of a seal.Two things follow, and it is worth keeping them apart:
ImportSegmentFilecommits a manifest per adopted segment (segstore.go:1561), so a sync of N segments pays the same O(N²).Asks, in the order they pay
Both are independent of #30's tiered merge, which reduces how many segments exist; this is about what each seal costs regardless. #32 is the special case where the seal has nothing to seal at all.