Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
65 changes: 64 additions & 1 deletion docs/plans/tally-design.md
Original file line number Diff line number Diff line change
Expand Up @@ -123,6 +123,41 @@ the host reading it.
where the documents say one thing and the numbers are another, and nothing in
it is wrong enough to notice.

## It is fed sequentially, by a follower

The fold reads its source from beginning to end, once, in one process — the follower's `--run` consumer. The follower is not
merely the delivery mechanism: it holds the position durably, holds the
source's retention back while the tally is behind it, supplies the registry
and the systemd lifecycle, and restarts into exactly the re-fold `safe_offset`
makes correct.

⚠ **Parallel workers are not the first thing to reach for**, for two
reasons:

* **It needs additive partials**, which is their THIRD use after spilling and
repair — two workers can both contribute to one bucket, and with a merge
that replaces a cell the second silently erases the first. And it must
partition by SOURCE POSITION rather than by day: a day's entries are not
contiguous in the source (late arrivals are why a citation exists at all),
and there is no per-chunk logline range to find them by
([logline-order.md](logline-order.md)). Partitioning by chunk range falls
out for free, a worker's consumed range being the partial identity already
wanted.
* **The gain is small.** Measured over 300,000 real lines: reading records
0.52 s, the fold 4.53 s, writing blocks 0.05 s — so blocks are 1% and the
fold is 90%. A day is ~47 s single-threaded, and the largest backfill that
can ever be asked for is the SOURCE's retention (weeks, not years, since
nothing can tally what was dropped), so a 30-day rebuild is ~24 minutes
once.

⚠ **Nor is sharing the extraction, obvious as it looks.** The four metrics of
the measured document each run their own extract regex over every line, and
cost is spread evenly
and the metric with the SMALLEST regex (17 characters) is joint-most expensive
at 1.31 s, because it is a histogram of 22 buckets and turns one line into 22
samples. The cost is per sample produced and per metric evaluated, not per
regex byte, so sharing the extraction wins much less than it looks.

## Recovery is re-derivation

A follower's position is precious because it shipped bytes it cannot un-ship.
Expand All @@ -135,6 +170,29 @@ and not a correctness one.
nothing else. Nothing consults a watermark, so a re-derivation does not depend
on where the read started.

⚠ **And a record stream is forward-only, which is fine for a CRASH and not
for a repair.** The tally is fed a stream and cannot ask for bytes again, so
the two cases separate:

* **A crash needs no seek**, because of `Roller::safe_offset` — "the oldest
source byte any OPEN bucket still depends on. A consumer may not report past
this". The position is therefore always behind every unfinished bucket, so a
restart re-sends from before that bucket's FIRST entry, re-folds it whole,
and the complete total replaces the partial one. ⚠ **The block writer's
`How::Merge` depends on this, invisibly**: advance `safe_offset` any further
and blocks are corrupted by code that never mentions them. ⚠ That dependency
is a consequence of a merge REPLACING a cell — under additive partials
([tally-partials.md](tally-partials.md)) the position may advance freely.
* **A repair does need to go back, and a REWIND is the wrong way to do it.**
A tally is not only a consumer: `query --records --from X --to Y | tally`
is a bounded DIRECT read of the source, and it already exists. So "the regex
was wrong, recompute last week" is a one-shot pass over that window while
the follower keeps going forward — no position is moved, and the live path
is never interrupted. ⚠ What it needs is `How::Regenerate` rather than
`Merge`, which the writer does not expose: merging a recomputation into
what is there would replace cell by cell and leave any series the new
definition no longer produces standing.

## Copying is a file sync or a bundle

Blocks that are immutable once past the floor, under a manifest with a crc32
Expand Down Expand Up @@ -188,7 +246,12 @@ a working set that saturates rather than drifting.
for a partial — and it is also the atomicity defect below;
* **compaction's schedule**, and whether a query merges or refuses;
* **the write-batching mechanism** — a WAL for samples in the `.sap` shape is
the candidate, since one late sample otherwise rewrites a whole day block.
the candidate, since one late sample otherwise rewrites a whole day block;
* **a repair pass** — a bounded read with `How::Regenerate`, which is what
makes "fix the definition and recompute" an operation rather than a plan.
⚠ Not a rewind of the follower: the direct read already exists, and moving
a live position to recompute history would stop the live path to fix the
past.

⚠ **A defect that exists today:** `commit` renames a block into place and THEN
saves the manifest, so a rewrite at the same `(t0, generation)` leaves a window
Expand Down
1 change: 1 addition & 0 deletions packaging/timberfs-completion.bash
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,7 @@ _timberfs() {
--deadline | --positions | --batch-size | --follow-from | --delete-empty | \
--look-in | --etc | --extractor | --metric | --width | --grace | \
--provision | --run | --pack | --unpack | --block-buckets | \
--blocks | --block-flush | \
--query | --series | --since | --until)
COMPREPLY=($(compgen -f -- "$cur"))
return 0
Expand Down
36 changes: 36 additions & 0 deletions packaging/timberfs.1
Original file line number Diff line number Diff line change
Expand Up @@ -1332,6 +1332,42 @@ shows the tape as written, both lines included, which is what a tape
viewer should do. See LATENESS below for when a provisional line is
written.
.PP
.B WRITING THE GRID (EXPERIMENTAL, A HARNESS)
.PP
.BI \-\-blocks " DIR"
writes the numbers as columnar BLOCKS into DIR instead of as tally lines
on stdout \(em the storage the design settles on, where a line is the
INTERCHANGE form and not the store. Verified against 300,000 real log
lines: the blocks render back to exactly what the line path wrote.
.IP
⚠ A HARNESS rather than an interface. Nothing in a provisioned deployment
runs this: it is how the writer is exercised against a real store until
.B \-\-run
can write blocks itself. And it names a DIRECTORY where every other
timberfs argument names a store: the manifest carries an id, but nothing
searches for block stores by it yet, so a path is the only address there
is.
.IP
.BI \-\-block\-flush " N"
is how many samples are buffered before a commit (default 50000). ⚠ It is
a WRITE\-AMPLIFICATION control and not a latency one: a block is a day, so
each commit rewrites up to a megabyte, and one real day of a busy store is
~268,000 samples. Measured on 300,000 lines, the same numbers cost 43, 5
or 1 block writes at 1000, 10000 and 50000. The cost of buffering is
VISIBILITY \(em a sample is not in a block until it is flushed \(em and the
cost of losing the buffer to a crash is a re\-read, the position not
having moved. The durable form of it would be a write\-ahead log for
samples.
.IP
⚠ ONE store in, one directory out, which is why it is here and not on
.BR \-\-run :
a provisioned run serves a SELECTION, one sink per source store, and
where each one's blocks go is a provisioning question rather than a
writer one.
.IP
⚠ Nothing is written to stdout, the unit announcement included: a block
store carries its own definitions, and a unit is a property of one.
.PP
.B PACKING THE GRID (EXPERIMENTAL, A MEASUREMENT)
.PP
.BI \-\-pack " DIR"
Expand Down
27 changes: 27 additions & 0 deletions src/main.rs
Original file line number Diff line number Diff line change
Expand Up @@ -667,6 +667,28 @@ enum Command {
/// format --fold takes
#[arg(long)]
observations: bool,
/// EXPERIMENTAL, and a HARNESS rather than an interface: write
/// the numbers as columnar BLOCKS into DIR instead of as tally
/// lines on stdout, which is how the block writer is exercised
/// against a real store until a provisioned `--run` can write
/// them.
///
/// ⚠ It names a DIRECTORY where every other timberfs argument
/// names a store: the manifest carries an id, but nothing
/// searches for block stores by it, so a path is the only
/// address there is.
///
/// ⚠ One store in, one directory out. Not on `--provision`'s
/// `--run`, which serves a SELECTION with a sink per source
/// store: where each one's blocks go is a provisioning question.
#[arg(long, value_name = "DIR", conflicts_with_all = ["fold", "provision", "run", "try_it", "check", "pack", "unpack", "query", "observations"])]
blocks: Option<PathBuf>,
/// With --blocks: samples buffered before a commit. A block is a
/// day, so each commit rewrites up to a megabyte — this is the
/// write-amplification control, and a sample is not in a block
/// until it is flushed
#[arg(long, value_name = "N", default_value_t = timberfs::tally::DEFAULT_BLOCK_FLUSH, requires = "blocks")]
block_flush: usize,
/// EXPERIMENTAL, and a measurement rather than a feature: read
/// tally lines on stdin and write them as columnar BLOCKS into
/// DIR — the grid of series x buckets, one file per range,
Expand Down Expand Up @@ -1910,6 +1932,8 @@ fn main() -> anyhow::Result<()> {
grain::cmd_reindex(&file)?;
}
Command::Tally {
blocks,
block_flush,
extractors,
try_it,
check,
Expand Down Expand Up @@ -1958,6 +1982,9 @@ fn main() -> anyhow::Result<()> {
.map(append::parse_duration_ms)
.transpose()?;
tally::cmd_tally(&tally::TallyOpts {
blocks,
block_flush,
block_buckets,
extractors,
etc,
try_it,
Expand Down
Loading
Loading