Skip to content

feat(writing): keep file structure together, front-loaded by default - #180

Open
Blackclaws wants to merge 13 commits into
Apollo3zehn:devfrom
Blackclaws:feat/metadata-placement
Open

Blackclaws wants to merge 13 commits into
Apollo3zehn:devfrom
Blackclaws:feat/metadata-placement

Conversation

@Blackclaws

@Blackclaws Blackclaws commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

Structure is currently allocated in encode order, so it ends up spread evenly through the file. A reader that fetches ranges rather than seeking freely therefore pays for the whole file to read the structure: on a 620 MB reporting file, walking the 8.2% of bytes that are structure touched 591 of 591 one-megabyte ranges.

H5WriteOptions.MetadataPlacement adds two alternatives to the layout the writer produces today, which stays the default:

Placement Allocation Effect
Interleaved in encode order the default, and the layout the writer produces today, byte for byte
Aggregated from blocks that double in size, forming a few clusters needs no estimate of the total
FrontLoaded from one region reserved at the front a reader fetches all structure in one range

FrontLoaded sizes its reservation by measuring, so callers do not have to — and because the measurement is exact, it costs no file size at all.

Measured in RangeReadCostTests, over a block-fetching stream that counts what an HTTP range client would have transferred (600 groups of deflated measurement series, 256 kB blocks):

interleaved    fetched 4,050,543 B of 4,050,543 B   100.0% of file   16 requests
aggregated     fetched   786,432 B of 4,052,705 B    19.4% of file    3 requests
front-loaded   fetched   262,144 B of 4,050,543 B     6.5% of file    1 request

And in MetadataLayoutTests, on 2,000 groups each with a contiguous dataset and a string attribute, at 1 MiB ranges:

interleaved       16,878,158 bytes   structure in 17/17 ranges
front-loaded      16,878,158 bytes   structure in  1/17 ranges

Same byte count, one range instead of seventeen.

How it works

The writer allocates every address before writing and then seeks to it, so segregating regions needed no change to any encoding logic — only the allocator, plus classifying the allocation sites as structure or payload. Three of those are not what they look like:

  • An implicit chunk index is not an index. The chunks sit contiguously from the address in the layout message and a chunk's address is arithmetic on it, so what is allocated there is the whole payload of an unfiltered chunked dataset.
  • The global heap serves both kinds. An attribute's variable-length value is structure — it is what a viewer reads while browsing — while a dataset's own variable-length elements are payload. One classification for the whole heap has to be wrong for one of them, so the manager keeps a collection open per kind.
  • A compact dataset's payload lives inside its object header and makes no allocation of its own, so for datasets under 64 kB payload is structure and there is nothing to separate. Files that lean on PreferCompactDatasetLayout see little from this.

The reservation is sized by a pass that encodes the file against a stream discarding everything and reads the metadata total off the allocator. It is exact rather than estimated because it is the same encoder and the same allocator, which is what lets it account for what an estimate cannot: a chunk index's size scales with chunk count, and global heap collections are allocated at a 4 kB minimum each and left partly empty. An earlier estimator missed both and over-reserved threefold, so it was removed rather than kept as a fallback.

The pass does not compress, and for fixed-size element types it does not touch the data at all — no selection walk, no chunk buffers, no byte copying. Best of seven runs on a 64 MB file, against a full write of the same file:

unfiltered    34.9 ms -> 1.3 ms    43% of a write -> 2%
filtered      29.4 ms -> 0.9 ms   2.6% of a write -> 0.1%

MetadataLayoutTests.TheSizingPassMeasuresExactlyWhatAWriteAllocates asserts per shape — contiguous, compact, chunked filtered and unfiltered, single-chunk, variable-length strings, Nullable<T>, attributes, object references, many groups — that the pass measures exactly what a write allocates. That equality is what both shortcuts rest on, and an inequality either way is a defect: over-counting wastes the reservation's tail, under-counting spills.

Set MetadataReservation to skip the pass. Nothing else requires it — including a deferred write through BeginWrite, since nothing a caller writes later changes how much structure the file has.

Worth flagging

  • There is no free list, so a region's unused remainder is abandoned rather than reclaimed. A measured reservation abandons nothing, so this is only a cost when you set MetadataReservation yourself, or when a reservation comes up short and spills into blocks — which loses locality for the remainder rather than failing.
  • Metadata blocks grow rather than being fixed. A block is claimed in full, so a fixed size is a floor on file size: at the 8 MB default, 33 kB of metadata produced an 8.4 MB file. Blocks now start at 64 kB and double up to MetadataBlockSize, which bounds what is claimed at roughly twice what is used. MetadataBlockSize is therefore a ceiling, and must be positive for any placement that uses blocks.
  • The end-of-file address covers allocated rather than written space, because an address can be referenced without ever being written to: a dataset declared through BeginWrite and never written has its payload allocated and named by its layout message. Under a clustered placement nothing extends the stream over it, and h5dump then reports "invalid dataset size, likely file corruption". The stream is extended to the allocator's mark instead — translated by the base address, since the allocator counts from the superblock and the stream from the start of the file.
  • Overlaps feat(writing): size a string attribute to its own value #179 only in H5WriteOptions, where both add a property to the record body; whichever lands second takes a one-line conflict.
  • No CHANGELOG.md entry here, on the assumption you would rather write that yourself — happy to add one.

Changes since the first review

Everything below came out of your two comments and of reviewing the result again afterwards.

  • Interleaved is the default, as requested. The feature is now purely additive: files written with default options are byte-for-byte what they were.
  • The sizing pass no longer walks fixed-size data — your suggestion, implemented; see the reply on that thread for the two corrections it needed.
  • Corrected the BeginWrite claim, twice. It was never chunk indexes, and it is no longer variable-length data either. NothingDeferredOutrunsTheSizingPass asserts the shortfall is zero for all three.
  • Front loading now costs nothing. The reservation added 12 kB of slack, justified on a reading of the superblock accounting that was backwards — it left the region long, not short. Subtracting instead makes it exact.
  • A dataset's variable-length values are no longer counted as structure. They shared the global heap with attribute values, so a variable-length string dataset with no attributes measured 97.3% structure and a front-loaded region swallowed its payload. Now 0.0%.
  • Fixed a corruption bug of my own: the extension above compared an allocator address against a stream length without translating for UserBlockSize, so a user block plus a never-written deferred dataset produced a file h5dump rejects.
  • Fixed a misclassification of my own: the implicit chunk index, per the bullet above. Worth knowing that the measurements in the first version of this description therefore did not cover unfiltered chunked data.
  • Made three assertions able to fail. TestUtils.DumpH5File strips h5dump's error stack so the fixture comparisons do not depend on the installed version — which also strips every trace of failure, so DoesNotContain("error") always passed. On lzf.h5, which h5dump rejects with exit code 1 and 2,910 bytes of error stack, the stripped output contains "error" zero times. There is now a RunH5Dump that reports the exit code, and it also returns null when h5dump is absent so that Skip.If(dump is null) skips instead of failing — it was dead code, and those tests failed on a machine without h5dump.
  • Reworked the RangeReadCostTests payload from a strided integer ramp to a measurement series. The ramp is pathological for a fast match finder: .NET 9 replaced the bundled zlib with zlib-ng, and that shape compressed 73% worse there, moving every figure the test reports. The figures above now land within 0.2% on both runtimes with identical request counts — and the request count is now asserted rather than merely printed.
  • MetadataBlockSize <= 0 and a negative MetadataReservation now throw instead of silently turning Aggregated into Interleaved.
  • Added a Metadata placement section to doc/writing/index.md, since the feature is opt-in and would otherwise only be discoverable from XML docs.

Two pre-existing issues turned up along the way, neither related to placement and neither touched here: #184 (deflate at compression-level: 1 expands data on .NET 9+) and #183 (any UserBlockSize above 512 produces a file PureHDF cannot reopen). #183 is why the new UserBlockSize coverage uses 512 and no larger value.

@Apollo3zehn

Copy link
Copy Markdown
Owner

Thanks for this PR as well! This is a nice feature for high latency environments :-)

I started the review and put some time into figuring out why it is necessary to still write the data into the Stream that measures the metadata size because preparing the data for write does not come for free and it would be good to skip unnecessary heavy work. The answer is two fold: There are cases where we can skip that and some where it is not possible, i.e. when writing the data itself causes metadata modifications which is the case for variable length data. So I asked GPT for a plan. Please tell me if you agree to its content or if you have any doubts. If the plan is fine I will add a commit with the proposed changes. The plan:

Suggested improvement: skip raw data writes during metadata sizing for fixed-size datasets

The metadata sizing pass currently avoids compression for filtered chunks, but it still walks dataset data, allocates chunk buffers, copies bytes, and registers raw chunk data. For fixed-size unmanaged datasets, that work appears unnecessary because it does not affect metadata size.

Why this should be safe

MeasureMetadataSize already performs the metadata-relevant work:

  • It encodes the file structure with MetadataPlacement = Interleaved and SizeOnly = true.
  • It creates each dataset layout and H5D_Base.
  • It stores each dataset in Context.DatasetToInfoMap.
  • It disposes the writer, which disposes each H5D and finalizes chunk index metadata.
  • It encodes global heap collections before writing the final superblock.

For fixed-size unmanaged dataset payloads, WriteData only copies raw data into the target layout. That raw data copy does not change the amount of metadata that must be reserved.

Chunked layout details

For chunked datasets, skipping raw data writes during the sizing pass should still preserve metadata sizing:

  • Filtered multi-chunk datasets use fixed-array indexing.
  • The fixed-array metadata size depends on WriteChunkInfos.Length, ChunkSizeLength, and page information derived from the entry count.
  • It does not depend on the actual chunk address, compressed chunk size, or filter mask values.
  • If data writing is skipped, chunk entries may remain at default values, but their encoded width and count remain unchanged.
  • Unfiltered multi-chunk datasets use implicit indexing; their data region is allocated during layout creation, so WriteData does not add metadata.
  • Single-chunk datasets store the relevant indexing information in the layout message.
  • Contiguous datasets allocate payload as AllocationKind.RawData, not metadata.

This matches the existing observation that compressed output size is not metadata-size-relevant: fixed-array entries are fixed-width, and ChunkSizeLength is derived from the uncompressed chunk size.

Keep data encoding where it has metadata side effects

The sizing pass must still run WriteData for datasets whose element encoding can allocate or alter metadata:

  • Compact datasets, because the encoded data is stored in the object header.
  • Variable-length strings and variable-length sequences, because encoding allocates global heap metadata.
  • Object references, because encoding may recursively encode referenced groups or datasets.
  • Nullable/reference-containing element types, because they are routed through the per-element encoding path.

Proposed code shape

Add a small guard in the write path, after dataset encoding/layout creation but before WriteData:

if (context.SizeOnly && !RequiresSizeOnlyDataWrite<T>(dataset))
    return;

A possible predicate:

private static bool RequiresSizeOnlyDataWrite<T>(H5Dataset dataset)
{
    return
        dataset.Layout is DataLayoutMessage4
        {
            Properties: CompactStoragePropertyDescription
        }
        || DataUtils.IsReferenceOrContainsReferences(typeof(T))
        || Nullable.GetUnderlyingType(typeof(T)) is not null;
}

The exact location would likely be InternalWriteDataset<T, TElement>, since InternalEncodeDataset should still run in the sizing pass to create layouts, datatype messages, dataspace messages, H5D instances, and deferred metadata structures.

Expected benefit

This should avoid unnecessary work during metadata measurement for fixed-size datasets:

  • no source-data traversal
  • no chunk-buffer allocation
  • no raw byte copying
  • no raw chunk registration/allocation for filtered chunks
  • no selection walking beyond what is needed for metadata

Compression is already skipped, so this targets the remaining fixed-size raw-data overhead in the measurement pass.

@Apollo3zehn

Copy link
Copy Markdown
Owner

Apart from that, due to the complexity involved and the limitations being imposed ("BeginWrite needs an explicit MetadataReservation") I would like to keep Interleaved the default.

BeginWrite needs an explicit MetadataReservation. The measuring pass cannot know what a caller will write later, so it under-counts those chunk indexes and the shortfall is placed as Aggregated would place it.

I am not sure I fully understand this statement. Why is it under-counting chunk indexes? The number of chunks are always known beforehand and will not change during BeginWrite.

Blackclaws and others added 9 commits September 1, 2026 17:21
Structure is allocated in encode order, so it ends up spread evenly through the
file. A reader that fetches ranges rather than seeking freely therefore pays for
the whole file in order to read the structure: on a 620 MB reporting file,
walking the 8.2% of bytes that are structure touches 591 of 591 one-megabyte
ranges.

H5WriteOptions.MetadataPlacement adds two alternatives to the encode-order
layout, which stays the default:

  Interleaved  in encode order. The default, and what the writer has always
               produced, byte for byte.
  Aggregated   structure is allocated from large blocks, so it forms a few
               clusters. Needs no estimate of the total.
  FrontLoaded  structure is allocated from one region reserved at the front, so
               a reader can fetch all of it in one range.

Measured on 2,000 groups, each with a contiguous dataset and a string attribute,
at 1 MiB ranges:

  interleaved       16,878,158 bytes   structure in 17/17 ranges
  front-loaded      16,890,494 bytes   structure in  1/17 ranges   +0.07%

The reservation is sized by measurement rather than by estimate: a pass encodes
the file against a stream that discards everything and reads the metadata total
off the allocator. That is exact because it is the same encoder and the same
allocator, which is what lets it account for the two things an estimate cannot -
a chunk index's size scales with chunk count, and global heap collections are
allocated at a 4 kB minimum each and left partly empty. Set MetadataReservation
to skip the pass.

The pass does not compress. No metadata size depends on compressed output - a
filtered chunk index entry's size field is sized from ChunkByteSize, the
UNCOMPRESSED chunk size - and compression is around 97% of a filtered write, so
the pass costs at most 1.6% of one. On an unfiltered write the share is larger,
but such a write is cheap to begin with: 264 MB in 324 ms.

The writer allocates every address before writing and then seeks to it, so
segregating regions needs no change to any encoding logic - only the allocator,
plus classifying the allocation sites as structure or payload. Two of those
sites are not what they look like:

- An implicit chunk index is not an index at all. The chunks sit contiguously
  from the address in the layout message and a chunk's address is arithmetic on
  it, so what is allocated there is the whole payload of an unfiltered chunked
  dataset. Counted as structure it would front-load the payload along with
  everything else and dissolve the separation, silently - the file stays correct
  and the same size.
- A compact dataset's payload lives inside its object header and makes no
  allocation of its own, so for datasets under 64 kB payload IS structure and
  there is nothing to separate. Files that lean on PreferCompactDatasetLayout
  will see little from this.

Two more things worth knowing:

- There is no free list, so a region's unused remainder is abandoned rather than
  reclaimed; that is the file-size cost above. A reservation that comes up short
  spills into blocks rather than failing.
- The end-of-file address must cover allocated rather than written space, because an
  address can be referenced without ever being written to: a dataset created through
  BeginWrite has its payload allocated and named by its layout message whether or not
  the caller ever writes it. Nothing extends the stream over that payload, and the
  placement decides whether anything else does - an interleaved file usually has
  metadata allocated after it, whose writes extend the stream by accident, while a
  front-loaded one serves metadata from the front and ends before the payload the
  layout points at. h5dump then rejects the dataset: "invalid dataset size, likely
  file corruption". So the stream is extended to the allocator's mark rather than the
  address being trimmed to the stream - trimming would be smaller, since an abandoned
  region tail is usually the last thing in the file and referenced by nothing, but it
  cannot be told apart from the payload case and getting it wrong truncates data.

Writing variable-length data through BeginWrite is the one case a measurement
cannot reach: how much global heap such data needs follows from the values, which
the pass has not seen. Everything else about a deferred dataset is covered, since
a chunk index is sized from the chunk count and that comes from the dimensions.
Set MetadataReservation there.

Not a breaking change: nothing is removed, renamed or altered in signature,
H5WriteOptions gains properties, the placement enum is new, and the default
places exactly as before.
Asserts the locality of the placement from the reader's side rather than from the
written bytes: a stream that serves reads in fixed-size blocks and counts them
stands in for an HTTP range-request client, and the same tree is written both
ways so the only variable is where the structure went.

Front loading has to at least halve the transfer and hold fewer blocks. At a
256 kB block over 600 groups, each with a deflated dataset and two attributes:

  interleaved    3,666,125 B fetched of 3,666,125 B   100.0% of file   14 requests
  front-loaded     262,144 B fetched of 3,678,461 B     7.1% of file    1 request

Deflate is not incidental - without a filter these datasets are stored compact,
payload inside the object header, and there is no separation left to measure.

Co-Authored-By: Claude <noreply@anthropic.com>
The pass exists to total up metadata, so it only has to do work that decides how
much metadata there is. Copying fixed-size payload into a layout decides nothing:
a contiguous region and an implicit chunk index are raw data, a fixed-array index
is sized from the chunk count, and every index's encoded width comes from the
dataspace rather than from a value. So for fixed-size elements the pass now skips
the data path entirely - no selection walk, no chunk buffers, no byte copying, no
raw chunk allocation - on top of already skipping compression.

The pass still encodes data whose encoding allocates:

  compact          payload lives inside the object header, so here data IS
                   metadata
  strings, jagged  variable-length data takes global heap space sized from the
  arrays, structs  values themselves
  holding either
  object refs      encoding one encodes the object it points at, header and all
  Nullable<T>      encoded as a variable-length sequence of one, so it takes
                   global heap space too - and holds no reference, so the
                   reference test does not see it

The test is on the element type rather than on the dataset type, because an array
is itself a reference: asked of the dataset type it answers yes for every
array-backed dataset and nothing would be skipped.

Sizing pass against a full write of the same 64 MB file, best of seven runs each -
the fastest run being the one least perturbed by whatever else the machine is doing:

  unfiltered    34.9 ms -> 1.3 ms    43% of a write -> 2%
  filtered      29.4 ms -> 0.9 ms   2.6% of a write -> 0.1%

MetadataLayoutTests asserts per shape that the pass still measures exactly what a
write allocates, which is what the shortcut rests on - an over-count wastes the
reservation's tail and an under-count spills.
A block is claimed in full, so a fixed block size is a floor on file size rather
than a proportional cost. At the 8 MB default, 33 kB of metadata produced an
8.4 MB file - 205 times the interleaved layout of the same content. Aggregation
paid it on any file whose metadata was under 8 MB, and a front-loaded write paid
it too whenever its reservation was exhausted, which deferred variable-length
data guarantees: the global heap such data needs is sized from values the sizing
pass never saw.

Blocks now start at 64 kB and double up to MetadataBlockSize, which bounds what
is claimed at roughly twice what is used. Deferred strings, 33 kB of metadata:

  interleaved      40,974 bytes
  aggregated       73,584 bytes   was 8,396,656
  front-loaded     86,078 bytes   was 8,409,150

The cost is a handful of clusters rather than one, since the ramp adds at most
log2(MetadataBlockSize / 64 kB) blocks before reaching the ceiling - a constant,
after which the steady state is what it was. Measured over the range-fetching
stream from RangeReadCostTests, 3,000 groups with deflated datasets at 256 kB
blocks:

                          file          fetched   requests
  interleaved       18,343,706       18,343,706         70
  aggregated        18,378,700        1,835,008          7   was 25,784,268 / 4
  front-loaded      18,356,042        1,048,576          4

So aggregation trades three range requests for not inflating the file by 41%, and
front loading - the placement to reach for when round trips matter - is untouched.
MetadataBlockSize becomes a ceiling rather than a size; nothing about it changes
for a file with enough metadata to reach it.
MetadataAbandoned counted a region's remainder only when the region was replaced.
A region is replaced only when a request does not fit it, so those remainders are
small by construction, while the tail of the region left open at the end - the
part that is actually large - was never counted at all. It reported 48 bytes of
waste on a write that had claimed 8 MB and used 33 kB of it.

It now includes the open region's tail, which makes it agree exactly with what the
file grew by. Those are independently measurable: growth is observed from outside,
abandonment reported from inside, and since a region is claimed in full and there
is no free list, the two are the same number. AbandonedSpaceAccountsForTheFileGrowth
asserts that equality rather than a bound.

The figure is live rather than final while a write is in progress - it answers what
would be abandoned if the write ended now. Nothing in the library reads it; it is
diagnostic.
The comparison covered interleaving against front loading only, which leaves the
choice between the two clustered placements unmeasured. They answer different
questions - a reservation is sized by measuring the file first and gets the
structure into one range, aggregation needs no such pass and settles for a handful
- so a caller picking between them wants the gap, and a caller who cannot afford
the sizing pass wants to know aggregation is still worth having.

600 groups, deflated datasets, 256 kB blocks, on .NET 8:

                    file          fetched   requests
  interleaved  3,666,125        3,666,125         14
  aggregated   3,673,087          524,288          2
  front-loaded 3,678,461          262,144          1

Aggregation must halve what a walk transfers, and a measured reservation must not
do worse than not measuring. It must also not buy locality with file size: blocks
are claimed in full, so a block of fixed size is a floor rather than a proportional
cost, and at the 8 MB default this file came out three times larger while still
walking in one request. That is bounded here from the reader's side, where the
trade is visible, as well as in MetadataLayoutTests from the writer's.
The payload was a strided integer ramp, chosen only so that it would not compress
away to nothing. That turns out to be pathological for a fast match finder: .NET 9
replaced the bundled zlib with zlib-ng, whose match finder differs below maximum
effort, and this shape compressed 73% worse there than on .NET 8. Every figure the
test reports moved with it, which made the numbers quoted for the placements a
property of the target framework as much as of the layout.

It is now a periodic signal with a slow drift, rounded to two decimals - which is
what these files actually hold. It compresses to about 39%, so the file is neither
dominated by payload nor too small for locality to be expressible, and it lands
within 0.2% on both runtimes:

                       .NET 8      .NET 10
  interleaved       4,050,543    4,058,242
  aggregated        4,052,705    4,060,404
  front-loaded      4,062,879    4,070,578

Fetched bytes and request counts are now identical on both: 16 requests and the
whole file interleaved, 3 and 786,432 B aggregated, 1 and 262,144 B front-loaded.

The rounding is what buys the compression - full-mantissa doubles differ in their
low bytes at every sample and barely compress - and 2,048 of them is the same 16 kB
per group as before, so the file stays the size the block arithmetic assumes.
The guidance for deferred writes is that MetadataReservation is worth setting for
variable-length data and unnecessary for fixed-size data. That rested on reading the
code rather than on a measurement, and the first version of it was wrong in the
other direction - it claimed chunk indexes were under-counted, which they are not.

  filtered chunked int (fixed array)        measured   495   allocated     495   shortfall       0
  unfiltered chunked int (implicit index)   measured   192   allocated     192   shortfall       0
  contiguous strings (global heap)          measured   200   allocated 295,112   shortfall 294,912

A chunk index is fully accounted for however the chunks are indexed, because its
size follows from the chunk count and that follows from the dimensions, which are
fixed when the dataset is declared: WriteChunkInfos is sized to TotalChunkCount in
H5D_Chunk4.Initialize, and writing data later fills entries in an array whose length
was settled then. Global heap space is not accounted for, because how much a
variable-length value needs follows from the value.

So the shortfall is zero for one and 288 kB for the other, which is the distinction
the guidance draws, now asserted rather than reasoned.
The allocator counts addresses from the superblock; the driver's Length and
SetLength count from the start of the file. A user block puts those a
UserBlockSize apart, and the extension compared them untranslated - so with a user
block the file looked long enough already and the extension was skipped.

That matters for exactly the case the extension exists for: a dataset declared
through BeginWrite and never written has its payload allocated and named by its
layout message, and nothing else extends the stream over it under a clustered
placement. h5dump rejects the result:

  H5D__contig_check(): invalid dataset size, likely file corruption

for both Aggregated and FrontLoaded at UserBlockSize = 512, where the same file
with no user block is clean. Interleaved survived by accident, as it does without a
user block - metadata allocated after the payload extends the stream on its own.

EndOfFileAddressTests now runs each placement with and without a user block. The
zero case cannot stand in for the nonzero one, since the whole defect is the
translation between the two coordinate systems.

512 rather than a larger user block because that is the only size this library
round-trips today: the superblock search in NativeFile.InternalOpenAsync doubles
its step while seeking relatively, so it probes 512, 1536, 3584 rather than the
512, 1024, 2048 the format specifies. That is a reader bug of its own, unrelated to
placement, and is reported separately.
@Blackclaws

Copy link
Copy Markdown
Contributor Author

I'm on it. I've taken your proposal and Claude implemented it after fixing a couple of things it found in there that didn't work. Also found a number of other edge cases, an update to the PR is coming either today or tomorrow.

The reservation was the measured total plus 3 * 4096 bytes of slack, justified on
the grounds that the measurement counts the superblock while the superblock is
served from outside the region, "leaving the region short by exactly that much".
It leaves the region LONG by that much. Subtracting rather than adding makes the
reservation the exact number of bytes that will be served from it.

Front loading therefore costs nothing at all rather than a little. Measured across
five shapes - contiguous, filtered chunked, variable-length strings, Nullable<T>,
200 groups - the cost over the interleaved layout of the same content:

  measured + 3 * 4096      12,336 bytes abandoned
  measured                     48 bytes abandoned
  measured - superblock          0 bytes abandoned

One region opened in every case, at every setting: the slack was never keeping a
final allocation from spilling, because nothing was spilling. On the 2,000-group
fixture a front-loaded file is now 16,878,158 bytes against an interleaved
16,878,158, with structure in 1 of 17 ranges rather than 17.

Nothing caught this. AutoSizedFrontLoadingCostsNothing now asserts the sizes are
equal rather than within half a percent, and asserts MetadataRegionsOpened is 1
with nothing abandoned - the property that distinguishes a measurement that was
short (spills, losing the locality) from one that was long (pays for a tail nothing
uses). MetadataRegionsOpened existed to report exactly that and was checked nowhere.

Also in this file, two claims that did not survive being measured:

- MetadataLayoutTests.Options reserved 3 MB and said 1 MB would spill into a second
  region. The measured need is 878,158 bytes and 1 MB opens one region, so the
  reservation is now 1 MB and the comment says what it is for - exercising the path
  that skips the sizing pass, at the price an explicit reservation pays, which is
  its unused tail. That price was 13.44% of the file and is now 1.01%.
- MeasureBaseline wrote with default options, so it ran with compact layout ON,
  which the class remark calls out as the one thing that must be off for these
  files to measure anything.
…a heap

Global heap collections were all allocated as structure, justified on the grounds
that the heap holds attribute payload - "what a viewer reads while browsing rather
than while reading a dataset". That is true of an attribute's value and false of a
dataset's own variable-length elements, which go through the same manager. A
dataset of variable-length strings with no attributes at all measured 97.3%
structure, so a front-loaded region swallowed the payload and there was nothing
left to separate. Silently: the file stays valid and the same size, and only the
locality quietly fails to materialise.

The manager now keeps one collection open per allocation kind and allocates each
with its own kind. Per kind rather than per object, because a collection is
allocated whole - one shared collection would put both kinds at whichever address
the first object claimed.

  vlen strings as dataset, no attributes    97.3% -> 0.0%
  jagged arrays as dataset, no attributes   97.3% -> 0.0%
  one large string as an attribute          100%  -> 100%
  vlen dataset plus a string attribute              0.4%

Which kind is being encoded cannot be passed in: the delegates that call AddObject
come from DatatypeMessage.Create, shared by both paths. So the writer scopes it at
the two points where it can - a dataset's data write and an attribute's encode -
and both restore it, so nesting stays correctly attributed. An object reference
among a dataset's elements encodes the object it points at, and that object's
attributes are still attributes.

Two consequences worth knowing:

- A deferred write now needs no reservation at all. Variable-length data written
  after BeginWrite was the one case a measurement could not reach; its heap is
  payload now, so the shortfall is zero and the file comes out the size an
  interleaved one does. NothingDeferredOutrunsTheSizingPass asserts that, and is
  the third version of this guidance - it said the opposite twice.
- GlobalHeapManager's flush now clears every open collection rather than one. A
  flush encodes and drops the whole map, so a collection left open afterwards would
  take objects nothing would ever encode again.
…rate options

Three assertions that could not fail, and one option that meant nothing.

DumpH5File strips h5dump's HDF5-DIAG error stack so the 134 fixture comparisons do
not depend on the locally installed version. That also strips every trace of
failure, so asserting the dump does not contain "error" always passed. On a file
h5dump rejects outright - lzf.h5, exit code 1 with 2,910 bytes of error stack - the
stripped output contains the word "error" zero times.

RunH5Dump now reports the exit code and both streams unmodified, and the placement
and end-of-file tests ask it whether the C library accepted the file. DumpH5File
keeps its current output for the fixture comparisons.

It also returns null when h5dump is not installed rather than throwing
Win32Exception, which makes Skip.If(dump is null) do what it says. It was dead
code: these tests failed rather than skipped on a machine without h5dump, which
would have hit CI or a contributor and not me.

MetadataBlockSize of zero or less silently turned Aggregated into Interleaved -
the block path is gated on a positive size, so a caller asked for clustering, got
none, and had nothing to tell them why. That and a negative MetadataReservation now
throw. Interleaved still ignores the block size, because it genuinely uses no
blocks.

And RangeReadCostTests now asserts the request count. It is the headline figure of
this feature and the unit every claim about it is stated in, and it was printed and
never checked: front loading must take exactly one request, and aggregation a
handful against interleaving's one per block.
…ttering it

Classifying a dataset's variable-length values as payload fixed what a reader
fetches and broke what a writer can stream, and both matter.

A global heap collection is allocated when it is needed and written much later,
when it flushes. A writer uploading to object storage buffers the file in fixed
parts and can only ship a part once every byte in it is written, so such a
collection is a hole that pins its part until the flush. Scattered through the
file, those holes pin every part and the whole object stays in memory. Gathered,
they pin a few.

Measured over that upload model - 2,000 groups, a fixed-size dataset and one small
variable-length value each, 1 MB parts, so payload dominates and the values pack
into shared collections:

  variable-length values as ATTRIBUTES   parts held   shipped during the write
    Interleaved                                  33                          0
    Aggregated                                    6                         28
    FrontLoaded                                   3                         31

  variable-length values as DATASETS
    Interleaved                                  33                          0
    Aggregated                    was 7, 33 after the reclassification, now 9
    FrontLoaded                   was 3, 33 after the reclassification, now 6

So the reclassification alone cost 3 parts to 33 - 3.1 MB resident against 34.6 MB,
scaling with file size rather than staying flat.

AllocationKind.DeferredRawData is therefore payload that is clustered: its own
blocks, never the metadata region, so a reader walking the structure still does not
fetch it while a writer still sees parts complete. Six rather than the original
three is the honest price of two clusters instead of one, which is what keeping
payload out of the structure region costs.

Chunk and contiguous payload needs none of this - it is written at the address it
was allocated and leaves no hole.

The clustering logic is now shared, since two independent streams of allocations
need it, and PayloadAbandoned reports the second region's unused tail.
MetadataAbandoned and MetadataRegionsOpened stay structure-only: a reservation
should not be sized against payload. Together the two figures still account for
every byte a clustered file carries that an interleaved one does not, which
AbandonedSpaceAccountsForTheFileGrowth asserts.
@Blackclaws

Copy link
Copy Markdown
Contributor Author

Thanks — and yes, I agree with the premise, and with where the plan draws the line. I have implemented it (perf(writing): skip fixed-size data in the metadata sizing pass) rather than leaving you to write the commit, since checking it turned up two things worth flagging. Drop that commit and take your own if you would rather.

Where the plan is right: nothing about a fixed-size dataset's metadata footprint depends on its data. I checked every encode size involved — SingleChunkIndexingInformation.GetEncodeSize and ImplicitIndexingInformation.GetEncodeSize ignore their values entirely, and the chunk-index case is the one you raise on the other thread. The list of things that must still be encoded is right too: compact, variable-length, object references, Nullable<T>.

Two corrections it needed:

The guard has to sit at the WriteData call inside InternalEncodeDataset, not in InternalWriteDataset. InternalWriteDataset is only the deferred path — what writer.Write(dataset, data) reaches. The sizing pass never calls it: it calls writer.Write() and then Dispose, so all of its data encoding goes through InternalEncodeDataset's own WriteData call. Guarding InternalWriteDataset would have skipped nothing at all, and nothing would have failed to say so.

The predicate has to test TElement, not T. T is the dataset type, and an array is itself a reference, so IsReferenceOrContainsReferences(typeof(int[])) is true — asked of T the predicate answers "needs data" for every array-backed dataset, which is very nearly all of them. Same silent no-op. On TElement it answers as intended, and it also covers the object-reference case for free, since H5ObjectReference is a class.

One addition: Nullable<T> is worth keeping in the predicate for a reason the plan does not give. It is not that nullables go through the per-element path — it is that GetTypeInfoForNullableValueType encodes them as a variable-length sequence of one and calls GlobalHeapManager.AddObject, so they take global heap space. Nullable<int> contains no reference, so the reference test does not see it. MetadataLayoutTests has a nullable ints shape that pins this: drop the clause and it measures 200 bytes against the 12,488 a write allocates.

What I did not do is make the compact case conditional. It is safe to skip too — CompactStoragePropertyDescription.Data is allocated by size in DataLayoutMessage4.Create, so the object header's size is already settled — but a compact dataset is capped at 64 kB, so there is nothing to win, and keeping it documents that payload and structure are the same bytes there.

Measured on a 64 MB file, sizing pass against a full write of the same file, best of seven runs each — the fastest run being the one least perturbed by whatever else the machine is doing:

unfiltered    34.9 ms -> 1.3 ms    43% of a write -> 2%
filtered      29.4 ms -> 0.9 ms   2.6% of a write -> 0.1%

The safety net is TheSizingPassMeasuresExactlyWhatAWriteAllocates, which asserts across ten shapes that the pass reports exactly what a write allocates. Equality, not a bound: over-counting wastes the reservation's tail and under-counting spills, so the number has to be the number. That is what makes both shortcuts — no compression, and now no fixed-size data — checkable rather than argued, and it went on to catch two allocation bugs of mine that had nothing to do with this commit.

@Blackclaws

Copy link
Copy Markdown
Contributor Author

Interleaved is the default now. Agreed on the reasoning, and it makes the change purely additive — default-option files are byte-for-byte what they were.

On the under-counting: you are right, and my statement was wrong. I checked it rather than argued it, and the shortfall for a deferred chunk index is exactly zero:

filtered chunked int (fixed array)        measured 495   allocated 495   shortfall 0
unfiltered chunked int (implicit index)   measured 192   allocated 192   shortfall 0

Exactly as you say: the chunk count is known before anything is written. WriteChunkInfos is sized to TotalChunkCount in H5D_Chunk4.Initialize, the fixed-array data block is sized from WriteChunkInfos.Length, and Initialize runs during the structure write, which the sizing pass performs in full. Writing the data later fills entries in an array whose length was already fixed.

I then replaced that wrong claim with a second wrong claim, so it is worth being explicit about where it ended up. I said the real shortfall was variable-length data, which was true at the time and is no longer: variable-length values were being allocated on the global heap as structure, and reviewing this again showed that is wrong for a dataset's own elements. An attribute's value is structure — it is what a viewer reads while browsing a file — but a dataset's variable-length elements are its payload, and both went through the same manager. A dataset of variable-length strings with no attributes at all measured 97.3% structure, which meant a front-loaded region swallowed the payload and there was nothing left to separate. Silently, too: the file stays valid and the same size, and only the locality quietly fails to materialise.

With that fixed, a variable-length value's heap space is payload, so writing more of it later does not move the metadata total either:

contiguous strings (global heap)          measured 200   allocated 200   shortfall 0

So the honest answer to your question is better than either of my attempts at it: nothing a caller writes after BeginWrite changes how much structure the file has, and MetadataReservation is not needed for a deferred write at all. A deferred variable-length write under FrontLoaded now opens one region, abandons nothing, and produces a file the same size as the interleaved one. NothingDeferredOutrunsTheSizingPass asserts all three shortfalls are zero — this guidance has been wrong twice, so it seemed worth pinning rather than restating.

The same review turned up a second misclassification of the same kind — an implicit chunk index is not an index, so what is allocated there is an unfiltered chunked dataset's whole payload — plus a corruption bug and a few other things. Those are in the rewritten description rather than here. The pattern they share is worth one sentence though: all of them were invisible from outside, because a misclassified allocation still produces a valid file of the same size and only the locality quietly fails to appear. It follows that the measurements in the first version of this description did not cover unfiltered chunked data or variable-length datasets, and the ones now in it do.

@Blackclaws

Copy link
Copy Markdown
Contributor Author

Force-pushed, and the description is rewritten rather than amended — worth a fresh read, since editing it does not notify.

The two replies above answer your comments. Everything else came out of reviewing the result again afterwards, and is listed under Changes since the first review rather than repeated here. Four of those are defects in code you had already read, so they are the ones I would not want buried:

  • The extension that covers allocated-but-unwritten space compared an allocator address against a stream length without translating for UserBlockSize, so a user block plus a never-written deferred dataset produced a file h5dump rejects.
  • The reservation added 12 kB of slack that was never needed. Front loading now costs zero bytes rather than nearly zero.
  • A metadata block was claimed in full, so the 8 MB default was a floor on file size: 33 kB of metadata produced an 8.4 MB file. Blocks now grow.
  • Three h5dump assertions could not fail, because TestUtils.DumpH5File strips the error stack the assertions were looking for.

Two unrelated pre-existing bugs turned up on the way and are filed separately: #183 and #184. Neither is touched by this PR.

Two things still open for you rather than for me: whether you want a CHANGELOG.md entry from me or would rather write it, and whether Aggregated is worth keeping at all — front loading now dominates it on every axis I can measure except needing no sizing pass, so if you would rather carry one option instead of two, dropping it is a smaller surface to maintain.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants