Conversation
|
Do you have evidence this is faster? there are some benchmarks (that unfortunately we keep forgetting about) that may be useful as a starting point |
|
Sure, this is what I was able to measure: The strongest evidence is the allocation reduction: all three scenarios allocate materially less memory per operation, and the wall-clock timings move in the same direction. |
|
@twsouthwick Would you like additional performance testing? |
|
@twsouthwick Any chance we can review this? |
twsouthwick
left a comment
There was a problem hiding this comment.
there are two changes going on here - can you explain why each are needed? I'd like to see each as their own change to understand if they are needed.
|
There are two separate optimizations that work together to improve the combined benchmark performance during loading. The first changes how child elements are found. Every lookup used to create a temporary ElementFactory object just so Array.BinarySearch could search for it. The new implementation compares the element name directly, avoiding that allocation. It uses a simple scan when there are four or fewer possible children because checking a few entries directly is cheaper. Larger collections still use binary search so they remain efficient. The second change reduces repeated work while loading children and attributes. Values that remain the same during loading, such as:
It also skips markup-compatibility tracking entirely when processing is disabled. |
|
@twsouthwick Still looking to merge this if possible as it improves loading performance in very large spreadsheets. |
Improves DOM population performance by reducing per-element overhead in the OpenXML SDK hot path: child element factory lookup no longer allocates a temporary lookup object for every XML node, tiny child factory lists use a cheaper linear scan while larger lists keep binary search, Populate reuses cached namespace resolver and markup-compatibility state, skips no-op MC processing work in the default NoProcess mode, and attribute loading avoids repeated/cold lookups by caching stable values. These changes are internal-only and preserve existing load semantics while making large worksheet/document materialization cheaper.