Context
pdfTextLayout.groupLines() (#110) orders items by baseline top-to-bottom, then by x left-to-right, and puts every item that shares a baseline on the same line. That is correct for single-column documents, but a classic two-column paper becomes:
LEFT COLUMN RIGHT COLUMN parsed as
A line 1 B line 1 A line 1 B line 1
A line 2 B line 2 A line 2 B line 2
A line 3 B line 3 A line 3 B line 3
and a paragraph bbox can span both columns.
KnowNote is a research tool, so two-column academic PDFs are a primary input, not an edge case. This does not need to be solved before citations ship — the page number and block text are still correct — but the highlight geometry (#72) and the reading order would be wrong on the papers users care most about.
Proposal
Add column detection before line grouping:
- Estimate column boundaries from the horizontal projection of text-item coverage (or a gutter of consistent whitespace across many lines).
- Segment lines into columns, then read columns top-to-bottom, left-to-right within the detected order.
- Keep the single-column path untouched; fall back to it when no stable gutter is found.
Do not turn this into a general PDF layout engine. Scope it to "detect N columns and read them in the right order".
Fixtures
Golden fixtures with known reading order, at minimum:
Acceptance
Verification
Synthetic-item unit tests for the column splitter plus golden assertions on the committed fixtures.
Part of the v1.4 — Trusted Research Loop epic: #82
Follow-up to #110 / PR #111. Feeds #72 (citation → exact paragraph highlight).
Context
pdfTextLayout.groupLines()(#110) orders items by baseline top-to-bottom, then by x left-to-right, and puts every item that shares a baseline on the same line. That is correct for single-column documents, but a classic two-column paper becomes:and a paragraph bbox can span both columns.
KnowNote is a research tool, so two-column academic PDFs are a primary input, not an edge case. This does not need to be solved before citations ship — the page number and block text are still correct — but the highlight geometry (#72) and the reading order would be wrong on the papers users care most about.
Proposal
Add column detection before line grouping:
Do not turn this into a general PDF layout engine. Scope it to "detect N columns and read them in the right order".
Fixtures
Golden fixtures with known reading order, at minimum:
Acceptance
Verification
Synthetic-item unit tests for the column splitter plus golden assertions on the committed fixtures.
Part of the v1.4 — Trusted Research Loop epic: #82
Follow-up to #110 / PR #111. Feeds #72 (citation → exact paragraph highlight).