Early-stage geometry-first PDF → HTML converter for text-based PDFs — verbatim text, source typography, zero images.
One command in. One elegant, dependency-free HTML file out.
pdf-html is an early-stage, geometry-first Python CLI that converts text-based PDFs into a single self-contained HTML file. Instead of NLP or OCR, it infers document structure purely from font metadata and bounding-box geometry — rebuilding heading hierarchies, nested lists, ruled tables (including rows that span page breaks), and multi-column reading order, while keeping every character of text verbatim. The output mirrors the source document's typography with CSS derived from its own fonts, sizes, and colors, and never embeds images. Deterministic, offline, MIT-licensed, and installable with pip install pdf-html.
-
✅ Best for: text-based PDFs with real text layers, headings, lists, tables, and multi-column layouts
-
⚠️ Not a universal OCR-first converter: scanned/image-heavy PDFs still need an OCR pre-pass or a dedicated workflow -
🏷️ Heading hierarchy — font-size tiers become real
<h1>–<h6> -
🎨 Typography & color — the page CSS is derived from the document's own fonts, sizes, and palette
-
📊 Tables — ruled tables are rebuilt as real
<table>elements, including lists inside cells and rows that continue across page breaks -
📝 Lists — bullets and numbered items become nested
<ul>/<ol>(indent decides nesting) -
🧭 Reading order — multi-column layouts are re-linearized column by column
-
✂️ Page furniture — repeated running headers/footers and page numbers are stripped
-
🔒 Text is verbatim — never summarized, reworded, or reordered within a block
-
🚫 No images, ever — image content is dropped by design; text alone carries the document
This is an early-stage, geometry-first conversion tool for text-based PDFs. It is designed to be reliable and deterministic where the PDF has a usable text layer, but it is not a universal OCR-first converter for scanned documents, forms, or arbitrary image-heavy PDFs.
- text-based PDFs with real text layers
- styled HTML with verbatim text
- heading hierarchy, lists, and multi-column layout reconstruction
- ruled table reconstruction with cell-aware content flow
- OCR and scanned-document support are planned as separate, opt-in paths
- arbitrary image-heavy PDFs or forms with weak text layers
- pixel-perfect visual reconstruction of every layout edge case
# install from PyPI (Python 3.10+)
pip install pdf-html
# or: uv pip install pdf-html
# convert
pdf-html report.pdf -o report.html
# open report.html in any browser — no external assets neededIf you are working from a cloned repository instead of the published package, install the local checkout with pip install -e . or uv pip install -e ..
pdf-html INPUT.pdf -o out.html [--extractor pymupdf|pdftotext]
[--style auto|default] [--paginate] [--no-tables] [--no-callouts]
[--keep-headers] [--allow-scanned]| Flag | Default | What it does |
|---|---|---|
-o, --output |
(required) | Path of the HTML file to write |
--extractor |
pymupdf |
Text extraction backend (pdftotext fallback planned) |
--style |
auto |
auto derives CSS from the document's own fonts/sizes/colors; default uses a clean built-in theme |
--paginate |
off | Wrap each PDF page in a <section class="sheet"> |
--no-tables |
off | Disable table reconstruction (table text flows as paragraphs) |
--keep-headers |
off | Keep repeated running headers/footers |
--allow-scanned |
off | Convert scanned/image PDFs instead of exiting with an OCR hint |
A deterministic pipeline — every stage is a small, pure, individually testable module:
🔍 Pipeline diagram — click to enlarge (click again to close) · open full screen ↗
---
config:
layout: elk
theme: neutral
---
flowchart LR
subgraph S1["📥 1 · Extract"]
direction TB
PDF@{ shape: doc, label: "📄 PDF" }
EX@{ shape: rect, label: "Extractor<br/>spans + font metadata" }
PDF --> EX
end
subgraph S2["🧹 2 · Clean & Profile"]
direction TB
HF@{ shape: rect, label: "Header/Footer<br/>stripping" }
SP@{ shape: hex, label: "Style Profiler<br/>body size · h1–h6 · palette" }
HF --> SP
end
subgraph S3["🧭 3 · Layout"]
direction TB
RO@{ shape: rect, label: "Reading Order<br/>column clustering" }
TR@{ shape: fr-rect, label: "Table Reconstructor<br/>per-cell pipeline" }
end
subgraph S4["🏗️ 4 · Structure"]
direction TB
SD@{ shape: div-rect, label: "Structure Detector<br/>headings · paragraphs · lists" }
LP@{ shape: rect, label: "List Parser<br/>nested ul/ol" }
AST@{ shape: bow-rect, label: "AST<br/>Document → Page → Blocks" }
SD --> LP --> AST
end
subgraph S5["🎨 5 · Render"]
direction TB
RN@{ shape: rect, label: "Renderer<br/>semantic HTML5 + inline CSS" }
OUT@{ shape: tag-doc, label: "🌐 out.html" }
RN --> OUT
end
EX --> HF
SP --> RO
SP --> TR
RO --> SD
TR --> SD
AST --> RN
| Module | Responsibility |
|---|---|
extractor.py |
TextExtractor ABC; PyMuPDF backend reads per-span size/weight/color/bbox and detects table regions |
header_footer.py |
Strips spans repeating on ≥ 60% of pages in the top/bottom 10% bands |
style_profiler.py |
Character-weighted font-size histogram → body size, heading tiers, color palette |
reading_order.py |
Column detection via x-gap clustering, with a card-grid fallback and straddle guard |
table_reconstructor.py |
Assigns spans to detected cells, runs the full pipeline inside each cell, merges cross-page rows |
structure.py |
Classifies lines into headings/paragraphs/list items from geometry + font cues |
list_parser.py |
Indent-based nesting; glyph style only picks ul vs ol; markers stripped, text verbatim |
ast.py |
Typed document model — Document → Page → Block, runs carry inline style |
renderer.py |
Single-file HTML5 with one <style> block and CSS variables from the profile |
The hardest part of PDF → HTML is tables. pdf-html:
- 🔍 Detects ruled tables geometrically (PyMuPDF
find_tables()) at extraction time - 📌 Assigns the page's styled spans to cells by bounding box — inline bold/color/size survive
- 🔄 Runs the normal line → paragraph → list pipeline inside every cell, so bullets in cells become real nested lists
- 🧵 Merges rows that continue across page breaks (empty-first-cell fragments) back into one row — even resuming mid-list-item
- 🏷️ Promotes a bold-only first row to a
<th>header row
| Principle | Meaning |
|---|---|
| 🧮 Pure geometry, no AI | All structure is inferred from font metadata and bounding boxes. No NLP, no LLM, no document-specific regexes |
| 🔒 Text is sacred | Output text is verbatim; only & < > are escaped |
| 🚫 No images | Spans overlapping image rects are dropped; <img> is never emitted |
| 🪂 Graceful degradation | Heuristic failures only affect styling — never text content or order |
| 🔧 Tunable & testable | Every heuristic threshold is a named module-level constant with focused unit tests |
Every PDF converter picks a trade-off. pdf-html optimizes for semantic, reflowable, styled HTML with a verbatim-text guarantee — a square none of the established tools occupy:
| Tool | Output | Semantic structure | Keeps typography | Deterministic | Footprint |
|---|---|---|---|---|---|
| pdf-html | Self-contained HTML5 | ✅ h1–h6, ul/ol, table |
✅ CSS derived from the source | ✅ | ~30 MB (PyMuPDF only) |
| pdf2htmlEX | Pixel-faithful HTML | ❌ positioned glyphs | ✅ visually | ✅ | C++ toolchain |
| Poppler pdftohtml | Positioned divs / bare text | ❌ | ✅ | system package | |
| pymupdf4llm | Markdown for LLM ingestion | ❌ discarded | ✅ | ~30 MB | |
| marker-pdf | Markdown/JSON via ML | ✅ | ❌ discarded | ❌ model-dependent | GB-scale models, GPU-friendly |
| docling | Markdown/HTML/JSON via ML | ✅ | ❌ discarded | ❌ model-dependent | GB-scale models |
| unstructured | Element JSON for RAG | ❌ | heavy optional deps | ||
| Adobe PDF Services | Structured JSON/HTML | ✅ | ❌ | cloud API, paid |
When to choose pdf-html — you want a readable, reflowable document that still looks like the original, produced offline, reproducibly, with text you can trust character-for-character (text-based PDFs, tables included, even across page breaks).
When to choose something else — you need pixel-perfect visual replicas (pdf2htmlEX), OCR-heavy scanned document conversion (marker, docling), or RAG-oriented element JSON (unstructured).
56 tests cover every pipeline stage plus end-to-end CLI runs over deterministic fixture PDFs (report, two-column paper, slide deck, brochure, ruled table):
uv sync # dev deps (pytest, reportlab)
uv run pytest # run the suite
uv run python tests/fixtures/make_fixtures.py # regenerate fixture PDFsVerbatim-ness is asserted mechanically: every source string drawn into a fixture must appear in the rendered HTML.
pdf-html/
├── src/pdf_html/
│ ├── cli.py # argparse CLI → pipeline → HTML
│ ├── extractor.py # PyMuPDF span + table-region extraction
│ ├── header_footer.py # repeated page-furniture stripping
│ ├── style_profiler.py # font-size histogram → style profile
│ ├── reading_order.py # column clustering & span ordering
│ ├── table_reconstructor.py # cell assignment, per-cell pipeline, row merging
│ ├── structure.py # heading / paragraph / list classification
│ ├── list_parser.py # nested list folding
│ ├── ast.py # typed document model
│ └── renderer.py # semantic HTML5 + derived CSS
├── .github/workflows/
│ └── publish.yml # CI: test → build → publish to PyPI on release
├── tests/ # 56 tests + deterministic PDF fixtures
├── README.md
└── TUTORIAL.md # step-by-step usage guide
- Ruled-table reconstruction with cross-page row merging
- Column-aware reading order with card-grid detection
- Repeated header/footer stripping
- Borderless-table detection (whitespace-gap heuristic)
- Callout/aside detection (
--no-calloutsflag already reserved) - Dependency-free
pdftotextfallback extractor -
colspan/rowspanfrom merged-cell geometry
Contributions welcome! Ground rules:
- 🐍 Python 3.10+, type hints throughout
- 📦 PyMuPDF is the only hard runtime dependency
- 🧩 Keep heuristics small, pure, and individually testable; thresholds as named constants
- ✅ One feature per commit (
feat|fix|docs|refactor|chore: ...); update README/TUTORIAL with any user-facing change - 🧪
uv run pytestmust stay green — fixtures are the contract
This project is intentionally published as an early-stage v0.x tool: the core pipeline is solid for text-based PDFs, but it is not a universal PDF converter for scanned pages, forms, or OCR-heavy corpora.
Releases are automated with GitHub Actions and PyPI trusted publishing — no API tokens involved. Publishing a GitHub release triggers publish.yml, which runs the test suite, builds the sdist + wheel, and uploads to PyPI:
# 1. bump version in pyproject.toml and src/pdf_html/__init__.py
# 2. commit, tag, and push
git tag -a vX.Y.Z -m "Release vX.Y.Z"
git push origin main --tags
# 3. publish the GitHub release — this triggers CI → tests → build → PyPI
gh release create vX.Y.Z --generate-notesManual fallback: uv build && twine upload dist/*.
MIT — see pyproject.toml.
Built with 🐍 + 📐 — early-stage geometry over guesswork.
If this project helped you, consider giving it a ⭐!