Skip to content
OmarMWarraichPublic

About

Convert any text-based PDF into beautiful, self-contained HTML — verbatim text, source typography, zero images.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

pdf-html logo

pdf-html

Early-stage geometry-first PDF → HTML converter for text-based PDFs — verbatim text, source typography, zero images.

PyPI Publish Python PyMuPDF License: MIT Tests Typed No LLM

One command in. One elegant, dependency-free HTML file out.


🎯 What it does

pdf-html is an early-stage, geometry-first Python CLI that converts text-based PDFs into a single self-contained HTML file. Instead of NLP or OCR, it infers document structure purely from font metadata and bounding-box geometry — rebuilding heading hierarchies, nested lists, ruled tables (including rows that span page breaks), and multi-column reading order, while keeping every character of text verbatim. The output mirrors the source document's typography with CSS derived from its own fonts, sizes, and colors, and never embeds images. Deterministic, offline, MIT-licensed, and installable with pip install pdf-html.

  • ✅ Best for: text-based PDFs with real text layers, headings, lists, tables, and multi-column layouts

  • ⚠️ Not a universal OCR-first converter: scanned/image-heavy PDFs still need an OCR pre-pass or a dedicated workflow

  • 🏷️ Heading hierarchy — font-size tiers become real <h1>–<h6>

  • 🎨 Typography & color — the page CSS is derived from the document's own fonts, sizes, and palette

  • 📊 Tables — ruled tables are rebuilt as real <table> elements, including lists inside cells and rows that continue across page breaks

  • 📝 Lists — bullets and numbered items become nested <ul>/<ol> (indent decides nesting)

  • 🧭 Reading order — multi-column layouts are re-linearized column by column

  • ✂️ Page furniture — repeated running headers/footers and page numbers are stripped

  • 🔒 Text is verbatim — never summarized, reworded, or reordered within a block

  • 🚫 No images, ever — image content is dropped by design; text alone carries the document

🧭 Early-stage scope

This is an early-stage, geometry-first conversion tool for text-based PDFs. It is designed to be reliable and deterministic where the PDF has a usable text layer, but it is not a universal OCR-first converter for scanned documents, forms, or arbitrary image-heavy PDFs.

What it does well

  • text-based PDFs with real text layers
  • styled HTML with verbatim text
  • heading hierarchy, lists, and multi-column layout reconstruction
  • ruled table reconstruction with cell-aware content flow

What it does not yet do well

  • OCR and scanned-document support are planned as separate, opt-in paths
  • arbitrary image-heavy PDFs or forms with weak text layers
  • pixel-perfect visual reconstruction of every layout edge case

⚡ Quick start

# install from PyPI (Python 3.10+)
pip install pdf-html
# or: uv pip install pdf-html

# convert
pdf-html report.pdf -o report.html

# open report.html in any browser — no external assets needed

If you are working from a cloned repository instead of the published package, install the local checkout with pip install -e . or uv pip install -e ..

🖥️ CLI reference

pdf-html INPUT.pdf -o out.html [--extractor pymupdf|pdftotext]
    [--style auto|default] [--paginate] [--no-tables] [--no-callouts]
    [--keep-headers] [--allow-scanned]
Flag Default What it does
-o, --output (required) Path of the HTML file to write
--extractor pymupdf Text extraction backend (pdftotext fallback planned)
--style auto auto derives CSS from the document's own fonts/sizes/colors; default uses a clean built-in theme
--paginate off Wrap each PDF page in a <section class="sheet">
--no-tables off Disable table reconstruction (table text flows as paragraphs)
--keep-headers off Keep repeated running headers/footers
--allow-scanned off Convert scanned/image PDFs instead of exiting with an OCR hint

🏗️ Architecture

A deterministic pipeline — every stage is a small, pure, individually testable module:

🔍 Pipeline diagram — click to enlarge (click again to close) · open full screen ↗
---
config:
  layout: elk
  theme: neutral
---
flowchart LR
    subgraph S1["📥 1 · Extract"]
        direction TB
        PDF@{ shape: doc, label: "📄 PDF" }
        EX@{ shape: rect, label: "Extractor<br/>spans + font metadata" }
        PDF --> EX
    end
    subgraph S2["🧹 2 · Clean &amp; Profile"]
        direction TB
        HF@{ shape: rect, label: "Header/Footer<br/>stripping" }
        SP@{ shape: hex, label: "Style Profiler<br/>body size · h1–h6 · palette" }
        HF --> SP
    end
    subgraph S3["🧭 3 · Layout"]
        direction TB
        RO@{ shape: rect, label: "Reading Order<br/>column clustering" }
        TR@{ shape: fr-rect, label: "Table Reconstructor<br/>per-cell pipeline" }
    end
    subgraph S4["🏗️ 4 · Structure"]
        direction TB
        SD@{ shape: div-rect, label: "Structure Detector<br/>headings · paragraphs · lists" }
        LP@{ shape: rect, label: "List Parser<br/>nested ul/ol" }
        AST@{ shape: bow-rect, label: "AST<br/>Document → Page → Blocks" }
        SD --> LP --> AST
    end
    subgraph S5["🎨 5 · Render"]
        direction TB
        RN@{ shape: rect, label: "Renderer<br/>semantic HTML5 + inline CSS" }
        OUT@{ shape: tag-doc, label: "🌐 out.html" }
        RN --> OUT
    end
    EX --> HF
    SP --> RO
    SP --> TR
    RO --> SD
    TR --> SD
    AST --> RN
Loading
Module Responsibility
extractor.py TextExtractor ABC; PyMuPDF backend reads per-span size/weight/color/bbox and detects table regions
header_footer.py Strips spans repeating on ≥ 60% of pages in the top/bottom 10% bands
style_profiler.py Character-weighted font-size histogram → body size, heading tiers, color palette
reading_order.py Column detection via x-gap clustering, with a card-grid fallback and straddle guard
table_reconstructor.py Assigns spans to detected cells, runs the full pipeline inside each cell, merges cross-page rows
structure.py Classifies lines into headings/paragraphs/list items from geometry + font cues
list_parser.py Indent-based nesting; glyph style only picks ul vs ol; markers stripped, text verbatim
ast.py Typed document model — Document → Page → Block, runs carry inline style
renderer.py Single-file HTML5 with one <style> block and CSS variables from the profile

📊 Table reconstruction highlights

The hardest part of PDF → HTML is tables. pdf-html:

  1. 🔍 Detects ruled tables geometrically (PyMuPDF find_tables()) at extraction time
  2. 📌 Assigns the page's styled spans to cells by bounding box — inline bold/color/size survive
  3. 🔄 Runs the normal line → paragraph → list pipeline inside every cell, so bullets in cells become real nested lists
  4. 🧵 Merges rows that continue across page breaks (empty-first-cell fragments) back into one row — even resuming mid-list-item
  5. 🏷️ Promotes a bold-only first row to a <th> header row

🧭 Design principles

Principle Meaning
🧮 Pure geometry, no AI All structure is inferred from font metadata and bounding boxes. No NLP, no LLM, no document-specific regexes
🔒 Text is sacred Output text is verbatim; only & < > are escaped
🚫 No images Spans overlapping image rects are dropped; <img> is never emitted
🪂 Graceful degradation Heuristic failures only affect styling — never text content or order
🔧 Tunable & testable Every heuristic threshold is a named module-level constant with focused unit tests

⚖️ How it compares

Every PDF converter picks a trade-off. pdf-html optimizes for semantic, reflowable, styled HTML with a verbatim-text guarantee — a square none of the established tools occupy:

Tool Output Semantic structure Keeps typography Deterministic Footprint
pdf-html Self-contained HTML5 ✅ h1–h6, ul/ol, table ✅ CSS derived from the source ✅ ~30 MB (PyMuPDF only)
pdf2htmlEX Pixel-faithful HTML ❌ positioned glyphs ✅ visually ✅ C++ toolchain
Poppler pdftohtml Positioned divs / bare text ❌ ⚠️ partial ✅ system package
pymupdf4llm Markdown for LLM ingestion ⚠️ headings & lists ❌ discarded ✅ ~30 MB
marker-pdf Markdown/JSON via ML ✅ ❌ discarded ❌ model-dependent GB-scale models, GPU-friendly
docling Markdown/HTML/JSON via ML ✅ ❌ discarded ❌ model-dependent GB-scale models
unstructured Element JSON for RAG ⚠️ element types ❌ ⚠️ heavy optional deps
Adobe PDF Services Structured JSON/HTML ✅ ⚠️ ❌ cloud API, paid

When to choose pdf-html — you want a readable, reflowable document that still looks like the original, produced offline, reproducibly, with text you can trust character-for-character (text-based PDFs, tables included, even across page breaks).

When to choose something else — you need pixel-perfect visual replicas (pdf2htmlEX), OCR-heavy scanned document conversion (marker, docling), or RAG-oriented element JSON (unstructured).

🧪 Testing

56 tests cover every pipeline stage plus end-to-end CLI runs over deterministic fixture PDFs (report, two-column paper, slide deck, brochure, ruled table):

uv sync                                        # dev deps (pytest, reportlab)
uv run pytest                                  # run the suite
uv run python tests/fixtures/make_fixtures.py  # regenerate fixture PDFs

Verbatim-ness is asserted mechanically: every source string drawn into a fixture must appear in the rendered HTML.

📁 Project structure

pdf-html/
├── src/pdf_html/
│   ├── cli.py                  # argparse CLI → pipeline → HTML
│   ├── extractor.py            # PyMuPDF span + table-region extraction
│   ├── header_footer.py        # repeated page-furniture stripping
│   ├── style_profiler.py       # font-size histogram → style profile
│   ├── reading_order.py        # column clustering & span ordering
│   ├── table_reconstructor.py  # cell assignment, per-cell pipeline, row merging
│   ├── structure.py            # heading / paragraph / list classification
│   ├── list_parser.py          # nested list folding
│   ├── ast.py                  # typed document model
│   └── renderer.py             # semantic HTML5 + derived CSS
├── .github/workflows/
│   └── publish.yml             # CI: test → build → publish to PyPI on release
├── tests/                      # 56 tests + deterministic PDF fixtures
├── README.md
└── TUTORIAL.md                 # step-by-step usage guide

🛣️ Roadmap

  • Ruled-table reconstruction with cross-page row merging
  • Column-aware reading order with card-grid detection
  • Repeated header/footer stripping
  • Borderless-table detection (whitespace-gap heuristic)
  • Callout/aside detection (--no-callouts flag already reserved)
  • Dependency-free pdftotext fallback extractor
  • colspan/rowspan from merged-cell geometry

🤝 Contributing

Contributions welcome! Ground rules:

  • 🐍 Python 3.10+, type hints throughout
  • 📦 PyMuPDF is the only hard runtime dependency
  • 🧩 Keep heuristics small, pure, and individually testable; thresholds as named constants
  • ✅ One feature per commit (feat|fix|docs|refactor|chore: ...); update README/TUTORIAL with any user-facing change
  • 🧪 uv run pytest must stay green — fixtures are the contract

🚀 Releases

This project is intentionally published as an early-stage v0.x tool: the core pipeline is solid for text-based PDFs, but it is not a universal PDF converter for scanned pages, forms, or OCR-heavy corpora.

Releases are automated with GitHub Actions and PyPI trusted publishing — no API tokens involved. Publishing a GitHub release triggers publish.yml, which runs the test suite, builds the sdist + wheel, and uploads to PyPI:

# 1. bump version in pyproject.toml and src/pdf_html/__init__.py
# 2. commit, tag, and push
git tag -a vX.Y.Z -m "Release vX.Y.Z"
git push origin main --tags

# 3. publish the GitHub release — this triggers CI → tests → build → PyPI
gh release create vX.Y.Z --generate-notes

Manual fallback: uv build && twine upload dist/*.

📄 License

MIT — see pyproject.toml.


Built with 🐍 + 📐 — early-stage geometry over guesswork.

If this project helped you, consider giving it a ⭐!

About

Convert any text-based PDF into beautiful, self-contained HTML — verbatim text, source typography, zero images.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages