A robust CLI tool for converting PDF documents into structured JSON data. This tool extracts text, layouts, font styles, and images with high precision using PyMuPDF.
- Component-Level Extraction: Extracts text, lines, rectangles, and images as individual JSON components.
- Font Intelligence: Identifies and extracts embedded font styles (Family, Full Name, PostScript) for accurate rendering.
- Visual Accuracy: Captures image positions, dimensions, and transformations.
- Smart Cleanup: Automatically removes overlaps, hidden texts, and redundant background elements.
- Local Storage: Completely standalone with local asset storage (no cloud dependencies).
- Docker Ready: Fully containerized for easy deployment and scaling.
- Clone the repository:
git clone https://github.com/x-eight/pdf-to-json.git cd pdf-to-json - Install dependencies:
pip install -r requirements.txt
docker build -t pdf-to-json .Run the CLI by providing the path to your PDF file:
python src/cli.py path/to/your/document.pdf-o, --output-dir: Specify where to save the JSON and extracted assets (default:./output).
Example with custom output:
python src/cli.py input.pdf --output-dir ./my_resultsdocker run -v $(pwd)/output:/app/output pdf-to-json input.pdfThe tool generates a structured output in the designated folder:
output/
├── input.json # Structured data of the PDF
├── images/ # Extracted portrait and inline images
└── fonts/ # Extracted embedded fonts
This project utilizes:
- PyMuPDF (fitz): For core PDF parsing and SVG extraction.
- Pillow: For image processing and optimization.
- FontTools: For deep inspection of embedded font metadata.
- Font Mapping: Uses
src/resources/fonts.jsonto accurately map embedded and standard fonts (Arial, Calibri, etc.) even when obfuscated.
This project is licensed under the MIT License.