把中文 PDF 课件提炼为结构化 Markdown 学习笔记的 Agent Skill,含零截断图例裁剪脚本(pdfplumber + pypdfium2 + Pillow)
-
Updated
Sep 20, 2026 - Python
把中文 PDF 课件提炼为结构化 Markdown 学习笔记的 Agent Skill,含零截断图例裁剪脚本(pdfplumber + pypdfium2 + Pillow)
High-fidelity PDF to DOCX with editable layout, Google Docs support, and render-back verification.
PDF tables, word boxes, form fields & page render for DuckDB (Python, pdfplumber)
Preflight checks for document extraction pipelines — validate, render, and screen PDFs before they reach your LLM. Pure-Python wheel, in-memory only.
Mi herramienta de escritorio para comprimir archivos PDF
Preprocessing pipeline for Colombian legal documents that cleans and structures PDF/HTML content for RAG systems, preserving legal formatting while removing UI noise and metadata.
Turn invoices, receipts and scanned documents into a balanced, audit-ready ledger — automatically.
PDF layout-surgery engine on a fully permissive stack - read, erase, and redraw text and vector art by coordinate.
A Python tool for extracting highlighted text from PDF files while preserving formatting attributes (headers, bold, italic) and removing unwanted line breaks and page breaks. Perfect for integrating with content management systems.
20 PDF classifiers, one verdict matrix: should this PDF go through fast text extraction, or do we need OCR?
Le couteau suisse PDF 100% offline - 0 tracker - CircaFrax - Astra
To associate your repository with the pypdfium2 topic, visit your repo's landing page and select "manage topics."