Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

DocU — PDF → DOCX High-Fidelity Converter / 高保真 PDF 转 Word 工具

English | 中文 | AI Tools

Powered by MinerU for document parsing. Renders structured parse results into native, editable .docx with proper tables, formulas, images, and text formatting.

基于 MinerU 文档解析引擎,将结构化解析结果渲染为原生可编辑的 .docx 文件,支持表格、公式、图片和文本格式。


English

Features

  • Native DOCX tables — HTML tables with colspan/rowspan → real Word tables
  • Formula rendering — LaTeX equations → high-quality embedded images
  • Mixed inline content — text, inline math, symbols in proper OOXML runs
  • Font fallback — CJK / Latin / Math / Monospace per-character-script mapping
  • Three modes — full auto (PDF→DOCX), from MinerU cache, or pure data-driven

Quick Start

pip install docu[mineru]
from docu import DocU

# One-liner
DocU(backend="pipeline").convert("paper.pdf", "paper.docx")

Usage

Full auto (PDF → DOCX)

from docu import DocU
DocU(backend="pipeline", lang="en").convert("input.pdf", "output.docx")

From existing MinerU output (no mineru at runtime)

DocU().convert_from_parse_dir("./mineru_output/paper/auto/", "output.docx")

From structured data

DocU().convert_from_middle_json(pages_list, "./images/", "output.docx")

Batch conversion

from docu import batch_convert
batch_convert(["p1.pdf", "p2.pdf"], "./output/", backend="pipeline")

Custom fonts

from docu import DocUFontConfig
DocU(font_config=DocUFontConfig(
    font_cjk="微软雅黑", font_latin="Times New Roman",
    font_math="Cambria Math", font_mono="Consolas", base_size_pt=11,
)).convert("in.pdf", "out.docx")

CLI

docu -p input.pdf -o output.docx -b pipeline -l en
docu --font-cjk "微软雅黑" --font-latin "Times New Roman" -p in.pdf -o out.docx

Architecture

PDF → [MinerU] → content_list_v2 → [DocU] → .docx
                       │
                       ├── _adapt.py    — middle_json compat
                       ├── _block.py    — block-level renderer
                       ├── _table.py    — HTML → native DOCX table
                       ├── _inline.py   — inline mixed-content renderer
                       ├── _fonts.py    — font fallback engine
                       └── _symbols.py  — Unicode script detection

Requirements

  • Python ≥ 3.10
  • python-docx, beautifulsoup4, lxml, Pillow, click
  • mineru ≥ 3.0 (optional, for PDF parsing)

中文

特性

  • 原生 DOCX 表格 — HTML 表格(含 colspan/rowspan)转为真实 Word 表格
  • 公式渲染 — LaTeX 公式以高质量图片嵌入
  • 混合内联内容 — 文本、行内公式、特殊符号在同一 OOXML run 中正确排版
  • 字体回退 — 中日韩 / 拉丁 / 数学 / 等宽字体按字符脚本自动映射
  • 三种模式 — 全自动(PDF→DOCX)、基于 MinerU 缓存、纯数据驱动

快速开始

pip install docu[mineru]
from docu import DocU

# 一行搞定
DocU(backend="pipeline").convert("论文.pdf", "输出.docx")

使用方式

全自动转换(PDF → DOCX)

from docu import DocU
DocU(backend="pipeline", lang="zh").convert("输入.pdf", "输出.docx")

基于已有 MinerU 输出(无需运行 MinerU)

DocU().convert_from_parse_dir("./mineru_output/论文/auto/", "输出.docx")

基于结构化数据

DocU().convert_from_middle_json(pages_list, "./images/", "输出.docx")

批量转换

from docu import batch_convert
batch_convert(["论文1.pdf", "论文2.pdf"], "./输出/", backend="pipeline")

自定义字体

from docu import DocUFontConfig
DocU(font_config=DocUFontConfig(
    font_cjk="微软雅黑", font_latin="Times New Roman",
    font_math="Cambria Math", font_mono="Consolas", base_size_pt=11,
)).convert("in.pdf", "out.docx")

命令行

docu -p 输入.pdf -o 输出.docx -b pipeline -l zh
docu --font-cjk "宋体" --font-latin "Times New Roman" -p in.pdf -o out.docx

架构

PDF → [MinerU] → content_list_v2 → [DocU] → .docx
                       │
                       ├── _adapt.py    — middle_json 适配层
                       ├── _block.py    — 块级渲染器
                       ├── _table.py    — HTML → 原生 DOCX 表格
                       ├── _inline.py   — 内联混合内容渲染
                       ├── _fonts.py    — 字体回退引擎
                       └── _symbols.py  — Unicode 脚本检测

依赖

  • Python ≥ 3.10
  • python-docx, beautifulsoup4, lxml, Pillow, click
  • mineru ≥ 3.0(可选,用于 PDF 解析)

🤖 AI Coding Tools

DocU is designed to work well with AI-assisted development workflows.

Claude Code

Add to your project's CLAUDE.md:

# DocU project conventions

- Run with: `python -m docu` or `docu` CLI
- Test commands: `pytest tests/`
- Install deps: `pip install -e ".[dev,mineru]"`
- Key entry points: `src/docu/_core.py::DocU.convert()`, `src/docu/cli.py::main()`
- Font config: `src/docu/_fonts.py::DocUFontConfig`
- Render pipeline: `_block.py` → `_table.py` → `_inline.py`
- MinerU integration: `src/docu/_pdf_parser.py`
- All internal modules prefixed with `_` (private API)

Codex / Copilot

The codebase follows these conventions that help LLMs navigate:

  • Entry point: DocU.convert() is the main public API
  • Internal modules: _block.py, _table.py, _inline.py, _fonts.py, _symbols.py
  • Data flow: PDF → MinerU → content_list_v2 JSON → BlockRenderer → TableRenderer → InlineRenderer → .docx
  • Type hints: Used throughout for better IDE/LLM comprehension
  • Private-by-convention: _ prefix = internal, no _ = public

Recommended Prompts

"Convert this PDF to DOCX using DocU with Chinese font support"
"Add a new font fallback rule for Arabic script in _fonts.py"
"Fix the table colspan rendering in _table.py"
"Add OMML formula support as an alternative to image-based formulas"

License

Apache 2.0

About

PDF to DOCX high-fidelity converter powered by MinerU — native tables, formulas, images, and CJK font support

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages