Skip to content

About

把中文 PDF 课件提炼为结构化 Markdown 学习笔记的 Agent Skill,含零截断图例裁剪脚本(pdfplumber + pypdfium2 + Pillow)

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

10 Commits

Folders and files

Repository files navigation

agent-skill-pdf-lecture-notes

把中文 PDF 课件(大学课程讲义/幻灯片)提炼为结构化中文 Markdown 学习笔记的 Agent Skill(skill 名:pdf-lecture-notes-zh),在对应知识点下嵌入完整裁剪的图例。也可按 SKILL.md 的流程纯手工执行。

特性

  • 确定性图例裁剪算法:图形对象联合 bbox → 文字吸收到不动点 → 截断词校验到零截断;对整块位图型图例、固定角标压住图例两类情况有专门处理路径(精确框 + 局部遮盖)。
  • 可程序化校验:图片路径与可加载性检查、孤儿文件检查、逐行/逐列墨迹扫描。
  • 固定笔记结构:「重点回顾与易混淆点」+「课后自检清单」,可直接用于复习自测。
  • 依赖极简:pdfplumber + pypdfium2 + Pillow 三个,无需 OCR、numpy 或云端服务。

安装

git clone https://github.com/wonder37-debug/agent-skill-pdf-lecture-notes.git
cd agent-skill-pdf-lecture-notes

python3 -m venv .venv
.venv/bin/pip install -r requirements.txt          # Linux / macOS
# .venv\Scripts\pip install -r requirements.txt    # Windows

作为 Agent Skill 使用(Claude Code / WorkBuddy 等):把整个目录放进对应的 skills 目录即可,例如 ~/.claude/skills/pdf-lecture-notes-zh/。

快速开始

# 1) 提取全文 + 图形统计,渲染所有页面,诊断固定角标(页眉/页脚校徽)
python scripts/extract_and_render.py 课件.pdf -o _tmp --scale 3 --diagnose-deco

# 2) 翻看 _tmp/pages/*.png,按 SKILL.md 第 2 步的标准挑出值得截的图例页

# 3) 编写裁剪配置并批量裁剪(auto = 矢量图例,bbox = 整块位图,masks = 局部遮盖)
python scripts/crop_figures.py 课件.pdf figures.json -o figures --scale 3

# 4) 交付前校验
python scripts/check_figures.py 笔记.md
python scripts/check_figures.py 笔记.md --ink figures/fig_a.png --y0 69 --scale 3

工作流

① 抽文本 + 图形统计 + 整页渲染 ──► ② 挑图例页(只截真图)
                                          │
                                          ▼
      ⑤ 写 Markdown 笔记 ◄── ③④ 裁剪图例 + 目视复核(零截断校验)
                                          │
                                          ▼
                  ⑥ 校验图片引用与完整性 ──► ⑦ 清理中间文件

每一步的判据与操作细节见 SKILL.md。

脚本

脚本 作用
scripts/extract_and_render.py 提取每页文本与 images/curves/rects/lines 计数,整页渲染;--diagnose-deco 诊断页眉/页脚固定角标,输出可直接填入配置的 bbox 列表
scripts/crop_figures.py 配置驱动裁剪:auto 用确定性 bbox 算法(含通栏装饰线剔除、角标护栏、重叠告警),bbox 用精确框,masks 做局部遮盖;单项出错跳过、不中断整批
scripts/check_figures.py 图片路径与可加载性校验、孤儿文件检查、逐行(--ink)/逐列(--col)墨迹扫描

裁剪配置

figures.json 为数组,每项一张图,完整示例见 examples/figures.example.json:

[
  {"page": 29, "name": "fig_example_auto", "mode": "auto", "margin": 40,
   "deco": [[17.8, 20.2, 103.4, 60.7], [599.3, 368.4, 697.0, 399.6]]},

  {"page": 11, "name": "fig_example_bbox", "mode": "bbox",
   "box": [388, 148, 652, 368]},

  {"page": 29, "name": "fig_example_mask", "mode": "bbox",
   "box": [58, 206, 650, 386],
   "masks": [
     {"box": [608.5, 366, 650, 386], "fill": "bg"},
     {"box": [601, 382, 608.5, 386], "fill": [0, 112, 192]}
   ]}
]
  • mode: "auto":图例由矢量线条 + 文字标签绘制时用,脚本执行确定性 bbox 算法。deco 用 --diagnose-deco 的输出填全,漏填会使裁剪边界扩到页边。
  • mode: "bbox":图例是一整块嵌入位图时用(PPT 导出型课件常见),直接给精确框。
  • masks:裁剪后局部遮盖,用于清除压在图例角上的校徽。fill: "bg" 自动取图中出现次数最多的颜色作为背景色(这类课件背景常是 (242, 242, 242) 而非纯白)。遮盖框只能落在图例本体之外,否则会啃掉图例边框。

适用范围

  • 适用:有文本层的课件 PDF(PPT / LaTeX 导出),图例由位图或矢量图形构成。已在部分大学课程课件上验证,验证页数跨度 30~150 页。
  • 不适用:纯扫描件(无文本层,需先 OCR);图例本身是照片/自然图像(没有"完整边框"的概念)。

License

MIT

About

把中文 PDF 课件提炼为结构化 Markdown 学习笔记的 Agent Skill,含零截断图例裁剪脚本(pdfplumber + pypdfium2 + Pillow)

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages