Skip to content

About

CLI tool to extract comments from PDFs and generate todo lists and AI context files for implementing feedback

Resources

Stars

4 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

PDF Feedback Todo Extractor

A CLI tool that extracts comments and annotations from PDFs (especially academic papers) and generates:

  1. A markdown todo list for tracking feedback implementation
  2. Structured context for AI assistants (like Cursor) to help implement the feedback

Quick Install

Install from GitHub Repository

Option 1: Using uv (Recommended)

# Clone the repository
git clone https://github.com/finnoh/pdf-feedback-extractor.git
cd pdf-feedback-extractor

# Install dependencies and set up the environment
uv sync

# Install as a CLI tool (makes it available system-wide)
uv pip install -e .

After installation, use pdf-feedback from anywhere:

pdf-feedback your_paper.pdf

Option 2: Using pip

# Clone the repository
git clone https://github.com/finnoh/pdf-feedback-extractor.git
cd pdf-feedback-extractor

# Install the package
pip install -e .

After installation, use pdf-feedback from anywhere:

pdf-feedback your_paper.pdf

Option 3: Use without installation (uv run)

# Clone the repository
git clone https://github.com/finnoh/pdf-feedback-extractor.git
cd pdf-feedback-extractor

# Set up the environment
uv sync

# Run the tool (no installation needed)
uv run pdf-feedback your_paper.pdf

Features

  • Extracts comments/annotations from PDFs with page numbers and authors
  • Extracts highlighted text associated with each comment
  • Generates markdown todo lists organized by page and author
  • Creates structured context files for AI-assisted implementation
  • Source file matching: Provide your .md or .tex source file to get line numbers for each comment

Usage

Basic usage

pdf-feedback paper_review.pdf --output-dir ./feedback

With source file (for line number matching)

If you have the source .md or .tex file that was used to generate the PDF, you can provide it to get line numbers in the context file:

pdf-feedback paper_review.pdf --source-file paper.tex --output-dir ./feedback

This will:

  • Extract comments from the PDF
  • Search for highlighted text in the source file
  • Provide line numbers where the text appears in the source file
  • Include this information in the context file for easier implementation

Note on text matching: The tool uses fuzzy matching to find text in the source file. It first tries exact matches (normalized for whitespace), then falls back to matching lines containing at least 50% of the key words from the highlighted text. This helps handle cases where PDF text extraction might differ slightly from the source file.

Command-line Options

  • pdf_path: Path to the PDF file (required)
  • --output-dir: Directory to save output files (default: current directory)
  • --context-only: Only generate context file, skip todo list
  • --source-file: Path to source .md or .tex file used to generate the PDF (optional, for line number matching)

Output Files

  • feedback_todos.md: Markdown todo list with all comments organized by page
  • feedback_context.md: Structured context file for AI assistants with:
    • Comment details and highlighted text
    • Line numbers in source file (if --source-file is provided)
    • Implementation instructions

Examples

Example 1: Extract comments from a PDF

pdf-feedback review.pdf

Creates:

  • feedback_todos.md - Your todo list
  • feedback_context.md - Context for AI assistants

Example 2: With source file and custom output directory

pdf-feedback paper.pdf --source-file paper.tex --output-dir ./review_feedback

This will create the output files in ./review_feedback/ and include line numbers from paper.tex in the context file.

Example 3: Only generate context file

pdf-feedback paper.pdf --context-only

Only generates feedback_context.md, skipping the todo list.

Requirements

  • Python 3.8+
  • PyMuPDF (fitz) - automatically installed with the package
  • uv (optional, for environment management) or pip

Project Structure

.
├── pdf_feedback_extractor.py  # Main CLI script
├── pyproject.toml             # Project configuration (dependencies, entry points)
├── requirements.txt           # Alternative pip requirements
├── README.md                  # This file
└── .venv/                     # Virtual environment (created by uv sync, gitignored)

Development

If you want to contribute or modify the code:

# Clone and set up
git clone https://github.com/finnoh/pdf-feedback-extractor.git
cd pdf-feedback-extractor
uv sync

# Make changes, then test with:
uv run pdf-feedback your_paper.pdf

About

CLI tool to extract comments from PDFs and generate todo lists and AI context files for implementing feedback

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages