A CLI tool that extracts comments and annotations from PDFs (especially academic papers) and generates:
- A markdown todo list for tracking feedback implementation
- Structured context for AI assistants (like Cursor) to help implement the feedback
# Clone the repository
git clone https://github.com/finnoh/pdf-feedback-extractor.git
cd pdf-feedback-extractor
# Install dependencies and set up the environment
uv sync
# Install as a CLI tool (makes it available system-wide)
uv pip install -e .After installation, use pdf-feedback from anywhere:
pdf-feedback your_paper.pdf# Clone the repository
git clone https://github.com/finnoh/pdf-feedback-extractor.git
cd pdf-feedback-extractor
# Install the package
pip install -e .After installation, use pdf-feedback from anywhere:
pdf-feedback your_paper.pdf# Clone the repository
git clone https://github.com/finnoh/pdf-feedback-extractor.git
cd pdf-feedback-extractor
# Set up the environment
uv sync
# Run the tool (no installation needed)
uv run pdf-feedback your_paper.pdf- Extracts comments/annotations from PDFs with page numbers and authors
- Extracts highlighted text associated with each comment
- Generates markdown todo lists organized by page and author
- Creates structured context files for AI-assisted implementation
- Source file matching: Provide your
.mdor.texsource file to get line numbers for each comment
pdf-feedback paper_review.pdf --output-dir ./feedbackIf you have the source .md or .tex file that was used to generate the PDF, you can provide it to get line numbers in the context file:
pdf-feedback paper_review.pdf --source-file paper.tex --output-dir ./feedbackThis will:
- Extract comments from the PDF
- Search for highlighted text in the source file
- Provide line numbers where the text appears in the source file
- Include this information in the context file for easier implementation
Note on text matching: The tool uses fuzzy matching to find text in the source file. It first tries exact matches (normalized for whitespace), then falls back to matching lines containing at least 50% of the key words from the highlighted text. This helps handle cases where PDF text extraction might differ slightly from the source file.
pdf_path: Path to the PDF file (required)--output-dir: Directory to save output files (default: current directory)--context-only: Only generate context file, skip todo list--source-file: Path to source .md or .tex file used to generate the PDF (optional, for line number matching)
feedback_todos.md: Markdown todo list with all comments organized by pagefeedback_context.md: Structured context file for AI assistants with:- Comment details and highlighted text
- Line numbers in source file (if
--source-fileis provided) - Implementation instructions
pdf-feedback review.pdfCreates:
feedback_todos.md- Your todo listfeedback_context.md- Context for AI assistants
pdf-feedback paper.pdf --source-file paper.tex --output-dir ./review_feedbackThis will create the output files in ./review_feedback/ and include line numbers from paper.tex in the context file.
pdf-feedback paper.pdf --context-onlyOnly generates feedback_context.md, skipping the todo list.
- Python 3.8+
- PyMuPDF (fitz) - automatically installed with the package
- uv (optional, for environment management) or pip
.
├── pdf_feedback_extractor.py # Main CLI script
├── pyproject.toml # Project configuration (dependencies, entry points)
├── requirements.txt # Alternative pip requirements
├── README.md # This file
└── .venv/ # Virtual environment (created by uv sync, gitignored)
If you want to contribute or modify the code:
# Clone and set up
git clone https://github.com/finnoh/pdf-feedback-extractor.git
cd pdf-feedback-extractor
uv sync
# Make changes, then test with:
uv run pdf-feedback your_paper.pdf