Skip to content

[Artifact Forest 5/9] Recover screenplay structure from PDF and unstructured text #426

Description

@szmyty

Parent: #421
Depends on: #423
Coordinates with: #424, #425

Outcome

Support an explicitly lossy, review-required path from screenplay-shaped PDF or text into a recoverable screenplay representation and then into Fountain/other script derivatives.

Pipeline

PDF
 ├─ text-bearing → deterministic text/layout extraction
 └─ image-only   → bounded OCR fallback
                     ↓
              screenplay recovery
                     ↓
          renderflow.screenplay/v1
             ├─ Fountain
             ├─ FDX
             ├─ FadeIn
             └─ OSF

Provider candidates

  • Poppler pdftotext / equivalent local deterministic extractor.
  • Existing Tesseract provider for image-only pages.
  • Optional reviewed AI transform only when deterministic recovery is insufficient and policy explicitly allows it.

Requirements

Acceptance criteria

  • Text PDF → reviewed screenplay candidate works on a synthetic fixture.
  • Image-only PDF takes the OCR path or reports provider unavailability.
  • Recovery evidence names extraction method and confidence/ambiguity.
  • Fountain/FDX/etc. derivatives are marked as recovered/lossy lineage.
  • Source PDF is unchanged and not regenerated by an exhaustive forest unless explicitly requested.
  • Malformed/encrypted/unsupported PDFs produce typed outcomes.
  • Optional AI cannot upgrade confidence or approval without explicit evidence/review.

Non-goals

Claiming exact Final Draft reconstruction or visual-layout preservation from arbitrary PDFs.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions