Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
43 changes: 43 additions & 0 deletions .github/workflows/python-publish.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
# This workflow will upload a Python package using uv when a release is published
# Adapted from: https://docs.astral.sh/uv/guides/integration/github/#publishing-to-pypi

name: Upload Python Package

on:
# normal behavior: run when a new release is published
release:
types: [published]
# allow running manually on main (restriction within job)
workflow_dispatch:

concurrency:
# Cancel existing job(s) for workflow when a new one is queued
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true

jobs:
pypi-publish:
name: Upload release to PyPI
runs-on: ubuntu-latest
environment:
name: pypi
url: https://pypi.org/project/spanerr/
permissions:
id-token: write # IMPORTANT: this permission is mandatory for trusted publishing
contents: read
if: github.event_name == 'release' || github.ref_name == 'main'
steps:
- name: Checkout repository
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6.0.3
with:
persist-credentials: false
- name: Install uv
uses: astral-sh/setup-uv@08807647e7069bb48b6ef5acd8ec9567f424441b # v.8.2.0
with:
enable-cache: false
- name: Install Python 3.12
run: uv python install 3.12
- name: Build package
run: uv build
- name: Publish package
run: uv publish
100 changes: 100 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
@@ -1 +1,101 @@
# spanerr

`spanerr` is a Python library for evaluating span-level text annotations using customizable alignment and scoring strategies.
`spanerr` operates explicitly over span text boundaries (i.e., text indices) rather than over the annotated text itself.

Many span-level annotation tasks diverge significantly enough from named-entity recognition (e.g., long text spans, large label set) that the typical formulations for the evaluation metrics of precision, recall, and $F_1$ scores become insuitable.
`spanerr` addresses this issue by not only supporting customized scoring of (partial) span matches, but also customizing how the spans within document-level annotation sets are aligned for evaluation.
`spanerr` is designed for maximal flexibility generally leaving it to the user to determine what assumptions and restrictions are required in their use case.
Comment thread
laurejt marked this conversation as resolved.

[![unit tests](https://github.com/Princeton-CDH/spanerr/actions/workflows/unit-tests.yml/badge.svg)](https://github.com/Princeton-CDH/spanerr/actions/workflows/unit-tests.yml)
[![codecov](https://codecov.io/gh/Princeton-CDH/spanerr/graph/badge.svg?token=Wd3vZ38Bxz)](https://codecov.io/gh/Princeton-CDH/spanerr)

## Basic Usage

### Installation

Use pip to install as a Python package directly from GitHub.
Use a branch or tag name, e.g. `@develop` or `@0.1.0` if you need to install a specific version

```sh
pip install git+https://github.com/Princeton-CDH/spanerr.git#egg=spanerr
```
Comment thread
laurejt marked this conversation as resolved.

### Core Data Types

`spanerr` has three first-class objects:

- `Span`: An individual span annotation.
- `DocSpans`: A set of span annotations for a document.
- `SpanAlignment`: A set of aligned span annotations (reference, system) within a single document.

All three of these data types are immutable, but `SpanAlignment` does not currently support hashing.

#### Binarization

`spanerr` provides functionality for "removing" span labels from `Span` and `DocSpans` by setting them to a default label (empty string).
For `DocSpans`, overlapping spans will be merged and optionally neighboring spans can be merged.

#### Loading from dictionaries

`Spans` can be loaded from dictionaries with the following fields:

- `start` (int): starting text index (inclusive)
- `end` (int): ending text index (exclusive)
- `label` (str): optional span label (defaults to empty string)
Comment thread
laurejt marked this conversation as resolved.

`DocSpans` can be loaded from dictionaries with the following fields:

- `doc_id` (str): optional document id (defaults to empty string)
- `spans` (list[dict]): list of spans (in dictionary form, see above)

### Core Functionality

In `spanerr` there are two core components to evaluating span annotations: (1) how spans are aligned and (2) how aligned spans are scored.
`spanerr` is intentionally designed so that these two components can be heavily customized.

#### Aligning span annotations

An alignment strategy is represented as a function that takes two sets of annotations (`DocSpans`) as input and returns an alignment (`SpanAlignment`) which will then be used for scoring.
There are little restrictions on the alignments themselves: the resulting `SpanAlignment` may contain transformed versions of the input `DocSpans` and no restrictions are made on its mapping between reference and system spans.
The idea is to allow for the creation of whatever alignment is useful for scoring.

The following alignment strategies are provided in `spanerr.align`:

- Select First : Select the first (sequential) matching system span for each reference span.
By default, spans match if they overlap and have the same label.
- Select Best : Select the best matching system span for each reference span.
By default, given spans that overlap and have the same label, the best match is the span pair with the highest jaccard similarity.
- Corppa : The alignment strategy used by [`corppa`](https://github.com/Princeton-CDH/corppa).
See `corppa`'s [evaluation documentation](https://github.com/Princeton-CDH/corppa/tree/main/src/corppa/poetry_detection/evaluation) for more detail.

For additional flexibility, alignment strategies may take additional inputs to further customize their behavior (e.g., use different span matching and span scoring strategies), but these will generally need to be set to a specific value (e.g., via a lambda function) before they can be used within `spanerr`'s evaluation workflow.
See the `spanerr.align.construct_aligner` for an example.

### Scoring Aligmnents

Scoring an alignment (`SpanAlignment`) entails computing an alignment's *relevance score* (i.e., true positive, numerator for precision and recall).
So, a scoring strategy is represented as a function that takes an alignment (`SpanAlignment`) as input and returns its score (`float`).
This allows for both alignment-independent scoring strategies in which scoring depends solely on the individual scores of each reference-system span pair as well as those that don't.

Currently, `spanerr` provides the building blocks for constructing alignment-independent strategies using `spanerr.eval.relevance_score` and `spanerr.span_utils.composite_match_score`.
See `spanerr.compute_metrics.get_scorer` for an example of constructing scoring strategy functions.

### Span Annotations File Format

A set of span annotations can be loaded into `spanerr` by providing a JSONL file with each line corresponding to a different document's annotations (i.e. `DocSpans`).
For examples see the files in `tests/test_data`.

### Scripts

Installing `spanerr` currently provides access to the following command line script:

- `spanerr-metrics`: For calculating entity- or document-level aggregated precision, recall, and F-1 scores for given reference and system span annotation sets.
(Corresponds to `spanerr.compute_metrics.py`)

## License

This project is licensed under the [Apache 2.0 License](LICENSE)

(c)2026 Trustees of Princeton University.
Permission granted for non-commercial distribution online under a standard Open Source license.
3 changes: 3 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,9 @@ test = [
"pytest-cov>=7.1.0",
]

[project.scripts]
spanerr-metrics="spanerr.compute_metrics:main"

[tool.uv]
exclude-newer= "1 week"

Expand Down
65 changes: 25 additions & 40 deletions src/spanerr/align.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,13 +25,14 @@
def select_first_match(
ref: DocSpans,
sys: DocSpans,
is_match: CheckSpanPair,
is_match: CheckSpanPair = partial_overlap,
exclusive: bool = True,
) -> SpanAlignment:
"""
Builds a span alignment using a select first match strategy. Each reference
span is aligned with the first matching system span as determined by the provided
`is_match` method. By default, alignments are exclusive.
span is aligned with the first matching system span. By default, a match
corresponds to spans that overlap and have the same label. By default,
alignments are exclusive.
"""
align_map = {}
sys_span_pool = dict.fromkeys(sys.spans)
Expand All @@ -50,14 +51,16 @@ def select_first_match(
def select_best_match(
ref: DocSpans,
sys: DocSpans,
is_match: CheckSpanPair,
score_match: ScoreSpanPair,
is_match: CheckSpanPair = partial_overlap,
score_match: ScoreSpanPair = Span.jaccard,
exclusive: bool = True,
) -> SpanAlignment:
"""
Builds a span alignment using a greedy select best match strategy. Each reference
span is aligned with its best matching system span as defined by the provided
`is_match` and `score_match` methods. By default, alignments are exclusive.
span is aligned with its best matching system span. By default, a match
corresponds to spans that overlap and have the same label and the best match
corresponds to the match with the highest jaccard similarity.
By default, alignments are exclusive.
"""
mapping = {}
sys_span_pool = dict.fromkeys(sys.spans)
Expand Down Expand Up @@ -142,52 +145,34 @@ def construct_aligner(
- select_best: corresponds to select_best_match
- corppa: corresponds to corppa_align
"""
## Validate inputs and get alignment method
align_method = None
match strategy:
case "select_first":
# Validate input parameters
if is_match is None:
raise ValueError(f"Strategy {strategy} requires is_match parameter")
if score_match is not None:
raise ValueError(
f"Strategy {strategy} does not use score_match parameter"
)
# Construct aligner
if exclusive is None:
return lambda r, s: select_first_match(r, s, is_match)
else:
return lambda r, s: select_first_match(
r, s, is_match, exclusive=exclusive
)
align_method = select_first_match
case "select_best":
# Validate input parameters
if is_match is None or score_match is None:
raise ValueError(
f"Strategy {strategy} requires is_match and score_match parameters"
)
# Construct aligner
if exclusive is None:
return lambda r, s: select_best_match(r, s, is_match, score_match)
else:
return lambda r, s: select_best_match(
r, s, is_match, score_match, exclusive=exclusive
)

align_method = select_best_match
case "corppa":
# Validate input parameters
if exclusive is not None:
raise ValueError(
f"Strategy {strategy} does not use exclusive parameter"
)
# Construct aligner
if is_match is not None and score_match is not None:
return lambda r, s: corppa_align(
r, s, is_match=is_match, score_match=score_match
)
elif is_match is not None:
return lambda r, s: corppa_align(r, s, is_match=is_match)
elif score_match is not None:
return lambda r, s: corppa_align(r, s, score_match=score_match)
else:
return lambda r, s: corppa_align(r, s)
align_method = corppa_align
case _:
raise ValueError(f"Unknown alignment strategy: {strategy}")
# Determine optional args
options = {}
if is_match is not None:
options["is_match"] = is_match
if score_match is not None:
options["score_match"] = score_match
if exclusive is not None:
options["exclusive"] = exclusive
# Construct aligner
return lambda r, s: align_method(r, s, **options) # ty: ignore[invalid-argument-type]
Loading
Loading