Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 28 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -82,6 +82,34 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
`CorrectionEntry.severity` sets it and is rejected on non-accept actions,
on acceptances of other checks, and with any value other than `hard`.
Hard acceptances are highlighted in the Correction Log β€” #298
- **Backcheck targets**: New `checks/backchecks/coverage.py`.
`compute_backcheck_coverage` returns `BackcheckCoverage`: eligible unique
survey IDs (after duplicate handling and the optional eligibility filter),
how many have a matching backcheck, the on-track %, points vs the target, and
expected backchecks, `ceil(survey_target Γ— target% / 100)`, when
`survey_target` is set. `compute_staff_coverage` returns per-enumerator
coverage, including enumerators with no backchecks, or backchecks done per
backchecker. Neither needs comparison columns. The Backchecks Summary has a
new Targets section showing coverage against the target % and backchecks
done against expected, each with its deviation as a delta. The settings
panel gains an eligibility filter (`eligibility_column`,
`eligibility_values`).
`settings_from_page_config` builds `BackcheckSettings` from the page config,
where a target of 0 means not set β€” #318

### Changed

- **Breaking**: `BackcheckSettings.backcheck_target_percent` is now
`float | None` (0–100), defaulting to None instead of 10;
`effective_target_percent` applies the 10% default. The target resolves from
the settings panel, then the page config, then 10%. A cleared panel value, or
a saved value outside 0–100, falls back to the page config.
`BackcheckSettings` gains `survey_target` β€” #318
- **Breaking**: `compute_enumerator_backchecker_stats` no longer returns the
"Surveys" and "Backchecks" columns (both were the count of compared survey
KEYs). The Enumerator Backchecker Error Statistics table now takes them from
`compute_staff_coverage` and renders before comparison columns are
configured β€” #318

### Fixed

Expand Down
29 changes: 26 additions & 3 deletions docs/USER_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -994,9 +994,29 @@ Configure validation:
- **Enumerator**: Original data collector
- **Back Checker**: QC validator
- **Date**: Back check date
- **Target %**: Target back check rate (e.g., 10%)
- **Backcheck target (%)**: Share of surveys to back check (e.g., 10%).
Pre-filled from the page configuration. A value you enter here is saved and
overrides the page configuration; clear it to fall back. If neither is set,
10% is used and a warning is shown.
- **Eligibility Filter**: Optional survey column and values that mark a survey
eligible for back checks (e.g., `consent` in `1`). Only eligible surveys
count towards coverage.
- **Handle Duplicates**: Include or exclude duplicates

##### Backchecks Summary

A row of counts (survey observations, back check observations, enumerators and
back checkers), then a **Targets** section with two metrics side by side:

- **Backcheck Coverage**: Share of eligible unique survey IDs (after duplicate
handling) with at least one matching back check, and how many points it is
above or below the target. It is calculated before any back check columns
are configured.
- **Backchecks vs Expected**: When the page configuration sets the target
number of survey responses, back checks done against the back checks
expected (survey target Γ— target %, rounded up), and how many back checks
above or below that it is. Values over 100% are shown as is.

**Add Back Check Columns**:
Click "Add a back check column" (+ button):

Expand Down Expand Up @@ -1039,7 +1059,10 @@ Detailed column-level validation:
Performance by original enumerator:

- Enumerator ID
- \# surveys back checked
- Surveys: eligible unique submissions
- Backchecks: how many of those were back checked
- Coverage % and points vs target (coverage below target is highlighted;
enumerators with no back checks show 0%)
- \# values compared
- \# different values
- Error rate (%)
Expand All @@ -1049,7 +1072,7 @@ Performance by original enumerator:
Performance by validator:

- Back Checker ID
- \# surveys validated
- Backchecks: unique surveys back checked
- \# values compared
- \# discrepancies
- Error rate (%)
Expand Down
12 changes: 7 additions & 5 deletions src/datasure/checks/backchecks/compute.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,9 +29,8 @@ def load_default_backchecks_settings(

Loads previously saved backcheck report settings from the settings file
and merges them with the provided default configuration. Saved settings
take precedence over defaults.

Cached for 60 seconds to reduce file I/O operations.
take precedence over defaults, except a cleared or invalid backcheck
target, which falls back to the configured one.

Parameters
----------
Expand All @@ -46,6 +45,11 @@ def load_default_backchecks_settings(
Merged settings combining saved and default configurations.
"""
saved_settings = load_check_settings(settings_file, TAB_NAME)
# A cleared target falls back to the configured one, as does a value saved
# by the old count-based input that is not a valid percentage.
saved_target = saved_settings.get("backcheck_target_percent")
if not isinstance(saved_target, int | float) or not 0 <= saved_target <= 100:
saved_settings.pop("backcheck_target_percent", None)

default_settings: dict = dict(config)
default_settings.update(saved_settings)
Expand Down Expand Up @@ -919,8 +923,6 @@ def _calculate_staff_statistics(
# Initialize stats dict
staff_stats = {
staff_col: staff_name,
"Surveys": staff_data[survey_key].n_unique(),
"Backchecks": staff_data[survey_key].n_unique(),
"Avg Days": _calculate_average_days(staff_data, survey_date, backcheck_date),
}

Expand Down
213 changes: 213 additions & 0 deletions src/datasure/checks/backchecks/coverage.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,213 @@
"""Backcheck coverage against the backcheck target.

The target % resolves from the settings panel, then the page config, then
`DEFAULT_TARGET_PERCENT`. Coverage counts unique survey IDs, after duplicate
handling and the optional eligibility filter, that have a matching backcheck.
"""

import math
from dataclasses import dataclass

import polars as pl

from datasure.checks.backchecks.compute import _prepare_data_for_merge
from datasure.checks.backchecks.models import BackcheckSettings

DEFAULT_TARGET_PERCENT: float = 10.0

# Column `backchecked_surveys` adds to flag surveys with a matching backcheck.
BACKCHECKED: str = "_backchecked"


def settings_from_page_config(config: dict) -> BackcheckSettings:
"""Build backcheck settings from the page configuration.

The page configuration stores 0 for a target left blank, so 0 means
"not set" for both the backcheck target % and the survey target.
"""
return BackcheckSettings(
**{
**config,
"backcheck_target_percent": config.get("backcheck_target_percent") or None,
"survey_target": config.get("survey_target") or None,
}
)


def effective_target_percent(settings: BackcheckSettings) -> float:
"""Return the backcheck target %, or the default when none is set."""
if settings.backcheck_target_percent is None:
return DEFAULT_TARGET_PERCENT
return settings.backcheck_target_percent


@dataclass(frozen=True)
class BackcheckCoverage:
"""Backcheck progress against the target %.

`on_track_percent` and `points_vs_target` are None when there are no
eligible surveys. The expected-total fields are None when the survey
target is not set.
"""

eligible: int
backchecked: int
target_percent: float
on_track_percent: float | None
points_vs_target: float | None
expected_backchecks: int | None
expected_progress_percent: float | None


def backchecked_surveys(
survey_data: pl.DataFrame,
backcheck_data: pl.DataFrame,
settings: BackcheckSettings,
) -> pl.DataFrame | None:
"""Return one row per eligible survey ID, flagged if it was backchecked.

Both datasets go through the duplicate-handling option first, as in the
backcheck comparison. Matching is by survey ID only.

Returns
-------
pl.DataFrame | None
The deduplicated eligible survey rows plus a boolean `BACKCHECKED`
column, or None if either dataset lacks the survey ID column.
"""
survey_id = settings.survey_id
if (
not survey_id
or survey_id not in survey_data.columns
or survey_id not in backcheck_data.columns
):
return None

option = settings.drop_duplicates_option
surveys = _prepare_data_for_merge(survey_data, survey_id, option).unique(
subset=[survey_id], keep="first", maintain_order=True
)
backchecked_ids = _prepare_data_for_merge(backcheck_data, survey_id, option)[
survey_id
].drop_nulls()

surveys = surveys.filter(pl.col(survey_id).is_not_null())
eligibility_column = settings.eligibility_column
if (
eligibility_column
and eligibility_column in surveys.columns
and settings.eligibility_values
):
surveys = surveys.filter(
pl.col(eligibility_column).cast(pl.Utf8).is_in(settings.eligibility_values)
)

return surveys.with_columns(
pl.col(survey_id).is_in(backchecked_ids.implode()).alias(BACKCHECKED)
)


def compute_backcheck_coverage(
survey_data: pl.DataFrame,
backcheck_data: pl.DataFrame,
settings: BackcheckSettings,
) -> BackcheckCoverage | None:
"""Compute overall backcheck coverage against the target.

Returns None if either dataset lacks the survey ID column.
"""
surveys = backchecked_surveys(survey_data, backcheck_data, settings)
if surveys is None:
return None

eligible = surveys.height
backchecked = int(surveys[BACKCHECKED].sum())
target_percent = effective_target_percent(settings)
on_track_percent = backchecked / eligible * 100 if eligible else None

expected = None
if settings.survey_target:
# Round away float noise first so an exact product doesn't round up.
expected = math.ceil(round(settings.survey_target * target_percent / 100, 9))

return BackcheckCoverage(
eligible=eligible,
backchecked=backchecked,
target_percent=target_percent,
on_track_percent=on_track_percent,
points_vs_target=(
on_track_percent - target_percent if on_track_percent is not None else None
),
expected_backchecks=expected,
expected_progress_percent=backchecked / expected * 100 if expected else None,
)


def compute_staff_coverage(
survey_data: pl.DataFrame,
backcheck_data: pl.DataFrame,
settings: BackcheckSettings,
staff_type: str = "enumerator",
) -> pl.DataFrame:
"""Compute backcheck coverage per enumerator, or backchecks per backchecker.

Parameters
----------
survey_data : pl.DataFrame
Survey dataset.
backcheck_data : pl.DataFrame
Backcheck dataset.
settings : BackcheckSettings
Backcheck settings.
staff_type : str
Either "enumerator" or "backchecker".

Returns
-------
pl.DataFrame
For enumerators: the enumerator column, "Surveys" (eligible unique
submissions), "Backchecks" (how many of those were backchecked),
"Coverage %" and "vs target" (points). Enumerators with no backchecks
appear at 0%. For backcheckers: the backchecker column and
"Backchecks" (unique backchecked survey IDs). Empty if the survey ID
or staff column is missing.
"""
surveys = backchecked_surveys(survey_data, backcheck_data, settings)
if surveys is None:
return pl.DataFrame()

if staff_type == "enumerator":
staff_col = settings.enumerator
if not staff_col or staff_col not in surveys.columns:
return pl.DataFrame()
target_percent = effective_target_percent(settings)
return (
surveys.filter(pl.col(staff_col).is_not_null())
.group_by(staff_col, maintain_order=True)
.agg(
pl.len().alias("Surveys"),
pl.col(BACKCHECKED).sum().cast(pl.Int64).alias("Backchecks"),
)
.with_columns(
(pl.col("Backchecks") / pl.col("Surveys") * 100).alias("Coverage %")
)
.with_columns((pl.col("Coverage %") - target_percent).alias("vs target"))
.with_columns(pl.col("Surveys").cast(pl.Int64))
)

staff_col = settings.backchecker
if not staff_col or staff_col not in backcheck_data.columns:
return pl.DataFrame()
survey_id = settings.survey_id
backchecked_ids = surveys.filter(pl.col(BACKCHECKED))[survey_id]
return (
_prepare_data_for_merge(
backcheck_data, survey_id, settings.drop_duplicates_option
)
.filter(
pl.col(survey_id).is_in(backchecked_ids.implode())
& pl.col(staff_col).is_not_null()
)
.group_by(staff_col, maintain_order=True)
.agg(pl.col(survey_id).n_unique().cast(pl.Int64).alias("Backchecks"))
)
16 changes: 14 additions & 2 deletions src/datasure/checks/backchecks/models.py
Original file line number Diff line number Diff line change
Expand Up @@ -76,8 +76,20 @@ class BackcheckSettings(BaseModel):
)
enumerator: str | None = Field(None, description="Column containing enumerator")
backchecker: str | None = Field(None, description="Column containing back checker")
backcheck_target_percent: int = Field(
10, description="Target percentage of backchecks"
backcheck_target_percent: float | None = Field(
None,
ge=0,
le=100,
description="Target percentage of surveys to backcheck; None if not set",
)
survey_target: int | None = Field(
None, ge=0, description="Target number of survey responses"
)
eligibility_column: str | None = Field(
None, description="Survey column that marks a survey eligible"
)
eligibility_values: list[str] | None = Field(
None, description="Values of eligibility_column that mark a survey eligible"
)
drop_duplicates_option: str = Field(
"drop", description="How to handle duplicate entries"
Expand Down
Loading
Loading