Skip to content
35 changes: 34 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
`src/datasure/processing/correction_log.py`, shared with the replication
package, which now exports legacy logs with these columns) β€” #296
- **Accept action**: `CorrectionProcessor.accept_value` records that a flagged
value (outliers, constraints, backchecks, duplicates, GPS) was reviewed and
value (outliers, constraints, duplicates, GPS) was reviewed and
is correct, with a required reason. An acceptance is rejected if the data
no longer holds the value being accepted. `get_active_acceptances` returns the
acceptances whose recorded value still matches the data (for GPS, both
Expand Down Expand Up @@ -102,9 +102,42 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
sidebar, saved in `cache/user_settings.json`, else the OS login
(`getpass.getuser()`). Existing logs load with a null `user`. The
Correction Log table and `correction_log.csv` include it β€” #321
- **Backcheck mismatch attribution**: New Streamlit-free
`checks/backchecks/attribution.py`. Reviewers attribute mismatches in
Comparison Results Details to an `ErrorSource` (Enumerator, Backchecker,
Respondent, Unattributed) from a pinned Review button that opens a dialog;
selecting several mismatch rows and clicking Review on one of them
attributes them together. Only `match_status == "mismatch"` rows can be
attributed, and Backchecker and Respondent need a note. Entries are appended
to `bc_attribution_{page_name_id}` in the `logs` db with the survey and
backcheck KEYs, column, both values (as text), source, note, user
(`get_reviewer_name`) and date. `mark_error_sources` adds an `error_source`
column: the latest entry per KEY pair and column applies only while both
values still equal the stored ones, otherwise the mismatch is Unattributed.
An Attribution log expander lists the history. Attribution never changes
the data, the mismatch counts or the regular error rate β€” #301
- **Adjusted error rate**: `compute_enumerator_backchecker_stats` adds
"Adjusted Error Rate % (Cat n)" and "(Total)": for enumerators
(mismatches βˆ’ Backchecker βˆ’ Respondent) Γ· values compared, for backcheckers
(mismatches βˆ’ Enumerator βˆ’ Respondent) Γ· values compared. Unattributed
mismatches always count. `compute_column_stats` adds "Enumerator / Backchecker
/ Respondent / Unattributed Mismatches" counts. The Backchecks Summary shows
"Mismatches Attributed" (% of mismatches with a source). New optional
`BackcheckSettings.error_rate_target_percent` ("Error rate target (%)" in
Tracking Options): each regular and adjusted rate column above it is
highlighted in both the enumerator and backchecker views β€” #301
- **Overall backcheck error rates**: `compute_overall_error_rates` returns an
`OverallErrorRate` (compared, mismatches, error rate, enumerator and
backchecker adjusted rates) for the total over categories 1–3 and for each
category. The Backchecks Summary gains an Error Rates section below Targets
with one card each, the enumerator adjusted rate as a grey, arrowless delta
and the backchecker adjusted rate in the help β€” #301

### Changed

- **Breaking**: `backchecks` is no longer in `ACCEPT_CHECK_TYPES`, so a
backcheck mismatch can't be accepted; it can only be attributed β€” #301

- **Breaking**: `BackcheckSettings.backcheck_target_percent` is now
`float | None` (0–100), defaulting to None instead of 10;
`effective_target_percent` applies the 10% default. The target resolves from
Expand Down
6 changes: 4 additions & 2 deletions docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -96,8 +96,10 @@ Per project (UUID-keyed):

- `cache/{project_id}/data/` β€” the DuckDB databases `raw.duckdb`,
`prep.duckdb`, `corrected.duckdb`
- `cache/{project_id}/settings/` β€” `logs.duckdb` (import/prep logs), JSON
settings, and credential metadata
- `cache/{project_id}/settings/` β€” `logs.duckdb` (import/prep logs, and the
append-only backcheck attribution logs `bc_attribution_{page_name_id}`
written by `checks/backchecks/attribution.py`), JSON settings, and
credential metadata
- `cache/projects.json` β€” the project registry
- `cache/user_settings.json` β€” per-user preferences, currently the
"Reviewer name" recorded as `user` in correction logs
Expand Down
72 changes: 69 additions & 3 deletions docs/USER_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -1001,6 +1001,10 @@ Configure validation:
Pre-filled from the page configuration. A value you enter here is saved and
overrides the page configuration; clear it to fall back. If neither is set,
10% is used and a warning is shown.
- **Error rate target (%)**: Optional. The highest acceptable error rate. In
the enumerator and back checker statistics, every error rate and adjusted
error rate above it is highlighted, each column on its own. Leave blank to
highlight nothing.
- **Eligibility Filter**: Optional survey column and values that mark a survey
eligible for back checks (e.g., `consent` in `1`). Only eligible surveys
count towards coverage.
Expand All @@ -1009,7 +1013,9 @@ Configure validation:
##### Backchecks Summary

A row of counts (survey observations, back check observations, enumerators and
back checkers), then a **Targets** section with two metrics side by side:
back checkers) and **Mismatches Attributed**, the share of mismatches given an
error source (see [Attributing Mismatches](#attributing-mismatches)). Then a
**Targets** section with two metrics side by side:

- **Backcheck Coverage**: Share of eligible unique survey IDs (after duplicate
handling) with at least one matching back check, and how many points it is
Expand All @@ -1020,6 +1026,14 @@ back checkers), then a **Targets** section with two metrics side by side:
expected (survey target Γ— target %, rounded up), and how many back checks
above or below that it is. Values over 100% are shown as is.

Below the targets, an **Error Rates** section has one card for the total error
rate and one for each category. Each card shows mismatches Γ· values compared
over every back check. Its grey delta is the enumerator adjusted error rate,
and the card's help gives the back checker adjusted error rate (see
[Attributing Mismatches](#attributing-mismatches)). A category with no values
compared shows N/A, with "No values compared" in place of the delta. The cards
appear once back check columns are configured.

**Add Back Check Columns**:
Click "Add a back check column" (+ button):

Expand Down Expand Up @@ -1056,6 +1070,8 @@ Detailed column-level validation:
- \# surveys, backchecks, compared
- \# different values
- Error rate (%)
- \# mismatches attributed to the enumerator, the back checker and the
respondent, and \# still unattributed (there is no adjusted rate per column)

##### Enumerator Statistics

Expand All @@ -1068,7 +1084,7 @@ Performance by original enumerator:
enumerators with no back checks show 0%)
- \# values compared
- \# different values
- Error rate (%)
- Error rate (%) and adjusted error rate (%), by category and in total

##### Back Checker Statistics

Expand All @@ -1078,7 +1094,7 @@ Performance by validator:
- Backchecks: unique surveys back checked
- \# values compared
- \# discrepancies
- Error rate (%)
- Error rate (%) and adjusted error rate (%), by category and in total

##### Comparison Details

Expand All @@ -1090,8 +1106,58 @@ Record-level validation results:
- Survey value
- Back check value
- Comparison result
- Error source, for mismatches
- Column name

##### Attributing Mismatches

Back check results measure how well data was collected, so the Back Checks
page never changes survey or back check data. There is no way to accept a
mismatch as valid or to replace a survey value with the back check value.
Expected differences, such as "Don't know" against "Refused", belong in the
exclude and no-differences lists in the settings. The comparison uses the
corrected survey data, so corrections made on the Correct Data page are
reflected.

What you can record is who caused each mismatch:

- **Enumerator**: the survey value is wrong.
- **Backchecker**: the back check value is wrong. A note is required.
- **Respondent**: the respondent gave different answers. A note is required.
- **Unattributed**: no source yet. Every mismatch starts here, and choosing it
clears an earlier attribution.

Click **Review** on a mismatch in the Comparison Details table to open the
attribution dialog. It shows the survey and back check values, which you
can't edit. To attribute several mismatches at once, select their rows, then
click **Review** on one of the selected rows. Only mismatches can be
attributed.

Each attribution records your reviewer name and the date. The latest
attribution for a mismatch applies only while the survey and back check values
are the ones you attributed. If either value changes, the mismatch is
Unattributed again; if the values now match, there is no mismatch to
attribute. The **Attribution log** expander below the table lists every
attribution, newest first, including the ones later replaced.

Attribution never changes the mismatch counts or the regular error rate. It
only affects the **adjusted error rate**, which has the same denominator as
the error rate (values compared):

- Enumerator adjusted error rate = (mismatches βˆ’ Backchecker βˆ’ Respondent) Γ·
values compared
- Back checker adjusted error rate = (mismatches βˆ’ Enumerator βˆ’ Respondent) Γ·
values compared

Unattributed mismatches always count against both. For example, an enumerator
with 10 values compared and 4 mismatches, 1 attributed to the respondent and 1
to the back checker, has an error rate of 40% and an adjusted error rate of
20%.

Duplicate and unmatched back check IDs are not handled here. Until they can be
resolved from the Duplicates page, the **Handle Duplicates** setting decides
which duplicates are compared.

---

### 9. GPS Checks Report
Expand Down
Loading
Loading