Skip to content

feat(backchecks): attribute mismatches to an error source and report an adjusted error rate - #323

Open
iabaako wants to merge 7 commits into
feat/321-correction-log-userfrom
feat/301-backcheck-attribution
Open

iabaako wants to merge 7 commits into
feat/321-correction-log-userfrom
feat/301-backcheck-attribution

Conversation

@iabaako

@iabaako iabaako commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Pull Request Summary 🚀

Closes #301.

Stacked PR. This branches from feat/321-correction-log-user (#322) and
targets it. Retarget to main once #322 merges. #306 branches from this
branch.

What does this PR do? 📝

Reviewers can now record who caused each backcheck mismatch, and the
Backchecks page reports an adjusted error rate next to the regular one. The
page never changes survey or backcheck data.

From the issue (#301):

  • No more accepting backchecks. backchecks is removed from the check
    types that can be accepted, so a backcheck mismatch can't be accepted as
    valid. It can only be attributed.
  • Attribution in Comparison Results Details.
    • Each mismatch row has a pinned Review button that opens a dialog. The
      dialog shows the survey and backcheck values read-only and offers
      Enumerator, Backchecker, Respondent or Unattributed.
    • To attribute several mismatches at once, select their rows and click
      Review on one of the selected rows.
    • Only mismatches can be attributed.
    • A note is required for Backchecker and Respondent.
    • A new Error Source column shows the source; every mismatch starts as
      Unattributed.
  • Storage.
    • Each attribution is appended to a new table,
      bc_attribution_{page_name_id}, in the logs database, with the reviewer
      (from Corrections: record who made each correction log entry #321) and the date.
    • The latest entry for a mismatch wins, and earlier entries are kept as
      history.
    • An attribution lapses if either value changes. If the values now match,
      there is no mismatch to attribute.
    • An Attribution log expander lists the history, newest first.
  • Adjusted error rate. It uses the same denominator as the regular rate
    (values compared), and Unattributed mismatches always count.
    • Enumerators: (mismatches − Backchecker − Respondent) ÷ values compared.
    • Backcheckers: (mismatches − Enumerator − Respondent) ÷ values compared.
  • Statistics tables.
    • The enumerator and backchecker tables gain "Adjusted Error Rate %" for
      each category and the total.
    • The variable table gains Enumerator, Backchecker, Respondent and
      Unattributed mismatch counts.
    • The summary gains a "Mismatches Attributed" metric.
  • After a save, the whole app reruns and a toast confirms it.
  • Docs. The User Guide documents the policy and the two rates.

Not in the issue — added during review with the requester:

  • Error rate target (%). The issue asks for each rate column to be
    "highlighted against the target", but no error-rate target existed; the only
    target was backcheck coverage. A new optional setting under Tracking Options
    sets one. Each regular and adjusted rate column above it is highlighted, in
    both the enumerator and backchecker views. With no target set, nothing is
    highlighted.
  • Error Rates cards in the Backchecks Summary. A new section below Targets
    has one card for the total error rate and one for each category.
    • Each card's grey delta is the enumerator adjusted rate.
    • The card's help shows the backchecker adjusted rate and the counts.
    • A category with no values compared shows N/A with "No values compared",
      so all cards stay the same height.
  • Comparison table layout. Review and Error Source are the first two
    columns and stay pinned while scrolling. The Review column has a fixed
    width that fits its label.
  • Faster table interaction. Comparison Results Details now reruns on its
    own: selecting rows, changing its filters or opening Review reruns only that
    section, not the whole report. Saving an attribution still reruns the whole
    app so the rates refresh.
  • Stable row order. The comparison table is now sorted by column name and
    keys, so a Review click always resolves to the row that was clicked. This
    changes the default row order users see.

Why is this change needed? 🤔

Backcheck results measure data collection quality, so the page must not offer
a way to make poor results disappear. Before this PR, backchecks could be
accepted through the correction log. Yet some mismatches are not the
enumerator's fault: the backchecker recorded a wrong value, or the respondent
changed their answer. Attribution records that without hiding anything:

  • Every mismatch stays in the regular error rate.
  • Attribution only feeds the adjusted rate.

Known gap, expected: duplicate and unmatched backcheck IDs are out of
scope (moved to #306 and #303; #302 is closed). Until #303 lands, the
existing Handle Duplicates setting keeps deciding which duplicates are
compared. #303 closes this gap.

How was this implemented? 🛠️

  • New Streamlit-free module checks/backchecks/attribution.py. It holds
    the error sources, the note rule, building log entries (mismatch-only),
    marking each comparison with its source (latest entry wins, lapse rule),
    counts, both adjusted-rate formulas, the attributed share, and log storage.
    It is tested without a running app, like checks/outliers/review.py.
  • Matching. Keys and values are stored and compared as text, so the lapse
    check doesn't depend on column types. When the survey KEY is also the merge
    ID, the survey and backcheck KEYs are the same column, and that case is
    handled.
  • compute.py.
    • The staff statistics take the staff type to pick the right adjusted
      formula.
    • The column statistics add the per-source counts.
    • A new function computes the overall error rates for the summary cards.
  • report_ui.py.
    • The page loads the attribution log once per run and marks the comparison
      results, and every section reads from those marked results.
    • The Review button and dialog follow the outliers pattern.
    • Comparison Results Details runs as a fragment.
  • Settings. BackcheckSettings gains error_rate_target_percent (0–100).
    A saved value outside that range falls back to the default, as the coverage
    target does.

How to test or reproduce ? 🧪

Automated. just test gives 3634 passed, 6 skipped, and
just pre-commit-run passes. The new tests cover:

  • the mismatch-only restriction
  • the required note
  • the lapse rule
  • both adjusted-rate formulas
  • the variable-table counts
  • the attributed % metric
  • the error rate target
  • the summary cards

Manual, on a project with survey and backcheck data and at least one
configured backcheck column:

  1. Open the Backchecks page. Every mismatch in Comparison Results Details shows
    Unattributed, and the adjusted rates equal the regular rates.
  2. Click Review on a mismatch and choose Respondent. Save stays
    disabled until you type a note. Save it. A toast appears, the row shows
    Respondent, and the enumerator adjusted rate drops. The regular rate and
    the mismatch counts don't change.
  3. Select several mismatch rows, click Review on one of them, and attribute
    them all to Enumerator. All selected rows update.
  4. Open Attribution log. Every entry is listed with your reviewer name and
    the date, including earlier ones that were replaced.
  5. Correct one of the attributed survey values on the Correct Data page, then
    come back. That mismatch is Unattributed again.
  6. Set Error rate target (%) in Tracking Options. Rates above it are
    highlighted, and each rate column is checked against the target on its own.
  7. Check the new Error Rates cards under Targets.
  8. Select and unselect rows. Only Comparison Results Details reruns, not the
    whole page.
  9. On the Correct Data page, confirm backchecks can't be accepted.

Screenshots (if applicable) 📷

Correction log shows who made the change. User is shown in the sidebar to the left.
Screenshot 2026-10-06 100616

Each error can be attributed. Clicking "Review" shows a dialog box.
Screenshot 2026-10-06 120957

Attribution log shows a history of changes including Error Source, User and Datetime
Screenshot 2026-10-06 121021

Backcheck Summary now includes overall Error Rates for backchecks with a delta showing adjusted rates.
Screenshot 2026-10-06 120909

Checklist ✅

  • I have run and tested my changes locally (full test suite, plus each new
    widget rendered in Streamlit's test runner. The manual checks above
    still need a browser run.)
  • I have limit this PR to less than 1000 lines of code change (if not,
    explain why): 2054 lines changed in total, of which 922 are source, 1009
    tests and 110 docs. The source change alone is under 1000. It runs over
    because the issue's acceptance criteria call for tests of every rule,
    and because of the extras listed above.
  • I have updated/added tests to cover my changes (if applicable)
  • I have updated/added requirements to cover my changes (if applicable):
    N/A, no new dependencies.
  • I have run linting and formatting on any code changes (if applicable)
  • I have updated the documentation (README, etc.) accordingly. Changed:
    docs/USER_GUIDE.md (policy, both rates, attribution, new settings and
    cards), docs/ARCHITECTURE.md (new table), CHANGELOG.md.
  • I have reviewed and resolved any merge conflict: none against the base
    branch, feat/321-correction-log-user.

Reviewer Emoji Legend

:code: Meaning
😃👍💯 :smiley: :+1: :100: I like this...

...and I want the author to know it! This is a way to highlight positive parts of a code review.
⭐⭐⭐ :star: :star: :star: Important to fix before PR can be approved...

And I am providing reasons why it needs to be addressed as well as suggested improvements.
⭐⭐ :star: :star: Important to fix but non-blocking for PR approval...

And I am providing suggestions where it could be improved either in this PR or later.
⭐ :star: Give this some thought but non-blocking for PR approval...

...and consider this a suggestion, not a requirement.
❓ :question: I have a question.

This should be a fully formed question with sufficient information and context that requires a response.
📝 :memo: This is an explanatory note, fun fact, or relevant commentary that does not require any action.
⛏ :pick: This is a nitpick.

This does not require any changes and is often better left unsaid. This may include stylistic, formatting, or organization suggestions and should likely be prevented/enforced by linting if they really matter
♻️ :recycle: Suggestion for refactoring.

Should include enough context to be actionable and not be considered a nitpick.

🤖 Generated with Claude Code

iabaako and others added 5 commits October 6, 2026 11:24
Reviewers can attribute each backcheck mismatch to the enumerator,
backchecker or respondent from a pinned Review button on Comparison
Results Details; selecting several mismatch rows attributes them together.
A note is required for Backchecker and Respondent. Attributions are
appended to bc_attribution_{page_name_id} in the logs db with the user and
date, and lapse when either value changes. An Attribution log expander
lists the history.

The enumerator and backchecker tables gain adjusted error rates, the
variable table gains mismatch counts per source, and the summary shows the
share of mismatches attributed. A new optional error rate target
highlights each regular and adjusted rate column above it.

Backcheck mismatches can no longer be accepted: backchecks is dropped from
ACCEPT_CHECK_TYPES. The Backchecks page never changes survey or backcheck
data.

Closes #301

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Error Source now comes right after the Review button, and both stay
pinned while scrolling. The Review column has a fixed width that fits its
label instead of the default, which was too wide.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comparison Results Details now renders in an st.fragment, so selecting
rows, changing its filters or opening the Review dialog reruns just that
section instead of the whole report. Saving an attribution still reruns
the whole app (st.rerun(scope="app")) so the rates are refreshed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Backchecks Summary gains an Error Rates section below Targets: one
card for the total error rate and one for each category. Each card shows
the enumerator adjusted error rate as a grey delta, and its help gives the
backchecker adjusted rate and the counts behind the rate. A category with
no values compared shows N/A.

New compute_overall_error_rates returns the rates as OverallErrorRate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A category with no values compared had no delta line, so its card was
shorter than the others. It now shows "No values compared" as a grey
delta, and the cards use bordered columns, which stretch to the tallest
card in the row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@iabaako
iabaako requested a review from a team as a code owner October 6, 2026 12:16
@iabaako
iabaako added this pull request to stack #313 October 6, 2026 12:27
@iabaako
iabaako requested a balanced review from Copilot October 6, 2026 12:27
Comment thread tests/checks/backchecks/test_report_ui_attribution.py Fixed
Comment thread tests/checks/backchecks/test_report_ui_attribution.py Fixed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Clearing a configured error-rate target does not persist across reruns, causing highlighting to return unexpectedly.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
What changed in this PR

Adds backcheck mismatch attribution, adjusted error rates, and related reporting while preventing backcheck acceptance.

Changes:

  • Adds append-only mismatch attribution and review UI.
  • Adds adjusted rates, targets, summary cards, and source counts.
  • Updates tests and documentation.
File Description
src/​datasure/​checks/​backchecks/​attribution.py Implements attribution logic and storage.
src/​datasure/​checks/​backchecks/​compute.py Computes adjusted rates and source counts.
src/​datasure/​checks/​backchecks/​models.py Adds the error-rate target setting.
src/​datasure/​checks/​backchecks/​report_ui.py Adds review UI, metrics, and highlighting.
src/​datasure/​checks/​backchecks/​settings_ui.py Adds error-rate target controls.
src/​datasure/​processing/​correction_log.py Disallows backcheck acceptance.
tests/​checks/​backchecks/​test_attribution.py Tests attribution rules and storage.
tests/​checks/​backchecks/​test_compute.py Tests rates, counts, and settings.
tests/​checks/​backchecks/​test_report_ui_attribution.py Tests attribution UI behavior.
tests/​checks/​backchecks/​test_settings_ui.py Tests target settings UI.
tests/​processing/​test_corrections.py Tests rejection of backcheck acceptance.
docs/​USER_GUIDE.md Documents attribution and adjusted rates.
docs/​ARCHITECTURE.md Documents attribution-log storage.
CHANGELOG.md Records user-facing changes.
docs/​changelog_guide.md Formats example code.
.claude/​subagents/​py-format-lint/​SUBAGENT.md Formatting-only cleanup.
.claude/​skills/​surveycto-api/​SKILL.md Formatting-only cleanup.
.claude/​skills/​duckdb/​SKILL.md Formatting-only cleanup.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/datasure/checks/backchecks/compute.py Outdated
iabaako and others added 2 commits October 6, 2026 13:13
Clearing the optional error rate target saved None, but loading the
settings treated None as invalid and fell back to the configured target,
so highlighting could come back after a rerun. A saved None now stays;
only values outside 0-100 fall back. Also tidies two test assertions
flagged by code scanning.

Addresses review comments on #323.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Running pre-commit on all files reformatted code examples in three
.claude skill files and docs/changelog_guide.md. Those changes are
unrelated to #301, so they are restored to match the base branch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@sonarqubecloud

sonarqubecloud Bot commented Oct 6, 2026

Copy link
Copy Markdown

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Backchecks: attribute mismatches to an error source and report an adjusted error rate

2 participants