Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
37 changes: 37 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,43 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Added

- **Correction log**: New `source` column records which page made each entry;
existing logs load with `source = corrections_page`. A new `check_type`
column goes with the new `accept` action
(`CORRECTION_LOG_SCHEMA`, `ensure_log_columns` in the new Streamlit-free
`src/datasure/processing/correction_log.py`, shared with the replication
package, which now exports legacy logs with these columns) β€” #296
- **Accept action**: `CorrectionProcessor.accept_value` records that a flagged
value (outliers, constraints, backchecks, duplicates, GPS) was reviewed and
is correct, with a required reason. An acceptance is rejected if the data
no longer holds the value being accepted. `get_active_acceptances` returns the
acceptances whose recorded value still matches the data (for GPS, both
latitude and longitude). Replay and the generated `4_corrections.do` skip
`accept` rows; `correction_log.csv` keeps them, and the README's correction
counts exclude them. Values compare by value, not text: missing matches
missing (None or NaN), numbers compare numerically, and every row with the
KEY must match β€” #296
- **Atomic apply**: `CorrectionProcessor.apply_corrections` applies a list of
`CorrectionEntry` objects all or nothing; if the log save fails, the
corrected data is restored β€” #296
- **Shared correction form**: `src/datasure/utils/correction_form.py`
(`render_correction_form`, `render_correction_inputs`,
`apply_correction_entries`) renders the action, new-value and reason inputs
for a prefilled KEY/column/current value, with namespaced widget keys. The
Correct Data page now uses it β€” #296

### Fixed

- **Corrections cache**: `CorrectionProcessor`'s cached reads were keyed only on
`alias`, so two projects sharing an alias shared cached corrected data and
logs. The processor is now hashed by `project_id` β€” #296
- **Apply button**: A new value of `0` no longer disables Apply. An empty
string still does; use "remove value" to blank a cell β€” #296
- **Correction log schema**: Removing the last correction entry now leaves an
empty log with the full schema, including status columns β€” #296

## [1.1.0] - 2026-09-21

### Added
Expand Down
4 changes: 3 additions & 1 deletion docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,13 +35,15 @@ src/datasure/
β”‚ └── local.py # Local file import (csv/xlsx/xls/json/dta/parquet)
β”œβ”€β”€ processing/
β”‚ β”œβ”€β”€ prep.py # Data preparation operations (Polars)
β”‚ └── corrections.py # Data correction application
β”‚ β”œβ”€β”€ corrections.py # Corrections and accept entries (CorrectionProcessor)
β”‚ └── correction_log.py # Correction log schema/backfill (no Streamlit)
β”œβ”€β”€ replication/ # Stata/Python replication package export
β”œβ”€β”€ models/
β”‚ β”œβ”€β”€ schemas.py # Pydantic models
β”‚ └── enums.py # Prep action/method enums
β”œβ”€β”€ utils/ # Shared utilities (DuckDB, cache, config, charts,
β”‚ # credentials, SurveyCTO API, UI helpers, ...)
β”‚ └── correction_form.py # Shared correction form used across pages
└── views/ # Streamlit pages (top-level page scripts)
β”œβ”€β”€ start_view.py # Project selection/creation
β”œβ”€β”€ import_view.py # Credentials + data import
Expand Down
12 changes: 12 additions & 0 deletions docs/USER_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -460,6 +460,18 @@ All corrections are tracked with:
- Action type
- Reason for correction
- Timestamp
- Status of the last reapply, with the reason if it failed
- Source: the page that made the entry (`corrections_page` for this page)

The log can also contain **accept** entries. An accept entry records that a
flagged value was reviewed and is correct. It never changes the data, and the
Check type column shows which check it applies to. It stays in effect only
while the value is unchanged. Accept entries are kept in `correction_log.csv`
in the replication package but are not part of the corrections script. You can
remove an accept entry with "Remove correction step" like any other entry.

To blank a cell, use "remove value": "modify value" needs a non-empty new
value (`0` is valid).

#### Verifying Corrections

Expand Down
63 changes: 63 additions & 0 deletions src/datasure/processing/correction_log.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
"""Schema and vocabulary of the correction log (`corr_log_{alias}`).

Kept free of Streamlit so the replication package can read logs with the
same schema and backfill rules as `CorrectionProcessor`.
"""

import polars as pl

CORRECTIONS_PAGE_SOURCE = "corrections_page"

# Actions that change the data when a correction is applied or replayed.
MODIFY_VALUE_ACTION = "modify value"
REMOVE_VALUE_ACTION = "remove value"
REMOVE_ROW_ACTION = "remove row"
CORRECTION_ACTIONS = (MODIFY_VALUE_ACTION, REMOVE_VALUE_ACTION, REMOVE_ROW_ACTION)

# An "accept" entry records that a flagged value was reviewed and is correct.
# It never changes the data; check pages use it to stop flagging the value.
ACCEPT_ACTION = "accept"
ACCEPT_CHECK_TYPES = ("outliers", "constraints", "backchecks", "duplicates", "gps")

# Full schema of a persisted correction log (`corr_log_{alias}`), in column order.
CORRECTION_LOG_SCHEMA: dict[str, pl.DataType] = {
"date": pl.Datetime("us"),
"KEY": pl.String,
"ID": pl.String,
"action": pl.String,
"column": pl.String,
"current_value": pl.String,
"new_value": pl.String,
"reason": pl.String,
"status": pl.String,
"status_reason": pl.String,
"source": pl.String,
"check_type": pl.String,
}

# Values given to columns that were added to the log after some logs were
# already persisted. Every legacy entry came from the Corrections page and
# was applied successfully when it was logged.
_LOG_BACKFILL_DEFAULTS: dict[str, str | None] = {
"status": "Successful",
"status_reason": None,
"source": CORRECTIONS_PAGE_SOURCE,
"check_type": None,
}


def ensure_log_columns(df: pl.DataFrame) -> pl.DataFrame:
"""Backfill columns missing from logs persisted before those columns existed."""
if df.width == 0:
return df
for column, default in _LOG_BACKFILL_DEFAULTS.items():
if column not in df.columns:
df = df.with_columns(
pl.lit(default, dtype=CORRECTION_LOG_SCHEMA[column]).alias(column)
)
return df


def empty_correction_log() -> pl.DataFrame:
"""Return a correction log with no entries and the full schema."""
return pl.DataFrame(schema=CORRECTION_LOG_SCHEMA)
Loading
Loading