Repository navigation
Conversation
…as_row The rebase of 9b83901 onto the src/ -> foresight/ refactor left ingest/ under backend/src/, so foresight.ingest was unimportable and broke model and Alembic imports. as_row was misspelled mas_row; the tests and the documented contract both use as_row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Settings now also reads the repo-root .env, which docker-compose already uses, so commands run from backend/ pick up the API credentials. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolves by (source, source_ref) before dedupe_key so a start-date revision updates the existing row instead of orphaning it. Structured sources outrank text-extracted ones, and a revision that lands on another row's key merges the two. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Geo-filters around the hotel, pages through results, and fetches incrementally on updated.gte. Requests deleted states explicitly, since the API returns only active events by default. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
python -m foresight.ingest.run --source predicthq --since 7d --report prints fetched/new/updated/unchanged/rejected plus market totals. Records failing validation are logged and counted as rejected rather than aborting the run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PredictHQ often lists one event several times, typically one copy active and one deleted. The first run correctly ignored the deleted copy, but on every later run its sighting already pointed at the event, so it counted as the event's own record and its deletion was applied (20 live Dublin events). Live copies that differed slightly also overwrote each other on every run, so an identical rerun reported 238 updated instead of 0. Each observation now stores the status it reported, and every sighting of an event is ranked the same way: source precedence, then live over withdrawn, then latest upstream revision, with source_ref as a tie-break. Only the top-ranked sighting sets the row, so reruns settle on one winner. Migration backfills status for existing PredictHQ sightings from their payload. Repository tests now clear both tables inside their rolled-back transaction, so locally ingested events no longer break their counts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
services/features/events.py turns the events table into the expected_attendance feature the models train on: same-night events add up, multi-day events are spread over their nights, withdrawn events are dropped, provisional ones count by confidence, and an optional radius drops far venues. scripts/forecast_events.py prices upcoming nights with and without the real events. generate_mock_data.generate can take attendance as input so the mock competitor, fare and flight placeholders respond to real events, as they do in the training data. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Five PredictHQ ingest runs (Dublin and Boston, including same-query reruns), the pre-fix runs that exposed the duplicate-listing bug, a timed-out run, the model forecast, and a findings page summarising them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t CLI docker compose up already ran the migrate service, but `just db-up` started only the database, and the ingest CLI run from the host never ran migrations at all. Pulling the new observation status column then broke the ingest with a missing-column error until someone ran `just migrate` by hand. `just db-up` now starts the migrate service with the database, and the ingest CLI upgrades to head before fetching (a no-op when current). Alembic gets no ini file there, so it leaves the run report's logging alone. The findings page drops the now-unneeded manual migration step. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…as_row backend/src/ is outside the hatch package and the ruff src root, so foresight.ingest was not importable and foresight/database/models/external.py could not load. Same move Heather made locally, so the two land identically.
env_file also reads ../.env so commands run from backend/ pick up the repo-root file.
Text sources write dates as people do. Resolves day ranges (including ones crossing a month or a new year), month-year and season/quarter spans, reporting the precision it achieved rather than rounding a vague date into a confident one. Relative forms are refused outright.
Articles in, provisional events out, with every drop recorded on IngestReport as a named reason. collect_events is the single call a scheduled worker makes; hours_back is the scheduling knob.
Replays recorded batches by default so the pipeline is demonstrable without credentials; --live goes through the same collect_events call.
Brings in the Windows justfile fix. pyproject.toml conflicted only because both sides appended a runtime dependency; keeps both httpx and scikit-learn. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An unscoped run rejected 188 of 198 articles on city alone -- paid-for noise. One search per city, filtered to that city's country, cuts that to 60 and raises the keep rate from 2% to 38% on three calls instead of four.
Keep rate is not quality. The kept set still contains historical entities ("Irish Civil War", "midterms"), wrong-city assignments (Reading Festival -> Dublin), and one Tribeca Lisboa run split across four rows. entities.Event is not a scheduled-event list, and confidence scores extraction certainty rather than event-ness, so it cannot yet be used as a quality filter. Precision work still outstanding.
…ternal-factors-events # Conflicts: # .env.example # backend/foresight/config.py # backend/pyproject.toml
…ternal-factors-events
Replaces the in-memory EventStore stand-in, which existed only because repository.py was not yet on this branch. One transaction per run; new/updated/unchanged now come from repository.Outcome rather than a local dict.
CityTarget.slug was never called. horizon_days and the plan_searches queries override were knobs nothing turned outside tests. No behaviour change.
Fetch and print without writing, so PredictHQ output can be inspected the way scripts/run_asknews.py --dry-run already allows. With --lat/--lon/--city it touches no database and skips the automatic migration; --hotel-id still needs one to resolve the market.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds event ingestion from two external sources — PredictHQ (structured listings) and AskNews (news articles) — behind one shared
NormalizedEventcontract that future event sources should also target. Events merge onto canonical rows and feed the pricing model as per-night expected attendance.Type of change
feat— new featuredocs— documentation onlytest— adding or updating testschore/build/ci— tooling, deps, or pipeline changesChanges
ingest/normalize.py—NormalizedEvent,dedupe_key(normalized title + city + local start date), and the shared enums. Validates on construction.ingest/predicthq.py— structured listings near a point, paged byupdated.gte. Requestsdeletedstate so upstream cancellations propagate.ingest/asknews.py+ingest/textdates.py— infers events from articles: per-city scoped search, prose date resolution, and a named reject reason for everything dropped.ingest/repository.py—upsert_event, source-agnostic. Resolves by(source, source_ref)beforededupe_keyso a revised start date updates rather than orphans. PredictHQ outranks AskNews on contested fields.ingest/run.py— ingest CLI, with--dry-runto fetch and print without writing.services/features/events.py— stored events to the per-nightexpected_attendancefeature.docs/event-extraction.md— architecture, file inventory, and how to test it.httpx, promotesscikit-learnto a runtime dependency.Database migrations
just migrate-create "...")just migrate) and rolls back (just migrate-down)Two migrations: canonical
events+event_observations, andstatuson observations. Verified down and back up with data in the table.Checklist
just lint) and code is formatted (just format).env.examplejust dev) and verified the API at/docs— N/A, no API surface in this PR; ingestion runs via CLI and scriptsHow to test
Full instructions in
docs/event-extraction.md→ "How to test it".Quickest check — no credentials, no database, no API credits:
Then the suite:
just test(132 passing;test_repository.pyandtest_event_features.pyneed Postgres viajust db-up).Only
--livespends AskNews credits — use--runs 1if you try it.Notes for reviewers
Please sanity-check
NormalizedEvent. It is the contract every future event source has to fit, and the expensive thing to change later. Specifically:dedupe_keyexcludes category on purpose (sources disagree, and that must not split one event in two), and coarse dates are kept atMONTH/QUARTERprecision rather than rounded away.Known gaps, not blockers:
entities.Eventis not a scheduled-event list.confidencemeasures extraction certainty, not event-ness — a live run scored "midterms" at 0.95. Not safe as a quality filter. AskNews events carry no attendance so they contribute nothing to pricing today, butnightly_attendancemultiplies by confidence, so this must be fixed before they ever do.python -m foresight.ingest.run(PredictHQ) andscripts/run_asknews.py(AskNews). Both now support--dry-run; folding AskNews in as--source asknewswould collapse them.Unverified:
--dry-runon the PredictHQ path has no live token behind it on my side — it is covered by a unit test and confirmed to touch no database, but nobody has run it against real PredictHQ data.Heads-up for anyone running this locally: if you already have Postgres on your machine it binds
127.0.0.1:5432and wins over Docker, sojust db-upsucceeds while the app talks to a different database.