Fix enrichment re-arm and migration backfill gaps
Three crash- and migration-recovery gaps in the unified enrichment model left rows stuck or misclassified: - Upsert re-armed a blank, newly-linked tracker pattern without clearing its prior enrichment payload. A crash between the worker's claim and persist then left the row with a stale payload, so the stale-recovery sweep (which only catches rows with a null payload) skipped it forever. Clear enrichment on re-arm so the row reads as not-yet-completed again, and pin the behavior with a test. - The migration added last_enrichment_attempt_at to common_third_parties without seeding it. Rows with prior attempts kept a NULL clock and could never satisfy the stale-reset predicate. Backfill from updated_at, the historical claim-time proxy. - The migration switched the tracker-pattern enriched-state source to the enrichment payload without backfilling rows previously marked by enriched_at, making already-enriched rows read as unenriched. Seed a provenance sentinel for rows that carried the old done-flag. Signed-off-by: Émile Ré <emile@probo.com>
This commit is contained in:
@@ -30,8 +30,30 @@ ALTER TABLE common_tracker_patterns
|
||||
ALTER TABLE common_tracker_patterns
|
||||
RENAME COLUMN enriched_at TO last_enrichment_attempt_at;
|
||||
|
||||
-- Backfill the enriched-state marker for rows that were terminally enriched
|
||||
-- under the old model. "Enriched" previously meant enriched_at IS NOT NULL (a
|
||||
-- terminal done-flag); it now means a non-null enrichment payload. The rename
|
||||
-- above moved the old done-flag into last_enrichment_attempt_at, so seed a
|
||||
-- provenance payload for every row that carried it. Without this, rows already
|
||||
-- enriched before the deploy would read as permanently unenriched. The payload
|
||||
-- carries no per-field provenance, so completeness checks treat it as fully
|
||||
-- enriched.
|
||||
UPDATE common_tracker_patterns
|
||||
SET enrichment = '{"status": "migrated"}'::jsonb
|
||||
WHERE last_enrichment_attempt_at IS NOT NULL;
|
||||
|
||||
-- common_third_parties already carries enrichment/enrichment_attempts; give
|
||||
-- it the same explicit last-attempt timestamp (previously only recorded in
|
||||
-- the enrichment JSON) so the stale-recovery clock has a dedicated column.
|
||||
ALTER TABLE common_third_parties
|
||||
ADD COLUMN last_enrichment_attempt_at TIMESTAMP WITH TIME ZONE;
|
||||
|
||||
-- Seed the new clock for rows that already spent an attempt. The claim path
|
||||
-- stamps last_enrichment_attempt_at and updated_at together, so updated_at is
|
||||
-- the historical proxy for the last attempt. Without this, a row claimed but
|
||||
-- never completed before the migration keeps a NULL clock, and the stale-reset
|
||||
-- predicate (last_enrichment_attempt_at < stale_before) can never be true for
|
||||
-- NULL, so it would never be re-queued. Never-attempted rows keep a NULL clock.
|
||||
UPDATE common_third_parties
|
||||
SET last_enrichment_attempt_at = updated_at
|
||||
WHERE enrichment_attempts > 0;
|
||||
|
||||
Reference in New Issue
Block a user