Skip to content

fix(scraper): visible scrape diagnostics + dispatchable GH scraper path - #62

Merged
KSaifStack merged 13 commits into
mainfrom
fix/scraper-survey
Sep 29, 2026
Merged

KSaifStack merged 13 commits into
mainfrom
fix/scraper-survey

Conversation

@KSaifStack

Copy link
Copy Markdown
Owner

Root-cause path for the frozen scraper (last_seen_at stuck 12:01Z):\n\n- GH Actions DATABASE_URL secret = 127.0.0.1:5432 (garbage, proven by run 36520376480) - the old workflow could never have connected\n- prod /update 500s deterministically ~3.2s while /count (warm pool) works - consistent with a stale/rotated DATABASE_URL on the Railway service (fixed on the dashboard)\n\nChanges:\n- /update returns exception detail instead of bare 500\n- /health gains last_scrape {ran_at, result, error}\n- scrape.yml: hourly + workflow_dispatch GH scrape (needs a real DATABASE_URL secret to function)\n\nDeploy boots, then /health shows last_scrape.error and we see the exact failure without guessing.\n\nNEED YOU (dashboard): (1) searchtern-production Variable DATABASE_URL vs the actual Postgres string; (2) repo Actions secret DATABASE_URL = the public Railway PG string so the GH scraper works.

@vercel

vercel Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
search-tern Ready Ready Preview Sep 29, 2026 11:55am UTC

Restores a visible scrape path (the Railway scheduler runs silently and
last_seen_at is frozen since 12:01Z). Runs the same scraper against the
DATABASE_URL secret so a broken/stale connection string — the suspect —
shows up in Actions logs instead of an in-process 500.
…spatchable GH scrape

- /update now returns the exception detail (was a bare 500), and /health
  gains last_scrape {ran_at, result, error} so a silently failing scheduled
  scrape is diagnosable by HTTP instead of Railway logs.
- Reintroduce scrape.yml (hourly + workflow_dispatch) against DATABASE_URL,
  restoring a visible scrape path as fallback.
The scrape aborted with
  psycopg2.errors.UniqueViolation: duplicate key value violates unique
  constraint "internships_fingerprint_key"
on the first run with a working DATABASE_URL.

Prod carries two unique constraints and the upsert named only one.
_ensure_schema creates internships_unique_job (company, role, location,
link); internships_fingerprint_key on `fingerprint` was added by hand in
the SQL editor and exists nowhere in the repo. So ON CONFLICT
(company, role, location, link) routed around its own arbiter for any
listing that shared a fingerprint with a different link, and the
UniqueViolation killed the run.

- Drop the conflict target. Both constraints now route to the same
  DO UPDATE, and since the SET list omits company/role/location/link the
  pre-existing row keeps its link.
- Key `seen` on job_fingerprint() instead of the raw lowercase triple.
  The two disagreed: "Acme, Inc." and "Acme Inc." have distinct lowercase
  keys but one fingerprint, so both could land in a single
  execute_values page and raise "ON CONFLICT DO UPDATE command cannot
  affect row a second time". Making the in-memory dedup use the same
  normaliser as the constraint also tightens cross-source dedup.

test_fingerprint_dedup.py covers both directions: lookalikes collapse,
distinct jobs survive. It fails against the old dedup key.
With the upsert fixed, the run got further and failed in
_backfill_fingerprints() with the same UniqueViolation. That UPDATE filled
every NULL-fingerprint row in one batch, but prod already contains rows whose
content normalises to the same fingerprint, so a single batch can assign the
same value twice and no arbiter is involved.

Assign greedily against a set seeded from the already-populated fingerprints,
and skip rows whose value is taken. Postgres treats NULLs as distinct, so a
skipped row neither violates the index nor blocks later runs. The corrected
`seen` dedup stops new collisions from appearing.

Test covers it: two mutation checks confirm the assertions fail against the
old blind dedup key and a blind backfill.
The previous commit wrote "ON CONFLICT DO UPDATE" while the next line
already read "DO UPDATE SET", producing
    ON CONFLICT DO UPDATE DO UPDATE SET
which is a psycopg2 SyntaxError. py_compile passed because the statement is
a bare string literal; it only failed against prod.

Reduced to "ON CONFLICT / DO UPDATE SET". Added assertions over _UPSERT_SQL
so this class of typo fails locally instead: exactly one DO UPDATE, a conflict
target never reintroduced, and the SET list never overwrites company, role,
location or link -- those are the columns a conflict must preserve.
…constraint

Dropping the conflict target does not work: Postgres requires an inference
specification for DO UPDATE and rejects a bare "ON CONFLICT" with
  ON CONFLICT DO UPDATE requires inference specification or constraint name
Only DO NOTHING may omit the target.

Arbitrate on (fingerprint) instead. fingerprint is
norm(company)|norm(role)|norm(location), so an identical
(company, role, location) necessarily yields an identical fingerprint --
every four-column conflict is also a fingerprint conflict, and one target
covers both. This is the same fix the previous two commits were reaching for,
stated in a form Postgres accepts.

_ensure_schema now creates internships_fingerprint_key when absent, after the
backfill. It previously existed only because someone added it by hand, so a
fresh or restored database failed every scrape with "no unique or exclusion
constraint matching the ON CONFLICT specification". A UniqueViolation there is
caught and reported instead of killing the run: `seen` still dedupes by
content in Python.

Test asserts the target is (fingerprint), that one DO UPDATE clause is
present, and that the four-column conflict is a fingerprint conflict, so both
the bare form and a return to the four-column target fail locally.
…D CONSTRAINT

_prod's internships_fingerprint_key is a unique index, not a table
constraint, so the pg_constraint guard found nothing to add and the
subsequent ADD CONSTRAINT died with
  psycopg2.errors.DuplicateTable: relation "internships_fingerprint_key"
already exists -- the constraint collided with the index of the same name.

ON CONFLICT (fingerprint) infers from a unique index just as well as from a
constraint, so CREATE UNIQUE INDEX IF NOT EXISTS is both sufficient and
idempotent against either form, since a constraint is backed by an index of
the same name.
Sizes the blast radius before deciding how to reconcile them. Rows left
NULL are invisible to ON CONFLICT (fingerprint), so a later insert matching
one on the four-column constraint is unarbitrated and aborts the scrape.
…em NULL

The 212 rows left NULL by the previous commit are why the scrape then failed
on internships_unique_job instead. A NULL is invisible to ON CONFLICT
(fingerprint), so an insert matching one of those rows on
(company, role, location, link) reaches no arbiter and aborts the run. One
error traded for another.

Deleting them is the same reconciliation _ensure_schema already does when
internships_unique_job's definition changes: keep the lowest id, drop the
later copy. The survivor is the row the frontend tracker already resolves to,
since a missing fingerprint is a 404 on every tracked job. Nothing references
internships by foreign key; job references live in agent_proposals.payload.

The SELECT is now ORDER BY id so "keep the oldest" is deterministic instead of
depending on heap order.

This deletes 212 rows on the next run. They are the same job listed twice by
the app's own definition of a job -- identical normalised company, role and
location -- which the unique index exists to forbid.
…each run

Guessing at the arbiter from a UniqueViolation three layers up wasted four
runs. The fingerprint index is the load-bearing piece of the write path; print
it and the remaining NULL count once per run.
…bypassed

A constraint firing means ON CONFLICT (fingerprint) matched nothing, which
for an identical 4-tuple can only be true if the stored row's fingerprint
differs from the computed one. Print it instead of spending a run guessing.
…e normaliser

Non-NULL is not the same as correct. The upsert arbitrates on (fingerprint)
and nothing else, so a row whose stored fingerprint is not what the current
normaliser computes for its own columns is invisible to that arbiter while
still holding (company, role, location, link) -- an insert matching it there
aborts the run with UniqueViolation on internships_unique_job.

Also replaces the write-failure diagnostic's DETAIL regex with a verbatim
print: the pattern was anchored on "link=(", which never appears, so it
matched nothing on the one run it existed for and the diagnostic was silent.
Non-NULL is not the same as correct, and a wrong fingerprint is a hard
failure rather than a cosmetic one. The upsert arbitrates on (fingerprint)
and nothing else, so a row whose stored value is not what job_fingerprint
computes for its own columns is invisible to that arbiter while still
holding (company, role, location, link) -- an insert matching it there
aborts the whole run on internships_unique_job. That is what killed every
scrape after the 212 NULL rows were reconciled: 52 more rows hold a
normaliser's value that exists nowhere in this repo, which kept punctuation
as a space and collapsed NYC to ny.

The pass is now seeded with nothing and walks every row in id order, so
"already taken" means an earlier row rather than a pre-existing value. A
correct row is compared and skipped, so the extra cost is one comparison.
@KSaifStack
KSaifStack changed the base branch from dev to main September 29, 2026 11:56
@KSaifStack
KSaifStack merged commit 124a748 into main Sep 29, 2026
3 checks passed
@KSaifStack
KSaifStack deleted the fix/scraper-survey branch September 29, 2026 11:56

This branch was successfully deployed

1 active deployment
Preview — d023fba2 Deployed Sep 29, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant