fix(scraper): visible scrape diagnostics + dispatchable GH scraper path - #62
Merged
Merged
Conversation
Contributor
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Restores a visible scrape path (the Railway scheduler runs silently and last_seen_at is frozen since 12:01Z). Runs the same scraper against the DATABASE_URL secret so a broken/stale connection string — the suspect — shows up in Actions logs instead of an in-process 500.
…spatchable GH scrape
- /update now returns the exception detail (was a bare 500), and /health
gains last_scrape {ran_at, result, error} so a silently failing scheduled
scrape is diagnosable by HTTP instead of Railway logs.
- Reintroduce scrape.yml (hourly + workflow_dispatch) against DATABASE_URL,
restoring a visible scrape path as fallback.
KSaifStack
force-pushed
the
fix/scraper-survey
branch
from
September 29, 2026 04:20
bdcd5a5 to
a7d9163
Compare
The scrape aborted with psycopg2.errors.UniqueViolation: duplicate key value violates unique constraint "internships_fingerprint_key" on the first run with a working DATABASE_URL. Prod carries two unique constraints and the upsert named only one. _ensure_schema creates internships_unique_job (company, role, location, link); internships_fingerprint_key on `fingerprint` was added by hand in the SQL editor and exists nowhere in the repo. So ON CONFLICT (company, role, location, link) routed around its own arbiter for any listing that shared a fingerprint with a different link, and the UniqueViolation killed the run. - Drop the conflict target. Both constraints now route to the same DO UPDATE, and since the SET list omits company/role/location/link the pre-existing row keeps its link. - Key `seen` on job_fingerprint() instead of the raw lowercase triple. The two disagreed: "Acme, Inc." and "Acme Inc." have distinct lowercase keys but one fingerprint, so both could land in a single execute_values page and raise "ON CONFLICT DO UPDATE command cannot affect row a second time". Making the in-memory dedup use the same normaliser as the constraint also tightens cross-source dedup. test_fingerprint_dedup.py covers both directions: lookalikes collapse, distinct jobs survive. It fails against the old dedup key.
With the upsert fixed, the run got further and failed in _backfill_fingerprints() with the same UniqueViolation. That UPDATE filled every NULL-fingerprint row in one batch, but prod already contains rows whose content normalises to the same fingerprint, so a single batch can assign the same value twice and no arbiter is involved. Assign greedily against a set seeded from the already-populated fingerprints, and skip rows whose value is taken. Postgres treats NULLs as distinct, so a skipped row neither violates the index nor blocks later runs. The corrected `seen` dedup stops new collisions from appearing. Test covers it: two mutation checks confirm the assertions fail against the old blind dedup key and a blind backfill.
The previous commit wrote "ON CONFLICT DO UPDATE" while the next line
already read "DO UPDATE SET", producing
ON CONFLICT DO UPDATE DO UPDATE SET
which is a psycopg2 SyntaxError. py_compile passed because the statement is
a bare string literal; it only failed against prod.
Reduced to "ON CONFLICT / DO UPDATE SET". Added assertions over _UPSERT_SQL
so this class of typo fails locally instead: exactly one DO UPDATE, a conflict
target never reintroduced, and the SET list never overwrites company, role,
location or link -- those are the columns a conflict must preserve.
…constraint Dropping the conflict target does not work: Postgres requires an inference specification for DO UPDATE and rejects a bare "ON CONFLICT" with ON CONFLICT DO UPDATE requires inference specification or constraint name Only DO NOTHING may omit the target. Arbitrate on (fingerprint) instead. fingerprint is norm(company)|norm(role)|norm(location), so an identical (company, role, location) necessarily yields an identical fingerprint -- every four-column conflict is also a fingerprint conflict, and one target covers both. This is the same fix the previous two commits were reaching for, stated in a form Postgres accepts. _ensure_schema now creates internships_fingerprint_key when absent, after the backfill. It previously existed only because someone added it by hand, so a fresh or restored database failed every scrape with "no unique or exclusion constraint matching the ON CONFLICT specification". A UniqueViolation there is caught and reported instead of killing the run: `seen` still dedupes by content in Python. Test asserts the target is (fingerprint), that one DO UPDATE clause is present, and that the four-column conflict is a fingerprint conflict, so both the bare form and a return to the four-column target fail locally.
…D CONSTRAINT _prod's internships_fingerprint_key is a unique index, not a table constraint, so the pg_constraint guard found nothing to add and the subsequent ADD CONSTRAINT died with psycopg2.errors.DuplicateTable: relation "internships_fingerprint_key" already exists -- the constraint collided with the index of the same name. ON CONFLICT (fingerprint) infers from a unique index just as well as from a constraint, so CREATE UNIQUE INDEX IF NOT EXISTS is both sufficient and idempotent against either form, since a constraint is backed by an index of the same name.
Sizes the blast radius before deciding how to reconcile them. Rows left NULL are invisible to ON CONFLICT (fingerprint), so a later insert matching one on the four-column constraint is unarbitrated and aborts the scrape.
…em NULL The 212 rows left NULL by the previous commit are why the scrape then failed on internships_unique_job instead. A NULL is invisible to ON CONFLICT (fingerprint), so an insert matching one of those rows on (company, role, location, link) reaches no arbiter and aborts the run. One error traded for another. Deleting them is the same reconciliation _ensure_schema already does when internships_unique_job's definition changes: keep the lowest id, drop the later copy. The survivor is the row the frontend tracker already resolves to, since a missing fingerprint is a 404 on every tracked job. Nothing references internships by foreign key; job references live in agent_proposals.payload. The SELECT is now ORDER BY id so "keep the oldest" is deterministic instead of depending on heap order. This deletes 212 rows on the next run. They are the same job listed twice by the app's own definition of a job -- identical normalised company, role and location -- which the unique index exists to forbid.
…each run Guessing at the arbiter from a UniqueViolation three layers up wasted four runs. The fingerprint index is the load-bearing piece of the write path; print it and the remaining NULL count once per run.
…bypassed A constraint firing means ON CONFLICT (fingerprint) matched nothing, which for an identical 4-tuple can only be true if the stored row's fingerprint differs from the computed one. Print it instead of spending a run guessing.
…e normaliser
Non-NULL is not the same as correct. The upsert arbitrates on (fingerprint)
and nothing else, so a row whose stored fingerprint is not what the current
normaliser computes for its own columns is invisible to that arbiter while
still holding (company, role, location, link) -- an insert matching it there
aborts the run with UniqueViolation on internships_unique_job.
Also replaces the write-failure diagnostic's DETAIL regex with a verbatim
print: the pattern was anchored on "link=(", which never appears, so it
matched nothing on the one run it existed for and the diagnostic was silent.
Non-NULL is not the same as correct, and a wrong fingerprint is a hard failure rather than a cosmetic one. The upsert arbitrates on (fingerprint) and nothing else, so a row whose stored value is not what job_fingerprint computes for its own columns is invisible to that arbiter while still holding (company, role, location, link) -- an insert matching it there aborts the whole run on internships_unique_job. That is what killed every scrape after the 212 NULL rows were reconciled: 52 more rows hold a normaliser's value that exists nowhere in this repo, which kept punctuation as a space and collapsed NYC to ny. The pass is now seeded with nothing and walks every row in id order, so "already taken" means an earlier row rather than a pre-existing value. A correct row is compared and skipped, so the extra cost is one comparison.
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Root-cause path for the frozen scraper (last_seen_at stuck 12:01Z):\n\n- GH Actions DATABASE_URL secret = 127.0.0.1:5432 (garbage, proven by run 36520376480) - the old workflow could never have connected\n- prod /update 500s deterministically ~3.2s while /count (warm pool) works - consistent with a stale/rotated DATABASE_URL on the Railway service (fixed on the dashboard)\n\nChanges:\n- /update returns exception detail instead of bare 500\n- /health gains last_scrape {ran_at, result, error}\n- scrape.yml: hourly + workflow_dispatch GH scrape (needs a real DATABASE_URL secret to function)\n\nDeploy boots, then
/healthshows last_scrape.error and we see the exact failure without guessing.\n\nNEED YOU (dashboard): (1) searchtern-production VariableDATABASE_URLvs the actual Postgres string; (2) repo Actions secretDATABASE_URL= the public Railway PG string so the GH scraper works.