Skip to content

Fold yoheinakajima/glance (local VLM logit harness) - #73

Merged
24601 merged 2 commits into
mainfrom
cursor/fold-glance-vlm-logit-140d
Sep 21, 2026
Merged

24601 merged 2 commits into
mainfrom
cursor/fold-glance-vlm-logit-140d

Conversation

@24601

@24601 24601 commented Sep 21, 2026

Copy link
Copy Markdown
Owner

Summary

Catalogs yoheinakajima/glance (PyPI glance-vlm 0.3.1, https://glance.yohei.me) as a vision soft-score / logit-read peer in notes.md §147.

Glance is a harness around a frozen open VLM (default Qwen3-VL-4B). It reads answer-token logits in one forward pass and does not train weights. Typed noul / choice / score follow TypeSafe Jev (hosted text). Soft probabilities are for threshold, abstain, or rank. They are not a safety gate. Author metrics are marked theirs.

Shares the inference object with simple-jev, jev-visual, and LitJev. Distinct from the trained SmolVLM head already carded as hf:thaitea/laya-vision (§146) and from Open-Jev heads that train weights.

No Augustus call site. README still ends at License. Left as a draft for adversarial review.

Open in Web Open in Cursor 

Catalog the Apache-2.0 glance-vlm readout in notes §147: soft noul/choice/score for threshold, abstain, or rank, not a safety gate, with author metrics marked as theirs.

Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
@24601

24601 commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

ADV #73 glance (draft, not merged)

Reviewed cursor/fold-glance-vlm-logit-140d @ 5d6cb67fb13b0a84eb7f7aa9fc0670ef4bf68d75 against base main @ 86762804 (#72). Live recapture of yoheinakajima/glance at HEAD 8f36e063bffb76e71d1f4cea6792872fec25cfd6 (pushed 2026-09-21T17:37:40Z, tag v0.3.1 on that commit). README blob b57280394bc1ed537bc981e1c98a525d7c609c3e. LICENSE blob d645695673349e3947e8e5ae42332d0ac3164cd7. Description hash 980e4bcfb7b7. PyPI glance-vlm 0.3.1. Ledger: docs/CLAIMS.md. Positioning: docs/paper/RELATED_WORK.md.

ADV_FAIL

Honesty gates that hold

  1. Soft scores stay soft. The card says probabilities are for threshold, abstain, or rank, and are not a safety gate. No Augustus runtime call site. glance is not imported into the skill.
  2. README still ends at a clean ## License. No uniqueness-lock or HIGH digest after License. The glance lock lives in research/ overlays and uniqueness_gate.py only.
  3. Inserted prose adds no new em dashes. Dashes on the activation-triggers line were already on that line.
  4. Class framing holds. notes.md §147 names the decision-model class and TypeSafe Jev Choice/Score/Noul as the dominant exemplar. Glance is a vision peer beside simple-jev / jev-visual / LitJev. The class is not thinned to Jev-only.
  5. Harness vs trained Laya holds. Glance trains no weights. Catalogued Laya Vision stays hf:thaitea/laya-vision (§146), which matches the RELATED_WORK credit line laya-vision-smolvlm-256m. YOFO stays a paper name. No YOFO repo card.
  6. Fingerprints match the live commit: HEAD, README SHA, LICENSE SHA, size 51490, Apache-2.0, topics, created/pushed clocks, v0.3.1, PyPI 0.3.1, description hash. sources.json parses. One github row, one site row, one pypi row. The lock string appears once per overlay file (gate pattern, not a second synthesis card).
  7. Mechanism quotes that match the README lede: frozen Qwen3-VL-4B, answer-token logits, nothing generated, nothing trained, images stay local, noul / choice / score, POST /v1/decide on 127.0.0.1. glance/server.py does os.environ.setdefault("HF_HUB_OFFLINE", "1"). Repo ids for featherless-ai/simple-jev, hr98w/jev-visual, and zhengxuyu/litjev match RELATED_WORK.
  8. Author numbers that match the README zero-shot table: yes/no 0.939 (541), pick-one 0.933 (270) vs Gemini 3.1 Flash-Lite 0.933 and Pro 0.937, rating exact 0.669 vs Flash-Lite 0.763. Geometric probes match the README (diagonals 0.30, largest-of-four 0.52, eight objects 0.56, hidden-state stripe probe 0.99). Non-image-quality 0.55 at 300 labels and tilt 0.33 match. Explicit other at 7.7% matches CLAIMS ("knows when none of the options fits" is NOT supported as shipped). "Faster or cheaper than Jev" is Do not claim. WITHDRAWN "0.33 s" and "about five times cheaper" is recorded. Raw yes/no ECE 0.111 and 0.179 matches CLAIMS row "Probabilities out of the box are calibrated / NOT supported."

Star-noise, non-blocking: live REST is 4 stars and 1 fork. The card says 3 stars and 0 forks. pushed_at and HEAD are unchanged, so this is not a material revisit.

Spot-recapture failures (fixes)

  1. OpenJev v2 is not Zefan-Cai/Open-Jev. RELATED_WORK, "Where Glance sits": Glance trains no weights, unlike YOFO, Laya Vision, OpenJev v2, or OpenJev-Vision. OpenJev v2 is AlexWortega on the Hub (huggingface.co/AlexWortega/openjev): Qwen3.5 NLI cross-encoder, v2 is a trained 4B multimodal claim scorer. OpenJev-Vision is github.com/IamBusy/OpenJev-Vision (trained). razorback16's OpenJev is a different project, and RELATED_WORK says so. The card's parallel ("Catalogued Laya Vision is hf:thaitea/laya-vision. Catalogued Open-Jev with a trained head is Zefan-Cai/Open-Jev") assigns the author's OpenJev v2 to the wrong repo. Zefan-Cai/Open-Jev (§125) stays a separate trained-head namesake. AlexWortega/openjev is already a census name (§77). Do not mint a sibling card. Rewrite the namesake line so OpenJev v2, Zefan-Cai/Open-Jev, razorback16/openjev, and IamBusy/OpenJev-Vision stay four strings. Keep POST /v1/decide ≠ TypeSafe /v1/systemone ≠ IamBusy/OpenJev /v1/decide as a wire lock, not as the identity of OpenJev v2.

  2. Fitted rating numbers are under a zero-shot heading. The author-table label says "README, zero-shot, same items" and then lists unlabeled 0.67 to 0.76 and labeled 0.86 / ECE about 0.03. CLAIMS section F: do not say "Zero-shot" or "no training" for any number that used a fitted readout. The VLM is frozen. The readout is fit. Split the list. Zero-shot table only: 0.939 / 0.933 / 0.669. Setup table, marked fitted: unlabeled image-quality 0.67 to 0.76 (README; CLAIMS one-pass figure is 0.758 at 16 unlabeled images), labeled about 32 images 0.86 exact and ECE about 0.03, per rubric, does not transfer.

  3. "Overconfident until a fit" merges two different fits. CLAIMS: raw yes/no ECE 0.111 and 0.179. The measured remedy is a pooled two-number Platt map on those same two suites (ECE 0.061 and 0.053). Transfer to a new yes/no task is untested. glance fit is a per-rubric rating map, not that Platt map. judgment-class.md currently says "Raw yes/no stays overconfident until a fit." Name the Platt map, and say transfer is untested. Keep glance fit on the rating bullets.

  4. "One forward pass" is the default readout, not every rating path. README and glance/cli.py: unfitted ratings use jsondigits (1 pass, 0.669). Labeled glance fit uses ens4d (4 passes). Four-pass zero-shot rating is 0.570, behind the same model writing JSON (0.672). glance fit --unlabeled uses jsondigits. fast2 (2 passes) and digits (1 pass) remain available. Keep the one-pass quote for yes/no, pick-one, and the default unfitted rating. State the ens4d labeled path next to it.

  5. Rating limit the README states and the card skips. "When to use it, and when not to": Glance is not an image-quality metric. On KADID-10k it misses their own targets. Do not invent a KADID accuracy (CLAIMS: do not quote interim KADID numbers as results). A 0.9B model trained for image quality (Q-SiT-mini) is level under the same 32-label fit. Hand-built features beat it on low-level artifacts when labels are plentiful. The 0.76 / 0.86 image-quality figures need that caveat in the same paragraph.

Concurrent hourly 1203 (conflict note, not part of this verdict)

Draft PR #74 (cursor/hourly-1203-fold-9f39 @ 2d264909) is open off the same main @ 86762804 (#72). Its body claims notes.md §147 for hourly 1203: 73 novel cards plus 6 revisit densifies (79 items), composition 761-776, findings batch #129. The patch heading is ## 147. Hourly 1203 HIGH. This PR's heading is ## 147. User-provided glance HIGH. Both drafts edit the same overlay files (notes.md, uniqueness_gate.py, SKILL.md, judgment-class.md, ecosystem.md, changelog-hourly.md, refresh-log.md, revisit_fingerprints.json).

Merge order if this glance card reaches ADV_PASS: merge #73 first, so glance keeps §147. Then rebase #74 onto that main and move hourly 1203 to the next section id, with its uniqueness lock, gate, and fingerprints. Do not fold the 79 items into this branch, and do not ADV #74 here.

This review is still ADV_FAIL, so #73 is not ready to merge. Do not merge #73. Do not merge #74.

ADV_FAIL

Concrete fixes, in the card and in the one lock string: (1) split OpenJev v2 (AlexWortega) from Zefan-Cai/Open-Jev, razorback16/openjev, and IamBusy/OpenJev-Vision; (2) pull fitted 0.67 to 0.76 and 0.86 out of the zero-shot heading; (3) separate the yes/no Platt map from glance fit; (4) state jsondigits vs ens4d; (5) add the KADID / Q-SiT-mini / classical-feature rating limits without new KADID digits. Leave both drafts open.

Sources:

OpenJev v2 stays Hub AlexWortega/openjev, distinct from Zefan-Cai/Open-Jev, razorback16/openjev, and IamBusy/OpenJev-Vision. Zero-shot accuracy is only 0.939/0.933/0.669. Fitted image-quality numbers, the yes/no Platt map, jsondigits versus ens4d, and the KADID/Q-SiT-mini caveat are named separately. Glance remains notes.md §147.

Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
@cursor

cursor Bot commented Sep 21, 2026

Copy link
Copy Markdown

ADV corrections on 3090828. Draft left open. Not merged.

  1. OpenJev v2 is Hub AlexWortega/openjev (census §77, no sibling card). It stays distinct from Zefan-Cai/Open-Jev (§125), razorback16/openjev, and IamBusy/OpenJev-Vision. POST /v1/decide ≠ TypeSafe /v1/systemone ≠ IamBusy/OpenJev /v1/decide is a wire lock only.
  2. Zero-shot table is only yes/no 0.939, pick-one 0.933, rating exact 0.669. Fitted readout is separate: unlabeled image-quality 0.67 to 0.76 (CLAIMS one-pass 0.758 at 16 unlabeled) and labeled ~32 images 0.86 exact / ECE ~0.03. Those fitted numbers are not called zero-shot.
  3. Raw yes/no ECE 0.111 and 0.179. Remedy is the pooled two-number Platt map (ECE 0.061 and 0.053). Transfer to a new yes/no task is untested. glance fit stays the per-rubric rating map.
  4. Default unfitted rating is jsondigits (1 pass, 0.669). Labeled glance fit uses ens4d (4 passes). Four-pass zero-shot rating is 0.570, behind write JSON 0.672. glance fit --unlabeled uses jsondigits. fast2 and digits remain available.
  5. The 0.76 / 0.86 image-quality paragraph now says Glance is not an image-quality metric, KADID-10k misses their own targets (no interim KADID digits), Q-SiT-mini (0.9B) is level under the same 32-label fit, and hand-built features win on low-level artifacts when labels are plentiful.

Uniqueness lock updated to match. Live REST 4★ / 1 fork recorded as star-noise. Glance remains notes.md §147.

@24601

24601 commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

Re-ADV #73 glance (draft, not merged)

Reviewed cursor/fold-glance-vlm-logit-140d @ 3090828b1bb49b0ad50047f601eb79e764f62408 (PR head; parent 5d6cb67) against base main @ 86762804. Live recapture of yoheinakajima/glance is still HEAD 8f36e063bffb76e71d1f4cea6792872fec25cfd6, pushed 2026-09-21T17:37:40Z. Annotated tag v0.3.1 points at that commit. README blob b57280394bc1ed537bc981e1c98a525d7c609c3e. LICENSE blob d645695673349e3947e8e5ae42332d0ac3164cd7. Description hash 980e4bcfb7b7. Apache-2.0. Topics unchanged. PyPI glance-vlm 0.3.1. Repo size 51490. The fix summary is comment 5765699332.

ADV_PASS

Prior FAIL items, recaptured

  1. OpenJev v2 is Hub AlexWortega/openjev. RELATED_WORK lists it as a Qwen3.5 NLI cross-encoder whose v2 (4B) is trained and multimodal. CLAIMS calls that system a trained claim scorer. The card keeps it as a census name in notes.md §77 and does not mint a sibling section. It stays distinct from Zefan-Cai/Open-Jev (§125), razorback16/openjev, and IamBusy/OpenJev-Vision. POST /v1/decide ≠ TypeSafe /v1/systemone ≠ IamBusy/OpenJev /v1/decide is a wire lock, not OpenJev v2 identity. The same four-way split is in judgment-class.md and in the one lock string.

  2. Zero-shot table is only 0.939 / 0.933 / 0.669. Those are the README zero-shot rows (yes/no 541, pick-one 270, default unfitted rating). Fitted numbers sit in a separate paragraph that says not to call them zero-shot. Unlabeled image-quality 0.67 to 0.76 matches the README setup table. CLAIMS one-pass 0.758 at 16 unlabeled images is quoted there as fitted (CLAIMS: zero labels, not zero-shot). Labeled fit, about 32 images: 0.86 exact, ECE about 0.03, per rubric, does not transfer. That is the README setup row. CLAIMS also prints 0.856 and 0.857; the README table and the CLAIMS "do not say" line both round that figure to 0.86.

  3. Yes/no Platt is not glance fit. CLAIMS row "Probabilities out of the box are calibrated / NOT supported": raw ECE 0.111 and 0.179, pooled two-number Platt map on those suites to 0.061 and 0.053, transfer to a new yes/no task untested. The card, the how-to-apply lens, and judgment-class.md keep glance fit as the per-rubric rating map.

  4. Pass counts match the README and CLAIMS. Default unfitted rating is jsondigits (1 pass, exact 0.669). Labeled glance fit uses ens4d (4 passes). Four-pass zero-shot rating is 0.570, behind the same model writing JSON at 0.672. glance fit --unlabeled uses jsondigits. fast2 (2 passes) and digits (1 pass) remain available.

  5. The 0.76 / 0.86 paragraph carries the README limit. Glance is not an image-quality metric. On KADID-10k it misses their own targets. Interim KADID digits are not quoted (CLAIMS: do not quote interim numbers; the jsondigits row also carries KADID 0.347 / 0.328, and the card leaves those out). Q-SiT-mini (0.9B) is level under the same 32-label fit. Hand-built features beat it on low-level artifacts when labels are plentiful. Same paragraph as the fitted image-quality figures.

Other checks

  • Soft scores stay soft: threshold, abstain, or rank, not a safety gate. soft scores ≠ hard gates is in the card and the lock. No Augustus call site. glance is not imported.
  • README still ends at ## License. No uniqueness lock after License.
  • §147 prose and the judgment-class.md glance paragraph add no em dashes. Dashes on the long activation-triggers line were already on that line.
  • python3 .agents/skills/augustus/scripts/uniqueness_gate.py exits 0. The glance lock is one identical substring in the overlays and in UNIQ_GLANCE.
  • Fingerprints still match this live commit (HEAD, README SHA, LICENSE SHA, description hash, v0.3.1, size 51490, created and pushed clocks). Live REST is now 5 stars and 1 fork. The card records 4 stars and 1 fork from the correction pass. pushed_at and HEAD are unchanged, so that star move is star-noise, not a material revisit, and not a remaining fix.

Concurrent #74

Draft PR #74 (cursor/hourly-1203-fold-9f39 @ 2d264909) is still open off the same main @ 86762804 and still claims notes.md §147 for hourly 1203. Merge #73 first so glance keeps §147. Then rebase #74 onto that main and move hourly 1203 to §148, including its uniqueness lock, gate, and fingerprints. Do not merge #73 in this review. Do not merge #74.

Draft left open.

ADV_PASS

Sources:

@24601
24601 marked this pull request as ready for review September 21, 2026 18:52
@24601
24601 merged commit 8a096a8 into main Sep 21, 2026
1 check passed
cursor Bot pushed a commit that referenced this pull request Sep 21, 2026
Rebase onto post-#73 main (8a096a8). Glance keeps notes.md §147.
Hourly 1203 moves to §148. Composition 761-776 and findings batch #129 stay.
24601 added a commit that referenced this pull request Sep 21, 2026
… + REVISIT densify (#74)

Rebase onto post-#73 main (8a096a8). Glance keeps notes.md §147.
Hourly 1203 moves to §148. Composition 761-776 and findings batch #129 stay.

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants