Fold hourly 1110 HIGH: laya-vision same-species serving + novel cards + REVISIT densify - #72
Merged
Merged
Conversation
24601
commented
Sep 21, 2026
24601
left a comment
Owner
Author
There was a problem hiding this comment.
Adversarial review (hourly 1110, draft PR #72)
Do not merge from this comment. Catalog densify holds. No blockers.
Gates
- Occupancy. Current
origin/mainis6216705(merged #71, Release v0.5.1) and still ends atnotes.md§145 / composition items 729-744 / findings batch #127. This fold is §146 / 745-760 / batch #128. No collision with already-merged ranges. - Release pins. Diff vs base
6be01e04does not touch README, CHANGELOG, CITATION, marketplace, ordocs/release-notes-v0.5.1.md. A three-way merge onto6216705stays clean and keepsversion: 0.5.1, the release notes, and the slim skill, then adds the 1110 lock. Fold text "Does not bump 0.5.0" is this fold's non-bump, not a pin rewrite. - README / uniqueness.
README.mdstill ends at## License(MIT / SECURITY / CONTRIBUTING only).python3 .agents/skills/augustus/scripts/uniqueness_gate.pyprintsuniqueness-gate ok.invented_signal: falseinresearch/archive/hourly/2026-09-21T17/run_digest.json. Added human prose has no U+2014. - Class framing. TypeSafe Jev stays the default recommended path. The lane stays the wider decision-model class. laya-vision is same-species vision serving, not an 18th scoring-table row (no scoring-table edit). Act/escalate head is untrained (zero gradient, random init) and is not a gate.
- musubi-jev. README at
e943f21e4057is the kev tree (# Kev, jaredpalmer CI badge). The card says copied kev numbers are not a musubi bench and does not paste a kev table as a musubi measurement.musubi-labs/musubi-jev ≠ jaredpalmer/kev. - Soft scores.
act_above 0.8is still soft. JevFence: soft scores are not hard gates. Soft judgment is not a sole veto. Counts: 69 novel + 7 revisit.
theirs checked live
- thaitea/laya-vision sha
a2653db2831b, CC-BY-NC-SA-4.0, 13 likes. A-OKVQA 63.1% ECE 0.266 to 0.094 (n=1,138). ScienceQA image 89.0% ECE 0.080 to 0.034 (n=2,097). VQAv2 noul 73.2% ECE 0.085 to 0.042 (n=5,000, re-split, not published VQAv2). All n=8,235 75.9%, calibrated ECE 0.035 (raw ECE on that row is 0.108; the fold's 0.035 is the calibrated figure). About 71 ms. Option order varies 1.2 points. Train 94.7% vs val 63.1% on A-OKVQA. Usage loadsthaitea/laya-vision-smolvlm-256m(§87 card id, not this card). - 7Zenox/gemma-jev 144 authored decisions. Gemma 3 270M 0.293 below chance 0.333. Gemma 4 E2B-it JSON chat 0.807. Qwen3.5-4B JSON chat 0.813. bf16 vs fp32 argmax-agreement check not run. JSON chat and letter-slot softmax are not a calibrated Noul.
- AntonG87/codearia-sieve 6 of 6 and 11 of 11 are the judge table, marked not Harbor. Parse is not a decision.
- musubi-labs/musubi-jev kev copy, not a new bench.
- Pdbz199/local-decision-model independent from the public post. The 70 to 500 ms row is attributed to the Jev post, not measured here. Span start/end head is locate, not decide.
- gbesse/jev-rerank-server SciFact n=25 nDCG@10 0.616377 to 0.718260, Recall@10 0.84 unchanged. Ranking is not calibration.
- MarkChu-git/typesafe-mcp act/review/abstain is code. Default act_above 0.8, review_above 0.5, still soft.
- Adkid-Zephyr/chinese-workflow-decision-bench 64/64 and 63/64, synthetic, not a group-chat dump. 8 cases times 2 prompts, CPU and MPS labels matched, and that match is not a pretraining proof.
- pythongiant/laya-drift description is now monitor agent drift (was calculate agentic drift, §142). alignment 0.7 and plan_ref 0.3 are application policy. Densify, not a sibling.
- cloudbtl/JevRAG 20 branches times six hops is combinatorics, not a bench. walk, gather, file, watch. First card.
- quaeast/vllm2jev does not reproduce Jev calibration. wire-compat is not logit-equiv.
- malevrigns/agent-jev 79.25% (1585/2000), ECE 0.1687, marked theirs not Harbor.
- DjTaNg-404/Laya-2048 one seed-48 game 28,092, marked not Harbor. jast HEAD
c6588285208fand README SHAd9214ca31f92match §145; star 0 to 1 is star-noise.
Non-blocking
The Hub limitations also say score questions are untrained (no ordinal image data, outputs meaningless). The fold quotes the card lead (choice / score / noul) and the measured choice and noul rows, and it refuses the act head. It does not quote a score accuracy. Composition item 754 puts 64/64 on the line after the laya-drift sentence; the item title and the uniqueness lock attribute those fractions to the chinese bench, and §146 does too.
ADV_PASS
… + REVISIT densify Card hf:thaitea/laya-vision as a vision decision peer, not an 18th scoring row, and densify the seven revisit cards in place. Soft scores stay off the sole-veto path. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
24601
marked this pull request as ready for review
September 21, 2026 17:44
24601
force-pushed
the
cursor/hourly-1110-fold-4cda
branch
from
September 21, 2026 17:45
f022794 to
09addcb
Compare
This was referenced Sep 21, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Hourly 1110 fold off
mainat6be01e0429ed(merged #70). Notes §146, composition 745-760, findings batch #128. Does not bump 0.5.0. Does not reopen #23-#70.TypeSafe Jev stays the dominant exemplar and the default recommended path. The lane is still the broader decision-model class: classifiers, encoders and decoders, specialized AR and constrained heads, vision and listwise scorers, and what TypeSafe calls System One.
PRIMARY: hf:thaitea/laya-vision (13 likes, sha
a2653db2831b, CC-BY-NC-SA-4.0). SmolVLM-256M typed vision decisions (choice/score/noul, one forward pass, no text generation). Same-species serving, not an 18th scoring row. The act/escalate head is untrained (zero gradient, random init). Do not gate on it. A-OKVQA 63.1% ECE 0.266 to 0.094, ScienceQA 89.0%, VQAv2 noul 73.2%, all n=8,235 75.9% calibrated ECE 0.035 are theirs, not Harbor. The VQAv2 split is a re-split, not published VQAv2. Usage still loadsthaitea/laya-vision-smolvlm-256m(already §87).Also carded: AntonG87/codearia-sieve (page to typed fields; parse is not a decision; 6 of 6 and 11 of 11 theirs), 7Zenox/gemma-jev (letter-slot logits; JSON chat is not a calibrated Noul; Gemma and Qwen3.5 are not Archer), Pdbz199/local-decision-model (one pass from the public post; schema-valid is not the same as correct), musubi-labs/musubi-jev (README is a kev tree copy; copied kev numbers are not a musubi bench), quaeast/vllm2jev, IAmJSD/pg-laya, gbesse/jev-rerank-server (ranking is not calibration), MarkChu-git/typesafe-mcp (act_above 0.8 is still soft; soft judgment is not a sole veto).
REVISIT densify, marked theirs, prior section kept:
69 novel + 7 revisit.
invented_signal: false. Evidence archive:research/archive/hourly/2026-09-21T17/.Test plan
python3 .agents/skills/augustus/scripts/uniqueness_gate.py(includes fingerprint self-test)## Licenseand holds no uniqueness dump