Skip to content

Fold hourly 1110 HIGH: laya-vision same-species serving + novel cards + REVISIT densify - #72

Merged
24601 merged 1 commit into
mainfrom
cursor/hourly-1110-fold-4cda
Sep 21, 2026
Merged

24601 merged 1 commit into
mainfrom
cursor/hourly-1110-fold-4cda

Conversation

@24601

@24601 24601 commented Sep 21, 2026

Copy link
Copy Markdown
Owner

Summary

Hourly 1110 fold off main at 6be01e0429ed (merged #70). Notes §146, composition 745-760, findings batch #128. Does not bump 0.5.0. Does not reopen #23-#70.

TypeSafe Jev stays the dominant exemplar and the default recommended path. The lane is still the broader decision-model class: classifiers, encoders and decoders, specialized AR and constrained heads, vision and listwise scorers, and what TypeSafe calls System One.

PRIMARY: hf:thaitea/laya-vision (13 likes, sha a2653db2831b, CC-BY-NC-SA-4.0). SmolVLM-256M typed vision decisions (choice / score / noul, one forward pass, no text generation). Same-species serving, not an 18th scoring row. The act/escalate head is untrained (zero gradient, random init). Do not gate on it. A-OKVQA 63.1% ECE 0.266 to 0.094, ScienceQA 89.0%, VQAv2 noul 73.2%, all n=8,235 75.9% calibrated ECE 0.035 are theirs, not Harbor. The VQAv2 split is a re-split, not published VQAv2. Usage still loads thaitea/laya-vision-smolvlm-256m (already §87).

Also carded: AntonG87/codearia-sieve (page to typed fields; parse is not a decision; 6 of 6 and 11 of 11 theirs), 7Zenox/gemma-jev (letter-slot logits; JSON chat is not a calibrated Noul; Gemma and Qwen3.5 are not Archer), Pdbz199/local-decision-model (one pass from the public post; schema-valid is not the same as correct), musubi-labs/musubi-jev (README is a kev tree copy; copied kev numbers are not a musubi bench), quaeast/vllm2jev, IAmJSD/pg-laya, gbesse/jev-rerank-server (ranking is not calibration), MarkChu-git/typesafe-mcp (act_above 0.8 is still soft; soft judgment is not a sole veto).

REVISIT densify, marked theirs, prior section kept:

  • pythongiant/laya-drift §142 (description rewrite to monitor agent drift)
  • turenlabs/jast §145 (SHA unchanged; star 0 to 1 is star-noise)
  • Adkid-Zephyr/chinese-workflow-decision-bench §145 (64/64 and 63/64 theirs; synthetic, not a group-chat dump)
  • JabbaKadabra/SystemOneDotNet §143 (unofficial .NET client; wire-compat is not logit-equiv)
  • arnavm-codes/JevFence §145 (block table is policy; soft scores are not hard gates)
  • cloudbtl/JevRAG first card (revisit tag had no prior notes card; combinatorics is not a bench)
  • gbesse/decision-workbench §139 (human review stays separate from model output; demo scores are not accuracy measurements)

69 novel + 7 revisit. invented_signal: false. Evidence archive: research/archive/hourly/2026-09-21T17/.

Test plan

  • python3 .agents/skills/augustus/scripts/uniqueness_gate.py (includes fingerprint self-test)
  • Adversarial review of quoted theirs numbers against the Hub card and READMEs
  • Confirm README still ends at ## License and holds no uniqueness dump
Open in Web Open in Cursor 

@24601 24601 left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adversarial review (hourly 1110, draft PR #72)

Do not merge from this comment. Catalog densify holds. No blockers.

Gates

  • Occupancy. Current origin/main is 6216705 (merged #71, Release v0.5.1) and still ends at notes.md §145 / composition items 729-744 / findings batch #127. This fold is §146 / 745-760 / batch #128. No collision with already-merged ranges.
  • Release pins. Diff vs base 6be01e04 does not touch README, CHANGELOG, CITATION, marketplace, or docs/release-notes-v0.5.1.md. A three-way merge onto 6216705 stays clean and keeps version: 0.5.1, the release notes, and the slim skill, then adds the 1110 lock. Fold text "Does not bump 0.5.0" is this fold's non-bump, not a pin rewrite.
  • README / uniqueness. README.md still ends at ## License (MIT / SECURITY / CONTRIBUTING only). python3 .agents/skills/augustus/scripts/uniqueness_gate.py prints uniqueness-gate ok. invented_signal: false in research/archive/hourly/2026-09-21T17/run_digest.json. Added human prose has no U+2014.
  • Class framing. TypeSafe Jev stays the default recommended path. The lane stays the wider decision-model class. laya-vision is same-species vision serving, not an 18th scoring-table row (no scoring-table edit). Act/escalate head is untrained (zero gradient, random init) and is not a gate.
  • musubi-jev. README at e943f21e4057 is the kev tree (# Kev, jaredpalmer CI badge). The card says copied kev numbers are not a musubi bench and does not paste a kev table as a musubi measurement. musubi-labs/musubi-jev ≠ jaredpalmer/kev.
  • Soft scores. act_above 0.8 is still soft. JevFence: soft scores are not hard gates. Soft judgment is not a sole veto. Counts: 69 novel + 7 revisit.

theirs checked live

  • thaitea/laya-vision sha a2653db2831b, CC-BY-NC-SA-4.0, 13 likes. A-OKVQA 63.1% ECE 0.266 to 0.094 (n=1,138). ScienceQA image 89.0% ECE 0.080 to 0.034 (n=2,097). VQAv2 noul 73.2% ECE 0.085 to 0.042 (n=5,000, re-split, not published VQAv2). All n=8,235 75.9%, calibrated ECE 0.035 (raw ECE on that row is 0.108; the fold's 0.035 is the calibrated figure). About 71 ms. Option order varies 1.2 points. Train 94.7% vs val 63.1% on A-OKVQA. Usage loads thaitea/laya-vision-smolvlm-256m (§87 card id, not this card).
  • 7Zenox/gemma-jev 144 authored decisions. Gemma 3 270M 0.293 below chance 0.333. Gemma 4 E2B-it JSON chat 0.807. Qwen3.5-4B JSON chat 0.813. bf16 vs fp32 argmax-agreement check not run. JSON chat and letter-slot softmax are not a calibrated Noul.
  • AntonG87/codearia-sieve 6 of 6 and 11 of 11 are the judge table, marked not Harbor. Parse is not a decision.
  • musubi-labs/musubi-jev kev copy, not a new bench.
  • Pdbz199/local-decision-model independent from the public post. The 70 to 500 ms row is attributed to the Jev post, not measured here. Span start/end head is locate, not decide.
  • gbesse/jev-rerank-server SciFact n=25 nDCG@10 0.616377 to 0.718260, Recall@10 0.84 unchanged. Ranking is not calibration.
  • MarkChu-git/typesafe-mcp act/review/abstain is code. Default act_above 0.8, review_above 0.5, still soft.
  • Adkid-Zephyr/chinese-workflow-decision-bench 64/64 and 63/64, synthetic, not a group-chat dump. 8 cases times 2 prompts, CPU and MPS labels matched, and that match is not a pretraining proof.
  • pythongiant/laya-drift description is now monitor agent drift (was calculate agentic drift, §142). alignment 0.7 and plan_ref 0.3 are application policy. Densify, not a sibling.
  • cloudbtl/JevRAG 20 branches times six hops is combinatorics, not a bench. walk, gather, file, watch. First card.
  • quaeast/vllm2jev does not reproduce Jev calibration. wire-compat is not logit-equiv.
  • malevrigns/agent-jev 79.25% (1585/2000), ECE 0.1687, marked theirs not Harbor.
  • DjTaNg-404/Laya-2048 one seed-48 game 28,092, marked not Harbor. jast HEAD c6588285208f and README SHA d9214ca31f92 match §145; star 0 to 1 is star-noise.

Non-blocking

The Hub limitations also say score questions are untrained (no ordinal image data, outputs meaningless). The fold quotes the card lead (choice / score / noul) and the measured choice and noul rows, and it refuses the act head. It does not quote a score accuracy. Composition item 754 puts 64/64 on the line after the laya-drift sentence; the item title and the uniqueness lock attribute those fractions to the chinese bench, and §146 does too.

ADV_PASS

… + REVISIT densify

Card hf:thaitea/laya-vision as a vision decision peer, not an 18th scoring row, and densify the seven revisit cards in place. Soft scores stay off the sole-veto path.

Co-authored-by: Basit Mustafa <24601@users.noreply.github.com>
@24601
24601 marked this pull request as ready for review September 21, 2026 17:44
@24601
24601 force-pushed the cursor/hourly-1110-fold-4cda branch from f022794 to 09addcb Compare September 21, 2026 17:45
@24601
24601 merged commit 8676280 into main Sep 21, 2026
@24601
24601 deleted the cursor/hourly-1110-fold-4cda branch September 21, 2026 17:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants