Skip to content

Hand each reader the number measured for their own language - #49

Merged
peopleworks merged 1 commit into
mainfrom
teacher-package-says-what-the-tool-now-does
Aug 5, 2026
Merged

Hand each reader the number measured for their own language#49
peopleworks merged 1 commit into
mainfrom
teacher-package-says-what-the-tool-now-does

Conversation

@peopleworks

Copy link
Copy Markdown
Owner

Docs only. Found while checking what #32 made stale before writing the article about it.

Two problems, one caused by #32 and one much older

1. The syllabus promised something the tool no longer does.

I may run submitted work through an offline writing-analysis tool. It produces no verdict

True when written: the report's gate demanded a per-language threshold no language has, so it
withheld the verdict from every document ever analysed. #32 fixed the gate. Above the boundary the
tool now prints "Signs of AI writing".

That sentence is meant to be pasted into a course syllabus. A teacher would have been
contradicted by the tool in front of their own class. It now says what stays true — the tool reaches
no conclusion about who wrote anything, and nothing it produces decides anything.

2. The committee template handed over the pooled rate. This one is older, and worse.

The template gave the teacher under 4.1% overall to copy into a disciplinary finding — regardless
of the language of the work.

Language Texts Bound it supports
English 65 5.6%
Spanish 25 13.3%
pooled 90 4.1%

A committee judging a Spanish essay was being handed a tool three times better than the one
actually used
, in the teacher's own handwriting, in a document that gets read at appeal.

This is exactly the harm #36 removed from the report — sitting untouched in the package aimed at the
people who carry reports to committees. The template now points at the figure the report prints for
the work being judged, and says why copying a rate from a template overstates the tool.

The student sheet had the same defect quietly: both language editions quoted the pooled figure.
Each now carries its own, and says the pooled number is the more flattering one and that is why it is
not used. A Spanish-speaking student is the person this project is most likely to harm with a rate
measured mostly on English, and the sheet written for them was quoting it.

Every figure verified against published-calibration.json.

Two things went stale in the teacher package, and only one of them was #32's doing.

The syllabus wording said the tool "produces no verdict". That was true when it was written — the
report's gate demanded a per-language threshold that no language has, so it withheld the verdict
from every document ever analysed. #32 fixed the gate, and above the measured boundary the tool now
says "Signs of AI writing", which is a verdict by any reading. A teacher who pasted that sentence
into a syllabus would have been contradicted by the tool in front of their own class. It now says
what stays true: the tool reaches no conclusion about who wrote anything, and nothing it produces
is the basis of a decision.

The second is older and worse. The committee template handed the teacher the **pooled** false-
positive rate — "under 4.1% overall" — to write into a disciplinary finding, whatever language the
work was in. Spanish on its own supports 13.3%: a committee judging a Spanish essay was being handed
a tool three times better than the one actually used, in the teacher's own handwriting, in a
document that gets read at appeal. This is precisely the harm #36 removed from the report, sitting
untouched in the package aimed at the people who take reports to committees.

The template now points at the figure the report prints for the work being judged, and says plainly
why copying the pooled number from a template is the easiest way to overstate this tool at a hearing.

The student sheet had the same defect in a quieter form: both language editions quoted the pooled
figure. Each now carries its own — 6 in 100 for English over 65 texts, 13 in 100 for Spanish over
25 — and says that the pooled number is the more flattering one, which is why it is not being used.
A Spanish-speaking student was the person in this project most likely to be harmed by a rate
measured mostly on English, and the sheet written for them was quoting it.

No code changed. Every figure checked against published-calibration.json.
@peopleworks
peopleworks merged commit f20cf15 into main Aug 5, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant