Hand each reader the number measured for their own language - #49
Merged
Merged
Conversation
Two things went stale in the teacher package, and only one of them was #32's doing. The syllabus wording said the tool "produces no verdict". That was true when it was written — the report's gate demanded a per-language threshold that no language has, so it withheld the verdict from every document ever analysed. #32 fixed the gate, and above the measured boundary the tool now says "Signs of AI writing", which is a verdict by any reading. A teacher who pasted that sentence into a syllabus would have been contradicted by the tool in front of their own class. It now says what stays true: the tool reaches no conclusion about who wrote anything, and nothing it produces is the basis of a decision. The second is older and worse. The committee template handed the teacher the **pooled** false- positive rate — "under 4.1% overall" — to write into a disciplinary finding, whatever language the work was in. Spanish on its own supports 13.3%: a committee judging a Spanish essay was being handed a tool three times better than the one actually used, in the teacher's own handwriting, in a document that gets read at appeal. This is precisely the harm #36 removed from the report, sitting untouched in the package aimed at the people who take reports to committees. The template now points at the figure the report prints for the work being judged, and says plainly why copying the pooled number from a template is the easiest way to overstate this tool at a hearing. The student sheet had the same defect in a quieter form: both language editions quoted the pooled figure. Each now carries its own — 6 in 100 for English over 65 texts, 13 in 100 for Spanish over 25 — and says that the pooled number is the more flattering one, which is why it is not being used. A Spanish-speaking student was the person in this project most likely to be harmed by a rate measured mostly on English, and the sheet written for them was quoting it. No code changed. Every figure checked against published-calibration.json.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Docs only. Found while checking what #32 made stale before writing the article about it.
Two problems, one caused by #32 and one much older
1. The syllabus promised something the tool no longer does.
True when written: the report's gate demanded a per-language threshold no language has, so it
withheld the verdict from every document ever analysed. #32 fixed the gate. Above the boundary the
tool now prints "Signs of AI writing".
That sentence is meant to be pasted into a course syllabus. A teacher would have been
contradicted by the tool in front of their own class. It now says what stays true — the tool reaches
no conclusion about who wrote anything, and nothing it produces decides anything.
2. The committee template handed over the pooled rate. This one is older, and worse.
The template gave the teacher
under 4.1% overallto copy into a disciplinary finding — regardlessof the language of the work.
A committee judging a Spanish essay was being handed a tool three times better than the one
actually used, in the teacher's own handwriting, in a document that gets read at appeal.
This is exactly the harm #36 removed from the report — sitting untouched in the package aimed at the
people who carry reports to committees. The template now points at the figure the report prints for
the work being judged, and says why copying a rate from a template overstates the tool.
The student sheet had the same defect quietly: both language editions quoted the pooled figure.
Each now carries its own, and says the pooled number is the more flattering one and that is why it is
not used. A Spanish-speaking student is the person this project is most likely to harm with a rate
measured mostly on English, and the sheet written for them was quoting it.
Every figure verified against
published-calibration.json.