The calibration in #30 produced a list nobody had before, and this is what to do about it.
Every rule below fired on text published before generative models existed. Each hit is a false
positive by construction — there is no judgement call to make.
| Rule |
Texts it fired on (of 90) |
Share |
Total hits |
rhet.rule-of-three |
45 |
50% |
100 |
stat.burstiness |
25 |
27.8% |
25 |
lex.moreover |
23 |
25.6% |
67 |
lex.furthermore |
23 |
25.6% |
55 |
rhet.in-order-to |
22 |
24.4% |
60 |
lex.just |
19 |
21.1% |
30 |
lex.facilitate |
17 |
18.9% |
29 |
rhet.in-terms-of |
16 |
17.8% |
30 |
lex.comprehensive |
14 |
15.6% |
36 |
lex.ademas |
12 |
13.3% |
28 |
The problem
A rule firing on half of human academic prose is measuring the genre, not the machine. Scholars
were writing "furthermore" and grouping things in threes long before anyone trained a transformer.
What this is not
Not an argument for deleting them all. Some of these genuinely are AI tells and genuinely common in
human academic writing, which is exactly why they carry weight rather than certainty, and why the
catalog explains each one. The score is a reading of prose and prose is arguable — that was always the
deal.
But rhet.rule-of-three at 50% is not carrying information. It is background.
What would close this
For each rule near the top, one of:
- Reweight. Lower its contribution so it colours the score instead of driving it.
- Narrow it.
rhet.rule-of-three presumably fires on any three-item list; the AI tell is closer
to abstract triples in a summarising sentence, not "we measured height, weight and age".
- Demote to Info. Shown in the report, contributing nothing to the score — the catalog entry
still teaches the writer something.
- Retire it, and say so in the catalog.
The discipline this needs
Re-run the calibration before and after and put both numbers in the PR. Changing rules to make our
own number look better is exactly the failure mode this measurement exists to prevent, so the change
has to be argued on the rule's merits and the number reported either way.
Note also that lowering weights lowers the whole score distribution, which moves the recommended
threshold down with it. That is fine — the threshold is derived, not chosen — but it means "the
false-positive rate improved" is not by itself evidence of anything.
Related
The calibration in #30 produced a list nobody had before, and this is what to do about it.
Every rule below fired on text published before generative models existed. Each hit is a false
positive by construction — there is no judgement call to make.
rhet.rule-of-threestat.burstinesslex.moreoverlex.furthermorerhet.in-order-tolex.justlex.facilitaterhet.in-terms-oflex.comprehensivelex.ademasThe problem
A rule firing on half of human academic prose is measuring the genre, not the machine. Scholars
were writing "furthermore" and grouping things in threes long before anyone trained a transformer.
What this is not
Not an argument for deleting them all. Some of these genuinely are AI tells and genuinely common in
human academic writing, which is exactly why they carry weight rather than certainty, and why the
catalog explains each one. The score is a reading of prose and prose is arguable — that was always the
deal.
But
rhet.rule-of-threeat 50% is not carrying information. It is background.What would close this
For each rule near the top, one of:
rhet.rule-of-threepresumably fires on any three-item list; the AI tell is closerto abstract triples in a summarising sentence, not "we measured height, weight and age".
still teaches the writer something.
The discipline this needs
Re-run the calibration before and after and put both numbers in the PR. Changing rules to make our
own number look better is exactly the failure mode this measurement exists to prevent, so the change
has to be argued on the rule's merits and the number reported either way.
Note also that lowering weights lowers the whole score distribution, which moves the recommended
threshold down with it. That is fine — the threshold is derived, not chosen — but it means "the
false-positive rate improved" is not by itself evidence of anything.
Related
lex.ademasis the onlySpanish rule in the top ten and the Spanish side currently rests on encyclopedia prose alone.