Skip to content

Rules that fire on human academic writing: rule-of-three at 50% #31

Description

@peopleworks

The calibration in #30 produced a list nobody had before, and this is what to do about it.

Every rule below fired on text published before generative models existed. Each hit is a false
positive by construction — there is no judgement call to make.

Rule Texts it fired on (of 90) Share Total hits
rhet.rule-of-three 45 50% 100
stat.burstiness 25 27.8% 25
lex.moreover 23 25.6% 67
lex.furthermore 23 25.6% 55
rhet.in-order-to 22 24.4% 60
lex.just 19 21.1% 30
lex.facilitate 17 18.9% 29
rhet.in-terms-of 16 17.8% 30
lex.comprehensive 14 15.6% 36
lex.ademas 12 13.3% 28

The problem

A rule firing on half of human academic prose is measuring the genre, not the machine. Scholars
were writing "furthermore" and grouping things in threes long before anyone trained a transformer.

What this is not

Not an argument for deleting them all. Some of these genuinely are AI tells and genuinely common in
human academic writing, which is exactly why they carry weight rather than certainty, and why the
catalog explains each one. The score is a reading of prose and prose is arguable — that was always the
deal.

But rhet.rule-of-three at 50% is not carrying information. It is background.

What would close this

For each rule near the top, one of:

  1. Reweight. Lower its contribution so it colours the score instead of driving it.
  2. Narrow it. rhet.rule-of-three presumably fires on any three-item list; the AI tell is closer
    to abstract triples in a summarising sentence, not "we measured height, weight and age".
  3. Demote to Info. Shown in the report, contributing nothing to the score — the catalog entry
    still teaches the writer something.
  4. Retire it, and say so in the catalog.

The discipline this needs

Re-run the calibration before and after and put both numbers in the PR. Changing rules to make our
own number look better is exactly the failure mode this measurement exists to prevent, so the change
has to be argued on the rule's merits and the number reported either way.

Note also that lowering weights lowers the whole score distribution, which moves the recommended
threshold down with it. That is fine — the threshold is derived, not chosen — but it means "the
false-positive rate improved" is not by itself evidence of anything.

Related

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions