From 702d958d2faf3732fc2387250ffb160308a4b24e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Pedro=20Hern=C3=A1ndez?= Date: Mon, 10 Aug 2026 14:25:34 -0400 Subject: [PATCH] The page for somebody who will never read the README MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Closes #43. Everything this project has published is written for people who already know what a false-positive rate is: the README opens with badges, CALIBRATION.md opens with a confidence interval. Both are honest and both are unreadable to the person we most need to reach. It opens by answering the only question that matters, in the size of the question: **can software tell whether a machine wrote this? No.** Not this one, not the paid ones. Every competitor implies otherwise, and saying it first is the product. Then the three things it does, laid out so the argument is visible before a word is read: the score spans the row as an opinion about prose, and the two facts — characters nobody types, a bibliography against itself — sit paired underneath. A percentage is argued about for half an hour; the same identifier on two references is settled with "can you send me the paper?". Then the number nobody publishes, drawn rather than tabulated. Ninety circles, none of them filled, and the interval drawn around the zero so the reader sees why we do not claim zero. Below it, one bar per language, because handing the pooled figure to everyone would be convenient and false — 5.6% over 65 English texts, 13.3% over 25 Spanish ones, with the axis marked at 0, 5, 10 and 15 so the drawing can be checked against the number rather than trusted. The first draft of that chart could not be: it drew both bars on a 0-10% axis and labelled one of them 13.3%, so the Spanish bar showed 9.3% under a number that said otherwise. On the page whose entire argument is that we do not lie with numbers. Then what it does *not* tell you — how much machine writing it misses, never measured and deliberately so; that the ninety are published articles and not student essays; and that the signal firing most often on human writing is uniform sentence rhythm, at 27.8%. Then why trust this one: because on 10 August 2026 it was accusing people falsely and we published that. A tidy bibliography listing further reading was announced as "1 source contradiction" while the code held a field that knew the difference and a comment explaining it. Ask any paid detector when it last accused someone falsely. Then Monday morning: straight into Docs/Teaching, with the Spanish sheet carrying the Spanish figure. **It phones nowhere.** No CDN, no web font, no analytics, no image host, and no script at all — the two languages switch on a radio input and :has(). A page arguing that your students' work never leaves your machine cannot itself call somebody else's server, and the defect we shipped in the report this month was exactly that shape. 31 KB, opens from a downloads folder. Not linked from inside the app yet, on purpose: the interface is shared with the desktop host, where this file does not exist, so an in-app link would be broken there. Placement is a decision, not an oversight. Co-Authored-By: Claude Opus 5 --- README.md | 2 + src/SignsOfAI.Web/wwwroot/why.html | 530 +++++++++++++++++++++++++++++ 2 files changed, 532 insertions(+) create mode 100644 src/SignsOfAI.Web/wwwroot/why.html diff --git a/README.md b/README.md index 2c85f85..d1a1c97 100644 --- a/README.md +++ b/README.md @@ -12,6 +12,8 @@ [![NuGet MCP](https://img.shields.io/nuget/v/SignsOfAI.Mcp?logo=nuget&label=MCP%20server)](https://www.nuget.org/packages/SignsOfAI.Mcp) [![Available on CodeGuilds](https://img.shields.io/badge/Available_on-CodeGuilds-6366f1?logo=data:image/svg%2bxml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHZpZXdCb3g9IjAgMCAyNCAyNCI+PHBhdGggZmlsbD0id2hpdGUiIGQ9Ik0xMiAyTDIgN2wxMCA1IDEwLTV6TTIgMTdsMTAgNSAxMC01TTIgMTJsMTAgNSAxMC01Ii8+PC9zdmc+)](https://codeguilds.dev/packages/signsofai) +**[Are you a teacher? Start here →](https://peopleworks.github.io/SignsofAI/why.html)** — what this does and what it cannot do, in plain language, with the error rate drawn rather than tabulated. No badges, no interval notation, nothing to install. English & Spanish. + **[Try the live demo →](https://peopleworks.github.io/SignsofAI/)** — English & Spanish, runs in your browser. No signup, and the analysis uploads nothing. **[Download the Windows app →](https://github.com/peopleworks/SignsofAI/releases?q=desktop&expanded=true)** — the same tool in a window. Nothing to install alongside it: the .NET runtime is bundled. diff --git a/src/SignsOfAI.Web/wwwroot/why.html b/src/SignsOfAI.Web/wwwroot/why.html new file mode 100644 index 0000000..28b681a --- /dev/null +++ b/src/SignsOfAI.Web/wwwroot/why.html @@ -0,0 +1,530 @@ + + + + + + +¿Puede un programa saber si esto lo escribió una máquina? · Signs of AI Writing + + + + + + + + + + +
+ +

¿Puede un programa saber si un texto lo escribió una máquina?

+
No.
+

+ Ninguno puede. Ni este, ni los de pago, ni el que le enseñaron en una formación. + Si alguien le dice lo contrario, le está vendiendo algo. +

+

+ Esta página explica, sin una sola palabra técnica, qué hace entonces esta herramienta, + cuánto se equivoca —con el número medido, que casi nadie publica— y qué puede usted + hacer el lunes con todo eso. +

+ +

Entonces, ¿qué hace?

+

+ Tres cosas distintas. Y lo importante no son las tres: es que dos son hechos y una + es una opinión, y confundirlas es de donde salen las injusticias. +

+ +
+
+ Opinión sobre la prosa +

La puntuación. Señala palabras y giros que la escritura de máquina usa mucho, y le + enseña cada uno subrayado en el texto. Es un juicio sobre cómo suena. Se puede discutir, + y a veces se equivoca.

+
+
+ Hecho +

Caracteres que no se teclean. Las herramientas que «humanizan» texto dejan letras + de otro alfabeto que se ven idénticas. Le da el código, la línea y la columna. Está o no está.

+
+
+ Hecho +

La bibliografía contra sí misma. Una fuente citada que no aparece en su propia lista; + el mismo identificador en dos artículos distintos; un año que aún no ha llegado. Sin internet: + el documento se contradice solo.

+
+
+ +

+ Un porcentaje se discute en una reunión durante media hora. Que el mismo identificador + esté en dos referencias distintas se resuelve con una pregunta: «¿me manda el artículo?». + Por eso la herramienta enseña las dos cosas por separado y nunca suma la segunda a la primera. +

+ +

¿Cuánto se equivoca?

+

+ Esta es la pregunta que ningún detector contesta, y es la única que importa cuando hay + una persona delante. +

+

+ Lo medimos así: noventa textos publicados antes de que existieran los modelos + generativos. No los juzgó nadie: sabemos que son humanos por su fecha. Después los + pasamos por la herramienta y contamos a cuántos habría acusado. +

+ +
+ + 90 TEXTOS HUMANOS + + + + + + + + + + 0 + marcados + ninguno relleno, + ninguno acusado + + + LO QUE SE PUEDE AFIRMAR DE VERDAD + + + + 0 % + 4,1 % + 10 % + +
+ Ninguno de los noventa fue marcado. Pero cero de noventa no significa que la tasa + sea cero: con esa muestra, lo máximo que se puede afirmar honestamente es que se + equivoca menos del 4,1 % de las veces. Esa franja ámbar es la diferencia + entre una medición y una promesa. +
+
+ +

Y no es el mismo número para todos

+

+ Repartir la cifra global a todo el mundo sería cómodo y sería falso. El corpus tiene + sesenta y cinco textos en inglés y veinticinco en español, así que lo que se puede + afirmar sobre cada idioma es distinto — y en la herramienta cada lector ve el suyo, + no el promedio. +

+ +
+ + Inglés + 65 textos + + + 5,6 % + + Español + 25 textos + + + 13,3 % + + + + 0 % + 5 % + 10 % + 15 % + + +
+ Menos textos en español significa menos certeza, no más errores. La barra es más larga + porque sabemos menos, y decirlo es más útil que esconderlo detrás del promedio. +
+
+ +

Lo que no le dice

+

+ Cuánta escritura de máquina se le escapa. No está medido, y es a propósito: + para medirlo habría que reunir textos generados, y esa colección solo representa los modelos + que estaban de moda ese mes. Una herramienta que no marca nada tiene una tasa de error perfecta. +

+

+ Los noventa textos son artículos publicados, no trabajos de estudiantes. Un + ensayo de primero de carrera no se parece a un artículo revisado por pares, y esa diferencia + no está medida aquí. +

+
+ Y una que conviene saber antes de mirar ninguna puntuación: la señal que más salta en escritura + humana es el ritmo uniforme de las frases — aparece en el 27,8 % de los textos + humanos que medimos. Si la evidencia de un trabajo se apoya sobre todo en eso, vale menos. + La herramienta lo dice en cada informe, por su nombre. +
+ +

¿Por qué fiarse de este y no de otro?

+

+ No por la tecnología. Por esto: el 10 de agosto de 2026 esta herramienta acusaba + en falso, y lo publicamos. +

+

+ Un trabajo con la bibliografía en orden, que solo listaba una lectura adicional sin citarla, + se anunciaba en el informe como «1 contradicción de fuentes», encima + de la frase «el documento se contradice a sí mismo». En la página que un profesor imprime y + lleva a un comité. +

+

+ Lo peor no fue el fallo. Fue que el propio código sabía la diferencia: hay un + campo que distingue una contradicción real de un simple desorden, con un comentario que explica + que la gente lista lectura adicional con toda legitimidad. La línea que escribía el titular + nunca le preguntó. Estuvo cinco días publicado y no lo encontró ningún usuario. +

+

+ Está arreglado, está contado y el historial es público. Pregúntele a cualquier detector + de pago cuándo fue la última vez que acusó en falso. Esa respuesta, y no un porcentaje, + es lo que debería decidir de quién se fía. +

+ +

¿Y el lunes qué hago?

+

+ Nada de esto sirve si no se traduce en algo que decir en un aula o en una reunión. Eso está + escrito, en castellano, y no lleva software dentro: +

+ +

+ La hoja del estudiante lleva impresa la cifra medida para su idioma, no la global. + Si el trabajo está en español, el número que va al comité es el 13,3 %. +

+ +

+ Se ejecuta entera en su navegador. El texto no se sube a ningún sitio — tampoco a nosotros. + Esta página tampoco pide nada a ningún servidor: ni una fuente, ni una imagen, ni una analítica. +

+ +
+ + +
+ +

Can software tell whether a machine wrote this?

+
No.
+

+ None of them can. Not this one, not the paid ones, not the one demonstrated at a training day. + Anyone who tells you otherwise is selling something. +

+

+ This page explains, without a single technical word, what this tool does instead, how often it is + wrong — with the measured number, which almost nobody publishes — and what you can do with any of + it on Monday morning. +

+ +

So what does it do?

+

+ Three separate things. What matters is not the three: it is that two of them are facts and + one is an opinion, and confusing those is where the injustices come from. +

+ +
+
+ An opinion about prose +

The score. It names words and turns of phrase that machine writing leans on, and shows + you each one underlined in the text. It is a judgement about how the writing sounds. It can be + argued with, and sometimes it is wrong.

+
+
+ A fact +

Characters nobody types. Tools that "humanise" text leave behind letters from another + alphabet that look identical. It gives you the codepoint, the line and the column. Either it is + there or it is not.

+
+
+ A fact +

A bibliography against itself. A source cited in the text and missing from its own + list; the same identifier on two different papers; a year that has not happened. No internet + needed: the document contradicts itself.

+
+
+ +

+ A percentage can be argued about for half an hour in a meeting. The same identifier on two + different references is settled with one question: "can you send me the paper?" That is + why the tool shows the two apart, and never adds the second to the first. +

+ +

How often is it wrong?

+

+ This is the question no detector answers, and the only one that matters when there is a person + on the other side of it. +

+

+ Here is how it was measured: ninety texts published before generative models + existed. Nobody judged them human — their dates did. Then they went through the tool and + we counted how many it would have accused. +

+ +
+ + 90 HUMAN TEXTS + + + + + + + + + 0 + flagged + none filled in, + none accused + + WHAT CAN HONESTLY BE CLAIMED + + + + 0% + 4.1% + 10% + +
+ None of the ninety was flagged. But zero out of ninety is not a rate of zero: + with a sample that size, the most anyone can honestly claim is that it is wrong + less than 4.1% of the time. That amber band is the difference between a + measurement and a promise. +
+
+ +

And it is not the same number for everybody

+

+ Handing the pooled figure to everyone would be convenient and false. The corpus holds sixty-five + English texts and twenty-five Spanish ones, so what can be claimed about each language differs — + and in the tool, each reader is shown their own, never the average. +

+ +
+ + English + 65 texts + + + 5.6% + + Spanish + 25 texts + + + 13.3% + + + + 0% + 5% + 10% + 15% + + +
+ Fewer Spanish texts means less certainty, not more mistakes. The bar is longer because + we know less, and saying so is more use than hiding it behind an average. +
+
+ +

What it does not tell you

+

+ How much machine writing it misses. That is not measured, deliberately: measuring + it would mean collecting generated text, and any such collection represents whichever models were + fashionable that month. A tool that flags nothing has a perfect error rate. +

+

+ The ninety texts are published articles, not student work. A first-year essay + does not read like a peer-reviewed paper, and that difference is not measured here. +

+
+ And one worth knowing before you look at any score: the signal that fires most often on human + writing is uniform sentence rhythm — it appears in 27.8% of the human texts we + measured. If the evidence against a piece of work leans mostly on that, it is worth less. The + tool says so on every report, by name. +
+ +

Why trust this one and not another?

+

+ Not because of the technology. Because of this: on 10 August 2026 this tool was accusing + people falsely, and we published that. +

+

+ A piece of work with a perfectly good bibliography, which merely listed some further reading + without citing it, was announced in the report as "1 source contradiction", + above the sentence "the document disagrees with itself". On the page a teacher prints and carries + into a committee. +

+

+ The bug was not the worst part. The code itself knew the difference: there is a + field that separates a real contradiction from mere untidiness, with a comment explaining that + people legitimately list further reading. The line printing the headline never asked it. It was + published for five days and no user found it. +

+

+ It is fixed, it is written up, and the history is public. Ask any paid detector when it + last accused someone falsely. That answer, not a percentage, is what should decide who + you trust. +

+ +

What do I do on Monday?

+

+ None of this is worth anything until it becomes something you can say in a classroom or a + meeting. That part is written, and there is no software in it: +

+ +

+ The student sheet carries the figure measured for their language, not the pooled one. + If the work is in Spanish, the number that goes to the committee is 13.3%. +

+ +

+ It runs entirely in your browser. The text is not uploaded anywhere — not even to us. This page + requests nothing from any server either: no font, no image, no analytics. +

+ +
+ +
+ + + +