A detector called an 1859 book AI-generated while its own numbers said 85% human
The verdict sentence and the numbers directly beneath it disagreed on every run. The headline percentage also changed from 18% to 0% on identical text, while the breakdown underneath stayed exactly the same.
I pasted a paragraph published in 1859 into a free AI detector. Here is what it told me, in the order the page presents it:
18% · AI GPT Your input content seems to be AI generated
AI-generated 0% · Mixed 15% · Human-written 85%
Read that twice. The sentence says the text seems to be AI generated. The numbers immediately below it say zero percent AI-generated and eighty-five percent human-written.
Both are on screen at the same time, about a book published one hundred and sixty-seven years ago.
The contradiction was not a one-off
I ran the passage again, and then ran a variant. The verdict sentence never changed:
| Run | Headline | AI-generated | Mixed | Human-written | Verdict sentence |
|---|---|---|---|---|---|
| Original | 18% | 0% | 15% | 85% | “seems to be AI generated” |
| Original, again | 0% | 0% | 15% | 85% | “seems to be AI generated” |
| Punctuation variant | 6% | 0% | 35% | 65% | “seems to be AI generated” |
Three runs, three times the same sentence, including on a run scored 0% and 85% human-written.
If that sentence is what a hurried reader takes away, and it is the largest piece of prose in the result panel, then this tool tells you an 1859 text looks machine-written no matter what its own analysis found.
The headline number moved on identical input
Compare the first two rows. Same 1,050 characters, same tool, minutes apart. The breakdown is identical — 0% / 15% / 85% both times. The headline percentage went from 18% to 0%.
So the stable number is the one in small type, and the unstable one is the big number at the top. I only have two runs of it, so treat this as an observation rather than a measured variance, but the direction of the problem is clear: the figure most likely to be screenshotted and sent to a student is the figure that did not reproduce.
Changing only punctuation moved the human score by 20 points
The third row is the same paragraph with sentences joined at semicolons. No word was added, removed or altered, and no letter changed case — I asserted that programmatically before running it, and the tool independently counted 186 words for both versions.
Human-written fell from 85% to 65%, and Mixed rose from 15% to 35%.
That is a single observation and I am not going to build a rule on it. But it is consistent with the mechanism these tools are known to use: sentence structure feeds the score, so how you punctuate can move you toward an accusation without changing a single word you wrote.
The tool half-admits this itself. Under every result sits: “Borderline result — long sentences or complex words may be tipping the detector.”
The same page sells you the way out
Directly above the verdict, before you have even read it:
Still flagged as AI? GPTinf humanizes your text so it passes — and actually sounds human. [Not now] [Humanize now]
The product that flags your writing also sells the product that unflags it, on the same screen, in the same session. I am not going to speculate about intent. I will note that a detector with a humanizer attached has a structural reason for its verdict sentence to lean toward “AI generated”, and that the sentence did lean that way on every run including the 85%-human one.
The result panel also lists “Cross-checked with: Turnitin, Copyleaks, OriginalityAI, GPTZero, Crossplag, Sapling.ai, Gowinston.ai, ZeroGPT”. No detail is given about what that cross-checking is, and nothing in the result identifies which of them, if any, produced anything.
What this is and is not
It is one text and three runs on one tool. No accuracy rate is claimed, and none should be read into it. The contradiction between the sentence and the numbers is what reproduced, not a percentage.
It is not proof the underlying classifier is bad. On these runs the classifier’s own breakdown was right: 0% AI-generated for a passage from 1859 is the correct answer. The failure is in how that answer is presented, which is the part a teacher or an editor actually reads.
A fourth run did not complete. The tool sat on “Waiting for text …” and never returned a result, so it is not in the table.
The practical version
If you are handed a detector screenshot as evidence, the questions worth asking are: what did the breakdown say, as opposed to the sentence; does the number reproduce when the same text is submitted twice; and does the tool that produced it also sell a product for defeating it.
On this trial, those three questions had answers of 85% human, no, and yes.
Method
GPTinf’s free web detector, no account, no word limit, single browser session, in the same sitting. The 1859 passage is public domain and was chosen because its date makes machine authorship impossible. The punctuation variant was generated by joining sentence pairs at semicolons and verified to be token-identical to the original including capitalisation. Verdicts and percentages are quoted exactly as displayed, and each result was screenshotted as well as read, after a previous trial showed scraped numbers can be accurate and still completely misleading.