aieveryminute

I gave two AI detectors text from 1859 and 1996. Both called it human, but the newer one scored 23% AI

Four texts whose origin is not in doubt: two written long before language models existed, two generated this week. GPTZero got all four right. The interesting part is how much closer to the line the modern human writing sat.

Every few weeks someone is accused of using AI to write something they wrote themselves. The detectors that trigger those accusations advertise 99% accuracy. The people writing about them mostly quote that number back.

I wanted to see the thing work, so I gave two free detectors four pieces of text where the answer is not arguable.

The four texts

Two of them are human beyond dispute, because they were published decades before a language model could have written anything:

  • Darwin, 1859. A paragraph of flowing argumentative prose from a public domain book.
  • RFC 1958, 1996. Explanatory prose about internet architecture, written by a human standards author, twenty-six years before ChatGPT.

Two are machine-written, because I generated them minutes earlier and watched them appear:

  • A Claude essay, 250 words on why people resist changing their minds.
  • A DeepSeek V3 essay, the same prompt.

Everything went in unedited. No paraphrasing, no humanising, no mixing.

What GPTZero said

Text Origin Verdict AI / Human
Darwin, 1859 human “highly confident this text is entirely human” 0% / 100%
RFC 1958, 1996 human “moderately confident this text is entirely human” 23% / 77%
Claude essay AI “highly confident this text was AI generated” 100% / 0%
DeepSeek V3 essay AI “highly confident this text was AI generated” 100% / 0%

Four out of four correct. Both machine essays were caught at 100% with no hedging, and neither human text was flagged.

That is not the result I expected to be writing up. The prevailing story about these tools is that they misfire constantly, and on clean unedited samples this one did not misfire at all.

The margin is more interesting than the verdict

Look at the two human texts again. Darwin came back at 100% human, “highly confident”. The 1996 RFC came back at 77% human with 23% AI, and only “moderately confident”.

Both are unambiguously human. One was scored a quarter of the way toward a false accusation.

The difference between them is not authorship, it is register. The RFC is modern, plain, declarative technical prose: short sentences, measured claims, careful hedging, no ornament. That is also a fair description of how a language model writes. Darwin’s 1859 prose, with its long subordinate clauses and personal address to the reader, looks nothing like model output and the detector was certain about it.

Which suggests the risk is not “human writing gets flagged”, it is that a particular kind of human writing gets flagged: clear, plain, structured, unadorned. The kind taught in technical writing guides. The kind produced by people writing carefully in a second language.

This is one data point and I am not going to inflate it into a rate. But it lines up with why the complaints exist, and it is visible even in a test the detector passed.

The second detector agreed, and warns you not to trust it

I ran the same 1996 RFC text through QuillBot’s free detector, which needs no account either. It returned 0% AI, 100% human-written — more confident than GPTZero on the same passage.

Underneath the result panel, in small grey text, sits this:

Never rely on AI detection alone to make decisions that could impact someone’s career or academic standing.

That is the vendor’s own footer, on a product marketed to educators. It is the most honest sentence I read all day, and it is the one nobody quotes.

What this trial does not tell you

Four texts is four texts. This is a trial, not a benchmark. No accuracy rate is being claimed here and none should be inferred from it.

Everything was clean and unedited, which is the easy case. The hard case, and the one that produces the accusations, is AI text lightly rewritten by a human, or human text passed through a paraphraser. I have not tested that yet.

I only got one of the four texts through QuillBot. Its detector page redirected into a document editor partway through and stopped accepting programmatic input. Rather than report a partial matrix as if it were complete, that is where it stopped.

What I would want next

The question worth answering is not “do detectors catch raw model output”, because on this evidence they do, easily. It is how much human editing it takes before AI text reads as human, and how much plain technical writing it takes before human text reads as AI. That is where careers actually get damaged, and both directions are testable.

Method

Human samples are a paragraph from an 1859 public domain book and prose from RFC 1958, published 1996, both chosen because their date makes machine authorship impossible. AI samples were generated immediately before testing, one from Claude and one from DeepSeek V3, on an identical prompt, and pasted unedited. Both detectors were used through their free web interfaces with no account. Verdicts are quoted exactly as the tools displayed them.


Follow-up 2026-08-09. Every verdict above was produced on GPTZero’s Model 4.8b, which the result panel prints beside the score. Later the same day, once the free scan allowance ran out, the tool began answering with an older model and classified this same 1859 passage as 100% AI, “highly confident”. The readings in this post still stand as recorded, but the free tier does not reliably give you the same detector twice: what happened, and why the model label is the thing to check.

POSTaieveryminute.com#tool-trialbuilt 2026-08-31 17:47 UTC