A free AI detector switched to an older model when my scans ran out, then called an 1859 book 100% AI
The same 1,050 characters of Darwin scored "highly confident this text is entirely human" earlier in the session. After the free quota ran out, the identical text scored "highly confident this text was AI generated" — and the model label on the result had quietly changed.
I nearly published something badly wrong today, and the thing that stopped it is worth more than the thing I was trying to measure.
What happened
Earlier in the session I ran four texts of undisputed origin through GPTZero’s free detector. A passage from an 1859 public domain book came back “We are highly confident this text is entirely human”, 0% AI / 100% human. Correct, and unsurprising.
An hour later I pasted the identical 1,050 characters back in. Same tool, same browser, same free tier.
We are highly confident this text was AI generated. AI 100% · Mixed 0% · Human 0%
Same text. Opposite verdict. Both delivered with “highly confident”.
The text did not change. The model did.
The result panel prints a small model label next to the verdict. On the earlier, correct scans it read:
Model 4.8b
On the scan that called Darwin machine-written it read:
Model 2025-03-13-base (and elsewhere on the page: Model 3.3b)
Alongside it, a modal I had not seen before: “Sign up to Continue”, and in the results panel, “Create a free account to unlock more scans.”
So the free allowance had run out, and rather than refusing to answer, the tool answered anyway — using a different, older model, with no warning that the thing grading my text was not the thing it had been grading with ten minutes earlier.
Why this matters more than the false positive itself
A false positive on Victorian prose is a curiosity. A tool that silently changes the model behind a confident verdict is a different category of problem.
Nothing in the verdict tells you it happened. The wording is identical: “We are highly confident”. The percentage is a round 100%. The layout is the same. The only tell is a small grey model string that means nothing to anyone who has not been staring at it across two sessions.
If you are a teacher who ran a few essays this morning and a few more this afternoon, you have no way of knowing that the afternoon batch was scored by something else. The results look exactly as authoritative.
What I can and cannot claim
Established, and reproduced: the identical 1,050-character 1859 passage received “highly confident, entirely human” on Model 4.8b and “highly confident, AI generated” after the sign-up wall appeared, with a different model label on the result.
Correlated, not isolated: the quota exhaustion and the model change appeared together. I have not proven the first causes the second. It is the obvious reading and I am not going to state it as more than that.
Not a claim about GPTZero’s accuracy. On its current model, in the earlier trial, it got all four texts right. This is not “the detector is bad”. It is “the free tier can stop being the detector you think you are using, without saying so.”
I also could not continue testing once the wall appeared, so the sample here is small by force rather than by choice.
The near miss, which is the actual lesson
Between those two verdicts I also ran a longer passage from a 1996 internet standards document. It came back 100% AI, highly confident, twice.
For about ten minutes I believed I had a genuinely strong finding: a detector confidently flagging a document written twenty-six years before ChatGPT. It would have made a far better headline than this one.
It was almost certainly an artifact of the same degraded model, and I cannot now separate the two, because the runs happened after the quota was already gone and I had not been recording the model label per run.
The result that looks spectacular is the one most likely to be your harness failing. The check that caught it was not clever: it was screenshotting the page instead of scraping a number out of it. The percentages I was reading programmatically were correct and completely misleading. The model label and the sign-up wall were only ever visible in the picture.
Record the tool’s own version string alongside every result. I now have four scans I cannot interpret because I did not.
Method
GPTZero’s free web detector, no account, single browser session. Texts were pasted unedited; verdicts and model labels are quoted exactly as displayed. The 1859 passage is public domain, chosen originally because its date makes machine authorship impossible, which is also what made the second verdict obviously wrong rather than merely surprising.