aieveryminute

7 of 12 runs got the date right. Every miss that showed its working was arithmetic

Six models from four vendors, asked the same working-day question twice each in separate chats. All eight runs that showed a route to their answer set the problem up correctly. Five answers were still wrong, and four of the five announce it in their own text.

A project starts on 3 March 2026 and runs for 120 working days, Monday to Friday, holidays ignored. What date does it finish?

It is ordinary project arithmetic, and it has two defensible answers depending on whether the start date counts as day one. Both were computed and both were declared correct before a single model was asked:

Reading Answer
The start date is working day 1 Monday 17 August 2026
120 working days after the start Tuesday 18 August 2026

Six models are offered on duck.ai, DuckDuckGo’s free chat surface, and their labels name four different makers. Each was asked twice, in a fresh chat, with nothing scripted.

What came back

Model, as the UI labelled it Run 1 Run 2 Right
GPT-5.4 nano 2 July 2026 17 September 2026 0/2
GPT-5.4 mini 25 August 2026 22 August 2026 0/2
Claude Haiku 4.5 17 August 2026 17 August 2026 2/2
Mistral Small 4 27 August 2026 18 August 2026 1/2
gpt-oss 120B 18 August 2026 18 August 2026 2/2
Gemma 4 31B 18 August 2026 18 August 2026 2/2

7 of 12 runs correct. Three of the six models were right both times.

Nothing that showed its working misread the question

Eight of the twelve runs showed a setup, and all eight got the setup right. Nobody counted calendar days instead of working days, nobody forgot the weekends, nobody wandered off into holidays. The other four printed a date and nothing else.

Two methods appeared, and both are sound:

  • The shortcut. 120 working days is exactly 24 five-day weeks, so 120 working days after the start is 168 calendar days later, landing on the same weekday. Five runs used it.
  • Month by month. Count the working days in March from the 3rd, then April, May, June, July, and finish the remainder in August. Two runs used it, both from Claude Haiku 4.5, and its month counts are right every time: 21, 22, 21, 22, 23.

The eighth run, GPT-5.4 nano’s first, stated the setup correctly and then produced a date with no route to it at all.

Three of the five wrong answers showed working, and all three were a correct method executed wrongly. Mistral Small 4’s first run states 24 weeks × 7 days/week = 168 calendar days, which is right, then writes 3 March 2026 + 168 days = 27 August 2026. Three March plus 168 days is 18 August. The method got the model to the right doorstep and the addition walked it past.

The other two wrong answers showed nothing, so nothing can be said about them. GPT-5.4 mini replied with a bare date on both runs. Whether it misread the question is not decidable from this trial and no claim is made either way.

The same model, the same prompt, the same method, two different dates

Mistral Small 4 is the clearest artefact in the trial. Both of its runs pick the 24-week shortcut. Both state 168 calendar days. One lands on 27 August and one on 18 August.

The only visible difference is how the addition was carried out. The wrong run asserted the sum in a single step. The right one walked it: 28 days left in March, then April, May, June, July, leaving 18 days into August.

I am not claiming that decomposing the addition causes the correct answer; two runs cannot establish that. But it is the only thing that changed, and it is where the error was.

Four of the five wrong answers give themselves away

None of these requires knowing the right answer.

GPT-5.4 nano, run 1 answered “Friday, 2 July 2026”. The 2nd of July 2026 is a Thursday.

GPT-5.4 nano, run 2 opened with “the day-of-week for 3 March 2026 is Wednesday”. It is a Tuesday. It then added 168 days to 3 March and got 17 September, which is 198 days. Its two runs are 77 days apart from each other, the widest spread in the trial.

GPT-5.4 mini, run 2 answered “22 August 2026”. That is a Saturday. A project whose working days are Monday to Friday cannot finish on one, so this is refuted by the prompt itself.

Mistral Small 4, run 1 wrote: “3 March 2026 (Tuesday) + 168 days lands on 27 August 2026 (Thursday).” Both weekdays are correct in isolation. But 168 is a multiple of seven, so a shift of 168 days cannot change the weekday. The sentence contradicts itself, and you can see it without opening a calendar.

That leaves GPT-5.4 mini’s first run: 25 August 2026. It is a Tuesday, it is in the right month, it looks entirely reasonable, and it is six working days too many. It is the only wrong answer here with nothing at all to catch it on, and the whole reply was the date.

So the cheap check is to make it show three things

The catchable errors all lived in the working, which means the useful move is not “ask it to explain” in general but to demand the three quantities that constrain each other:

  1. The weekday of the start date.
  2. The weekday of the finish date.
  3. The number of calendar days between them.

The start weekday and the gap together fix the finish weekday exactly, so the three cannot all be right unless the arithmetic is. In particular: if the gap is a multiple of seven the two weekdays must match, and if it is not they must differ. And a finish date on a Saturday fails before you check anything at all.

Four of the five wrong answers here violate one of those constraints in their own text. I am not claiming a measured catch rate for the protocol, because I did not run the trial that way: three of those four volunteered the contradiction unprompted, and I cannot know what the silent runs would have said if asked. What I can say is that every failure that was catchable was catchable there.

The thing that looks like a rule and is not

In this trial, every model that gave the same answer twice gave the right answer, and every model that changed its answer was wrong at least once. Stability and correctness lined up perfectly across all six.

Do not take that as a correctness check. This site has already published the counterexample: a model that returned the same wrong answer eight times out of eight, stable and confidently incorrect on every run. Asking twice catches an unstable error. It tells you nothing at all about a stable one. Twelve runs agreeing on this occasion is a coincidence, not a finding, and it is recorded in the corpus as one.

What this does not settle

Twelve runs is not a benchmark. Two runs of six models on one question on one day. It cannot rank these models and no ranking is offered here. What twelve runs can show is that a single answer from any of them is not evidence, which is the only thing the tally is used for.

One question, one shape. Working-day arithmetic against a fixed calendar. The visible failures were arithmetic slips rather than confusion about the task, and that may well be specific to this shape of problem.

Silent runs cannot be scored for comprehension. Four of the twelve, two of them wrong, printed a date and no working. The finding about setups covers only the eight that showed one.

Free tier only. These are the models a public free surface offers, several of them small or fast variants. A vendor’s larger paid model is not represented and nothing here is a claim about one.

The reading split is real. Two runs landed on 17 August, five on 18 August, five on neither, and both readings count as correct here.

Only two runs said which reading they used, both from GPT-5.4 nano, and both of those runs were wrong. Everyone else left it to be inferred from their method. If you need one convention, put it in the prompt: none of the twelve asked which you wanted.

Controls

The rubric was fixed before running. Both readings were computed and accepted as correct before the first prompt was typed, so no answer could be graded after the fact to suit a conclusion.

Ground truth is computed, not asserted. The assembler derives both readings from a date library, and it also decides every weekday, every month’s working-day count and every date-offset claim quoted from a run. The scoring is not a second opinion; it is the same arithmetic the models were asked for, run by something that does not slip.

Four things per run were entered by hand: the finish date it gave, the weekday and offset claims it made, which method it showed, and whether it said in words that the start date counts. The verbatim answer sits beside all four in the corpus, so you can check my reading of every one. Where a run showed no working, its setup is stored as unknown rather than as correct, because a bare date is not evidence either way and an earlier draft of this post got that wrong.

Model labels come from the UI, screenshotted rather than scraped, because a product can change what sits behind a label without saying so.

One New Chat per run, so no previous answer sat in context.

Nothing scripted. Twelve prompts typed by hand into a public UI on its free tier.

All twelve runs, verbatim, with every derived verdict, are in working-day-arithmetic-12-runs.json.

POSTaieveryminute.com#model-behaviourbuilt 2026-08-31 17:47 UTC