aieveryminute

What is your setup costing you?

What a Claude Code session costs, at startup and across the whole run, from measured figures.

These figures predate the release you are running. Claude Code 2.1.251 is installed; these figures were last measured on an earlier version and have not been re-checked on this release, longest lag first: commands and subagent definitions (2.1.232, 19 patch releases behind), parallel subagents (2.1.238, 13 patch releases behind), path variation in the floor (2.1.241, 10 patch releases behind). Every other figure on this page was measured on 2.1.251. The floor has moved by as much as 5,594 tokens between two consecutive releases measured here, so treat the stale rows as indicative until they carry the current version.

then, across the session
tokens before your first message
tokens sent across the whole session

What each thing costs

These are the per-unit figures the calculator interpolates between. The two baselines were measured directly on Claude Code 2.1.251, the subagent line was re-measured on 2.1.238, the skills curve on 2.1.251, the MCP curves on 2.1.251, and the command and subagent-definition curves on 2.1.232:

The two baselines also depend on your project's directory path, which Claude Code puts in the system prompt. Measured on 2.1.241 across ten paths, the tool-search-on floor spanned 82 tokens. Those absolute readings, 22,765 to 22,847, no longer apply: the floor itself fell 5,594 tokens on 2.1.251, so treat the SPREAD as the transferable part and not the endpoints. It tracks the path's tokens rather than its length: two paths of identical 67-character length differed by 14 tokens (also2.1.241), and a 38-character path of many short segments cost 20 tokens more than a 47-character path of one repeated character. That last comparison wasre-measured on 2.1.251 across twelve rounds and the gap is still exactly 20, with both paths having fallen by exactly the same 5,727, so the release change shifted this term rather than rescaling it. The figures below were measured in a short sandbox path, so a deeper or more word-rich project directory will read a little higher. The whole range is about a third of one percent, so it matters for reproducing a measurement rather than for your bill. All 48 runs are in path-floor-2-1-241.json.

ItemCost
Claude Code itself, tool search on17,207 tokens
Claude Code itself, tool search off32,400 tokens
A user-level CLAUDE.md0.37 tokens per byte, always
CLAUDE.md106 tokens + 0.26 per byte, every session (see the note below on markdown-structured files)
One skill, below about 200 of them39 tokens, at a 90-character description
One skill, past about 400 of them3.9 tokens
A skill description, while the listing is small0.276 tokens per character, to 300
A skill description past 300 characters0.363 to 0.482 tokens per character, measured on real English descriptions
A skill's body, any size up to 229KB0 tokens
One slash command, below about 20040.0 tokens
One slash command, past about 4005.6 tokens
One subagent definition51 tokens, at every count measured
One MCP tool, deferred~16 tokens
One MCP tool, loaded upfront652 tokens, up from 304 on 2.1.224
One parallel subagent20,581 to 28,504 tokens + about 1.93x your configuration again

The CLAUDE.md rate is low for markdown-structured files

The 0.26 per byte above was fitted on flowing English prose. On 2026-08-24 a currently-trending CLAUDE.md was measured directly and unchanged at 2,357 bytes: theandrej-karpathy-skills file, which is markdown-structured with headers, bold runs, bullet lists and a fenced code block rather than paragraphs. Across 15 rounds and 120 runs it cost 982 tokens at the mode, on 38 of 60 paired differences, which is 0.417 tokens per byte. The row above predicts 719 tokens for that file, so the calculator is 263 low on it.

Two qualifications, both honest. The paired difference is multimodal: it takes 982 on most pairs and wanders on the rest, and the full observed range spans both this page's prediction and the 0.2759 rate measured on prose, so neither is refuted by that data. And one file does not re-fit a curve that was fitted from 2KB to 28KB. The rate above is left unchanged and this note is the disclosure. Receipts:the corpus andthe pre-registration, which recorded both predictions before anything ran.

One scope note, because it changes who this applies to. That repository is not a single file: it ships the same guidance as a Claude Code plugin, as a Cursor rule and as a project-root CLAUDE.md, and its README recommends the plugin first. The 2,357 bytes measured here are the CLAUDE.md install, which is the option the repository lists second. What the plugin path costs was not measured.

The skill description rate was fitted on English

The 0.276 tokens per character row above was fitted on English descriptions, and it does not transfer. On 2026-08-27 nine skills were installed one at a time on 2.1.247, four rounds each, every delta measured against a floor re-measured in the same round. Descriptions containing no CJK cost 0.363 to 0.482 tokens per description character. Descriptions that are majority Chinese cost0.887 to 1.048. The arms do not overlap, and the two character-length ranges do overlap (English 198 to 485, Chinese 167 to 518), so length is not carrying the difference.

Judged only on the cells inside the 30 to 300 character range this row was declared valid for, the row lands within 17 tokens either way on English (predicting 94 against 77 measured, and 108 against 120). On Chinese it is 90 low on a 167-character description and 180 low on a 289-character one. The cleanest single comparison sits inside one pack by one author, which holds style and source constant: its 455-character English description costs 185 tokens and its 518-character Chinese one costs 497, so 1.14x the characters cost 2.69x the tokens.

Three qualifications, all honest, and the first one matters most. Measured perbyte of UTF-8 rather than per character, the two armsoverlap: English 0.363 to 0.482, Chinese 0.459 to 0.513. A Chinese character is about 2.3 bytes, and that accounts for most of the per-character gap. So this is not evidence that the tokeniser handles Chinese badly per unit of encoded text. It still costs you, because the 1,024-character description cap countscharacters and this row prices per character, so a Chinese description reaches that cap having spent roughly twice the tokens. Second, six of the nine cells sit outside the declared 30 to 300 range, so the full arm-level miss is not attributable to the tokeniser alone, and only the four in-range cells above are quoted against the row. Third, no cause below the tokeniser is claimed. The rate is left unchanged and this note is the disclosure, the same way the markdown-structured note above is handled. Receipts:the corpus, which publishes both units per cell and every round, including the three that lost one to this harness's intermittent component.

And the practical version, measured. This page prices skills percharacter and CLAUDE.md per byte, so the same content can look cheaper or dearer depending only on which row you read. That gap was closed on 2026-08-27 by measuring it directly: one translation pair, the same project instructions in English and Chinese, byte sizes matching within 1.4% while character counts differ by 2.8x. Across 16 self-paired rounds the English file cost 604 tokens and the Chinese one678, which is 1.12x. So switching your CLAUDE.md between these two languages is not a cost lever in either direction, and both popular intuitions are wrong by a wide margin: a per-character reading suggests 2 to 3x, and a character-count reading suggests 0.36x against 678 measured. The arms overlap round to round, so that is a modal difference rather than a resolved effect.The corpus andthe pre-registration, which recorded all three predictions before anything ran.

What this does not cover is a skill's body. Bodies are free until the skill is invoked, and none was invoked here, but a body is typically far longer than a description, so the same per-character penalty would land on a much larger number. That is unmeasured.

Your baseline includes your global config, not just Claude Code

A session loads your user-level CLAUDE.md as well as the project one. On the machine these figures were measured, that global file was 72,783 bytes and cost26,763 tokens, measured directly and reproduced twice. It was making the "empty" baseline roughly 156% larger than Claude Code's own floor, against the 17,207-token figure in the table above.

So the two are separated here. The base figure is Claude Code by itself, and your global config is an input like any other. If you do not know how big yours is, check~/.claude/CLAUDE.md, or justmeasure your actual project, which captures everything at once.

Why the session total dwarfs the startup context

A tool call is two API calls: Claude Code sends the conversation, gets back a request to run a tool, then sends the whole conversation again with the result appended. So each round trip costs roughly your entire current context, and the context grows as results accumulate.

That makes session cost grow faster than linearly with the number of round trips. The model used here sums the context sent on every trip; it predicts themeasured runs to within 0.45%. It also means startup context is not paid once, it is paid on every round trip, which is the strongest reason to keep it small.

The numbers age differently, and it matters

The per-unit costs are paired deltas: measured by taking a baseline and a configuration back to back and subtracting. Baseline drift cancels out, which is why the MCP figures came back identical to the token across a version bump.

The two baselines are not deltas, so they move whenever the system prompt does, and unlike a delta nothing cancels out. They shifted by several hundred tokens between 2.1.223 and 2.1.224. They were also wrong here until 2026-08-13, and by far more than drift: both were about 9,000 tokens too high, because they had been derived by subtracting one user-level file from a loaded machine rather than measured on a machine with nothing user-level loaded at all. Both are now measured directly, twice per round across four rounds, with 0 to 4 tokens of noise. Treat the per-unit deltas as solid and the absolute totals as good to within a percent or so.

How the estimate is built

Linear interpolation between measured points, never a fitted curve, because a curve would invent precision the runs do not support. Beyond the largest measured point the final marginal rate is extended, and the output says so.

CLAUDE.md is the exception: it is linear rather than interpolated, because it measured 0.26 tokens per byte to three decimal places across a 14x range in file size. Seewhat CLAUDE.md costs. Below about 2KB the figure is derived from that rate rather than measured, since the baseline drift is larger than the signal at that size.

Measured points for skills: 10, 40, 100, 150, 200, 300, 400, 600, 800 and 1,000. For MCP tools: 5, 50 and 200, in both deferred and loaded modes. Working fromthe skills measurements andthe MCP measurements.

Subagent definitions never get cheaper. Skills and slash commands both cost about 40 tokens each and then collapse to under 6 once you pass a few hundred. A subagent definition costs 51 at every count measured, so it is the one part of a plugin where the count really is the thing to watch. See what a plugin's parts cost. Commands and subagent definitions are not included in the parallel-subagent term below: that model was fitted against a configuration containing neither, so folding them in would extrapolate it onto inputs it was never calibrated with.

The CLAUDE.md rate is the weakest constant on this page. It is one number, 0.26 tokens per byte, and token density is a property of the text rather than the file: the same size measured 0.2844 on ordinary prose and 0.368 on a real user-level config. Treat the CLAUDE.md rows as accurate to about 30%, not to the decimal.

Count every skill, not just your project's. A skill costs 39 tokens while your listing is small and 3.9 once it is large, with the change falling somewhere between 200 and 400 skills, so the number that decides your price is the total across your project, your user scope and every plugin you have installed. Enter that total. The curve is measured at 90-character descriptions; longer ones cost more, at 0.276 tokens per character, until the listing is large enough that they stop being charged at all.

Do not extend 0.276 past 300 characters. That rate was fitted between 30 and 300, and real descriptions routinely run longer. Past the bound, the measured figure is0.363 to 0.482 tokens per character on five real English descriptions each installed alone, spanning 198 to 485 characters. A398-character description measured on 2.1.251 cost 159 tokens, which is 0.400 per character and sits inside that band; extending 0.276 instead would have predicted 134 and under-budgeted by 25. A description in Chinese costs roughly twice as much per character again, because the cap counts characters rather than bytes.

MCP figures assume a tool definition of roughly 798 bytes. A tool with a much larger parameter schema costs proportionally more when loaded upfront, and about the same when deferred.

CALCaieveryminute.comcalculatorbuilt 2026-08-31 17:47 UTC