aieveryminute

The same forty skills cost 1,558 tokens, or 152

On a clean project a skill costs 39 tokens until you have about two hundred, then 3.9. On a machine with its own configuration loaded it costs 3.8 from the smallest block measured. Same version, same fixture, same round, 10.2x apart, and the obvious explanation for it turned out to be false.

This site has published, since 2.1.223 and re-verified on 2.1.224, that skills are almost free: 1,000 of them cost 1,461 tokens, about 1.5 each, and that description length is free, because forty skills with 1,500-character descriptions cost the same 80 tokens as forty with 90-character ones.

Measured on a clean project on 2.1.231, neither reproduces. Forty skills with 90-character descriptions cost 1,570 tokens, and the description is most of that.

I nearly published that as a version change. It is not one, and the thing that stopped me was a control I nearly did not run.

The same forty skills, twice, in the same round

The clean-project sweep found the marginal cost of a skill falling from 39.0 tokens to 3.9 somewhere between 200 and 400 skills. The earlier work ran without --setting-sources project, on a machine with its own configuration loaded, so I ran both conditions in the same round, on the same version, with the same fixture. The only difference is whether that configuration was allowed to load.

40 skills added Tokens Per skill
Clean project 1,558 38.95
Machine’s own config loaded 152 3.80

10.2x, from nothing but the flag. At 150 skills it is the same story: 5,848 against 579.

So the earlier figure was measured on a configured machine, which sits in the cheap regime. That is not a different release. It accounts for the order of magnitude and not for all of it: in that regime the same counts measure 152 and 3,875 here against the published 80 and 1,461, so those figures are still about twice too low, and that residual is unexplained.

The explanation I had, and why it is wrong

The obvious reading is that the machine already has enough skills of its own to be past the collapse, so its listing is compacted before you add anything. I wrote that down, and then counted.

It loads 65 skills: 29 in user scope and 36 from ten installed plugins. A raw find ~/.claude/plugins -name SKILL.md returns 213, which is the number I had, but most of that is marketplace clones and cached copies of versions that are not the installed one, and none of those load.

65 is not past the collapse. It is well below it, in the range where the clean sweep charges 39 tokens a skill, ten times what the configured arm actually measures.

So the hypothesis is refuted by its own machine. Loading a configuration moves the price of a skill by 10.2x, and the pre-existing skill count is not what does it. Dropping that flag also loads a user-level CLAUDE.md, plugin commands, plugin agents, MCP servers and hooks. Any of them could be responsible. I am not going to guess which, and I am leaving the refuted version here rather than deleting it, because it is the part of this a reader most needs to be able to distrust. (Measured the next day: slash commands are one of them, and this machine ships 82. What a plugin’s parts cost.)

On a clean project, the count really is the variable

Hold the description at 90 characters and sweep the count. Measured between adjacent points:

Between Tokens per extra skill
10 → 40 39.0
40 → 100 39.0
100 → 150 39.0
150 → 200 39.0 or 35.0
200 → 300 10.6
300 → 400 5.9
400 → 600 3.9
600 → 800 3.9
800 → 1000 3.9

Nothing varies down that table except the count. Flat at 39.0 to 200, then flat at 3.9 from 400 on.

The 150 → 200 row carries two values rather than one because the 200-skill cell splits two-two on a 198-token component that fires on some calls and not others, giving 39.0 in two rounds and 35.0 in the other two. Its median, 37.0, is a number no round produced, so it is not printed here as a measured dip. There is no evidence the cost departs from 39.0 before 200.

Those are rates across intervals of thirty to two hundred skills. No individual position is measured, and the single-skill cell is pure noise.

What a description costs, at forty skills

Hold the count at forty and vary the description:

Description length Tokens per skill Rate over the previous step
30 chars 21.4
90 chars 39.4 0.300 /char
300 chars 96.4 0.271 /char
800 chars 178.6 0.164 /char
1500 chars 169.8 −0.013 /char

A least-squares fit over the three points up to 300 characters gives 0.276 tokens per character on a fixed 13.7 per skill, R² 0.9997. It is linear only over that range. Between 300 and 800 the rate falls to 0.164, and between 800 and 1500 it goes negative: the 1,500-character arm lands 351 tokens below the 800-character arm and the two do not overlap across four rounds. Growth does not merely stop, it reverses slightly. I have no explanation and am not offering one.

Applying the fitted rate across the whole table would overstate by 31% at 800 characters and 152% at 1500, which is why the range is stated and not the coefficient alone.

Past the collapse, a longer description costs no more

The tempting story is that the price falls between 200 and 400 skills because descriptions stop being carried. That is a mechanism, and the sweep cannot test it: description length was only ever varied at forty skills. So I varied it at 600, with the forty-skill contrast running in the same rounds as a positive control.

30-char descriptions 300-char descriptions Difference
40 skills, the control 844 3,646 or 3,844 2,802 to 3,000 more
600 skills 10,462 9,684 778 less

At forty skills, 270 extra characters cost 70 to 75 tokens per skill. Charged that way at 600 skills, the long-description arm would be 42,000 to 45,000 tokens more expensive. It is 778 cheaper. Neither pair overlaps across four rounds.

Corrected 2026-08-22: this row originally read 3,745, a difference of 2,901, and 72.5 tokens per skill extrapolating to about 43,500. The four rounds of that cell came in at 3,646, 3,646, 3,844, 3,844, a two-two split with nothing between, so 3,745 is the midpoint of two modes and no round produced it. A median is robust to an outlier and not to bimodality at even n. The range replaces it. The finding does not move: the extrapolation and the measurement stay an order of magnitude apart under every value in the range.

So a 300-character description costs no more than a 30-character one once you have 600 skills, and the positive control is what makes that believable: without it, “the arms landed together” and “the harness stopped varying the description” look identical. This tests two lengths at one count. It does not map the boundary, and it does not say whether some intermediate length is charged at some intermediate count.

The reversal, longer descriptions coming in slightly cheaper rather than equal, is 778 tokens on a 10,462-token cell. Direction reported because it is measured, unexplained.

What this changes

If your machine carries a real configuration, skills are almost free. The configured arm here pays 3.8 tokens a skill. The old advice survives for you very nearly as written, for a reason nobody has established.

On a clean project they are not. Forty skills with 90-character descriptions cost about 1,570 tokens every session, roughly 7% of a 22,643-token floor. Small, but twenty times what this site told you.

Below the collapse, description length is the lever, not skill count. Forty skills at 300 characters cost 3,856 tokens against 856 at 30. Write descriptions long enough to match on, and know that down there you are paying for every character of them.

Reference material still belongs in a skill body. That is the one thing that did not move under either condition: the same 28,000 bytes costs 7,379 tokens in CLAUDE.md and stays free in a skill body.

The body is still free

The original work’s signature result is that a skill body is free: it loads only when the skill is invoked, which is the whole reason reference material belongs in one.

That holds:

40 skills, body size Tokens
200 B 1,574
4,000 B 1,576
28,000 B 1,576

The 200 B and 4,000 B arms overlap across rounds, so the honest statement is that no effect above roughly 200 tokens is detectable across a 140x range in body size. That still rules out the thing that matters: 28,000 bytes charged at CLAUDE.md’s rate would be about 7,300 tokens, and it is not there.

What shows the fixture is registering as skills at all is not this null but the level. 1,574 tokens for forty directories is not what nothing costs. The null itself has no positive control, since no cell here charges for a body, so on its own it could not tell a free body from a fixture that never loaded one.

What this does not settle

Why a configured machine is cheap. Measured, not explained here. Partially answered the next day: slash commands are drawn against the same thing skills are, and this machine ships 82 of them against its 65 skills, so the denominator was more than twice what it was first counted as. Subagent definitions and CLAUDE.md are not in it, and 147 entries is still under the 200 where the clean curve starts to bend, so that narrows the gap rather than closing it. See what a plugin’s parts cost.

The last factor of two. The published figures are 80 tokens for 40 skills and 1,461 for 1,000. The configured arm measures 152 and 3,875 for the same counts in the same regime. A residual of roughly 1.9 to 2.6 is unattributed: it could be version, fixture detail, a different configuration on the day, or a different usage field being read.

The transition is bracketed, not located. Between 200 and 400 skills on a clean project, with no mechanism claimed for it.

One skill is inside the noise. That cell measured -149 tokens with a 521-token spread, so there is no claim here about a single skill, and none about any single ordinal position.

One description shape. Every description is cut from one paragraph of ordinary English. A description of code or symbols tokenises differently and 0.276 per character will not transfer. A first pass built descriptions from a single repeated character and measured 106 tokens per skill instead of 39, because a run of one character tokenises far worse than prose. That pass was discarded.

90 characters is a fixture, not a typical description. Nothing here measures how long real descriptions are.

One machine, one model. Opus on 2.1.231, project-scope skills, and the configured arm is specific to this machine’s installed configuration.

Controls

Both conditions in the same round, same version, same fixture, with the isolation flag asserted present on one arm and asserted absent on the other, so the treatment is proved on both sides rather than only one.

A positive control on the one null that carries a claim: the 30-versus-300 contrast at 40 skills, run in the same rounds as the 600-skill cells, separating by 2,802 to 3,000 tokens.

Paired floors, re-measured every round, each arm against its own. Within each corpus the clean floor was identical in all four rounds. Across corpora it is not: the three runs published here measured the same clean floor, same version, same flag, at 22,643, 22,528 and 22,520, and the floor corpus behind the calculator puts it at 22,404. That 239-token spread is why nothing is measured against a floor from another session, and it is why the same cell reads 1,570 in one corpus and 1,558 in another.

Two intermittent components, both disclosed. A 198-token component fires on some clean-arm calls: it splits the 40, 200, 300 and 800-skill cells and the two smallest body cells. A 150-token component does the same on the configured arm, where in one round it was absent from the floor and the 1,000-skill cell while present in the other two, moving those two deltas by exactly 150. The estimator is a median throughout because of these.

Arrival, 132 of 132. Every reply had to be exactly OK.

The fixture is counted, not assumed: skills are counted back off disk on all 132 runs. Bytes written are recorded on the 80 clean-sweep runs, and the description-at-scale runs additionally read the description back off the written file and assert its length.

No exclusions. All 132 runs are scored.

Every run is published: the clean-project sweep, the regime control and the description-at-scale test.


Re-verified 2026-08-15 against Claude Code 2.1.233. The figure did not move, and the noise did.

Four rounds, same harness, same three arms. Ten of the eleven count cells returned the same value in all four rounds. On 2.1.232 not one of them did: every cell carried a spread of 198, 396 or 719 tokens, which is one, two or three firings of the intermittent component this site has been caveating since the first skills measurement. On 2.1.233 that component does not appear in this harness at all. The one cell still moving is 1,000 skills, at 521.

Every point on the curve is unchanged to the token except one. At 10 skills the mode went from 202 to 400, and it moved by exactly 198 — the component’s own size. On 2.1.232 that cell read [202, 202, 400, 598]; on 2.1.233 it reads [400, 400, 400, 400]. Whether the cell genuinely costs 400 now, or whether the component became constant there rather than disappearing, is not established by this and I am not going to guess.

The part worth taking away is that the per-skill cost and the startup floor moved independently over the same release. Between 2.1.232 and 2.1.233:

2.1.232 2.1.233 change
Per skill at 200 skills 39.05 39.05 none
Tool-search-off floor 40,192 36,941 −8%
Tool-search saving 17,664 14,644 −17%

So a release that moved the floor by 8% and the tool-search lever by 17% left what a skill costs completely untouched. If you are estimating, those are separate numbers with separate release histories, and re-checking one tells you nothing about the other.

Every run: skills-cost-2-1-233.json, and the previous release’s for comparison at skills-cost-2-1-232.json.

POSTaieveryminute.com#context-costbuilt 2026-08-31 17:47 UTC