aieveryminute

A trending 817-skill pack costs about 56,600 tokens before you type anything

mukul975/Anthropic-Cybersecurity-Skills was installed and measured 24 times on Claude Code 2.1.236. In an empty, isolated project it adds about 56,600 tokens of startup context, 69.30 per skill on the dedicated harness's ten rounds. This site's own synthetic skills curve predicts 13.46 per skill at that count. Two candidate explanations were tested here and refuted, and the cause is not established.

mukul975/Anthropic-Cybersecurity-Skills trended on 2026-08-20 and ships 817 Agent Skills mapped to MITRE ATT&CK, NIST CSF and four other frameworks. It is 7.4x larger than the biggest real pack this site had measured.

It is not an Anthropic repository. The name says otherwise and it is what a lot of people will find first, but the marketplace manifest names an individual owner.

What it costs

Population: the pack exactly as it ships, copied into an otherwise empty project’s .claude/skills/, on one machine, on Claude Code 2.1.236. Estimator: the mode of the per-round deltas from the dedicated harness, ten rounds, each against a floor re-measured in the same round.

Tokens Per skill
817 skills, as it ships 56,622 69.30

Ten rounds, spread zero: the absolute session context read 79,401 every single time against a floor of 22,779.

Two further harnesses re-measured the same pack for other reasons and got modes of 56,592 and 56,602. Twenty-one of the twenty-four readings sit inside a 48-token band, 56,592 to 56,640.

That is a repeatability check and not corroboration. All three harnesses were written within nine minutes of each other and all three take their number from one shared measurement call. It tells you three wrappers around one call reproduce within each harness, not that three instruments agree, and not that the instrument is reliable. On the identical 22,779 floor two of them read 79,401 and 79,381 absolute, so most of the 30-token mode spread is arm-side rather than floor-side. And twice the same fixture collapsed to 40,133, which is the anomaly below.

This is an isolated empty project. --setting-sources project is asserted onto every command line, so nothing from the machine’s own configuration is in it. It is not a claim about what the pack costs on top of whatever you already have installed.

The gap against this site’s own curve

This site publishes a synthetic skills curve, and the calculator interpolates it:

Skills Tokens Per skill
measured cells 800 / 1,000 10,927 / 11,711 13.66 / 11.71
interpolated to 817 817 10,994 13.46

Against the real pack that is 5.15x low. Note what that does and does not indict: 69.30 sits inside this site’s own real-pack range of 34.08 to 274.0, and google/skills at 81.26 per skill on 111 skills with a 397-character description median is the near-exact control. This site’s real-pack measurements never disagreed with this result. The synthetic curve is the outlier.

The curve is not stale. curve-stability-2-1-234.json re-ran its 1,000-skill cell on 2.1.234 and got a mode of 11,701 against a published 11,711, with a spread of 897 across eight rounds, about 8% of the mode. One caveat that belongs here: that re-measurement covered two cells, and the other one, at 10 skills, had its estimator refused for a spread 281% of its own mode. The 1,000-skill end reproduces. The small end does not.

The calculator now discloses this above two hundred skills.

Two explanations, tested and refuted

Every arm is 817 skill directories against its own round’s floor; the real-description arm registers at most 795 of them.

Arm Tokens Verdict
Synthetic, all descriptions the same string 9,952
Synthetic, distinct strings, 30-word vocabulary 11,047–11,108, estimator refused not uniqueness
Real descriptions, everything else synthetic 44,572 see below
Real files, extra frontmatter deleted 56,600 not frontmatter

None of these steps is a clean one-variable change, and I am not going to pretend otherwise. The two synthetic arms differ in repetition, in vocabulary and in syntax at once: the repeated arm is one 80-character noun phrase repeated about five times and cut mid-word, the distinct arm is random words from a disjoint thirty-word list. The step from row three to row four changes bodies, reference files, scripts, 22 descriptions and all 817 skill names together — and names matter here, because the whole delta is a name-and-description listing.

Description length is not varied anywhere in this corpus, so nothing here refutes it. The synthetic arms hold every description at a constant 396 characters, matching the pack’s median; that makes them a matched control, not a length test.

Description uniqueness is not the answer. Making all 817 descriptions distinct moved the reading from 9,952 to between 11,047 and 11,108. The guarded estimator refused a single figure there because no round repeated, so that cell is published as a range.

The extra frontmatter is not charged on these rounds, which is what Anthropic’s documentation says. These files carry sixteen frontmatter field types — domain, subdomain, tags, mitre_attack, nist_csf and more — with a 322-byte median for the extra keys alone (806 bytes for the whole frontmatter block). The docs say Level 1 loads name and description only. So I deleted every other key from the real files, 338,442 bytes, holding bodies and reference files byte-identical and verifying per round that both arms carried the same 817 names and the same 817 descriptions.

Stripped minus full is +8 tokens, and it is not noise: the full arm read 79,373 absolute and the stripped arm 79,381, each reproducing exactly in all four usable rounds. Six rounds ran; two were destroyed by the anomaly below, which landed on one side of this very comparison. Deleting 338KB of text moved the reading eight tokens in the wrong direction. That is a deterministic offset, not a cost.

What I cannot explain

The arm carrying the pack’s real descriptions in an otherwise minimal skill measured 44,572, against about 11,000 for synthetic descriptions at the same count and the same median length. The distributions are not the same: the synthetic arms hold a constant 396 characters, while as installed the real arm is 795 descriptions running 97 to 743 with a 395-character median, plus 22 at zero. As the pack ships they run 97 to 1,016 and none is zero; the 22 that break the fixture are the long ones. The obvious story is that real words cost more than repeated ones.

I am not publishing that, because neither fixture is carrying its descriptions whole, which makes the two arms useless as a test of what words cost.

The synthetic arms carry 817 × 396 characters: 40,850 words in the repeated arm and 43,684 to 43,773 across rounds in the distinct one. Any BPE tokeniser charges at least one token per word, so if that text were fully in the measured context neither arm could cost under 40,850 tokens. They measured 9,952 and 11,047–11,108. The synthetic descriptions cannot be fully in context, by a factor of about four.

Applied to the other arm the same bound is much tighter, and it took two drafts to get it right. The 795 real descriptions that survive that arm’s fixture are 300,742 characters and 38,829 words, so the floor is 38,829 tokens — and the arm measured 44,572, which is above it. On the word count alone, nothing is excluded.

It only bites once the listing overhead comes out, and this corpus supplies that number. The synthetic identical arm cost 9,952 at the same 817 skills, the same probe-skill-NNNN names and the same 200-byte bodies, with its descriptions provably not fully present, by about four times. So at most 9,950 of the 44,572 is overhead, which leaves at least roughly 34,600 tokens to carry 38,829 words.

So the real descriptions cannot be fully in context either — provided the listing overhead at 817 skills exceeds 5,743 tokens. No arm here measures that overhead directly, so this is a conditional bound rather than a proof. It is a mild condition: every synthetic cell this site has at that scale sits far above it, 10,994 interpolated at 817 and 10,462 at 600 skills. But it is a condition, and the margin is about 11% rather than the 4x the synthetic arms give.

So two fixtures matched on skill count measured 34,620 tokens apart, and neither can be carrying its descriptions whole. What that gap is made of is not established here. The control that would test vocabulary alone — synthetic descriptions drawn from a large vocabulary at constant length — was not run. That is the next measurement.

That arm also has a fixture defect, reported rather than buried. Rewriting each description as an unquoted single-line YAML scalar makes 22 of the 817 files invalid YAML, because those descriptions contain a colon followed by a space. Those 22 lose both name and description on a parse. They are the long ones, mean 553 characters, and they account for all 12,163 missing characters exactly. Whether Claude Code drops them from its listing was not measured. No share of the total is claimed from that arm.

The anomaly

Two of the twenty-four readings came in at 17,352 instead of about 56,600, both in one harness and both in its full arm — which is one half of the frontmatter comparison above, 2 of 6 there against 0 of 6 in stripped.

It is not floor drift: those rounds carried the normal 22,781 floor, and it is the arm’s absolute context that fell, 79,373 to 40,133. Arm order within a round is fixed and both landed in the first slot, but 2 of 6 against 0 of 6 does not establish a position effect and I am not claiming one.

A third reading, 57,489, is not an anomaly but is worth explaining, because it is a control failing. In that round neither arm moved at all — full read 79,373 and stripped 79,381, exactly as in the clean rounds — and only the floor dropped, by 897. The subtraction pushed that 897 into both published deltas. Pairing added variance there rather than removing it, which is the same failure mode this site reports from a different corpus, measured on 2026-08-18.

One clean negative

Zero of the 817 descriptions exceed the documented 1,024-character maximum. The longest is 1,016.

So the cut this site measured — where a description stops reaching context between 1,400 and 1,550 characters — cannot bite this pack, because nothing here is long enough to reach it. Worth stating, because Anthropic’s own repository exceeds the documented 1,024 limit in two of its five plugins and this third-party pack does not. That cut was located with one skill installed; at 10 and 40 skills only its lower bound was re-checked. It is not re-tested here.

What to actually do

Do not install 817 skills. The pack is one skill per directory. Copy the ones you need.

But do not price the subset by multiplying. What twenty of them cost was not measured, and count-times-rate does not transfer here: this site’s published per-skill rates run from about 11.71 (the 1,000-skill synthetic cell) to 274 (a one-skill pack), and they straddle 69.30 in both directions. Even at a fixed 300-character description they bracket it, 93.62 per skill at forty skills and 16.14 at six hundred.

Treat the calculator’s skills curve as a floor above a couple of hundred skills. That rests on one real pack at one count. The calculator says so.

Security scanned before installing

817 third-party files entering an agent’s context. No plugin hooks, no lifecycle or postinstall scripts. 1,093 Python scripts ship inside skills, all 1,093 unique rather than templated.

Every danger-grep hit was read in context and was a false positive: eval( matches are YARA rules and Splunk SPL query functions, curl | sh matches are grype and syft install lines plus detection rules that flag that pattern, base64 decoding is JWT parsing and PowerShell deobfuscation, /etc/shadow and keychain are forensic analysis targets, and every AWS key is AWS’s own documented example key, AKIAIOSFODNN7EXAMPLE, or a honeytoken decoy. The grep carried a positive control that fired 2 of 2, so the zero means clean rather than broken. No skill body is executed by the harness.

Controls

Paired, with a stated failure. Floor re-measured every round; every delta is against that round’s floor. In one round that pairing injected 897 tokens of floor noise into both arms, as above.

Interleaved, with a stated limit. Compared arms ran inside the same round, so between-round drift cannot land on one arm. Position within the round is not controlled, and the synthetic-versus-real comparison spans two harnesses and is not interleaved at all.

Arrival. Every reply had to be exactly OK, published per run, 80 of 80.

Isolation. --setting-sources project asserted onto the command line in every call.

Fixture, with a stated limit. Skills counted off disk every round, and the frontmatter arms hash both arms’ name and description sets. That count asserts 817 SKILL.md files exist; it does not assert 817 registrable skills, and it did not catch the 22 broken ones.

Nothing invoked. A trivial prompt; no skill body executed or read.

What this does not settle

The cause. Two candidates eliminated here, description string-uniqueness and extra frontmatter. Description length was refuted by an earlier corpus at 600 skills, not by these arms. The remaining candidate, the description text itself, is not supported by these arms.

Anything about other packs. The six real packs this site has measured run 34.08 to 274.0 tokens per skill and do not order by count. This is one more point, not a trend.

Anything on a configured machine. Every reading here is isolated and empty.

Small skill counts. Every cell is 817 skills.

Body size. No arm here isolates it: the arm with real 9,949-byte median bodies also carries real descriptions. This site’s earlier finding that bodies are free until invoked is carried forward unretired rather than re-verified.

Long descriptions. desc-limit-2-1-235.json measured 0.015 tokens per character above 500 characters, at 40 skills. Nothing here bears on it, and the arithmetic above says descriptions are not fully charged at all at 817 skills. Carried forward unretired and untested.

All 30 rounds, 80 runs and five arms are in cybersec-817-2-1-236.json.

POSTaieveryminute.com#context-costbuilt 2026-08-31 17:47 UTC