Skip to main content
All projects
Automation & AI
/projects/claude-superpack

Claude Superpack

  • 9 skills, -73% always-on context (2,753 -> 742 tokens)
  • 15 on-demand domain references, 4 deterministic scripts
  • 148/148 structure checks, 100% routing coverage
  • Published to npm with GitHub Actions OIDC provenance
project.claude-superpack.local
Claude Superpack project screenshot

Problem: Skill packs for Claude Code fail in one of two directions—stay small and carry no real engineering judgment, or load so much always-on context that every turn pays for domain knowledge used once a month, with no mechanism to verify a claimed fix actually works.

Solution: Rebuilt from a 33-skill always-loaded pack into a 9-skill risk-tiered engineering OS that sizes every task on two independent axes (scope and risk), pulls in 15 domain references only when a task needs them, and runs 4 deterministic Node scripts instead of asking the model to eyeball diffs, secrets, and build logs.

Impact: Cut always-on context 73% (2,753 -> 742 tokens, cross-checked against Claude Code's own estimator) while adding capabilities v4 never had—a debugging skill, an evidence gate before any completion claim, and a high-precision credential scanner—then shipped it as a public npm package with GitHub Actions OIDC trusted publishing.

Overview

claude-superpack is an engineering operating system for Claude Code: nine always-on skills classify every request on two independent axes—scope (how much work) and risk (what happens if it's wrong)—then route into fifteen domain references and four deterministic scripts that load only when the task actually needs them. One rule overrides everything else: no completion claim without evidence produced after the change. Published to npm with GitHub Actions OIDC trusted publishing so every version is verifiably built from the repository.

Before / after

Always-on context

Before: ~2,753 tokens across 33 always-loaded skills (v4)

After: ~742 tokens across 9 skills (v5), cross-checked against Claude Code's own estimator

Completion claims

Before: Model asserted a fix worked based on its own read of the diff

After: Evidence gate blocks any completion claim without tool output produced after the change

Codebase awareness

Before: Persistent JSON knowledge graph under ~/.claude/graphs/, which went stale silently

After: Blast radius computed fresh from the diff and grep every time—slower per call, never confidently wrong

Cross-session memory

Before: Five skills wrote to ~/.claude/memory/, duplicating dedicated memory plugins

After: Removed entirely; defers to claude-mem or Claude Code's native memory, keeps in-task state in durable plan files

Stack

Two-axis task sizing: Scope (S0-S3) x Risk (R0-R3) sets planning depth and gate count independently
9 skills: 4 content-activated (debugging, securing, shipping, requirements-grilling) + 5 phase-activated, reached from a core gate table
15 domain references (frontend, backend, data, ai-llm, cloud, devops, sre, security, performance, testing, architecture, quality, docs, parallel, routing), each structured Decide first / Build right / Failure modes / Evidence
4 zero-dependency Node scripts: recon.mjs, gates.mjs, diffstat.mjs, secrets.mjs
Published as @sahilbnsll/claude-superpack via GitHub Actions OIDC trusted publishing with signed provenance attestation
Three-tier benchmark suite (structure, routing vocabulary, paired live evaluation) gates every release via npm run bench

Decisions

Key trade-offs and design calls that shaped the final delivery.

Two independent axes instead of one A/B/C/D tier

Context: v4's single complexity tier conflated "big" with "dangerous"—a 1,000-line refactor of a test helper and a one-line IAM policy change got identical treatment

Decision: Split into Scope (S0-S3, sets planning depth) and Risk (R0-R3, sets gate count), so a trivial edit gets one tool call and a production/IaC change doesn't execute until blast radius is stated out loud and confirmed

Delete the persistent memory system and codebase graph

Context: v4's ~/.claude/memory/ and ~/.claude/graphs/ cost context every session, duplicated dedicated memory plugins, and depended on the model reliably self-reporting its own state—a paired evaluation found that compliance happened once in eighty trials despite the instruction reaching the model every time

Decision: Removed both. Blast radius is now computed fresh from the diff and grep on every task; cross-session recall defers explicitly to claude-mem or Claude Code's native memory instead of reimplementing it

npm publishing via GitHub Actions OIDC, not a stored token

Context: A long-lived npm token sitting in CI is a standing credential that can leak or go stale

Decision: Trusted publishing authenticates through GitHub Actions' OIDC identity—no npm token exists to leak—and every version carries a signed provenance attestation tying it to the exact workflow run that built it

Architecture

The primary system boundaries, runtime pieces, and how the project was structured in production.

superpack skill (SKILL.md + gate table)

Core Routing & Sizing

Classifies every request on scope and risk, then routes to phase-activated skills via an explicit gate table—the only way those skills can be reached, since nobody types "I would like a blast-radius check now."

15 on-demand markdown files

Domain References

92KB of engineering depth that costs zero always-on tokens and loads only when the router selects it, versus being baked into an always-loaded skill description.

recon.mjs / gates.mjs / diffstat.mjs / secrets.mjs

Deterministic Scripts

Replaces model guesswork with real tool output—secrets.mjs found 4 real credentials across 1,094 changed lines in 40ms with zero false positives; gates.mjs turned 15,247 bytes of build output into a 288-byte digest that preserved the evidence line.

Mermaid source. Paste into mermaid.live to visualize the diagram.

flowchart TD
  U[User request] --> C[superpack: classify scope x risk]
  C -->|content-activated| CA[debugging-systematically / securing-changes / shipping-safely / grilling-requirements]
  C -->|gate table| PA[codebase-recon -> planning-changes -> verifying-evidence -> reviewing-before-done]
  PA --> DR[15 domain references, loaded on demand]
  PA --> SC[recon.mjs / gates.mjs / diffstat.mjs / secrets.mjs]
  SC --> EV[Evidence gate]
  EV --> DONE[Completion claim]

Pipeline

How changes moved from development through validation and deployment.

1

Verify

npm run bench

Three-tier benchmark (148 structure checks, routing vocabulary coverage, paired live-evaluation harness) must pass before any release; prepublishOnly re-runs it with --no-write so published tarballs stay reproducible from their commit

2

Publish

GitHub Actions OIDC

Three independent jobs (verify, npmjs, github-packages), each skipping a version that already exists on that registry, so re-running a release is always safe

3

Attest

npm provenance

Every published version carries a signed attestation tying it to the workflow run that built it, verifiable via npm view @sahilbnsll/claude-superpack dist.attestations

4

Install

npx @sahilbnsll/claude-superpack install

No install scripts and no global install required—current npm blocks postinstall scripts by default, so a global install previously downloaded the package and silently skipped installing the skills

Incidents

Operational failures, rehearsals, or recovery moments that changed how the system was run.

Global install silently skipped installing the skills

P1

Resolution: Removed postinstall/preuninstall entirely as a breaking change (v5.0.2); install is now an explicit `npx ... install` step that exits non-zero on failure instead of a lifecycle script that swallowed errors

Lesson: A package that writes into the user's home directory merely by being downloaded is a supply-chain smell regardless of whether it currently works—npm blocking install scripts by default was the signal this needed fixing, not the cause of the bug

Every skill description loaded twice

P2

Resolution: v4 installed each skill both at ~/.claude/skills/<name>/ and again as a plugin-shaped copy underneath, roughly doubling always-on token cost to ~5,600 tokens/turn; the installer now installs one copy, deletes any leftover duplicate, and `claude-superpack doctor` detects the condition going forward

Lesson: An installer needs an install manifest, not an assumption that its own previous version installed cleanly

Published tarballs weren't reproducible from their commit

P3

Resolution: prepublishOnly ran the benchmark suite, which rewrote a results file with a fresh timestamp on every publish, so the packed tarball never matched the commit and left the working tree dirty; prepublishOnly now runs with --no-write

Lesson: A pre-publish hook that mutates the working tree undermines the exact guarantee—build reproducibility—that provenance attestation exists to provide