CompressionStatus / live
Compressionlive
Caveman
Why use many token when few token do trick
A skill and plugin for Claude Code, Codex, Gemini, Cursor, and 30+ agents. It removes filler from model replies while keeping code, commands, and errors byte-for-byte exact. The 65% benchmark covers output only; input and reasoning tokens are unchanged.
Product demo / Caveman
illustrative local workbenchCompression workbench local demo
intensity
156 est. tok
Output · caveman49 est. tok
Auth: validate email pre-insert; reject disposable domains; keep `oauth_callback`; log `organization_id` + `request_id`; omit query/secrets; test valid/malformed/duplicate; return `invalid_email`.
illustrative rewrite · not production engine output
Token map · first 48 kept removed code
01voiceactive
02structuralavailable
03CCRunused
04recoverynot needed
Capability ledger
06- 0165% average output-token reduction across 10 prompts (22–87% range)
- 02Output only: input and reasoning tokens are unchanged, and the skill adds prompt overhead
- 03Four intensity levels: lite, full, ultra, and wenyan (classical Chinese)
- 04Works across 30+ coding agents from one skill
- 05Byte-exact recovery for engine transformations through the CCR store
- 06Structural compressors for JSON, logs, ASTs and tool schemas
Product ledger
02- Language
- Claude Code skill
- License
- MIT
Product index
0501CaveGemmaCaveman compression baked into Gemma's weights.Compression / live02Caveman CodeTerminal coding agent. Half the tokens.Agent toolkit / live03CavememPersistent memory your agents recall over MCP.Agent toolkit / live04CavekitCompressed, spec-driven development.Agent toolkit / live05Caveman ProxyThe byte-safe LLM gateway.Cloud / in development