Caveman Cuts AI Output Tokens by 65%. It Saved Me Almost Nothing.
A tool that shrinks your agent's replies by 65% sounds like an obvious install. Then you look at which side of the bill it actually touches. Input tokens, output tokens, and why the same tool can be a win for one setup and a loss for another.
I run most of my life through an AI setup now. Job search, three side products, family health tracking, a memory system that remembers what I told it six weeks ago. All of it costs tokens, and tokens cost money, and money is a thing I watch closely these days.
So when I found Caveman, I actually got a little excited. The pitch is great: make your AI coding agent talk like a caveman, same answers, 65% fewer output tokens. One command to install, works across Claude Code, Cursor, Gemini, thirty-odd agents. No backend, no telemetry, nothing phoning home. I installed it the same evening.
Then I did the thing I do for a living, which is read past the headline number. The honest answer is that Caveman is a good tool that will do very little for me. Not because it is broken. Because of how I use my AI, which turns out to be the whole game.
What it actually is
Caveman is not a proxy or a compression engine sitting between you and the model. It is a skill. A prompt. It tells your agent to drop the filler, write in fragments, skip the “Great question! Let me help you with that” throat-clearing, and keep code, commands, and error messages exactly as they are.
That is it. And that restraint is the smart part. It does not touch what the model can do. It only touches how much it says.
The benchmarks back the claim. Across ten coding tasks the average output shrank 65%, with a range from 22% on the low end to 87% on the high end. A React re-render explanation went from an essay to four lines. That is a real result, and if you have ever watched an agent write six paragraphs to tell you to add a key prop, you know exactly why people want this.
Where the number gets slippery
Here is what made me put the calculator down and actually think.
A model turn has two sides. There is what you send in, which is your files, your context, the diff, the previous conversation. And there is what the model sends back. Input tokens and output tokens. On most pricing, input is cheaper per token, but there is usually a lot more of it.
Caveman only shrinks the output side. It says so itself, and I respect that the project spells this out in its own honest-numbers notes instead of hiding it. The input tokens do not move. The reasoning tokens do not move. And the skill adds roughly 1,000 to 1,500 input tokens every single turn, because the instructions have to ride along each time.
So the 65% is a real cut on one slice of the bill. The question is how big that slice is for you.
For a lot of coding work, the input slice is the big one. You are feeding the model a repo, a stack trace, a long back and forth. The reply is often the small part. Cut 65% off the small part, add a fixed cost to the big part, and the whole-session saving comes out modest. On workloads that are already terse, the project admits it can even go slightly negative.
Why my setup is the worst case for it
This part is specific to me, and probably to anyone running a personal AI OS instead of one-off coding sessions.
My system injects context on every prompt. Before I even ask a question, a hook pulls the five most relevant memories out of a vector database and loads them in. Then my instruction file loads. Then my memory index loads. Every turn carries this weight, and it is all on the input side, which Caveman does not touch at all.
The output, meanwhile, is usually the smallest thing in the exchange. I ask a sharp question, I want a sharp answer, I have already trained the thing to be blunt with me. Caveman would be squeezing juice from the one part of my setup that is already dry, while quietly adding its 1,500-token tax to the part that is actually heavy.
For me the math does not just underwhelm. It runs the wrong way.
The one feature I am still tempted by
There is a piece called caveman-compress that rewrites your saved memory or context files into a denser form, and claims around 46% smaller input on future sessions. That one is interesting to me, because it hits input, which is my actual cost center.
But I am holding off, and here is the caution if your memory is anything like mine. My files are not just prose. They are wired into a semantic search that runs on embeddings, they carry structured links between notes, and they are full of exact figures I cannot afford to have rounded or dropped. Embedding models also read natural sentences better than caveman fragments, so compressing the text can quietly make retrieval worse even when the file looks fine to a human. Compress for readability and you can lose fidelity you did not know you were spending. If I ever run it, it goes on one throwaway copy first, with a careful diff after.
Who should actually install this
I do not want this to read as a takedown, because it is not one. Caveman is well built, honest about its own limits, and solves a real annoyance. Let me just be clear about the fit.
Install it if you do a lot of interactive coding, you are on output-priced or per-message limits, and you are tired of scrolling through padding to find the one command you needed. The speed and readability alone might be worth it, forget the cost. That is a genuine everyday win.
Skip it, or at least run the numbers first, if your setup is input-heavy. Long context windows, RAG pipelines, big system prompts, a memory layer that reloads every turn. That is where the fixed per-turn tax eats the output savings, and where you will stare at your bill next month wondering why it looks the same.
The tool is not the variable. Your token profile is. Caveman just made me go and actually look at mine, which is the most useful thing a tool that saved me nothing has done in a while.