TokenMaestro
For people who build with AI

Make your AI coding plan go further. From the first request to the last token.

TokenMaestro conducts the work between you, your AI agent and your computer: a local model structures your request, it asks what’s missing before the agent starts, heavy work goes to tools installed on your machine, and each task gets the right model and effort. Every saving is measured on your machine, never estimated.

First release: the free X-ray, for Claude Code. Next: the full orchestra, and Codex, Gemini and other agents.

Free X-ray · No account · Nothing leaves your computer

Measured, never estimatedToken counts come from what the API returned, merged and de-duplicated correctly.
Cross-checkedOur count is compared with Claude Code’s own counters, and the difference is shown.
LocalRuns on your machine. No account, no upload.

The orchestra: fewer tokens from the start

Most token tools compress what the agent has already spent. TokenMaestro works before the spend: on the request, the understanding, the tools and the model.

Coming

A well-built request

A small open model runs on your computer and turns what you typed into a clear, structured request before it reaches the agent. It costs no tokens.

Coming

Questions before the work

When a request is ambiguous, TokenMaestro asks you first, so the agent doesn’t spend tokens guessing and redoing.

Coming

Local tools for the heavy lifting

Reading text in images, converting documents, cutting video: tools installed on your machine do it, set up by TokenMaestro with your permission. The agent only gets the result.

Coming

The right model and effort

Simple tasks go to a lighter model at low effort; the heavy model and high effort only when the task needs them.

Coming

Lean sessions

Compact or start fresh before the context balloons, and before a break lets the cache expire.

Free · first release

Every saving measured

The X-ray shows where every token went. Each fix is measured on your own work, on and off; if it doesn’t pay off, it is switched off.

Your agent, your choice: Claude Code first; Codex, Gemini and others next. TokenMaestro conducts; you choose the instruments.

Free

Start with the X-ray: what your plan was worth, and where it went

One command reads months of logs in seconds and writes a single dashboard file: value at API prices by day, model, effort level and project; days you hit the limit; and a cross-check against Claude Code’s own counters.

TokenMaestro dashboard: about 24 thousand dollars of API-equivalent value in five months, daily value chart with the days the plan limit was hit, and the cross-check against Claude Code’s counters.
The founder’s real dashboard: five months on a flat-rate plan, about $24,000 at API prices. The dashboard is in Portuguese today; other languages come next.

What it finds: the leaks, with a label on every number

“Measured” means the logs show it. “Ceiling” means the most a fix could recover — not a saving. In the founder’s logs:

53%

Huge contexts

of the value went to turns with more than 400k tokens of context. Each step re-reads everything before it.

7.5%

Cache rewritten after breaks

The cache expired 461 times during pauses, and the whole context was written again at full price.

~56k

Tokens before the first answer

Instructions, tools and memory every session starts with — re-read on every turn after that.

Leak cards in the dashboard: context above 200 thousand tokens, cache rewritten after breaks and without breaks, fixed weight re-read every turn, files re-read without changes, tools missing from the machine.

The weight map

What your agent read the most, by project and file. Area is the volume that entered the context; color shows how much of it was re-read with no change to the file.

Next: the same map drawn over your code’s structure, with the path your agent took in each session.

Treemap of the files the agent read the most, grouped by project.
Project and file names are replaced with neutral labels, as the X-ray lets anyone do before sharing.

Why you can trust the numbers

Most token tools report savings they estimate themselves. Independent measurements of popular ones found no savings, and sometimes higher cost. We started from the other end.

The log traps, handled

Claude Code writes several records per message, the early ones with partial counts; resumed sessions replay messages; subagents write separate files. Each case has a test.

Checked against Claude Code

On 8 comparable sessions, our cache-read count came out 4.9% below Claude Code’s own counters, never above. The numbers are a floor.

No savings claims until measured

No percentage goes on this page until a paired measurement is published with the raw data and a way to reproduce it.

Download

The X-ray is a small command-line tool. Paste one line in your terminal to install it, or download the file.

WindowsmacOSLinux

The first public release is being prepared. The download links for Windows, macOS and Linux go here when it is out.

Pricing

X-ray

Free
  • Value and usage dashboard
  • Leaks, with evidence labels
  • Weight map
  • Cross-check with Claude Code’s counters
  • Missing-tools check

Pro Coming

$79 / year

Founding members: $49 for the first year, first 500 people.

  • Structured requests from a local model
  • Questions before the agent starts
  • Local tools installed and tuned
  • The right model and effort for each task
  • Savings measured on your own work
  • Limit meter; Codex, Gemini and other agents as they arrive
Buy Pro

Purchase opens with the first public release.

Putting off one plan upgrade by a single month pays for a year of Pro.

FAQ

What does TokenMaestro read?

The session logs Claude Code already writes on your machine (~/.claude/projects): the token counts the API returned for each message, and which tools ran and how much they returned. Nothing is sent anywhere.

Is the dollar value what I paid?

No. It’s what your usage would have cost at API prices. On a flat-rate plan it shows how much your plan delivered, not a bill.

Claude Code already has /usage. Why this?

/usage tells you how much of your limit is gone. TokenMaestro tells you where it went, what it was worth, and what’s leaking.

How much will I save?

We don’t promise a number before measuring. The X-ray shows how much is leaking in your case. When fixes ship, each one is measured on your own work.

Do I need an account?

No. The X-ray needs no account and no email. Pro uses a license key, sent by email when you buy.

Does it work with Codex, Cursor or Gemini?

The X-ray reads Claude Code first. Codex and Gemini come next: you choose the agent, TokenMaestro conducts the rest.

Windows, macOS, Linux?

Yes. The X-ray runs on all three.

Do I need to install models or tools?

No. When a task calls for one, TokenMaestro shows what’s missing, explains why and installs it with your permission. The local model is small and open; if your computer can’t run it well, TokenMaestro works without it. (Pro, coming.)

Start with the free X-ray

Download the free X-ray