Home/Blog/Token inefficiency
Token inefficiency

Why your tokens
disappear so fast.

You did roughly the same work as last Tuesday and hit the cap by two o’clock. The cause is almost never the thing people blame, and four of the fixes cost nothing at all.

Tokens and usage · 4 min read

Anthropic publishes the list of what drives consumption, and it is worth reading before you upgrade anything: message length, file attachment size, current conversation length, tool usage, model choice, effort level, and artifact creation. Seven items. Most people assume the first one is the problem. It is usually the third.

01

Conversation length is the multiplier

A single message is cheap. A single message at the end of a four-hour thread is not, because the entire conversation is processed again on every turn. Anthropic notes that longer conversations which trigger automatic context management consume more of your usage limit. That compounding is why consumption feels fine and then falls off a cliff, rather than rising steadily. Two people can send exactly the same number of messages in a day and burn wildly different amounts, purely because one of them started fresh threads and the other did not.

02

Attachments get re-read, not remembered

The 80-page PDF you attached at ten o’clock is still in the window at four. It is not filed away somewhere cheap after the first read. If you need a document available all week, put it in a project, where uploaded content is cached and only the new or uncached portions count against your limits.

03

Tools and connectors are expensive by design

Anthropic describes tools and connectors as token-intensive, and that is not a criticism of them. Each tool the model can reach has to be described to it, and each result comes back into the window. Connect everything you might one day want and you pay for that inventory on every turn of every conversation. Switch on the ones a given piece of work actually needs, rather than running a permanently connected estate you use twice a month.

04

Effort level is a dial, and most routine work does not need it turned up

Higher effort produces more thinking, and more thinking is more tokens. On a genuinely hard problem that trade is worth making. On reformatting a list, tidying an email or answering something factual, you are paying a premium for deliberation the task does not require.

05

Retries are the cost nobody counts

A vague first prompt that produces a near-miss, followed by three corrections, costs several times what one properly specified request costs, and the whole exchange stays in the window afterwards. Anthropic’s own advice is to batch related questions into a single message and to review the prompt for clarity before sending it. Dull guidance, large effect. The prompt that takes an extra ninety seconds to write is routinely the cheapest one you will send all day, because the alternative is four exchanges that all stay in the window afterwards.

When you do hit the wall

Limits reset on a rolling five-hour window, with weekly limits on top for paid plans. At the cap you can wait for the reset, move to a higher plan, or switch on usage credits and keep working at standard API rates. Worth knowing before you upgrade in frustration: comparing what the $20, $100 and $200 plans actually give you is usually a better first move than assuming you need more.

See how coaching works
Questions

Things people
usually ask.

Get started

Same work, half the usage.
Usually two habits.

A specialist will watch how you actually use Claude for an hour and show you where the consumption is going. Most people leave with two changes and a noticeably longer day. We match you within 24 hours.

Book now →
Free · 5 minutes · no card · matched within 24 hours