Why your tokens
disappear so fast.
You did roughly the same work as last Tuesday and hit the cap by two o’clock. The cause is almost never the thing people blame, and four of the fixes cost nothing at all.
Anthropic publishes the list of what drives consumption, and it is worth reading before you upgrade anything: message length, file attachment size, current conversation length, tool usage, model choice, effort level, and artifact creation. Seven items. Most people assume the first one is the problem. It is usually the third.
Conversation length is the multiplier
A single message is cheap. A single message at the end of a four-hour thread is not, because the entire conversation is processed again on every turn. Anthropic notes that longer conversations which trigger automatic context management consume more of your usage limit. That compounding is why consumption feels fine and then falls off a cliff, rather than rising steadily. Two people can send exactly the same number of messages in a day and burn wildly different amounts, purely because one of them started fresh threads and the other did not.
Attachments get re-read, not remembered
The 80-page PDF you attached at ten o’clock is still in the window at four. It is not filed away somewhere cheap after the first read. If you need a document available all week, put it in a project, where uploaded content is cached and only the new or uncached portions count against your limits.
Tools and connectors are expensive by design
Anthropic describes tools and connectors as token-intensive, and that is not a criticism of them. Each tool the model can reach has to be described to it, and each result comes back into the window. Connect everything you might one day want and you pay for that inventory on every turn of every conversation. Switch on the ones a given piece of work actually needs, rather than running a permanently connected estate you use twice a month.
Effort level is a dial, and most routine work does not need it turned up
Higher effort produces more thinking, and more thinking is more tokens. On a genuinely hard problem that trade is worth making. On reformatting a list, tidying an email or answering something factual, you are paying a premium for deliberation the task does not require.
Retries are the cost nobody counts
A vague first prompt that produces a near-miss, followed by three corrections, costs several times what one properly specified request costs, and the whole exchange stays in the window afterwards. Anthropic’s own advice is to batch related questions into a single message and to review the prompt for clarity before sending it. Dull guidance, large effect. The prompt that takes an extra ninety seconds to write is routinely the cheapest one you will send all day, because the alternative is four exchanges that all stay in the window afterwards.
Limits reset on a rolling five-hour window, with weekly limits on top for paid plans. At the cap you can wait for the reset, move to a higher plan, or switch on usage credits and keep working at standard API rates. Worth knowing before you upgrade in frustration: comparing what the $20, $100 and $200 plans actually give you is usually a better first move than assuming you need more.
Things people
usually ask.
More input is more tokens, but a longer prompt that avoids three rounds of clarification is cheaper overall. Specificity usually pays for itself on the first exchange.
Check thread length, attachments still sitting in the window, how many tools or connectors were active, and the model and effort level you were on. Those four explain most day-to-day variation between two apparently similar sessions.
Sometimes, and it is worth knowing your own numbers first. Settings has a usage view. If you are consistently capped after habit changes rather than before them, a higher tier is a reasonable purchase; if you have not tried the habit changes, you will hit the new ceiling too.