Home/Blog/The Million-Token Window
Practice guide

Several novels, one request.
Now unlearn the chunking habit.

A one-million-token context window reached general availability for Opus 4.6 and Sonnet 4.6 on 13 March 2026, at standard pricing rather than the beta premium, with the media limit raised to 600 images or PDF pages per request. It has since become the default across the current range. This is what that scale changes in practice.

Tokens and usage · 4 min read

The number is abstract until translated: a million tokens is several long novels, a genuinely large codebase, or hundreds of pages of contracts and correspondence, readable in a single request without chunking or staged summarising. For the mechanics of tokens and what they cost, start with how Claude tokens work; this piece is about what the bigger desk changes in how you work.

01

A fair amount of engineering just became optional

Teams spent years working around smaller windows: splitting documents, building embeddings search, summarising before summarising again. For a meaningful share of use cases, that machinery is now unnecessary. Not all of it, and not for every scale of corpus, but the reflex to reach for retrieval infrastructure before trying the direct approach is worth re-examining.

02

The habit is the problem, not the ceiling

Teams that learned to chunk kept chunking after the limits rose, out of habit. A model reasoning over one contiguous document catches cross-references and inconsistencies that get lost when the same material arrives as pre-summarised fragments. If your process still splits everything, you are paying the old tax without the old reason.

03

Where whole-document context genuinely wins

Document-wide consistency checks. Full-contract review, where clause 4 quietly contradicts clause 87. Whole-codebase reasoning. Long correspondence histories where the relevant detail could be anywhere. The common property: value that lives in the connections across the material, which is exactly what chunking destroys.

04

Be honest about cost

A million tokens of input is not free, and reasoning over that much material is slower than a short prompt. The answer is not “always send everything” but knowing which tasks genuinely benefit from full context and reserving the big-window approach for those. Why Claude burns through tokens covers the usage side.

05

The retirement that forced the issue

The beta version of the large window on Sonnet 4.5 and Sonnet 4 was retired on 30 April 2026, pushing anyone still on those models to upgrade to keep it. Worth remembering as a pattern: capabilities migrate from beta to standard, and the beta path then closes. Workflows pinned to old models inherit old limits.

A quick self-test

Take your most recent multi-document task and count the manual steps that exist purely to fit material into a smaller window: splitting, summarising, re-assembling. Each one is a candidate for deletion, and each deletion removes a place where meaning quietly leaked.

Book a workflow review
Questions

Things people
usually ask.

Get started

Your documents fit now.
Does your process know that?

Bring one real multi-document workflow to a session and we will strip the steps the old limits forced on you. We match you within 24 hours.

Book now →
Free · 5 minutes · no card · matched within 24 hours