FRDownload

How do you make your Claude Code limits last?

Anthropic’s advice, sorted by impact. The first one alone often changes a whole day.

1. Start fresh between tasks

Anthropic is blunt: /clear “is the single most effective lever for both quality and cost”. Every message resends the whole context, so a conversation that drags on costs you on every turn. /clear costs nothing.

2. The right model for the task

“Opus costs several times more per turn than Sonnet, and Sonnet more than Haiku.” Anthropic recommends planning with Opus and executing with Sonnet, which the /model opusplan alias does. Our model guide covers which model suits which task.

3. Less useless context

4. Compact, at the right time

/compact summarizes the conversation to free up context. But “compacting a large context is itself a large request”. By default, Claude Code compacts on its own when the conversation nears the model’s limit (around 967K tokens on 1M-token models). /context shows what fills the window.

Caching helps too: on a subscription it lasts an hour. Switching models midway forces everything to be recomputed.

Frequently asked questions

/compact or /clear?

/clear when you move to another task: it is free and the most effective. /compact when you want to keep the thread of a long task, knowing the summary itself uses usage.

Does switching models mid-conversation cost anything?

Yes: Anthropic says switching models recomputes the whole request, even with identical content, because the cache does not carry over.

When does Claude Code compact on its own?

By default, when the conversation reaches the model’s context limit, around 967K tokens on 1M-token models. /autocompact lets you change that threshold.

Read next

Sources