# How to Track and Cut Claude Code Token Usage

> Run /usage to see session tokens, cost and plan limits — /cost is now just an alias. Then cut waste with /clear, subagents and the right model.

- Canonical: https://guides-ai.pages.dev/guides/claude-code-track-costs/
- Plate 03.06 · Topic: Claude Code commands (https://guides-ai.pages.dev/topics/commands/)
- Published: 21 Aug 2026 · Updated: 06 Sept 2026 · 3 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

One command tells you where the tokens went. Note that **`/cost` is now an alias for
`/usage`** — same screen, so use whichever you can remember.

## 1. Open /usage

```text
/usage
```

What it does: shows the Session block — total cost, API duration, and per-model input, output
and cache token counts. The dollar figure is computed locally at list price, so treat it as an
estimate rather than an invoice.

On Pro, Max, Team and Enterprise plans the same screen adds your plan usage bars. Press `d` or
`w` to switch between the last 24 hours and the last 7 days.

## 2. Read the attribution

The plan breakdown attributes recent usage to skills, subagents, plugins and individual MCP
servers, and flags any behavior — long context, cache misses — accounting for 10% or more. It's
the fastest way to discover that one forgotten MCP server is eating a fifth of your week.

## 3. Clear rather than compact

`/clear` starts a fresh context and costs nothing. `/compact` has to read everything it
summarizes, so it is itself a large request. Session totals reset when `/clear` starts a new
session. See [managing context](/guides/claude-code-manage-context/) for when each one wins.

## 4. Match the model to the job

Sonnet handles most coding work and costs less than Opus; reserve Opus for architecture and
multi-step reasoning. Switch mid-session with `/model`. For mechanical
[subagent](/guides/claude-code-subagents/) tasks, set `model: haiku` in the agent's frontmatter.

Extended thinking is on by default and its tokens bill as output. Lower it with `/effort` for
simple work.

## 5. Push verbose work out of your window

```text
Use a subagent to run the full test suite and report only the failures.
```

What it does: keeps thousands of lines of output in the subagent's context and returns a
summary to yours. A `PreToolUse` [hook](/guides/claude-code-hooks/) that greps a log before
Claude reads it does the same job even more cheaply.

## Verify it worked

Run `/usage` at the start of a task and again at the end. If the total climbed while you barely
typed, it's long context — every turn re-sends the whole conversation, priced in
[tokens](/glossary/#token). Your first message after an hour-long break also misses the prompt
cache and reprocesses everything.

Source: [Manage costs effectively](https://code.claude.com/docs/en/costs).
