Token Optimization
Token optimization reduces unnecessary input or output tokens while preserving the information and behavior an application needs. It includes selecting context, shortening repetitive instructions, controlling generated length and choosing suitable representations; the target is useful task completion per resource spent, rather than the shortest possible prompt.
What it is
Language models process text as tokens, whose count depends on the tokenizer rather than a simple word count. Input tokens consume context capacity and processing work, while generated tokens add decoding time and often dominate interactive latency. An application can remove irrelevant passages, avoid duplicating history, summarize completed work or request a bounded output format. These changes may alter behavior, so token savings are an engineering hypothesis to evaluate. Prompt caching addresses repeated computation, whereas token optimization changes how much material is processed or generated.
What the work involves
The practitioner profiles token use by call and identifies expensive patterns such as oversized retrieval batches or repeated tool schemas. A before-and-after evaluation records task success, omitted facts, latency and usage. It is useful to preserve a budget for essential exceptions instead of applying indiscriminate truncation. Deliverables include a context selection policy, output length rules and a token breakdown linked to task outcomes. For multi-call workflows, the total includes retries and verification calls, not merely the final response.
Illustrative example
An assistant summarizes a long meeting transcript. Sending every earlier conversation turn adds cost but little evidence, so the application keeps the user's requested decisions and the transcript, removes unrelated chat and asks for decisions with owners and unresolved questions. It then checks summaries against annotated meetings. If shorter inputs cause a disputed decision to disappear, the selection policy is revised even though the token count had improved. Savings remain subordinate to the required coverage.
Limits and common mistakes
Compression can erase qualifiers, provenance or rare but decisive facts. A short output limit can produce a polished yet incomplete answer, while overly narrow retrieval can force another costly call. Token counts also vary by language, content and model tokenizer. Optimization should compare end-to-end resource use at an acceptable quality level. A reduction measured on easy examples may fail when a task needs long evidence or several rounds of clarification.
Prerequisites
Related skills
- ← is part of: Prompt Caching
- → is part of: AI FinOps
- → is part of: Context Engineering
Sources and further reading
- Latency optimization
Discusses reducing generated tokens and improving application latency through request design.
- Effective context engineering for AI agents
Supports selecting and compressing context while preserving task-relevant information.
Last updated: 2026-10-10