Atlas · skill

Token Optimization

Token optimization reduces unnecessary input or output tokens while preserving the information and behavior an application needs. It includes selecting context, shortening repetitive instructions, controlling generated length and choosing suitable representations; the target is useful task completion per resource spent, rather than the shortest possible prompt.

conceptInference Efficiency

What it is

Language models process text as tokens, whose count depends on the tokenizer rather than a simple word count. Input tokens consume context capacity and processing work, while generated tokens add decoding time and often dominate interactive latency. An application can remove irrelevant passages, avoid duplicating history, summarize completed work or request a bounded output format. These changes may alter behavior, so token savings are an engineering hypothesis to evaluate. Prompt caching addresses repeated computation, whereas token optimization changes how much material is processed or generated.

What the work involves

The practitioner profiles token use by call and identifies expensive patterns such as oversized retrieval batches or repeated tool schemas. A before-and-after evaluation records task success, omitted facts, latency and usage. It is useful to preserve a budget for essential exceptions instead of applying indiscriminate truncation. Deliverables include a context selection policy, output length rules and a token breakdown linked to task outcomes. For multi-call workflows, the total includes retries and verification calls, not merely the final response.

Illustrative example

An assistant summarizes a long meeting transcript. Sending every earlier conversation turn adds cost but little evidence, so the application keeps the user's requested decisions and the transcript, removes unrelated chat and asks for decisions with owners and unresolved questions. It then checks summaries against annotated meetings. If shorter inputs cause a disputed decision to disappear, the selection policy is revised even though the token count had improved. Savings remain subordinate to the required coverage.

Limits and common mistakes

Compression can erase qualifiers, provenance or rare but decisive facts. A short output limit can produce a polished yet incomplete answer, while overly narrow retrieval can force another costly call. Token counts also vary by language, content and model tokenizer. Optimization should compare end-to-end resource use at an acceptable quality level. A reduction measured on easy examples may fail when a task needs long evidence or several rounds of clarification.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

Last updated: 2026-10-10