Atlas · skill

Prompt Engineering

Prompt engineering is the design and evaluation of instructions, examples and output requirements used to guide a language model on a task. It turns a vague request into an explicit interaction contract, then tests whether that contract produces useful behavior across representative inputs and failure cases.

conceptPrompt Design

What it is

A prompt influences generation by describing the task, supplying relevant material and indicating how an answer should be formed. It may include demonstrations, delimiters, a requested structure or rules for handling insufficient evidence. The model still uses its existing weights; changing a prompt is different from training or fine-tuning. Techniques such as few-shot examples and intermediate reasoning can help particular tasks, but they are choices to compare rather than mandatory ingredients. Context engineering is broader: it also decides which history, retrieved documents and tool results become available alongside the instructions.

What the work involves

The practitioner starts with success criteria and an evaluation set. A first prompt states the task plainly, defines important terms and makes uncertainty or escalation behavior explicit. Revisions follow observed failures: an example may clarify a boundary, while a missing input may require a tool rather than more prose. Prompts and model settings are versioned together. A useful deliverable includes the prompt, test cases, scoring criteria and a record of which changes improved performance without introducing unacceptable regressions.

Illustrative example

An assistant converts maintenance notes into structured repair summaries. The initial instruction produces fluent summaries but sometimes invents a replacement part. A revision requires every part to be supported by the note and permits an unknown value. Tests include incomplete notes and notes mentioning a part that was inspected but not replaced. The team compares results before deployment. If the remaining problem is missing access to the parts catalog, it adds a controlled lookup rather than repeatedly rewriting the wording.

Limits and common mistakes

A convincing example is weak evidence of general reliability. Prompts can overfit a small test collection, contain contradictory instructions or rely on model-specific behavior. They do not establish security permissions or guarantee factual correctness. Asking for extensive reasoning can increase latency without improving a task. Prompt quality is therefore measured against application outcomes, and meaningful model or input-distribution changes require renewed evaluation of the instruction contract.

Prerequisites

  • Chain-of-Thought and reasoning techniques exploit how Transformers process sequential tokens — understanding attention helps design better prompts

Related skills

Sources and further reading

Last updated: 2026-10-10