Atlas · skill

Supervised Fine-Tuning (SFT)

Supervised fine-tuning adapts a pretrained model using examples of the desired prediction or response. Practitioners choose targets, control which outputs contribute to the loss and check whether imitation generalizes to unseen inputs. Instruction-response training is one application of SFT, alongside classification, extraction and other labeled tasks.

conceptFine-Tuning Methods

What it is

Supervision supplies target outputs rather than only a preference between candidates or a reward for sampled actions. For an autoregressive language model, a common objective minimizes the negative log likelihood of target tokens given the context. Depending on the task, loss may cover a full text sequence, a completion or selected assistant turns. Other pretrained architectures use supervised task losses such as classification. Teacher forcing trains on reference prefixes, whereas deployment generates its own prefixes. SFT therefore teaches patterns represented by demonstrations, with behavior determined by their correctness, coverage and the loss assigned to different parts of each example.

What the work involves

Define the input-output contract and obtain reviewed targets that match it. Include realistic ambiguous and missing-input cases rather than only ideal answers. Inspect token-level masking, padding, packing and truncation to confirm that the intended target survives preprocessing. Split by underlying source or task family, reserve untouched evaluation examples and monitor overfitting. Measure task success on generated or predicted outputs, with format and factuality checks when appropriate. The deliverable is a model that reproduces the required behavior on unseen cases, accompanied by the exact target construction and loss configuration.

Illustrative example

An illustrative extraction model receives a note and must output a JSON object. The training record includes both the prompt and correct object, but the loss is applied to the object tokens. The developer checks examples where a field is absent and expects a null value. Evaluation uses notes from unseen reports and measures field correctness as well as parse validity. This is supervised fine-tuning even though the goal is a narrow extraction task.

Limits and common mistakes

The model can learn an answer's style without learning its correctness, and frequently occurring targets may dominate rarer requirements. Teacher forcing does not expose every error chain that occurs during generation. Incorrect masking can spend capacity predicting prompts, while truncation can remove the supervised answer. SFT does not directly express comparative preferences or guarantee safe behavior. Instruction tuning is a particular data and task formulation within the broader supervised adaptation approach, rather than a synonym for every SFT run.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

Last updated: 2026-10-10