BLEU
BLEU is a reference-based text generation metric built from modified n-gram precision and a penalty for overly short candidates. It originated in machine translation evaluation and remains a useful reproducible baseline, while its dependence on lexical overlap limits what it can say about open-ended answers, factuality or user usefulness.
What it is
BLEU counts candidate n-grams that appear in reference translations, clipping matches so repetition cannot receive unlimited credit. It combines precision across n-gram lengths and applies a brevity penalty when candidate output is too short relative to references. Corpus-level aggregation is central to its original use; sentence-level values depend strongly on smoothing and implementation. BLEU differs from ROUGE measures that emphasize overlap recall and from embedding-based semantic metrics. Tokenization, casing, reference selection and score scaling affect comparisons, so a reported number needs a clear calculation protocol rather than just the metric name.
What the work involves
The practitioner uses a documented implementation and records tokenization, smoothing, reference handling and whether scoring is sentence or corpus based. It checks examples where valid paraphrases receive low overlap and where lexical similarity hides a meaning error. Useful artifacts include the scoring configuration, comparable system outputs and human-reviewed samples. For translation, BLEU can complement expert adequacy and fluency review. For open-ended generation, task-specific factual and completeness checks are needed. Comparisons should not mix numbers produced by different preprocessing conventions or reference collections.
Illustrative example
A translation team evaluates two systems on the same held-out source texts and reference translations. BLEU provides a consistent corpus-level overlap baseline. Reviewers then inspect a sentence where one system changes a negation but retains most words, and another expresses the correct meaning with different phrasing. The metric alone cannot resolve that quality difference. The report uses human judgments and targeted error analysis alongside BLEU, rather than turning a small score change into a claim that every translation improved.
Limits and common mistakes
Valid wording can differ from a limited reference set, and high overlap can coexist with a critical factual or grammatical error. Short outputs make sentence scores unstable, while tokenization differences can distort comparisons. BLEU should be interpreted within its reference and calculation setting. It is a measurement of lexical overlap designed for a particular evaluation tradition, not a general correctness score for language-model answers or a replacement for task-relevant human assessment.
Prerequisites
Related skills
- → is subcategory of: LLM Evaluation Design
Sources and further reading
- BLEU: a Method for Automatic Evaluation of Machine Translation
Original definition of modified n-gram precision, brevity penalty and translation evaluation setting.
Last updated: 2026-10-10