Atlas · skill

BERTScore

BERTScore evaluates generated text by matching contextual token representations with those of reference text. It can recognize semantic similarity beyond exact word overlap, but it measures a relationship to the supplied references rather than directly proving factual correctness, completeness or suitability for a particular user task.

conceptBenchmarking

What it is

A contextual encoder represents tokens in candidate and reference text. BERTScore matches tokens through embedding similarity and aggregates the matches into precision, recall and an F-style score, with optional weighting or rescaling depending on configuration. This differs from BLEU or ROUGE variants based primarily on lexical overlap. The encoder, layer, tokenization and reference handling influence the result. BERTScore is a reference-based generation metric, not a retrieval index or a general hallucination detector. Semantically related wording can score well even when a critical number, name or negation makes the candidate wrong.

What the work involves

The practitioner selects and records the implementation and encoder configuration, then checks metric behavior on task-specific examples with human labels. References need sufficient coverage of valid answers, and multiple-reference handling should be documented. Useful artifacts include metric settings, score distributions and error examples where semantic similarity hides a decisive difference. BERTScore can complement deterministic entity or value checks and other quality measures. Cross-system comparisons keep configuration constant, while release decisions examine whether score changes correspond to improvements users or experts actually recognize.

Illustrative example

A summarization team compares two generated summaries against reference summaries. One uses different wording but preserves the main meaning, so BERTScore provides information missed by exact n-gram overlap. Another changes a delivery date while leaving most language intact. The team checks dates separately and inspects that case rather than assuming the semantic score detects every factual error. The final comparison reports semantic similarity alongside content coverage and factual review, preserving the distinct role of each measure.

Limits and common mistakes

The encoder may poorly represent specialist language, and similar token embeddings can obscure contradiction or entity errors. Reference incompleteness can penalize a valid answer or reward copying an incomplete one. Scores are also not directly comparable across configurations or tasks. BERTScore is useful as one evaluation signal after calibration, with independent checks for facts and requirements that contextual similarity alone cannot reliably establish.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

Last updated: 2026-10-10