Atlas · skill

Question Answering

Question answering produces an answer to a specific information request using a defined source of evidence. The competence is selecting extractive, generative or retrieval-based methods and testing answerability and support. A fluent answer or a plausible text span is insufficient when the available material does not actually contain the requested information.

Also searchable as: Question Answering (Q/A), question-answering

conceptText Understanding

What it is

Extractive QA selects an answer span from supplied context, often by predicting its start and end positions. Generative QA produces text and can combine or restate evidence. Open-domain systems additionally retrieve candidate sources before answering, while closed-context systems receive the material directly. These formulations have different failure sources: missing retrieval evidence, incorrect interpretation and unsupported generation. The question determines the requested relation, so locating a nearby entity is not enough. Unanswerable questions require a calibrated rejection or clarification path. Competence includes distinguishing an answer derived from the supplied evidence from information inferred or recalled by the model, and evaluating that distinction explicitly.

What the work involves

Define the question distribution, permitted sources and expected answer format. Build examples with reviewed evidence and deliberately unanswerable or ambiguous questions. Keep source documents and paraphrase families separate across training and evaluation. For extractive systems, inspect token offsets and context windows; for retrieval pipelines, measure evidence recall separately from answer quality. Assess correctness, support and appropriate abstention with task-specific matching and human review where needed. The deliverable is a QA workflow with reproducible source handling, linked evidence and a decision rule for questions that cannot be answered reliably from the available material.

Illustrative example

An illustrative manual assistant is asked which temperature triggers a shutdown. The manual states an operating range but omits the shutdown threshold. An extractive model selects the upper range value, which is a plausible number but answers a different relation. The evaluation treats this as unanswerable and tests abstention. A separate question with an explicit threshold should receive the supported value and source passage, demonstrating that rejection does not simply replace every numeric answer.

Limits and common mistakes

Span overlap metrics can reward an answer with the right words but wrong relationship. Retrieval may return topical material without the decisive fact. Generative systems can blend source evidence with unsupported recall, and long context can hide contradictions. Answerability scores need validation at the intended error cost. QA differs from broad information extraction, which fills a predefined structure across documents. Evaluate question interpretation, evidence availability, answer correctness and abstention separately, including cases where the safest correct answer is that the source does not say.

Prerequisites

  • mediumNLP

    Question interpretation and answer generation depend on language representation and evaluation.

Related skills

  • → is subcategory of: NLP

Sources and further reading

Last updated: 2026-10-10