Natural Language Processing & Computer Vision
41 skills · ontology graph below shows relations within this section.
What this domain covers
This edition groups 41 capabilities in Natural Language Processing & Computer Vision across 5 named categories. The inventory contains 29 concepts and 12 tools. Open an entry for its mechanism, practical workflow, example, limitations, and primary references.
Current category labels: Audio & Speech · Computer Vision · Multimodal Vision · NLP Foundations · Text Understanding
Frequent learning foundations
- NLP supports 6 mapped skills
- Linear Algebra supports 2 mapped skills
- Tokenization supports 2 mapped skills
- Text Classification supports 1 mapped skill
- Unsupervised Learning supports 1 mapped skill
Skills in this section
Audio AI applies learned models to sound, including speech, music and environmental events. The competence is choosing a meaningful audio task, representing recordings correctly and evaluating predictions or generated sound under realistic conditions. It spans several objectives rather than treating every recording as a speech transcription problem.
Computer vision extracts useful information from images or video through geometry, signal processing and learned models. Practitioners translate a visual question into an observable output, select suitable capture and annotation methods and test reliability under changing scenes. The competence extends from image preparation to evaluating complete perception systems.
Object detection identifies instances of target classes and estimates where they occur in an image. The skill is defining consistent object categories and boxes, selecting a suitable detector and evaluating both localization and missed or spurious detections. It produces instance locations rather than only a label for the whole scene.
OpenCV is a computer vision library for image and video operations, geometry and model inference. The competence is combining its functions into correct visual pipelines while managing array formats, coordinates and execution costs. Familiarity includes understanding the assumptions of each operation rather than only knowing how to display an image.
Vision-language models connect visual inputs with language representations or generated text. Practitioners select a model suited to retrieval, classification or visual dialogue and test whether its outputs are grounded in the image. The skill includes preprocessing, prompt design and evaluation that separates visual evidence from plausible linguistic guesses.
Natural language processing turns written or spoken language into representations, predictions or generated text for a defined task. The competence is identifying what linguistic information the application needs, selecting suitable data and methods and evaluating errors in context. It combines language-aware problem formulation with reproducible computational workflows.
Tokenization converts text into units that a language-processing system can represent and use. The skill is choosing or applying the correct segmentation scheme, preserving useful boundaries and checking model compatibility. Word tokens, subword tokens and byte-level units have different purposes and should not be treated as interchangeable linguistic objects.
Multilingual NLP builds language-processing systems that operate across multiple languages or transfer learning between them. Practitioners assess language coverage, representation quality and task behavior separately for each population. The competence includes scripts, morphology and code-switching, rather than assuming a model's multilingual label establishes equal performance everywhere.
Named entity recognition identifies text spans that refer to defined categories such as people, organizations or domain-specific items. The skill is designing an annotation scheme, locating exact boundaries and evaluating span and type errors. Recognizing a mention is separate from linking it to a database record or resolving every reference.
Semantic search retrieves material using learned representations intended to capture meaning beyond exact keyword overlap. The competence is selecting embeddings and similarity measures, constructing a useful index and evaluating retrieval against real information needs. Similar wording or a high vector score does not by itself establish that a result answers the query.
Audio processing prepares, transforms and measures sampled sound for analysis or playback. The competence is managing sampling, channels, levels and time-frequency representations while preserving the information required by the next stage. It supports audio AI workflows but also includes deterministic operations that do not recognize words or sound events.
ElevenLabs provides audio generation and transcription services with APIs and model-specific settings. The competence is selecting the appropriate service, managing voice and format configuration and checking the quality of returned audio or transcripts. Product familiarity includes current limits and reproducible requests, rather than assuming every model supports identical controls.
Librosa is a Python library for audio and music analysis, including loading, spectral features and timing estimates. The skill is selecting explicit sampling and frame parameters, interpreting outputs in physical units and validating results by listening and inspection. Its numerical features are useful inputs, rather than automatic evidence of musical or semantic understanding.
Speech recognition converts spoken language in audio into text. Practitioners select an acoustic and decoding approach, prepare recordings correctly and measure transcription errors under the expected voices and environments. The competence includes proper evaluation of names, numbers and timing, rather than assuming a fluent transcript faithfully captures the recording.
Text-to-speech converts written content into a spoken audio signal. The competence is controlling pronunciation, voice and delivery while preserving the intended words, then assessing intelligibility and playback behavior. Naturalness, timing and factual correctness are different qualities, so pleasing sound alone is insufficient evidence of a successful synthesis workflow.
Whisper is a family of speech models and an open-source implementation for transcription and related audio-language tasks. The competence is choosing the checkpoint and decoding mode, preparing audio correctly and testing long-recording behavior. Its outputs need comparison with the recording, especially for silence, names and unsupported or difficult speech.
Detectron2 is a framework for building and evaluating visual recognition models, especially detection and segmentation workflows. The competence is configuring models and datasets correctly, interpreting structured predictions and evaluating the complete pipeline. A working pretrained demo is a starting point for task adaptation, rather than proof of suitability for new images.
Emotion recognition assigns emotion-related labels to observations such as text, voice or facial behavior. The competence is defining exactly what the labels represent, validating the measurement and communicating uncertainty. Classifying an annotated expression is different from establishing a person's inner emotional state, intentions or psychological condition.
Facial recognition compares facial images to support identity verification or identification. The competence is designing the matching task, controlling capture quality and selecting thresholds using relevant error evidence. Detecting a face or estimating landmarks is a separate operation, and a similarity score alone does not establish a person's identity.
Image classification assigns one or more category labels to an image or defined crop. The skill is designing meaningful classes, selecting compatible input processing and evaluating mistakes under realistic visual variation. It answers what category the input belongs to, without necessarily locating every object or explaining the evidence used.
Image segmentation assigns regions or labels at pixel level, allowing a system to describe an object's shape or a scene's composition. Practitioners choose semantic, instance or panoptic outputs, define boundary conventions and evaluate masks. The competence includes preserving spatial detail and assessing errors that a whole-image label or box cannot reveal.
MMDetection is an OpenMMLab toolbox for configuring, training and evaluating object detection and related instance recognition models. The competence is navigating its configuration and data pipeline, selecting compatible dependencies and checking annotations and metrics. It enables reproducible experiments across model families rather than representing one particular detection algorithm.
MediaPipe provides components and task APIs for running machine learning perception in applications, including image and video workflows. The competence is selecting a supported task, managing model assets and runtime modes and interpreting results correctly. Application behavior also requires timing, coordinate and uncertainty handling beyond obtaining a prediction.
Object tracking associates observations of an object across time to estimate its continuing position or trajectory. The skill is combining detection, motion and appearance evidence while managing missed observations and identity changes. Tracking adds temporal association to perception; accurate detections in individual frames do not automatically yield reliable tracks.
YOLO refers to a family of visual recognition models and implementations originating in direct object detection. The competence is selecting a specific version and task, preparing its annotations and evaluating the exported model. Different YOLO releases and libraries have distinct architectures, supported outputs and operating requirements.
dlib is a C++ library with Python interfaces for machine learning, numerical operations and computer vision components. The competence is selecting the appropriate algorithm or pretrained asset, managing image and coordinate conventions and evaluating its output. A face-related example demonstrates a component, rather than establishing a complete identity or interpretation system.
Gensim is a Python library for corpus representations, vector models, similarity and topic-oriented text analysis. The competence is building consistent dictionaries and corpus transformations, selecting a representation suited to the question and interpreting results with independent checks. A learned topic or neighboring word is an analytical pattern rather than a verified semantic fact.
NLTK is a Python toolkit for language analysis, linguistic resources and educational NLP workflows. The competence is applying its tokenizers, corpora and linguistic algorithms with an explicit task and evaluation. Resource availability supports exploration, but the outputs still depend on language, preprocessing choices and the assumptions of each component.
spaCy is a library for building language-processing pipelines with tokenization, trained linguistic components and rules. The competence is selecting compatible language assets, configuring dependencies and inspecting structured annotations. It supports efficient application workflows, but each component and rule still needs evaluation on the language and domain where it will be used.
Information retrieval finds and ranks material that addresses an information need within a collection. The competence is defining relevance, indexing suitable units, choosing ranking methods and evaluating results with representative queries. Retrieval returns candidate evidence or documents; it does not automatically synthesize or verify an answer.
Intent detection identifies the action or goal expressed by a message so a system can choose an appropriate next step. The skill is defining an actionable intent inventory, separating similar requests and handling ambiguity or unsupported goals. Predicting an intent label is distinct from authorizing or executing the requested action.
Natural language understanding maps language to interpretations useful for a task, such as intent, semantic relationships or evidence-supported answers. The competence is specifying the intended meaning operation and testing it under ambiguity and context changes. A benchmark score or fluent response does not establish unrestricted understanding of language.
Summarization produces a shorter account of source material while preserving information needed by a particular reader or task. The skill is defining coverage and length priorities, choosing extractive or generative methods and checking faithfulness. A concise, fluent summary can still omit a decisive caveat or introduce an unsupported statement.
Text classification assigns a predefined category or set of categories to a text unit. Practitioners define labels, select representations and evaluate prediction errors in the application context. The competence includes class imbalance, ambiguous examples and unsupported categories, rather than assuming every text can be forced into one useful label.
Word2Vec learns dense word vectors from patterns of neighboring words in a corpus. The skill is choosing the context objective and corpus preparation, training or selecting embeddings and assessing their usefulness for a downstream task. Vector proximity reflects distributional usage and should not be treated as a complete definition of meaning.
TF-IDF weights text features using their frequency within a document and rarity across a corpus. The competence is building a stable vocabulary and weighting scheme, preserving sparse representations and evaluating their usefulness for search or classification. It highlights discriminative lexical evidence without learning a full semantic representation.
Topic modeling discovers recurring patterns of terms or document representations in a corpus to support exploration and organization. The competence is choosing a model, preparing meaningful document units and evaluating topic interpretability and stability. Learned topics are analytical constructs, rather than automatically correct labels or explanations of why people discuss a subject.
Text preprocessing prepares source text for a particular language-processing task through selected cleaning, normalization and structural transformations. The competence is deciding which changes preserve useful evidence and keeping them reproducible. It precedes or surrounds tokenization, but should not be reduced to a universal recipe that strips every text in the same way.
Sentiment analysis identifies evaluative polarity or opinion expressed in text, often about a particular target. The competence is defining what is being evaluated, labeling mixed or indirect opinions and measuring errors in context. It describes expressed appraisal, rather than establishing the author's emotional state or the truth of a statement.
Information extraction converts source content into structured entities, attributes, relations or events. Practitioners define a schema, connect outputs to evidence and evaluate both missing information and unsupported fields. The competence is broader than recognizing names, and differs from question answering because it produces a reusable structured representation under an extraction contract.
Question answering produces an answer to a specific information request using a defined source of evidence. The competence is selecting extractive, generative or retrieval-based methods and testing answerability and support. A fluent answer or a plausible text span is insufficient when the available material does not actually contain the requested information.