Image Segmentation
Image segmentation assigns regions or labels at pixel level, allowing a system to describe an object's shape or a scene's composition. Practitioners choose semantic, instance or panoptic outputs, define boundary conventions and evaluate masks. The competence includes preserving spatial detail and assessing errors that a whole-image label or box cannot reveal.
What it is
Semantic segmentation labels pixels by category without necessarily distinguishing separate objects of the same type. Instance segmentation produces a distinct mask for each object, and panoptic segmentation combines instance identities with scene-region categories. Neural models commonly encode visual context and recover spatial predictions through a decoder or task head. Training needs masks, polygons or other supervision appropriate to the output. Boundary ambiguity, resolution and ignored regions affect both learning and metrics. Overlap measures such as intersection over union assess region agreement, while boundary or instance measures answer additional questions. A segmentation mask is an estimate under an annotation scheme, not automatically a physical measurement in real-world units.
What the work involves
Decide whether the application needs class area, individual objects or precise boundaries. Write mask guidelines for occlusion, holes and uncertain pixels, then inspect annotation consistency. Preserve image-mask alignment during resizing and augmentation, using label-safe interpolation. Split by source scene or specimen. Evaluate class overlap, small-region performance and boundaries where relevant, and inspect masks over original images. Convert area to physical units only with appropriate geometry or calibration. The result is an evaluated mask generator and spatial processing contract, with evidence about which types of boundary and object separation it can support.
Illustrative example
An illustrative habitat-mapping project segments vegetation in aerial photographs. The developer defines how shadows and mixed boundary pixels are labeled and holds out entire survey areas. Evaluation compares overlap and errors along narrow vegetation strips. Resizing masks with ordinary image interpolation initially creates invalid class values, so the pipeline uses a suitable label-preserving operation. Final area summaries retain uncertain regions and account for image scale instead of treating every predicted pixel as an exact ground measurement.
Limits and common mistakes
Class imbalance can make background-heavy masks appear strong while missing small regions. Overlap metrics may conceal boundary errors or merged instances. Annotation disagreement sets practical limits on apparent precision. Aggressive downsampling removes detail, and mask area is not physical area without capture geometry. Detection boxes, semantic labels and instance masks solve different tasks. Evaluate the chosen output and its downstream use directly, including empty images, tiny objects and uncertain boundaries.
Prerequisites
Related skills
- → is subcategory of: Computer Vision
Sources and further reading
- U-Net: Convolutional Networks for Biomedical Image Segmentation
Encoder-decoder segmentation and preservation of spatial information.
- Microsoft COCO: Common Objects in Context
Instance masks and detailed localization as distinct visual annotations.
Last updated: 2026-10-10