Object Detection
Object detection identifies instances of target classes and estimates where they occur in an image. The skill is defining consistent object categories and boxes, selecting a suitable detector and evaluating both localization and missed or spurious detections. It produces instance locations rather than only a label for the whole scene.
What it is
A detector predicts class scores and spatial regions, usually bounding boxes, for a variable number of objects. Two-stage systems propose regions and classify or refine them; one-stage systems predict detections more directly from image features. Postprocessing filters low-confidence candidates and may suppress overlapping duplicate predictions. Training needs instance annotations whose coordinates follow a precise convention. Evaluation matches predictions with reference objects using overlap criteria and then measures precision and recall across confidence thresholds. Object detection differs from segmentation, which assigns detailed pixels, and from tracking, which associates instances over time. Its output supports counting or localization only within the defined categories and annotation rules.
What the work involves
Write class and box-boundary guidelines, including partially visible objects and difficult examples. Audit annotation completeness and convert coordinates carefully after resizing or cropping. Split data by scene, location or video sequence, not adjacent frames. Compare detector quality by object size, occlusion and class, using average precision plus application-specific false-positive and missed-object costs. Set thresholds on development data and measure latency with postprocessing included. Inspect overlays to distinguish localization errors from label errors. The deliverable is an evaluated detector and documented operating threshold whose spatial outputs match the downstream coordinate system.
Illustrative example
An illustrative shelf-auditing task detects individual packages. The developer labels boxes consistently when one package hides part of another and holds out entire stores. Two models have similar overall precision, but one misses small packages on distant shelves. Evaluation includes that subgroup and counts duplicate boxes after suppression. A coordinate test projects resized-image detections back onto the original photographs, catching a scaling error before the boxes are used for inventory review.
Limits and common mistakes
Detection confidence is not automatically a calibrated probability, and overlap metrics may not reflect downstream counting errors. Unannotated objects can make correct predictions appear false, while incomplete labels corrupt training. Small, occluded or unfamiliar objects are common failures. Suppression can remove legitimate nearby instances. A successful detector does not establish object identity over time or detailed shape. Evaluate class coverage, spatial accuracy and threshold behavior separately, including scenes that contain no target objects.
Prerequisites
Related skills
- → is subcategory of: Computer Vision
- ← is an instance of: YOLO
- ← is subcategory of: Object Tracking
- ← is an instance of: Detectron2
- ← is an instance of: MMDetection
Sources and further reading
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Two-stage detection through shared features and learned region proposals.
- Microsoft COCO: Common Objects in Context
Instance annotation and detection evaluation in complex scenes.
Last updated: 2026-10-10