01 — Understand the task
Segmentation produces a dense prediction at pixel level.
Classification assigns one category to a full image. Detection locates an object with a bounding box. Semantic segmentation assigns a class to every pixel, while instance segmentation also separates distinct objects from the same class. Two shoes in one scene therefore receive two different masks.
Fully Convolutional Networks (FCNs) established an architecture producing class maps at image resolution. Mask Region-based Convolutional Neural Network (Mask R-CNN) later added a mask branch to an object detector. YOLO instance-segmentation models pursue the same goal with an architecture designed for fast inference.
02 — Product hypothesis
Removing scenery can shift embeddings toward product shape and appearance.
An embedding model transforms the entire image into a vector. In fashion photography, that vector may encode the garment, model pose, studio, lighting and background. Such context is sometimes useful; it can also group products by art direction instead of cut, material or colour.
I therefore explored YOLOv8-seg on prototypes produced by clustering. The mask supports several preparations for comparison: full image, bounding-box crop, cut-out product on a neutral background and cut-out product with a margin of context.
03 — Protocol
Evaluate the mask and the retrieval engine within the same experiment.
Segmentation quality can be measured with Intersection over Union (IoU), the overlap between predicted and reference masks divided by the area of their union. Mean Average Precision (mAP) summarizes detection and mask quality across thresholds and categories. These metrics describe the model’s visual output.
The product decision requires a second set of measures: Recall@K, Precision@K or Normalized Discounted Cumulative Gain (NDCG) for neighbours, successful-segmentation coverage, per-category results, latency and cost. A highly accurate segmenter may still add little value when neighbours do not improve or too many catalogue items leave the pipeline.
04 — Decision
Use segmentation for categories where context demonstrably degrades retrieval.
The protocol is an ablation: the same catalogue, encoder and index are evaluated with and without masks. Differences are analyzed by category, photography style and segmentation confidence. This comparison identifies cases where scenery creates substantial unwanted variance.
In the project, segmentation remained in the experimental toolbox rather than entering the first service. Generic CLIP already supported raw visual similarity, while segmentation added another stage, possible failures and a cost that required justification. Final retrieval quality therefore remains the decision criterion.
05 — From concept to decision
Decide by photography style rather than for the whole catalogue.
A garment photographed in a busy setting may benefit from masking; styling may also reveal fit or use that helps retrieval. I would compare image preparations by photography style and retain the full image when segmentation fails.
The decisive test would examine retrieved neighbors and the proportion of products still served. Masks would be retained where their benefit outweighs lost context, failures and additional cost.