01 — Experimental question
Evaluate visual recommendation and semantic search separately.
For visual recommendation, every image becomes a query in turn. Its nearest neighbours count as relevant when they belong to the same category. For semantic search, the category name is used as a query while being deliberately removed from indexed text: the model must recover it from product names, brands and descriptions alone.
Metrics include Precision@K, Recall@K and Normalized Discounted Cumulative Gain at ten results (NDCG@10), which rewards relevant results placed early in the ranking. Experiments use roughly 1,880 Amazon Berkeley Objects products and a 1,000-product H&M subset.
02 — Visual results
FashionCLIP improves the ranking of visually similar products.
On Amazon Berkeley Objects, NDCG@10 increases from 0.618 with CLIP to 0.666 with FashionCLIP, a relative gain of about 7.9%. On H&M, it rises from 0.870 to 0.916, or roughly 4.6%.
The gap between catalogues is as informative as the gap between models. A dense catalogue with many items per category makes relevant neighbours easier to retrieve. A smaller, fragmented catalogue produces a more demanding benchmark.
03 — Text results
Fashion specialization narrows the gap without replacing a purpose-built text encoder.
On Amazon Berkeley Objects, FashionCLIP raises CLIP’s NDCG@10 from 0.076 to 0.296, yet the OpenAI text model reaches 0.622. On H&M, FashionCLIP reaches 0.638, compared with 0.556 for CLIP and 0.727 for OpenAI.
The most plausible explanation lies in the training objective: CLIP aligns short captions with images, while a dedicated text embedding organizes passages and queries by semantic proximity. FashionCLIP adds a fashion-specific vocabulary, reducing the gap on H&M.
04 — Architecture decision
Keep one encoder per use case until simplification is validated.
The results support a two-space architecture: CLIP or FashionCLIP for image similarity, and a dedicated text model for search. A shared space remains attractive for hybrid queries, but it needs its own benchmark.
Category membership is only a proxy for relevance and does not fully measure perceptual similarity. The next step should add human judgments, real multilingual catalogues and per-category analysis to verify that the measured gains match the intended product experience.