01 — Use case
Catalogue content supplies a signal as soon as a product arrives.
Collaborative recommendation learns from interactions between users and products. A new item has no such history. A multimodal approach immediately uses its image, name and description to position it near products with similar content.
Jina CLIP v2 encodes text and images into one shared space. Cosine distance measures the angle between normalized vectors: the closer their directions, the more related their content is considered.
02 — Incremental pipeline
Encode only new or modified fields.
The pipeline computes a fingerprint for each field’s content and compares it with the stored version. New or modified products are re-encoded; products removed from the catalogue are deleted. This logic sharply reduces the cost of subsequent runs.
Required fields determine product eligibility. If an optional field is missing, other available sources can compensate. Image-download failures are therefore handled through explicit rules instead of silent exclusion.
03 — Fusion
Combine modalities through readable, reversible weighting.
Each configuration gives a weight to the name, description or image. Embeddings are added with these coefficients and normalized. The operation can be recalculated quickly when product strategy changes, without rerunning model inference for unchanged fields.
Separating tables by model and field prevents incompatible vector spaces from being mixed. It also allows several versions to coexist during migration and keeps the origin of every representation identifiable.
04 — Search at scale
HNSW turns the catalogue into a navigable neighbourhood graph.
Hierarchical Navigable Small World (HNSW) organizes vectors in a multi-level graph. Search begins on a sparse level to approach the relevant region quickly, then refines neighbours through denser levels. The system returns the k closest products, removes the query product and applies a distance threshold.
The architecture makes recall, speed and memory trade-offs explicit. At 1,024 dimensions with 32-bit floating-point values, one million products require roughly 4 GB for vectors before index overhead. Batch size, embedding dimension and HNSW parameters must therefore be measured at the target scale.