A model becomes an operational service when tests, migrations, monitoring and fallbacks make its production behavior reliable.
Article context
Making a model work in a notebook opens the research; building a reliable service requires a different set of decisions.
CONTEXT
This post retraces a visual-search project for e-commerce: starting from a product image, representing the catalogue in vector space and returning visually related items. The work evolved from a collection of notebooks into an application programming interface (API) built with FastAPI and ClickHouse and designed to handle traffic.
SUMMARY
The research phase compared Contrastive Language–Image Pre-training (CLIP), FashionCLIP, clustering, generated descriptions, image–text projection and model specialization. The production phase retained a simpler chain: normalized embeddings, vector indexing, cosine-distance search, catalogue isolation, load testing, structured logging and end-to-end migration checks.
Key diagram
From exploratory notebook to a service you can trust
Each transition removes uncertainty: model relevance, pipeline reproducibility, output correctness and operational readiness.
A research-based explainer grounded in the sources listed below. Some mechanisms are simplified without changing the overall conclusions.
Contribution17
What must be added to a relevant model to create a reliable service?
I structured productionization around a reproducible pipeline, two measurable search strategies, an appropriate vector schema, output-correctness tests, verified migrations and usable observability.
Operational quality
Neighbour relevance
Latency and load
Migrations and fallback strategy
Takeaways
A machine-learning service becomes reliable through the systems around inference: schema, tests, versioning, security, observability and fallback strategies.
Keep experiments that improve results and leave the others outside the critical path.
Test neighbor content and ordering: a technically successful response can be functionally wrong.
Treat latency, migrations, versions and fallbacks as delivery criteria.
01 — Explore
Notebooks eliminate hypotheses before the architecture is fixed.
I compared CLIP ViT-B/32 and FashionCLIP, explored visual groups with K-means and principal component analysis (PCA), generated product descriptions, trained an image–text projection and tested several specialization and segmentation paths.
This phase identified the useful components for the target use case: generic CLIP was sufficient for raw visual similarity, image–text alignment was worth retaining, and several more complex components could remain in the toolbox without entering the first service.
02 — Build
A reproducible chain replaces a stack of experiments.
The pipeline loads images in batches, computes 512-dimensional vectors, applies L2 normalization and stores them in ClickHouse. A namespace isolates each catalogue. Two search strategies — database and in-memory computation — can be compared according to scale.
FastAPI exposes the service and components are deployed separately. The important decision was to preserve comparative measurement between strategies rather than choose early from an unverified traffic assumption.
03 — Harden
Silent errors require tests of outcomes, not only execution.
An incomplete migration had left the new tables outside the execution path. Another defect swapped metadata and distance columns: the query returned successfully, yet its ranking was wrong. Insertion tests, smoke tests — quick checks of essential functions — and output benchmarks exposed these cases.
Validation also covers latency, load, migrations and functional checks of returned neighbours.
04 — Operate
Security, observability and reproducibility make the service maintainable.
Credentials and service addresses were moved out of source code, software images were pinned to precise versions and debugging prints were replaced with structured logging. Dead code, unused dependencies and ambiguous configuration were removed.
A sub-200 ms latency target guided the vector schema and metadata denormalization. Decisions remain tied to their assumptions so they can be revisited as catalogue size and real traffic evolve.
Primary source
Visual Search project — research and productionization
Personal account of designing, auditing and hardening a similar-product search service.