Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
2026-08-20 · Hugging Face
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
Sentence Transformers is a Python library for using and training embedding and reranker models for applications like retrieval augmented generation, semantic search, and more. With the v6.0 update, it gains a fourth model type: MultiVectorEncoder, for ColBERT-style late interaction retrieval. Any PyLate checkpoint and any Stanford-NLP ColBERT checkpoint loads straight into it, and colpali-engine models for visual document retrieval can be used too, through the same familiar API you already use for dense, sparse, and reranker models.
What are Multi-Vector Models?
A dense embedding model reads a text and returns a single fixed-size vector. Everything the model noticed has to fit in those 384, 768, or 1024 numbers, and similarity is one dot product between two such summaries. This works remarkably well, but the compression is lossy in a specific way: a rare entity, an exact identifier, or one crucial clause in a long passage all have to compete for room in the same vector.
For the query "green sofa with wooden legs and rounded cushions", a single vector has to blend all four attributes into one point, so a green sofa with the wrong legs ends up sitting close to the one you actually asked for.
A multi-vector model (also called a late-interaction or ColBERT-style model) skips that compression. It runs the same transformer, but instead of pooling the token embeddings into one vector, it projects each token embedding down to a small dimension (classically 128) and keeps all of them. A 9-token document becomes a 9x128 matrix, not a 1x128 vector.
The interaction between query and document is deferred until scoring time, which is where the name "late interaction" comes from. Documents are still encoded independently and can be indexed offline, but scoring compares every query token against every document token. This sits between cross-encoders (early interaction, no precomputation) and bi-encoders (minimal interaction via one dot product, full precomputation).
The MaxSim Operator
Scoring uses MaxSim: for each query token, take its highest similarity against any document token, then sum those maxima across the query.
$$\text{MaxSim}(Q, D) = \sum_{Q_i \in Q} \max_{D_j \in D} Q_i \cdot D_j$$
Because the token embeddings are L2-normalized, each dot product is a cosine similarity. The whole sum lands within [-num_query_tokens, num_query_tokens].
The operator can be read as a soft alignment: every query token points at the one document token that best explains it. The alignment doesn't have to be lexical. Using lightonai/mLateOn, the query token "live" in "Where do penguins live?" finds its best match on "inhabit" at 0.94 similarity, despite sharing no characters.
This allows handling synonyms and paraphrases while still preserving exact token matches that a single-vector model would average away. It is not strictly one-to-one; several query tokens may align to the same document token.
What You Gain, and What It Costs
Multi-vector models deliver stronger retrieval quality, particularly on queries where one specific piece of a document determines relevance, on multi-requirement queries where each condition can find its own evidence, and on out-of-domain data where a dense model's learned compression may have dropped critical information.
The compression in dense models is learned from training queries, so the model keeps what those queries needed and drops everything else — which may include exactly what production queries require.
The trade-off is a larger index, since multiple vectors must be stored per document. The blog post demonstrates loading various checkpoint formats, encoding and scoring, integrating into a search stack, running on page images (state-of-the-art for visual document retrieval without OCR), and keeping the index affordable. All examples run after a simple `pip install -U sentence-transformers`.
Additional topics covered include semantic search, retrieve-and-rerank, token pooling, speeding up inference, evaluating models, and migration from PyLate or colpali-engine.