Last verified 5 min read Retrieval and embedding models

MultiVectorEncoder: what the 42x index figure compares

Sentence Transformers v6.0 added late interaction retrieval. Published index sizes, pooling reductions and compressed figures, separated by what each compares.

This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.

Conclusion: the 42x is measured against MiniLM, not DenseOn

According to the Hugging Face blog post “Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers” (18 August 2026, by Tom Aarsen, Antoine Chaffin and Raphael Sourty), encoding 4,874 Natural Questions passages with the multi-vector model lightonai/LateOn produced an uncompressed float32 index of 311.5MB, against 7.5MB for the dense all-MiniLM-L6-v2. That ratio is the “about 42x” quoted in the post.

The same post also reports a quality comparison, but against a different model. Every figure below comes from Hugging Face and LightOn; as of 24 August 2026 there is no third-party reproduction.

ComparisonThe two things comparedReported figures
Quality (NanoBEIR mean NDCG@10)LateOn vs DenseOn0.6868 vs 0.6764
Index sizeLateOn vs all-MiniLM-L6-v2311.5MB vs 7.5MB

DenseOn does not appear in the index size table at all. A sentence of the form “you pay 42x storage for 0.0104 NDCG” is therefore a reading the source does not support.

What MultiVectorEncoder is

sentence-transformers v6.0.0 was released on 18 August 2026; the GitHub release record and the PyPI upload record agree on that date. The development team describes MultiVectorEncoder as a fourth model class alongside SentenceTransformer, CrossEncoder and SparseEncoder, and states that PyLate checkpoints and Stanford-NLP ColBERT checkpoints load straight into it.

The quality figures come from two separate tables

ModelSource tableNanoBEIR mean NDCG@10
lightonai/LateOnLateOn vs DenseOn table0.6868
lightonai/DenseOnLateOn vs DenseOn table0.6764
lightonai/GTE-ModernColBERT-v1Supported Models table0.6720
lightonai/colbertv2.0Supported Models table0.6201
colbert-ir/colbertv2.0Supported Models table0.6053

The metric is the same in every row (mean NDCG@10 across the 13 NanoBEIR datasets), but the first two rows and the last three come from different tables. The post describes LateOn and DenseOn as trained on the same data with the same ModernBERT backbone and the same 149M parameters, differing only in whether they keep one vector per token or pool down to one per document, and notes that late interaction wins on 9 of the 13 datasets. Note also that colbertv2.0 appears twice under different repositories, so the repository prefix matters. The post itself adds that NanoBEIR is a small benchmark and is not a substitute for evaluating on your own data.

How large the index gets

The float32 index sizes reported for the 4,874 Natural Questions passages are below. LateOn produced 608,414 token vectors, an average of 124.8 per passage.

RepresentationVectorsDimensionsfloat32 size
Dense, all-MiniLM-L6-v24,8743847.5MB
Dense, gte-modernbert-base4,87476815.0MB
Multi-vector, LateOn608,414128311.5MB

The compressed figure differs between two primary sources

For the same 608,414 vectors stored as a fast-plaid (PLAID) index, the blog post says 92MB while the v6.0.0 release notes, published the same day, say 88MB. As of 24 August 2026 neither figure has been corrected publicly. The order of magnitude agrees in both documents, but there is no single confirmed value to quote.

Pooling: reduction and its quality cost

HierarchicalTokenPooling clusters each document’s token vectors with Ward linkage on cosine similarity, replaces each cluster with its mean, and keeps roughly 1 / pool_factor of the tokens. The post states that pooling applies to documents only by default, and that pooling all 608,414 vectors took about 6 seconds.

pool_factorToken vectorsReductionfloat32 index
1 (off)608,4141.00x311.5MB
2305,4381.99x156.4MB
3204,4072.98x104.7MB
4153,9363.95x78.8MB

The quality numbers have a different origin. The post cites the original token pooling paper (arXiv:2409.14683v1), which measured 100.6% of unpooled retrieval performance on average at pool_factor=2 and 99.0% at pool_factor=3 on BEIR. Those are not figures Hugging Face re-measured on Natural Questions. The post advises measuring the cost on your own corpus with an evaluator before settling on a factor.

Vector database timings, from one machine

The post lists the following timings for the same 4,874 documents and 608,414 token vectors. They were produced on a single machine (RTX 3090, i7-13700K) with no tuning beyond what the code shows, and the post presents them as an indication of the shape of the work rather than as a product benchmark.

ImplementationIngestionQuery
fast-plaid5s (indexing)11ms
Qdrant26.3s18ms
Weaviate41s17ms
Vespa~80s~75ms warm (~115ms first call)

On Qdrant the post states that the full scan costs 18ms and is exact at 4,874 documents, but that this does not extrapolate, and that Qdrant themselves suggest reserving late interaction for reranking a few hundred candidates rather than scanning a whole collection.

Details that differ per model

ItemWhat the source states
document_lengthTruncates, so anything past the cap never reaches the index. LateOn’s cap of 300 turns a 662-token passage into 273 vectors
Per-model capscolbert-ir/colbertv2.0 uses 180; lightonai/GTE-ModernColBERT-v1 uses caps of 48 and 300
attend=Falsecolbert-ir/colbertv2.0 and answerdotai/answerai-colbert-small-v1 reject Flash Attention at load time, so "sdpa" is required
MaxSim scoresThe score sums over query tokens, so scores are not comparable across models with different query recipes
Save compatibilityOne-way. PyLate, Stanford-NLP ColBERT and colpali-engine checkpoints load into MultiVectorEncoder, but MultiVectorEncoder.save_pretrained output is not loadable by any of them. The post adds that checkpoints in colpali-engine’s own format each need a small configuration added to their repository before they load

For a bounded scale the post points to similarity_fn_name="meanmaxsim", an average cosine similarity in [-1, 1], and it also mentions a way to raise the document length cap for a single call.

The same model family appears in our earlier piece on multilingual retrieval models trained without Japanese labels, and the question of when a public benchmark score is comparable at all is covered in NIST’s sequestered AITE evaluation. All figures and quotations here were checked against the primary sources on 24 August 2026.

Sources

  1. Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers huggingface.co published 2026-08-18 accessed 2026-08-24
  2. Releases · huggingface/sentence-transformers v6.0.0 github.com published 2026-08-18 accessed 2026-08-24
  3. sentence-transformers 6.0.0 · PyPI pypi.org published 2026-08-18 accessed 2026-08-24
  4. DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models (arXiv:2607.27178) 学術 published 2026-07-29 accessed 2026-08-24