Last verified 5 min read Retrieval and embedding models

Papers with Code search: what the 0.9955 recall measured

Recall 0.9955, p50 1.31ms, 75 papers/second: the corpus size, dimension and hardware Hugging Face measured each figure on, and what it never measured.

This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.

Conclusion: 0.9955 is a 5,000-paper number, not a 110,000-paper one

According to the Hugging Face blog post “How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code” (21 August 2026, by Niels Rogge and others), the reported Recall@20 of 0.9955 and the p50 1.31 ms / p95 2.21 ms latencies were measured on a 5,000-paper pilot with a 256-dimensional index.

The same post states that the production system maintains embeddings for more than 110,000 current papers sourced from arXiv and Daily Papers, but it does not report Recall@20 or HNSW latency at that scale.

Every figure below comes from Hugging Face itself. As of 26 August 2026 there is no third-party reproduction.

What each number was measured on

Performance and throughput

Reported figureValueMeasurement conditions
Recall@200.99555,000-paper pilot, 256 dimensions, ANN recall against exact search
p50 latency1.31 ms5,000-paper pilot, 256 dimensions, HNSW lookup
p95 latency2.21 ms5,000-paper pilot, 256 dimensions, HNSW lookup
Encoding throughput~75 papers/s5,000-paper pilot, 1024 dimensions, NVIDIA L4 GPU (l4x1, 24GB VRAM)
Storage~27%256-dim table and index relative to the 1024-dim version, in that test

The easiest condition to attach to the wrong number is the GPU. The post names the L4 as the runtime of the embedding generation Job, so it is bound to the throughput figure only. HNSW lookups happen in PostgreSQL with pgvector, and the post does not describe that runtime at all — no instance type, no memory, no PostgreSQL configuration.

Architecture and operational parameters

ItemAs stated in the post
Embedding modelQwen/Qwen3-Embedding-0.6B, pinned to an exact revision
Vectors256-dimensional, L2-normalized (Matryoshka representation truncated, then normalized)
Lexical branchWeighted PostgreSQL full-text search, up to 50 candidates
Semantic branchpgvector, up to 50 candidates
FusionWeighted reciprocal rank fusion, currently equal branch weights and k=60
Client timeoutOne second in production
EndpointMaximum of one replica, scales to zero when idle
Incremental runsHourly, at most 500 papers per run, batches of 16

According to the Qwen team’s model card (as recorded in April 2026), the model’s native embedding dimension is up to 1024, it supports Matryoshka Representation Learning, and user-defined output dimensions from 32 to 1024 are available. The blog post says the team chose 256 dimensions to make the search fast.

How to carry these numbers into your own estimate

  1. Read the conditional clause before the number. For the performance figures, the post states the conditions in a single sentence, and those conditions are the 5,000-paper pilot at 256 dimensions.
  2. Always quote corpus size, output dimension and hardware alongside the value. If any one of the three differs from your setup, the number is not directly transferable.
  3. Check whether the post reports a figure at your scale. For this source, recall and latency above 110,000 papers simply are not there.
  4. Do not extrapolate to fill the gap. Dividing a corpus size by “about 75 papers per second” produces a number the post never claims; the post does not state the wall-clock time taken to embed the full corpus.

Three things that are easy to misread

“About 27%” is what remained, not what was saved

The post states that the 256-dimensional table and index used about 27% of the storage of the 1024-dimensional version. That is the share remaining, which corresponds to a reduction of roughly 73%. Reading it as “27% smaller” understates the effect by a wide margin. The comparison also carries the qualifier that ANN recall stayed essentially the same in that test, and the post gives no production measurement of how the dimension reduction affected search quality.

Recall@20 is not end-user search accuracy

The 0.9955 figure describes how well the approximate nearest neighbour search reproduced the top results of an exact, brute-force search. It says nothing about whether those results are useful to a reader, so restating it as “99.55% search accuracy” changes the claim.

The one-second timeout belongs to the client

According to the post, the one-second production timeout is deliberate strict behaviour in the query client. What the post specifies for the endpoint itself is a maximum of one replica and scaling to zero when idle. If the endpoint is scaling up, times out, returns a malformed vector, or has no concurrency available, the semantic branch is skipped immediately and users still receive lexical results. On the write path, the source row is locked and its content hash rechecked before an embedding is written; a paper that changed during inference has its vector discarded and picked up by the next run.

For how an index size can look many times larger or smaller depending on the comparison, see our piece on what the “about 42x” figure in multi-vector retrieval compares. For the question of when a published benchmark number transfers to your own language and workload, see multilingual retrieval models trained without Japanese labels. All figures here were checked against the primary sources on 26 August 2026.

Sources

  1. How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code huggingface.co published 2026-08-21 accessed 2026-08-26
  2. Qwen/Qwen3-Embedding-0.6B model card huggingface.co published 2026-04-20 accessed 2026-08-26