Papers with Code search: what the 0.9955 recall measured
Recall 0.9955, p50 1.31ms, 75 papers/second: the corpus size, dimension and hardware Hugging Face measured each figure on, and what it never measured.
This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.
Conclusion: 0.9955 is a 5,000-paper number, not a 110,000-paper one
According to the Hugging Face blog post “How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code” (21 August 2026, by Niels Rogge and others), the reported Recall@20 of 0.9955 and the p50 1.31 ms / p95 2.21 ms latencies were measured on a 5,000-paper pilot with a 256-dimensional index.
The same post states that the production system maintains embeddings for more than 110,000 current papers sourced from arXiv and Daily Papers, but it does not report Recall@20 or HNSW latency at that scale.
Every figure below comes from Hugging Face itself. As of 26 August 2026 there is no third-party reproduction.
What each number was measured on
Performance and throughput
| Reported figure | Value | Measurement conditions |
|---|---|---|
| Recall@20 | 0.9955 | 5,000-paper pilot, 256 dimensions, ANN recall against exact search |
| p50 latency | 1.31 ms | 5,000-paper pilot, 256 dimensions, HNSW lookup |
| p95 latency | 2.21 ms | 5,000-paper pilot, 256 dimensions, HNSW lookup |
| Encoding throughput | ~75 papers/s | 5,000-paper pilot, 1024 dimensions, NVIDIA L4 GPU (l4x1, 24GB VRAM) |
| Storage | ~27% | 256-dim table and index relative to the 1024-dim version, in that test |
The easiest condition to attach to the wrong number is the GPU. The post names the L4 as the runtime of the embedding generation Job, so it is bound to the throughput figure only. HNSW lookups happen in PostgreSQL with pgvector, and the post does not describe that runtime at all — no instance type, no memory, no PostgreSQL configuration.
Architecture and operational parameters
| Item | As stated in the post |
|---|---|
| Embedding model | Qwen/Qwen3-Embedding-0.6B, pinned to an exact revision |
| Vectors | 256-dimensional, L2-normalized (Matryoshka representation truncated, then normalized) |
| Lexical branch | Weighted PostgreSQL full-text search, up to 50 candidates |
| Semantic branch | pgvector, up to 50 candidates |
| Fusion | Weighted reciprocal rank fusion, currently equal branch weights and k=60 |
| Client timeout | One second in production |
| Endpoint | Maximum of one replica, scales to zero when idle |
| Incremental runs | Hourly, at most 500 papers per run, batches of 16 |
According to the Qwen team’s model card (as recorded in April 2026), the model’s native embedding dimension is up to 1024, it supports Matryoshka Representation Learning, and user-defined output dimensions from 32 to 1024 are available. The blog post says the team chose 256 dimensions to make the search fast.
How to carry these numbers into your own estimate
- Read the conditional clause before the number. For the performance figures, the post states the conditions in a single sentence, and those conditions are the 5,000-paper pilot at 256 dimensions.
- Always quote corpus size, output dimension and hardware alongside the value. If any one of the three differs from your setup, the number is not directly transferable.
- Check whether the post reports a figure at your scale. For this source, recall and latency above 110,000 papers simply are not there.
- Do not extrapolate to fill the gap. Dividing a corpus size by “about 75 papers per second” produces a number the post never claims; the post does not state the wall-clock time taken to embed the full corpus.
Three things that are easy to misread
“About 27%” is what remained, not what was saved
The post states that the 256-dimensional table and index used about 27% of the storage of the 1024-dimensional version. That is the share remaining, which corresponds to a reduction of roughly 73%. Reading it as “27% smaller” understates the effect by a wide margin. The comparison also carries the qualifier that ANN recall stayed essentially the same in that test, and the post gives no production measurement of how the dimension reduction affected search quality.
Recall@20 is not end-user search accuracy
The 0.9955 figure describes how well the approximate nearest neighbour search reproduced the top results of an exact, brute-force search. It says nothing about whether those results are useful to a reader, so restating it as “99.55% search accuracy” changes the claim.
The one-second timeout belongs to the client
According to the post, the one-second production timeout is deliberate strict behaviour in the query client. What the post specifies for the endpoint itself is a maximum of one replica and scaling to zero when idle. If the endpoint is scaling up, times out, returns a malformed vector, or has no concurrency available, the semantic branch is skipped immediately and users still receive lexical results. On the write path, the source row is locked and its content hash rechecked before an embedding is written; a paper that changed during inference has its vector discarded and picked up by the next run.
Related reading
For how an index size can look many times larger or smaller depending on the comparison, see our piece on what the “about 42x” figure in multi-vector retrieval compares. For the question of when a published benchmark number transfers to your own language and workload, see multilingual retrieval models trained without Japanese labels. All figures here were checked against the primary sources on 26 August 2026.
Sources
この記事の日本語版: Papers with Code search: what the 0.9955 recall measured(日本語)