Last verified 5 min read Inference infrastructure and embedding models

TPU embedding parity: what the 0.999 was compared against

Google's TPU embedding post reports 0.999/0.995 cosine thresholds and 83,996 token/s. Which baseline and which configuration each number came from.

This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.

Conclusion: what the 0.999 was compared against depends on the table, and it is not a retrieval score

The Google Developers Blog post “Enterprise-Grade Precision for Long-Context Multimodal Embedding Inference on Cloud TPU”, dated 26 August 2026 and credited to Anthony Su and Injae Kwak, covers numerical parity when serving embedding models with vLLM on TPU. The stated pass thresholds are cosine similarity of at least 0.999 for text inputs and at least 0.995 for multimodal inputs.

Three things are worth separating before reusing any of these numbers.

  1. The thresholds describe agreement between vectors generated on TPU and reference vectors. They are not a comparison of retrieval quality such as Recall or nDCG.
  2. In Table 1 (Qwen3-Embedding-8B) the reference is a CPU baseline, not a GPU. Table 2 covers a different model and uses different references.
  3. The single throughput figure in the post was measured on a different configuration from the cosine similarity runs.

Everything below comes from Google’s own material: the Developers Blog post and the AI-Hypercomputer official recipes the post links to. As of 28 August 2026 there is no third-party reproduction.

What each number was measured on

Thresholds and similarity values

ItemValueSource and conditions
Text pass threshold0.999Body text and the Target Parity column of Table 1
Multimodal pass threshold0.995Quality Pass column of Table 2 (Qwen3-VL-Embedding-8B); the same column also shows ≥0.999 for its text rows
Table 1, Ironwood column (English long)0.99978456707K tokens, against a CPU baseline
Table 1, Trillium column (English long)0.9997664366Same run

Table 1 is headed “Qwen3-Embedding-8B (7K+ Tokens, TPU vs. CPU Baseline)” and its five columns are Input Corpus, Context Length, Ironwood Cosine Sim, Trillium Cosine Sim and Target Parity. It contains no throughput value and no tensor parallel size.

The eight values in that table match, digit for digit, the expected-value table in the official recipe’s README-correctness.md. There the columns are labelled v7x and v6e, the comparison target is a run with VLLM_TARGET_DEVICE="cpu", and the sequence length is stated as 7K tokens.

Performance figures

ItemValueSource and conditions
Total token throughput (blog)83,996bf16, 16K+ sequence length, TP=4, TPU Ironwood
Total token throughput (official recipe)84038.2716k workload
Request throughput5.13Identical in both documents

The requests per second agree; the tokens per second do not. Which figure is correct cannot be determined from the primary sources, and neither document explains the difference.

Tensor parallel size differs by purpose

PurposeTPU sideReference sideSource
Text parity validation21 (CPU)embed_script.py
Throughput measurement4—Deployment manifest (README.md)
Multimodal parity validation81 (GPU)embed_vl_script.py

The GKE deployment manifest that produced the 83,996 token/s figure uses --tensor-parallel-size=4, --max-model-len=16384 and --enable-chunked-prefill, on tpu7x with a 2x2x1 topology. The benchmark data was 1,000 randomly generated prompts of 16,384 input tokens driven through --backend=openai-embeddings, not real corpus data. The sample code printed in the blog post specifies tensor_parallel_size=2, so it belongs to the first row, the parity validation side.

In the official recipes, Ironwood corresponds to accelerator type v7x (TPU v7 / TPU7x) and Trillium to v6e. The parity validation runs on TPU VMs of type v7x-8 and v6e-4, while the throughput measurement runs on GKE with tpu7x 2x2x1, so the execution environments also differ.

How to read the published figures

  1. Determine whether a number belongs to Table 1 (cosine similarity) or to the caption sentence (throughput). TP=4 appears only in the caption.
  2. Check what the reference is. Table 1 is compared against a CPU baseline; Table 2 uses anonymised XPUs for both its text and multimodal test modes.
  3. If the question is retrieval quality, these numbers are not sufficient. The post contains no Recall or nDCG comparison.

Caveats

The primary source contradicts itself on sequence length

The heading of Table 1 says “7K+ Tokens”, while the caption attached to the same table says “16K+ sequence length”. The official recipe states 7K tokens for the parity validation.

The reference hardware cannot be identified from the post

Table 2 covers a different model, Qwen3-VL-Embedding-8B, and its Test Mode column lists both text and multimodal runs. It labels the reference machines XPU 1, XPU 2 and XPU 3. The official recipe’s README-correctness.md names Blackwell (B200) and Hopper (H200) explicitly, and the four values under XPU 2 match the H200 column while the four values under XPU 3 match the B200 column, digit for digit. Google never states this mapping, so the accurate statement is only that the values agree. No values corresponding to XPU 1 appear in the recipe tables.

The evaluation inputs live in the scripts, not the post

The post gives neither a sample count nor a dataset name. The linked scripts do specify the inputs: on the text side, four strings (English short, English long, Japanese, Korean) each repeated 180 times; on the multimodal side, a single input per test mode. So the information is absent from the post, not absent from the published material.

Sources

  1. Enterprise-Grade Precision for Long-Context Multimodal Embedding Inference on Cloud TPU developers.googleblog.com published 2026-08-26 accessed 2026-08-28
  2. Serve Qwen3-Embedding-8B with vLLM on TPU Ironwood (README.md) github.com published 2026-08-03 accessed 2026-08-28
  3. Qwen3-Embedding-8B E2E Correctness Validation on TPUs (README-correctness.md) github.com published 2026-08-03 accessed 2026-08-28
  4. AI-Hypercomputer/tpu-recipes embed_script.py (Qwen3-Embedding-8B) github.com published 2026-08-03 accessed 2026-08-28
  5. AI-Hypercomputer/tpu-recipes compare_precision.py (Qwen3-Embedding-8B) github.com published 2026-08-03 accessed 2026-08-28
  6. Qwen3-VL-Embedding-8B E2E Correctness Validation on TPU Ironwood (README-correctness.md) github.com published 2026-08-03 accessed 2026-08-28
  7. AI-Hypercomputer/tpu-recipes embed_vl_script.py (Qwen3-VL-Embedding-8B) github.com published 2026-08-03 accessed 2026-08-28