Last verified 4 min read Embedding models

mLateOn / mDenseOn: Is Japanese in the Training Data?

LightOn's multilingual retrieval models were retrieval-trained on 9 languages — Japanese is not one of them. What the reported MIRACL 72.1 actually means.

This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review. All benchmark figures below are LightOn’s self-reported numbers.

The short answer: retrieval training covers 9 languages, and Japanese is not one of them

For the multilingual retrieval models mDenseOn and mLateOn released by LightOn on July 30, 2026, the Hugging Face blog post explicitly lists the languages targeted by retrieval training: English plus eight machine-translated target languages (Arabic, French, German, Italian, Norwegian, Portuguese, Spanish, and Swedish). Japanese is not on that list.

That does not make these “models that never saw Japanese,” however. LightOn states that “unseen” refers specifically to its own retrieval training, and that those languages may still have appeared in the masked-language-modeling pretraining of the mmBERT backbone. The mmBERT-base model card describes pretraining on 1,833 languages, and its language metadata includes Japanese (jpn) as of October 2025. When you assess language support, you need to separate pretraining languages from retrieval-training languages.

Shared specifications

ItemmDenseOn / mLateOn
Release dateJuly 30, 2026
Parameters307M (both models)
BackbonemmBERT-base (jhu-clsp)
LicenseApache 2.0 (per model card)
Max context length8,192 tokens (query and document)
Retrieval stylemDenseOn: single vector; mLateOn: token-level late interaction

The companion paper, arXiv:2607.27178, was submitted on July 29, 2026 and only v1 exists as of August 2, 2026. The abstract says the models, datasets, and training code will be released.

What the reported numbers say

All figures below are LightOn’s self-reported values. As of August 2, 2026, we found no third-party reproduction.

The gap widens outside the training list

MetricmLateOnmDenseOn
BEIR average nDCG@1057.5656.70
MIRACL, trained languages avg.65.6159.61
MIRACL, 13 unseen languages avg.67.5957.42
MLDR, 6 unseen languages avg.66.5235.97

On trained languages the two models sit about 6 points apart. Averaged over languages absent from retrieval training the gap grows to about 10 points, and on long-document MLDR it exceeds 30 points. The paper’s abstract explains that even with a shared backbone, data, and objective, the dense model degrades outside the range covered by translate-train, while the late-interaction model generalizes better to unseen languages and scripts.

Japanese scores in isolation

Benchmark (Japanese)mLateOnmDenseOn
MIRACL ja (nDCG@10)72.163.0
MLDR ja (nDCG@10)73.550.4

Japanese is one of the 13 unseen languages in MIRACL and one of the 6 unseen languages in MLDR. Note that the Hugging Face blog lists per-language scores only for mLateOn; mDenseOn’s Japanese 63.0 and both MLDR values appear only in the paper’s appendix (Table 5 / Table 6).

72.1 is not the best Japanese score in that table

Reading down the Japanese column of appendix Table 5, BGE-M3 (73.1) and pplx-embed-v1-late-0.6b (72.7) — both larger models — sit above mLateOn’s 72.1. The primary sources claim leadership on BEIR averages, MIRACL trained-language averages, and MLDR — not on Japanese in isolation. The English-only predecessors are reported at 56.20 (DenseOn, 149M) and 57.22 (LateOn).

A checklist for Japanese RAG evaluation

  1. Check the model card or paper for the retrieval-training language list. A “multilingual” label alone cannot tell you whether a language was a retrieval-training target or merely passed through pretraining.
  2. If Japanese is not on the list, read the model’s Japanese score as generalization beyond the training list. In this case, the architectural difference produces a 72.1 vs 63.0 split on Japanese MIRACL.
  3. Re-compare with document length in mind. Japanese MLDR scores are 73.5 vs 50.4 — a much wider gap than on short-form MIRACL.
  4. Confirm license and context length. Both models are Apache 2.0 and are described as supporting context lengths up to 8,192 tokens (which does not mean always 8,192).
  5. Measure on your own data. LightOn itself advises validating performance per language, especially for low-resource languages and specialized domains.

Caveats

All figures are the developer’s own numbers

The models were released only three days before this article, and we could not find third-party reproduction or verification even in English-language searches. When citing these benchmarks, it is safest to attribute them: “according to LightOn.”

Questions the primary sources do not answer

  • Comparison against Japanese-specialized embedding models — the paper’s baseline list contains none.
  • Production quality on Japanese — the sources report benchmark values, not evaluations on business data.
  • Index size and latency for late-interaction retrieval — no concrete figures are given.

Limitations LightOn itself notes

Most of the multilingual pretraining pairs (roughly 2.8 billion query-document pairs in total) are machine translations of English data, and translation quality was not human-evaluated. LightOn also notes that MLDR overlaps with training languages, so baseline comparisons there are not fully controlled. Check the official paper and model cards for updated figures or corrections.

Sources

  1. mDenseOn with the mLateOn (Hugging Face Blog) huggingface.co published 2026-07-30 accessed 2026-08-02
  2. DenseOn with the LateOn (arXiv:2607.27178) 学術 published 2026-07-29 accessed 2026-08-02
  3. arXiv:2607.27178 HTML version, Appendix Table 5 / Table 6 学術 published 2026-07-29 accessed 2026-08-02
  4. lightonai/mLateOn model card huggingface.co published 2026-07-30 accessed 2026-08-02
  5. jhu-clsp/mmBERT-base model card huggingface.co published 2025-10-07 accessed 2026-08-02