MMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrieval
This paper introduces MMed-Bench-IR, a benchmark for multilingual medical retrieval in clinical settings, evaluating cross-lingual alignment, concept discrimination, and evidence retrieval.
First benchmark to evaluate multilingual medical retrieval capabilities in clinical settings, disentangling cross-lingual alignment, concept discrimination, and evidence retrieval.
Keywords
Before reading this…
Applications
- →Clinical settings
- →Retrieval-augmented generation (RAG)
To understand this paper, make sure you know these concepts first:
- Understanding of multilingual retrievalfind papers →
- Familiarity with medical terminologyfind papers →
Abstract
More Like ThisRetrieval-augmented generation (RAG) in clinical settings increasingly requires multilingual retrieval against predominantly English evidence corpora. Multilingual medical retrieval demands three capabilities: cross-lingual alignment, concept discrimination, and evidence retrieval. However, existing benchmarks evaluate these only in isolation, leaving the interaction between biomedical expertise and multilingual coverage unmeasured. We introduce MMed-Bench-IR, a benchmark designed to disentangle these axes across 6 languages and three structurally heterogeneous tasks: (1) cross-lingual medical QA retrieval with 6,127 queries grounded in the Unified Medical Language System (UMLS), (2) concept discrimination over 4,975 confusion sets at three difficulty tiers, and (3) multilingual evidence retrieval for RAG with 2,040 quality-assured queries. The three tasks share zero concept and query overlap by design, ensuring that aggregate scores reflect genuine capability breadth. Evaluation of ten systems across six paradigm families reveals severe cross-lingual failure: biomedical encoders that score 0.818 nDCG@10 in English drop to 0.056 in Japanese, a gap that English-only benchmarks cannot detect.