Medical LLM Benchmarks Miss Longitudinal EHR and Cross-Lingual Gaps

Most medical LLM benchmarks test static question answering with pre-selected evidence. They do not test whether a model can reason across months of longitudinal electronic health records, or whether it answers the same clinical question consistently when asked in English versus Mandarin. Two preprints published this week expose these gaps—and they matter acutely for Singapore hospitals deploying clinical AI in multilingual, EHR-heavy environments.

This post is for hospital CIOs, clinical informatics teams, and AI engineers evaluating medical LLMs for clinical AI services in Singapore and Asia.

Key takeaways

  • Static benchmarks hide longitudinal reasoning failures: Existing medical LLM benchmarks use pre-selected evidence snippets; new research shows models fail when reasoning across real longitudinal EHR timelines [1].
  • Cross-lingual consistency is untested: Medical LLMs produce inconsistent answers to the same clinical question in different languages, even when the medically correct answer should be identical [2].
  • Singapore hospitals face both gaps simultaneously: Multilingual patient populations and longitudinal EHR workflows mean both failure modes are production risks, not research curiosities.
  • Evaluation protocols must match deployment context: Benchmarks designed for static Q&A do not predict performance in longitudinal decision-making or multilingual clinical workflows.

Why static benchmarks miss longitudinal EHR reasoning

Most medical LLM benchmarks—MedQA, PubMedQA, MMLU-Medical—present a clinical question with a short vignette or pre-selected evidence. The model picks an answer. This setup tests recall and pattern matching, but it does not test whether a model can synthesize information across months of clinic visits, lab results, imaging reports, and medication changes.

ObGynLongBench, released this week, exposes this gap [1]. The benchmark uses real longitudinal obstetrics and gynecology EHRs and asks models to make clinical decisions—screening recommendations, diagnostic workups, treatment adjustments—based on patient timelines spanning multiple encounters. Early results show that models that score well on static medical Q&A benchmarks fail when required to track clinical context over time, miss relevant prior events, or conflate information from different visits.

For Singapore hospitals, this matters because clinical decision support, care coordination platforms, and ambient documentation systems all require longitudinal reasoning. A model that answers "What is the first-line treatment for gestational diabetes?" correctly in a multiple-choice exam may still recommend metformin to a patient whose EHR shows a documented sulfa allergy six months earlier—because it was never trained or evaluated on multi-visit reasoning.

We have seen this failure mode in pilot deployments: LLMs that perform well on static benchmarks produce clinically unsafe suggestions when integrated into EHR workflows, because they lack the architectural or training design to maintain coherent state across encounters. Evaluation must test the reasoning pattern the production system will actually use.

Cross-lingual consistency failures in multilingual clinical settings

Singapore hospitals serve multilingual patient populations. Clinical documentation, patient communication, and decision support often occur in English, Mandarin, Malay, or Tamil. A medical LLM deployed in this environment should answer the same clinical question consistently across languages—at least when the medically correct answer is language-invariant.

A preprint published September 7 challenges this assumption [2]. The authors tested multilingual medical LLMs on the same clinical questions posed in different languages and found significant cross-lingual inconsistency: models gave different answers to identical medical questions depending on input language, even when the correct answer should not vary by language. Existing multilingual medical benchmarks treat this variation as model error, but the paper argues that some cross-lingual variation may reflect legitimate cultural adaptation (e.g., dietary advice, end-of-life care preferences).

The distinction matters for deployment. If a model recommends different hypertension medications when queried in English versus Mandarin—not because of cultural context, but because of training data imbalance or tokenization artifacts—that is a safety failure. If it adjusts dietary counseling for a diabetic patient based on inferred cultural context, that may be appropriate clinical adaptation.

Singapore hospitals need evaluation protocols that distinguish these cases. Current multilingual medical benchmarks do not. We recommend:

  • Explicit cross-lingual consistency testing: Run the same clinical question through the model in English, Mandarin, and other deployment languages; flag inconsistencies for clinical review.
  • Cultural adaptation guidelines: Define which clinical domains permit language-conditioned responses (e.g., lifestyle counseling) and which require strict consistency (e.g., drug dosing, contraindications).
  • Audit trails for multilingual queries: Log input language and response for every clinical query; monitor for drift or inconsistency over time.

This is not a solved problem in the research literature. It is a deployment risk that Singapore hospitals must manage through evaluation design and operational monitoring, not by assuming that a model's English-language benchmark score generalizes to Mandarin or Malay.

Why Singapore hospitals face both gaps simultaneously

Singapore's healthcare AI deployment context is unusually demanding:

  • Longitudinal EHR workflows: Public hospital clusters maintain comprehensive longitudinal EHRs; clinical decision support must reason across years of patient history.
  • Multilingual clinical communication: Patients, families, and clinical staff communicate in multiple languages; documentation and decision support must support this without introducing inconsistency.
  • Regulatory scrutiny: The Health Sciences Authority (HSA) and Personal Data Protection Commission (PDPA) require evidence of safety and fairness; benchmarks that do not test deployment-relevant failure modes do not satisfy this burden.

Most medical LLM research originates in monolingual, single-encounter settings (e.g., US medical licensing exams, PubMed abstracts). Benchmarks reflect this context. Singapore hospitals deploying these models in multilingual, longitudinal EHR environments are operating outside the evaluated envelope.

We have worked with institutional partners to design custom evaluation protocols that test:

  • Longitudinal coherence: Does the model's recommendation change appropriately when new lab results arrive? Does it remember prior contraindications?
  • Cross-lingual consistency: Does the model give the same answer in English and Mandarin for language-invariant clinical questions?
  • Shortcut detection: Does the model rely on spurious correlations (e.g., language as a proxy for ethnicity) rather than clinical evidence?

These evaluations are not standardized. They require clinical domain expertise, multilingual validation, and access to representative EHR data. But they are necessary to de-risk deployment in Singapore's clinical AI context. For more on shortcut detection, see our earlier post on continuous monitoring.

What existing benchmarks do measure—and when they are useful

Static medical Q&A benchmarks are not useless. They test:

  • Medical knowledge recall: Does the model know first-line treatments, diagnostic criteria, contraindications?
  • Clinical reasoning on simplified cases: Can it apply guidelines to a short vignette?
  • Relative model comparison: Which of two models has better medical knowledge, holding evaluation context constant?

These are necessary but not sufficient for deployment. A model that fails MedQA is not ready for clinical use. A model that passes MedQA is not necessarily safe in longitudinal, multilingual EHR workflows.

We use static benchmarks as a first-pass filter, then layer on deployment-context evaluations:

  1. Static benchmark (MedQA, MMLU-Medical): Minimum threshold for medical knowledge.
  2. Longitudinal EHR simulation: Test reasoning across multi-visit timelines using de-identified institutional data.
  3. Cross-lingual consistency audit: Test language-invariant questions in English, Mandarin, and other deployment languages.
  4. Clinical review: Domain experts review model outputs for safety, appropriateness, and cultural sensitivity.

This is more expensive than running a Hugging Face leaderboard script. It is also the only way to generate evidence that the model will perform safely in the deployment environment. For RAG-based clinical systems, see our evaluation framework post for additional testing considerations.

Why this matters in Singapore and Asia

Singapore is a regional hub for healthcare AI deployment. Hospitals here are early adopters of clinical LLMs for documentation, decision support, and care coordination. But Singapore's multilingual, longitudinal EHR context is not well-represented in the benchmarks that medical LLM developers use to claim "clinical-grade" performance.

This creates a validation gap. A model trained and evaluated in the US or Europe may not generalize to Singapore's clinical workflows. Hospitals that deploy based on published benchmark scores—without testing longitudinal reasoning or cross-lingual consistency—are accepting unquantified risk.

The research published this week [1, 2] makes this gap explicit. It also provides a roadmap: hospitals need evaluation protocols that match their deployment context, not generic benchmarks designed for different clinical settings.

Across Asia, multilingual healthcare delivery is the norm, not the exception. The evaluation gaps identified here apply to hospitals in Malaysia, Hong Kong, India, and other multilingual markets. Singapore hospitals that develop robust evaluation protocols for longitudinal, multilingual clinical LLMs will be well-positioned to share those methods regionally—and to de-risk their own deployments.

What to do next

If you are evaluating medical LLMs for deployment in Singapore hospitals:

  • Do not rely solely on static benchmarks: MedQA and MMLU-Medical are necessary but not sufficient; layer on longitudinal EHR simulations and cross-lingual consistency tests.
  • Define language-invariant clinical domains: Identify which clinical questions should produce identical answers across languages (e.g., drug dosing) and which permit cultural adaptation (e.g., dietary counseling).
  • Build custom evaluation datasets: Use de-identified institutional EHR data to create longitudinal test cases that reflect your hospital's patient population, documentation patterns, and clinical workflows.
  • Monitor cross-lingual consistency in production: Log input language and model response for every clinical query; audit for inconsistency monthly.
  • Engage clinical domain experts in evaluation design: Longitudinal reasoning failures and cross-lingual inconsistencies are clinical safety issues, not just ML metrics; clinicians must define acceptable performance.

For hospitals building platform engineering infrastructure for clinical AI, evaluation is a platform capability, not a one-time pre-deployment check. Continuous monitoring, versioned evaluation datasets, and automated consistency testing should be part of your MLOps stack.

If you need help designing deployment-context evaluations for medical LLMs in Singapore hospitals, start a project with us.

FAQ

Why do static benchmarks dominate medical LLM evaluation if they miss deployment-relevant failures?

Static benchmarks are cheap to run, easy to compare across models, and align with academic publication incentives. Longitudinal EHR evaluation requires access to real patient data, clinical domain expertise, and custom infrastructure. Most research teams do not have these resources. But hospitals deploying clinical AI do—and should use them.

Should medical LLMs give different answers in different languages?

It depends. For language-invariant clinical facts (drug dosing, contraindications, diagnostic criteria), answers should be consistent across languages. For culturally sensitive domains (dietary advice, end-of-life care, mental health counseling), some adaptation may be appropriate. Hospitals must define these boundaries explicitly and test for unintended cross-lingual variation.

How do I test longitudinal reasoning without building a custom benchmark?

Start with ObGynLongBench [1] if your use case overlaps obstetrics/gynecology. For other specialties, create synthetic longitudinal cases by chaining real de-identified EHR encounters with known clinical decision points. Have domain experts review model outputs for coherence, safety, and appropriateness. This is lower-cost than a full benchmark but more rigorous than static Q&A.

What is the regulatory expectation for medical LLM evaluation in Singapore?

The Health Sciences Authority (HSA) requires evidence that AI-enabled medical devices are safe and effective for their intended use. If your intended use is longitudinal clinical decision support in a multilingual population, your evaluation evidence must test those conditions. Static benchmarks alone will not satisfy this burden. Work with clinical and regulatory affairs teams early to define evaluation scope.

Sources

[1] ObGynLongBench: Revealing the Evidence-to-EHR Gap in Longitudinal EHR Decision-Making. arXiv preprint cs.AI+health, September 7, 2026. https://arxiv.org/abs/2609.07601v1

[2] Perspectives on Cross-Lingual Consistency in LLMs for Medical Questions. arXiv preprint cs.CL+medical, September 7, 2026. https://arxiv.org/abs/2609.07687v1

[3] NIST AI Risk Management Framework. National Institute of Standards and Technology. https://www.nist.gov/itl/ai-risk-management-framework

[4] A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains. Wang S, Tang Z, Yang H. NPJ Digital Medicine, December 2, 2025. https://pubmed.ncbi.nlm.nih.gov/41454006/

[5] MedWER: A Reproducible, Model-Free Evaluation Protocol for Medical Speech Recognition. arXiv preprint cs.CL+medical, September 4, 2026. https://arxiv.org/abs/2609.05728v1