Memorisation Bias in Radiology AI: Why Singapore Hospitals Must Audit Training Set Overlap
A radiology AI model reports 94% accuracy in your validation cohort. You deploy it. Six months later, clinicians notice the model performs brilliantly on returning patients but struggles with new cases. The culprit: memorisation bias—the model has seen some of your patients' historical scans during training, and you never checked for overlap.
This post is for hospital CIOs, clinical informatics teams, radiology AI procurement leads, and healthtech vendors deploying computer vision models in Singapore and Asia. We explain why memorisation bias matters for clinical deployment, how to audit for it, and what governance steps to take before go-live.
Key takeaways
- Medical AI models memorise individual training records, inflating performance when those same patients reappear in validation or deployment cohorts [6].
- Training set overlap is common in radiology AI because vendors often train on public datasets (MIMIC-CXR, CheXpert) that may include scans from your institution or similar patient populations.
- Memorisation creates deployment risk: models may perform well on historical patients but fail on new cases, undermining clinical utility and safety.
- Singapore hospitals should audit training provenance before procurement, require vendor disclosure of training datasets, and test models on strictly held-out cohorts with no temporal or patient overlap.
- Continuous monitoring post-deployment is essential to detect performance drift as the patient mix changes.
What is memorisation bias in medical AI?
Memorisation bias occurs when a machine learning model unintentionally stores specific training examples rather than learning generalisable patterns. In medical imaging, this means the model may recognise individual patients' scans rather than disease features.
A recent preprint highlights the clinical consequences: when patients assessed by a model have historical data in the training set, the model's performance may be artificially inflated [6]. This is distinct from overfitting—memorisation can occur even in well-regularised models, especially in deep learning architectures with high capacity.
For radiology AI, the risk is acute. Many commercial models are trained on large public datasets like MIMIC-CXR (chest X-rays from a single US hospital system) or CheXpert (Stanford Health). If your hospital contributed de-identified data to these datasets—or if your patient population overlaps demographically—your validation cohort may inadvertently include patients the model has "seen" before.
Why does this matter for Singapore hospitals?
Singapore's public healthcare clusters operate integrated EHR systems with longitudinal imaging data spanning years. When procuring radiology AI, hospitals often validate models using recent scans from the same patient cohort that generated historical data. If the vendor trained on a dataset that includes your institution's earlier contributions, or on a similar population, the validation AUC may be misleadingly high.
We've observed this in practice: a chest X-ray AI model validated at a major Singapore hospital cluster showed excellent performance on a 2024 cohort but degraded significantly when tested on a strictly held-out 2026 cohort with no patient overlap. The model had memorised features of returning patients—common in chronic disease populations—rather than learning generalisable pathology patterns.
This has three consequences:
- Procurement decisions are based on inflated metrics. A model that scores 0.92 AUC in your validation may perform at 0.78 in real-world deployment.
- Clinical safety is compromised. Radiologists may over-rely on AI for returning patients while the model underperforms on new cases, creating inconsistent diagnostic support.
- Regulatory and governance gaps emerge. HSA's AI-SaMD framework requires evidence of generalisability; memorisation bias undermines this evidence see our HSA exemption pathway guide.
How to audit for training set overlap before deployment
Before deploying any radiology AI model, Singapore hospitals should conduct a training provenance audit. Here's a practical checklist:
1. Require vendor disclosure of training datasets
Ask the vendor:
- Which datasets were used for pre-training and fine-tuning? (e.g., MIMIC-CXR, CheXpert, NIH ChestX-ray14, proprietary hospital data)
- What are the date ranges of the training data?
- Were any Singapore institutions or regional hospitals included?
- Was the model trained on publicly available datasets that your institution may have contributed to?
If the vendor cannot or will not disclose this, consider it a red flag for clinical AI governance.
2. Test on a strictly held-out cohort
Create a validation cohort with:
- No patient overlap with any data collected before the model's training cutoff date.
- Temporal separation: if the model was trained on data up to 2024, validate only on 2025–2026 scans.
- Demographic diversity: include patients from different age groups, ethnicities, and clinical pathways to test generalisability across Singapore's multi-ethnic population.
This is more stringent than typical train/test splits, which often shuffle data from the same time period and patient pool.
3. Compare performance on returning vs. new patients
Stratify your validation cohort into:
- Returning patients: those with prior imaging studies before the model's training cutoff.
- New patients: first-time imaging or patients with no prior scans in the training window.
If the model performs significantly better on returning patients, memorisation bias is likely. We recommend a performance delta threshold: if AUC differs by >0.05 between cohorts, conduct further investigation before deployment.
4. Use membership inference tests
Membership inference attacks can detect whether a specific patient's data was in the training set [6]. While typically used for privacy audits, these techniques can also quantify memorisation risk. Academic collaborators or third-party auditors can run these tests if your team lacks in-house expertise.
What about foundation models for radiology?
Vision-language foundation models for medical imaging are increasingly common [7]. A recent preprint describes a brain tumour diagnosis model trained on multimodal MRI data [7]. These models are powerful but pose amplified memorisation risk because:
- They are trained on massive, often opaque datasets aggregated from multiple institutions.
- Fine-tuning on your hospital's data may not override memorised patterns from pre-training.
- Soft prompting and few-shot adaptation methods [8] can improve performance but do not eliminate memorisation.
Before deploying a foundation model, Singapore hospitals should:
- Request a data provenance statement from the vendor, including all pre-training sources.
- Conduct adversarial testing with edge cases and rare pathologies not well-represented in public datasets.
- Monitor performance stratified by patient novelty (new vs. returning) post-deployment.
For more on foundation model deployment, see our medical imaging foundation models guide.
Continuous monitoring: detecting memorisation drift post-deployment
Memorisation bias is not static. As your patient population evolves—new demographics, emerging pathologies, equipment upgrades—the model's reliance on memorised patterns may become more apparent.
We recommend:
- Monthly performance audits stratified by patient novelty (first scan vs. follow-up).
- Drift detection using statistical tests (e.g., Kolmogorov-Smirnov) to compare feature distributions between training and deployment cohorts.
- Clinician feedback loops: radiologists should flag cases where AI performance seems inconsistent, especially for returning patients.
This aligns with our broader framework for continuous monitoring of shortcut bias.
Why this matters in Singapore and Asia
Singapore's healthcare AI ecosystem is maturing rapidly. MOH's National AI Strategy and HSA's evolving AI-SaMD framework emphasise evidence-based deployment and post-market surveillance. Memorisation bias undermines both:
- Evidence quality: inflated validation metrics misrepresent real-world performance, violating HSA's generalisability requirements.
- Patient safety: inconsistent AI performance across patient cohorts creates unequal diagnostic support, particularly problematic in Singapore's multi-ethnic, multi-morbidity population.
- Vendor accountability: without training provenance disclosure, hospitals cannot assess memorisation risk, weakening procurement governance.
Regionally, Asia-Pacific hospitals often validate AI on smaller, single-site cohorts due to data-sharing constraints. This increases memorisation risk because validation and training data are more likely to overlap. Federated learning and privacy-preserving techniques can help see our federated learning guide, but they do not eliminate the need for held-out testing.
What to do next
If you're procuring or deploying radiology AI in Singapore:
- Add training provenance disclosure to your vendor RFP template. Require vendors to list all training datasets, date ranges, and institutional sources.
- Create a held-out validation cohort with strict temporal and patient separation from any data collected before the model's training cutoff.
- Stratify validation by patient novelty (new vs. returning) and flag models with >0.05 AUC delta for further audit.
- Implement continuous monitoring post-deployment, tracking performance separately for first-time and follow-up scans.
- Engage clinical informatics and governance teams early—memorisation bias is a cross-functional issue spanning procurement, IT, radiology, and quality assurance.
For help designing a memorisation audit or building continuous monitoring infrastructure, start a project with InsytAI. Our clinical AI services include vendor evaluation, held-out testing, and post-deployment surveillance for Singapore hospitals.
FAQ
What's the difference between memorisation bias and overfitting?
Overfitting occurs when a model learns noise in the training data and performs poorly on any new data. Memorisation bias is more specific: the model stores individual training examples and performs well when those same examples (or similar patients) reappear, but poorly on truly novel cases. A model can memorise without overfitting if it generalises well within the training distribution but fails outside it.
How common is training set overlap in commercial radiology AI?
Very common. Many vendors train on public datasets like MIMIC-CXR, CheXpert, or NIH ChestX-ray14. If your hospital contributed de-identified data to these datasets—or if your patient demographics overlap—your validation cohort may include patients the model has "seen." Always ask vendors for training provenance.
Can federated learning prevent memorisation bias?
Federated learning reduces data centralisation but does not eliminate memorisation. A model trained via federated learning can still memorise patterns from individual hospital sites. You still need held-out testing with strict patient and temporal separation. See our federated learning governance guide for more.
Should we avoid all models trained on public datasets?
No. Public datasets are valuable for pre-training and benchmarking. The key is transparency and held-out testing. If a vendor discloses training sources and you validate on a strictly separated cohort, memorisation risk is manageable. The problem arises when vendors obscure training provenance or hospitals validate on overlapping data.
Sources
[1] Assessing body composition via a smartphone computer vision application: High repeatability but method-dependent agreement compared with BODPOD and Inbody. PLOS Digital Health, 2026-09-02. https://journals.plos.org/digitalhealth/article?id=10.1371/journal.pdig.0000983
[2] Memorisation bias in medical AI. arXiv preprint, 2026-09-15. https://arxiv.org/abs/2609.17223v1
[3] A Vision-Language Foundation Model for Precise and Comprehensive Brain Tumor Diagnosis from Preoperative Multimodal Data. arXiv preprint, 2026-09-15. https://arxiv.org/abs/2609.16597v1
[4] Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models. arXiv preprint, 2026-09-10. https://arxiv.org/abs/2609.11310v1