AI-Assisted Systematic Reviews in Healthcare Research: What Singapore Institutions Need to Know

Systematic reviews are the backbone of evidence-based medicine, but they're notoriously time-intensive. A single review can take 12–18 months and cost upwards of SGD 100,000 in researcher time. This month, a peer-reviewed proof-of-concept in PLOS Digital Health demonstrated that prompt-driven generative AI can meaningfully accelerate systematic review workflows in digital psychiatry [1]. For Singapore healthcare researchers, clinical informatics teams, and research offices, this raises immediate questions: when is AI assistance appropriate, what governance is required, and how do we validate outputs without introducing bias?

This post unpacks the current state of AI-assisted systematic reviews, examines the new JAMA guidance on AI use in medical publishing [3], and provides a decision framework for Singapore institutions considering these tools.

Key takeaways

  • Proof-of-concept validated: A stage-matched comparative study shows prompt-driven LLMs can assist in screening, data extraction, and synthesis phases of systematic reviews in digital psychiatry [1]
  • Publisher guidance is tightening: JAMA's updated September 2026 guidance requires explicit disclosure of AI use, prohibits AI authorship, and mandates human verification of all AI-generated content [3]
  • Governance precedes adoption: Singapore institutions need clear protocols for AI tool selection, output validation, and disclosure before deploying these methods
  • Privacy and data residency matter: Many commercial LLM APIs process data outside Singapore; research involving patient data or unpublished findings requires privacy-preserving alternatives
  • Validation workload shifts but doesn't disappear: AI assistance changes the bottleneck from manual screening to systematic output verification

What changed in September 2026?

Two significant publications frame the current moment. First, the PLOS Digital Health study [1] provides stage-matched evidence that prompt-driven generative AI can assist across multiple systematic review phases—title/abstract screening, full-text screening, data extraction, and narrative synthesis—in digital psychiatry research. The authors used a structured prompt engineering approach and compared AI-assisted workflows against traditional manual methods.

Second, JAMA published updated guidance for author use of AI in medical publication [3]. The guidance acknowledges AI as "a tool that can be helpful to researchers and authors by reducing the burdens of some tasks and improving efficiency," but emphasizes that "as use of AI in scientific publishing grows exponentially, so do concerns about such use." The updated policy requires:

  • Explicit disclosure of AI tool use in methods sections
  • Prohibition of AI as author or co-author
  • Human verification and accountability for all AI-generated content
  • Transparency about which specific tasks involved AI assistance

For Singapore researchers, this creates a clear regulatory floor: any systematic review using AI assistance must document the tools, prompts, validation methods, and human oversight applied.

Why systematic reviews are a natural AI use case

Systematic reviews involve repetitive, rule-based tasks at scale—exactly where LLMs can provide leverage:

  1. Title/abstract screening: Applying inclusion/exclusion criteria to thousands of abstracts
  2. Data extraction: Pulling standardized fields (sample size, intervention, outcomes) from full-text papers
  3. Quality assessment: Applying structured checklists like PRISMA or Cochrane Risk of Bias tools
  4. Synthesis support: Drafting narrative summaries of findings across studies

The PLOS Digital Health study [1] demonstrated feasibility across all four stages in digital psychiatry, a domain with rapidly growing literature. However, the authors noted that "reliability score" (a measure of output consistency and accuracy) varied by task, with screening tasks showing higher reliability than synthesis tasks.

For Singapore institutions, this suggests a tiered adoption model: start with high-volume, low-ambiguity tasks (screening), validate extensively, then expand to more complex tasks (synthesis) only after establishing robust quality control.

Where governance gaps create risk

The JAMA guidance [3] addresses publication ethics but leaves operational questions unanswered:

Data residency and privacy

Most commercial LLM APIs (OpenAI, Anthropic, Google) process data in non-Singapore jurisdictions. If your systematic review involves:
- Unpublished patient-level data
- Proprietary hospital datasets
- Pre-publication manuscripts under review

...then default API use may violate PDPA, institutional data governance policies, or research ethics protocols. Singapore institutions need to evaluate:
- On-premise or Singapore-hosted LLM options (e.g., Azure OpenAI with Singapore region selection)
- Federated or privacy-preserving architectures for multi-site reviews [6, 11]
- Contractual data processing agreements with API providers

We've seen Singapore hospital research offices block AI-assisted review tools not because of accuracy concerns, but because the data governance paperwork wasn't in place.

Validation protocols

The PLOS Digital Health study [1] used stage-matched comparison—human reviewers independently completed the same tasks, and outputs were compared. This is rigorous but resource-intensive. For operational deployment, Singapore institutions need lighter-weight validation protocols:

  • Dual-review sampling: AI completes the full task; humans review a stratified random sample (e.g., 10% of excluded abstracts, 20% of extracted data fields)
  • Disagreement flagging: AI outputs confidence scores; human reviewers focus on low-confidence items
  • Checklist-driven audits: Structured quality checks (e.g., "Did the AI correctly apply exclusion criterion #3?") rather than full re-review

Without documented validation protocols, AI-assisted reviews risk rejection at journal submission or institutional review board (IRB) audit.

Prompt and model versioning

LLM outputs are non-deterministic and model-dependent. A systematic review conducted with GPT-4 in March 2026 may yield different results with GPT-4 in September 2026 if the model was updated. Singapore institutions need:
- Prompt libraries: Version-controlled, peer-reviewed prompts for common review tasks
- Model pinning: Use specific model versions (e.g., gpt-4-0613) rather than rolling gpt-4 endpoints
- Reproducibility documentation: Log prompts, model versions, temperature settings, and API timestamps

This aligns with broader clinical AI deployment practices we've implemented in Singapore hospitals: if you can't reproduce the output, you can't validate it.

Why this matters in Singapore and Asia

Singapore's healthcare research ecosystem is globally connected but resource-constrained. Academic medical centers produce high-quality systematic reviews, but researcher time is expensive and competing priorities (clinical duties, grant writing, teaching) create bottlenecks.

AI-assisted systematic reviews offer leverage, but adoption is uneven:

  • Public healthcare clusters have institutional review boards, data governance frameworks, and research offices that can standardize AI use policies
  • Smaller institutions and private hospitals often lack dedicated research infrastructure and may adopt tools ad hoc, creating compliance and quality risks
  • Regional collaboration (ASEAN, Asia-Pacific multi-site reviews) introduces additional data residency and governance complexity [6, 11]

The JAMA guidance [3] and PLOS Digital Health evidence [1] provide a foundation, but Singapore institutions need localized implementation frameworks that account for PDPA, MOH data-sharing guidelines, and institutional IRB requirements.

We've worked with Singapore health systems to develop AI tool evaluation checklists for research applications. The most common failure mode isn't technical—it's governance: tools adopted by individual researchers without institutional vetting, creating downstream compliance issues when manuscripts are submitted or audits occur.

What to do next

If your Singapore institution is considering AI-assisted systematic reviews:

  1. Establish institutional policy first: Work with your research office, IRB, and data governance team to define acceptable use, required disclosures, and validation standards before individual researchers adopt tools
  2. Start with screening tasks: Title/abstract screening has the highest reliability [1] and lowest risk; validate extensively before expanding to data extraction or synthesis
  3. Document everything: Prompts, model versions, validation protocols, and human oversight must be reproducible and auditable; this is non-negotiable for journal submission under the new JAMA guidance [3]
  4. Evaluate privacy-preserving options: If your reviews involve unpublished or sensitive data, assess Singapore-hosted LLM options or federated architectures [6, 11] before using public APIs
  5. Build prompt libraries: Develop and peer-review standardized prompts for common review tasks (PICO extraction, risk-of-bias assessment); version-control and share across your institution

For institutions ready to pilot AI-assisted reviews with robust governance, start a project with our team. We provide AI tool evaluation, validation protocol design, and governance framework implementation for Singapore healthcare research offices.

FAQ

Can AI replace human reviewers in systematic reviews?

No. The JAMA guidance [3] explicitly prohibits AI authorship, and the PLOS Digital Health study [1] demonstrates AI as an assistive tool, not a replacement. Human reviewers remain accountable for inclusion/exclusion decisions, data accuracy, and synthesis quality. AI shifts the workload from manual screening to systematic validation, but doesn't eliminate human judgment.

What about hallucinations and fabricated citations?

This is a critical risk. LLMs can generate plausible-sounding but non-existent citations, study details, or statistical results. Validation protocols must include:
- Citation verification: Cross-check every AI-extracted citation against the original source
- Data field audits: Verify extracted data (sample sizes, effect sizes, p-values) against full-text papers
- Synthesis review: Human experts must review and revise AI-generated narrative summaries

The PLOS Digital Health study [1] used stage-matched comparison to detect these errors; operational deployments need lighter-weight but still rigorous checks.

Do I need IRB approval to use AI in systematic reviews?

It depends. If your review involves only published literature (standard systematic review), IRB approval may not be required, but institutional research office notification is advisable. If your review involves unpublished data, patient records, or proprietary datasets, IRB review is likely required, and you'll need to document AI tool use, data processing locations, and validation methods in your protocol.

What's the current state of federated learning for multi-site reviews?

Federated learning enables collaborative analysis without sharing patient-level data—highly relevant for multi-site systematic reviews or meta-analyses involving unpublished datasets. Recent work [6, 11] demonstrates operational frameworks, but adoption in healthcare research remains limited. For Singapore institutions collaborating regionally, federated approaches may be necessary to satisfy cross-border data governance requirements. See our previous post on federated learning for hospital data governance for implementation considerations.

Sources

[1] PLOS Digital Health (2026). "Leveraging prompt-driven generative AI for systematic reviews in digital psychiatry: A stage-matched comparative proof-of-concept for healthcare researchers and clinicians." Published September 10, 2026. https://journals.plos.org/digitalhealth/article?id=10.1371/journal.pdig.0001666

[2] JAMA Network (2026). "The Organoid Era." Published September 15, 2026. https://jamanetwork.com/journals/jama/fullarticle/2853277

[3] JAMA Network (2026). "Updated Guidance for Author Use of AI in Medical Publication." Published September 15, 2026. https://jamanetwork.com/journals/jama/fullarticle/2852772

[6] arXiv (2026). "Multi-center Medical Data Mining with FL-Net - A One-stop Shop for Federated Learning." Preprint published September 17, 2026. https://arxiv.org/abs/2609.20650v1

[11] PubMed — Journal of the American Medical Informatics Association (2026). "PDA (Privacy-Preserving Distributed Algorithms) in action: ten principles for high-quality multi-site clinical evidence generation." Published September 1, 2026. https://pubmed.ncbi.nlm.nih.gov/42440280/