Gommers, J. et al. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study: a randomised, controlled, non-inferiority, single-blinded, population-based, screening-accuracy trial. Lancet 407, 505–514 (2026). Article PubMed Google Scholar Lu, M. Y. et al. A multimodal generative AI copilot for human pathology. Nature 643, 466–473
Gommers, J. et al. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study: a randomised, controlled, non-inferiority, single-blinded, population-based, screening-accuracy trial. Lancet 407, 505–514 (2026).
Google Scholar
Lu, M. Y. et al. A multimodal generative AI copilot for human pathology. Nature 643, 466–473 (2024).
Google Scholar
Singh, R., Bapna, M., Diab, A. R., Ruiz, E. S. & Lotter, W. How AI is used in FDA-authorized medical devices: a taxonomy across 1,016 authorizations. NPJ Digit. Med. 8, 388 (2025).
Google Scholar
Griot, M., Vanderdonckt, J. & Yuksel, D. Implementation of large language models in electronic health records. PLOS Digital Health https://doi.org/10.1371/journal.pdig.0001141 (2025).
Armitage, H. Clinicians can ‘chat’ with medical records through new AI software, ChatEHR. Stanford Medicine https://med.stanford.edu/news/all-news/2025/06/chatehr.html (2025).
Moor, M. et al. Foundation models for generalist medical artificial intelligence. Nature 616, 259–265 (2023).
Google Scholar
Clusmann, J. et al. The future landscape of large language models in medicine. Commun. Med. 3, 141 (2023).
Google Scholar
Bubeck, S. et al. Sparks of Artificial General Intelligence: Early experiments with GPT-4. Preprint at https://doi.org/10.48550/arXiv.2303.12712 (2023).
Singhal, K. et al. Toward expert-level medical question answering with large language models. Nat. Med. 31, 943–950 (2025).
Google Scholar
Gu, Y. et al. The illusion of readiness: stress testing large frontier models on multimodal medical benchmarks. Preprint at https://doi.org/10.48550/arXiv.2509.18234 (2025). This study reveals that medical LLM benchmarks may overestimate real-world readiness as benchmarks fail to capture brittleness and reasoning flaws.
Omar, M. et al. Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support. Commun. Med. 5, 330 (2025).
Google Scholar
Bedi, S., Jiang, Y., Chung, P., Koyejo, S. & Shah, N. Fidelity of medical reasoning in large language models. JAMA Netw. Open 8, e2526021 (2025).
Google Scholar
Goh, E. et al. GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial. Nat. Med. 31, 1233–1238 (2025).
Google Scholar
McDuff, D. Towards accurate differential diagnosis with large language models. Nature 642, 451–457 (2025).
Google Scholar
Tu, T. et al. Towards conversational diagnostic artificial intelligence. Nature 642, 442–450 (2025).
Google Scholar
Palepu, A., et al. Exploring large language models for specialist-level oncology care. NEJM AI 2, AIcs2500025 (2025).
Google Scholar
Singhal, K. et al. Large language models encode clinical knowledge. Nature 620, 172–180 (2023).
Google Scholar
Thirunavukarasu, A. J. et al. Large language models in medicine. Nat. Med. 29, 1930–1940 (2023).
Google Scholar
Van Veen, D. et al. Adapted large language models can outperform medical experts in clinical text summarization. Nat. Med. 30, 1134–1142 (2024).
Google Scholar
Tierney, A. A. et al. Ambient artificial intelligence scribes: learnings after 1 year and over 2.5 million uses. NEJM Catal. Innov. Care Deliv. https://doi.org/10.1056/CAT.25.0040 (2025).
Google Scholar
Heinz, M. V. et al. Randomized trial of a generative AI chatbot for mental health treatment. NEJM AI 2, AIoa2400802 (2025).
Comanici, G. et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. Preprint at https://doi.org/10.48550/arXiv.2507.06261 (2025).
Zakka, C. et al. Almanac — Retrieval-Augmented Language Models for Clinical Medicine. NEJM AI 1, AIoa2300068 (2024).
Google Scholar
Wiest, I. C. et al. Large language models for clinical decision support in gastroenterology and hepatology. Nat. Rev. Gastroeneterol. Hepatol. 22, 773–787 (2025).
Google Scholar
Ferber, D. et al. Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology. Nat. Cancer 6, 1337–1349 (2025). This study presents one of the first modular LLM-agent frameworks for healthcare, integrating biomedical tools and knowledge bases for clinical decision-making.
Google Scholar
Ge, J. et al. Development of a liver disease-specific large language model chat interface using retrieval-augmented generation. Hepatology 80, 1158–1168 (2024).
Google Scholar
Kresevic, S. et al. Optimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation-based framework. NPJ Digit. Med. 7, 102 (2024).
Google Scholar
Masanneck, L., Meuth, S. G. & Pawlitzki, M. Evaluating base and retrieval augmented LLMs with document or online support for evidence based neurology. NPJ Digit. Med. 8, 137 (2025).
Google Scholar
Zakka, C. et al. Almanac Copilot: towards autonomous electronic health record navigation. Preprint at https://doi.org/10.48550/arXiv.2405.07896 (2024).
Guo, D. et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, 633–638 (2025). This study shows that reinforcement learning can markedly improve LLM reasoning performance and introduces DeepSeek-R1 as a major open-weight model and one of few LLMs in peer-reviewed literature.
Google Scholar
Angus, D. C. et al. AI, health, and health care today and tomorrow: the JAMA Summit report on artificial intelligence. JAMA 334, 1650–1664 (2025).
Google Scholar
García-García, D., León-Gómez, I., Pérez-Marín, L. & Gómez-Barroso, D. Exploring all-cause mortality surveillance during the Iberian Peninsula power outage, Spain, 28 April 2025. Eurosurveillance 30, 2500405 (2025).
Google Scholar
Larsen, E., Fong, A., Wernz, C. & Ratwani, R. M. Implications of electronic health record downtime: an analysis of patient safety event reports. J. Am. Med. Inform. Assoc. 25, 187–191 (2018).
Google Scholar
Martin, G., Ghafur, S., Kinross, J., Hankin, C. & Darzi, A. WannaCry—a year on. BMJ 361, k2381 (2018).
Google Scholar
Law, R. Cyberattacks on healthcare: Russia’s tool for mass disruption. Medical Device Network https://www.medicaldevice-network.com/features/cyberattacks-on-healthcare-russias-tool-for-mass-disruption/ (2024).
Cartwright, A. J. The elephant in the room: cybersecurity in healthcare. J. Clin. Monit. Comput. 37, 1123–1132 (2023).
Google Scholar
Gordon, W. J. et al. Assessment of employee susceptibility to phishing attacks at US health care institutions. JAMA Netw. Open 2, e190393 (2019).
Google Scholar
Li, S., Surineni, K. & Prabhakaran, N. Cyber-attacks on hospital systems: a narrative review. Am. J. Ger. Psychiatry 7, 30–39 (2025).
Bowsher, G., Sullivan, R. & Lentzos, F. Tackling health disinformation in conflict settings. Lancet 405, 1052 (2025).
Google Scholar
Menz, B. D., Modi, N. D., Sorich, M. J. & Hopkins, A. M. Health disinformation use case highlighting the urgent need for artificial intelligence vigilance: weapons of mass disinformation: weapons of mass disinformation. JAMA Intern. Med. 184, 92–96 (2024).
Google Scholar
Perakslis, E. D., Ranney, M. L. & Goldsack, J. C. Characterizing cyber harms from digital health. Nat. Med. 29, 528–531 (2023).
Google Scholar
Nweke, L. O. Using the CIA and AAA Models to Explain Cybersecurity Activities. Preprint at https://pmworldlibrary.net/wp-content/uploads/2017/05/171126-Nweke-Using-CIA-and-AAA-Models-to-explain-Cybersecurity.pdf (2017).
Nasr, M. et al. Scalable Extraction of Training Data from Algined, Production Language Models. ICLR 2025 https://openreview.net/forum?id=vjel3nWP2a (2025).
Jha, R., Zhang, C., Shmatikov, V. & Morris, J. X. Harnessing the universal geometry of embeddings. NeurIPS 2025 Conference https://neurips.cc/virtual/2025/loc/san-diego/poster/116441 (NeurIPS, 2025).
Finlayson, S. G., Chung, H. W., Kohane, I. S. & Beam, A. L. Adversarial attacks against medical deep learning systems. Preprint at https://doi.org/10.48550/arXiv.1804.05296 (2018). This early and influential work identifies adversarial attacks as a fundamental risk for medical machine learning in healthcare.
Chakraborty, A., Alam, M., Dey, V., Chattopadhyay, A. & Mukhopadhyay, D. Adversarial attacks and defences: a survey. Preprint at https://doi.org/10.48550/arXiv.1810.00069 (2018).
Xu, H. et al. Adversarial attacks and defenses in images, graphs and text: a review. Int. J. Autom. Comput. 17, 151–178 (2020).
Google Scholar
Barreno, M., Nelson, B., Joseph, A. D. & Tygar, J. D. The security of machine learning. Mach. Learn. 81, 121–148 (2010).
Google Scholar
OWASP Top 10 for LLM Applications 2025 https://genai.owasp.org/resource/owasp-top-10-for-llm-applications-2025/ (OWASP, 2024). This resource provides a continuously updated overview of the most critical safety and security risks in LLMs.
CVSS v4.0 Specification Document. FIRST https://www.first.org/cvss/specification-document (2023).
Kaviani, S., Han, K. J. & Sohn, I. Adversarial attacks and defenses on AI in medical imaging informatics: A survey. Expert Syst. Appl. 198, 116815 (2022).
Google Scholar
Nagaraja, N. & Bahsi, H. Cyber threat modeling of an LLM-based healthcare system. In Proc. 11th International Conference on Information Systems Security and Privacy 325–336 (SCITEPRESS, 2025).
Li, M. Q. & Fung, B. C. M. Security concerns for large language models: a survey. Journal of Information Security and Applications, Volume 95, 2025 https://doi.org/10.1016/j.jisa.2025.104284 (2025).
Reason, J. Human error: models and management. Brit. Med. J. 320, 768–770 (2000).
Google Scholar
Joint Task Force Transformation Initiative. Guide for Conducting Risk Assessments https://doi.org/10.6028/nist.sp.800-30r1 (US Department of Commerce, 2012).
Ashkenazy, S. Why time-to-market is overtaking device security as a top priority. Cybellum https://cybellum.com/blog/why-time-to-market-is-overtaking-device-security-as-a-top-priority-in-2025/ (2024).
Piao, Y., Li, J. & Woods, D. W. Measuring the vulnerability disclosure policies of AI vendors. Preprint at https://doi.org/10.48550/arXiv.2509.06136 (2025).
Bommasani, R. et al. The 2024 Foundation Model Transparency Index. Preprint at https://doi.org/10.48550/arXiv.2407.12929 (2024).
Anthropic’s responsible scaling policy. Anthropic https://www.anthropic.com/rsp-updates (2026).
Lindsey, J. et al. On the biology of a large language model. Anthropic https://transformer-circuits.pub/2025/attribution-graphs/biology.html (2025). This work provides an extensive analysis of internal mechanisms in LLMs and introduces open-source tools to advance mechanistic interpretability.
Petri: an open-source auditing tool to accelerate AI safety research. Anthropic https://www.anthropic.com/research/petri-open-source-auditing (2025).
Joglekar, M. et al. Training LLMs for honesty via confessions. Preprint at https://doi.org/10.48550/arXiv.2512.08093 (2025).
Villalobos, P. et al. Position: will we run out of data? Limits of LLM scaling based on human-generated data. ICML 235, 49523–49544 (2024).
Alber, D. A. et al. Medical large language models are vulnerable to data-poisoning attacks. Nat. Med. 31, 618–626 (2025). This study demonstrates that even minimal data poisoning can induce clinically harmful behaviour in medical LLMs.
Google Scholar
Hubinger, E. et al. Sleeper agents: training deceptive LLMs that persist through safety training. Preprint at https://doi.org/10.48550/arXiv.2401.05566 (2024). This study characterizes difficult-to-detect and persistent ‘sleeper agent’ backdoor behaviours in LLMs.
Souri, H., Fowl, L. H., Chellappa, R., Goldblum, M. & Goldstein, T. in Advances in Neural Information Processing Systems 35 (eds Koyejo, S. et al.) 19165–19178 (NeurIPS, 2022).
Carlini, N. et al. Poisoning Web-Scale Training Datasets is Practical. in 2024 IEEE Symposium on Security and Privacy Vol. 29, 407–425 (IEEE, 2024).
Clusmann, J. et al. Incidental prompt injections on vision–language models in real-life histopathology. NEJM AI 2, AIcs2500078 (2025).
Google Scholar
Han, T. et al. Medical large language models are susceptible to targeted misinformation attacks. npj Digit. Med. 7, 288 (2024). This study demonstrates the feasibility of targeted model weight manipulation to introduce incorrect biomedical information into LLMs.
Google Scholar
Rao, P. S. B., Šćepanović, S., Jayagopi, D. B., Cherubini, M. & Quercia, D. The AI model risk catalog: what developers and researchers miss about real-world AI harms. In Proc. AAAI/ACM Conference on AI, Ethics, and Society Vol. 8, 2163–2150 (AAAI/ACM, 2025).
Carlini, N. et al. Poisoning web-scale training datasets is practical. ar5iv https://ar5iv.labs.arxiv.org/html/2302.10149 (2024).
Clusmann, J. et al. Prompt injection attacks on vision language models in oncology. Nat. Commun. 16, 1239 (2025). This study demonstrates prompt injection attacks on LLMs through hidden prompt instructions on medical imaging data.
Google Scholar
Zhang, Z., Qadir, M.I., Carstens, M. et al. Prompt injection attacks on vision-language models for surgical decision support. npj Digit. Surg. 1, 15 https://doi.org/10.1038/s44484-026-00014-6 (2026).
Lapuschkin, S. et al. Unmasking Clever Hans predictors and assessing what machines really learn. Nat. Commun. 10, 1096 (2019).
Google Scholar
Hou, G. et al. Evaluating robustness of large audio language models to audio injection: an empirical study. Preprint at https://doi.org/10.48550/arXiv.2505.19598 (2025).
Rajeev, M. et al. Cats confuse reasoning LLM: query agnostic adversarial triggers for reasoning models. Preprint at https://doi.org/10.48550/arXiv.2503.01781 (2025).
Alizadeh, M., Samei, Z., Stetsenko, D. & Gilardi, F. Simple prompt injection attacks can leak personal data observed by LLM agents during task execution. Preprint at https://doi.org/10.48550/arXiv.2506.01055 (2025).
Meincke, L. et al. Persuading large language models to comply with objectionable requests. Proc. Natl Acad. Sci. USA 123, e2535868123 (2025).
Zeng, Y. et al. How Johnny can persuade LLMs to jailbreak them: rethinking persuasion to challenge AI safety by humanizing LLMs. In Proc. 62nd Meeting of the Association for Computational Linguistics Vol. 1, 14322–14350 (ACL, 2024).
Schoene, A. M. & Canca, C. ‘For argument’s sake, show me how to harm myself!’: Jailbreaking LLMs in suicide and self-harm contexts. In 2025 IEEE International Symposium on Technology and Society https://doi.org/10.1109/ISTAS65609.2025.11269647 (IEEE, 2025).
Mehrotra, A. et al. Tree of Attacks: Jailbreaking Black-Box LLMs Automatically. In Advances in Neural Information Processing Systems 37 (eds Globerson, A. et al.) 61065–61105 (NeurIPS, 2024).
Jiang, F. et al. ArtPrompt: ASCII art-based jailbreak attacks against aligned LLMs. Preprint at https://doi.org/10.48550/arXiv.2402.11753 (2024).
Zhang, H., Lou, Q. & Wang, Y. Towards safe AI clinicians: a comprehensive study on large language model jailbreaking in healthcare. Preprint at https://doi.org/10.48550/arXiv.2501.18632 (2025).
Sharma, M. et al. Constitutional classifiers: defending against universal jailbreaks across thousands of hours of red teaming. Preprint at https://doi.org/10.48550/arXiv.2501.18837 (2025).
Yuan, Y. et al. From hard refusals to safe-completions: toward output-centric safety training. Superintelligence https://doi.org/10.70777/si.v2i6.15625 (2025).
Google Scholar
Zhang, Y., Juels, A., Reiter, M. K. & Ristenpart, T. Cross-VM side channels and their use to extract private keys. In Proc. 2012 ACM Conference on Computer and Communications Security https://doi.org/10.1145/2382196.2382230 (ACM, 2012).
Karim, H., Gupta, D. & Sitharaman, S. Securing LLM workloads with NIST AI RMF in the internet of robotic things. IEEE Access 13, 69631–69649 (2025).
Google Scholar
Huang, H., Meng, T. & Jia, W. Joint optimization of prompt security and system performance in Edge-Cloud LLM systems. Preprint at https://doi.org/10.48550/arXiv.2501.18663 (2025).
Schmotz, D., Abdelnabi, S. & Andriushchenko, M. Agent skills enable a new class of realistic and trivially simple prompt injections. Preprint at https://doi.org/10.48550/arXiv.2510.26328 (2025).
Narajala, V. S. & Habler, I. Enterprise-grade security for the model context protocol (MCP): frameworks and mitigation strategies. Preprint at https://doi.org/10.48550/arXiv.2504.08623 (2025).
Thoonsen, A. C. et al. Stimulating implementation of clinical practice guidelines in hospital care from a central guideline organization perspective: A systematic review. Health Policy 148, 105135 (2024).
Google Scholar
Clusmann, J. et al. The barriers for uptake of artificial intelligence in hepatology and how to overcome them. J. Hepatol. 83, 1410–1426 (2025).
Google Scholar
Wong, E. Y. T. et al. ESMO guidance on the use of large language models in clinical practice (ELCAP). Ann. Oncol. 36, 1447–1457 (2025).
Google Scholar
Sarthi, P. et al. RAPTOR: recursive abstractive processing for tree-organized retrieval. ICLR https://iclr.cc/virtual/2024/poster/19034 (2024).
Greshake, K. et al. Not what you’ve signed up for: compromising real-world LLM-integrated applications with indirect prompt injection. Preprint at https://doi.org/10.48550/arXiv.2302.12173 (2023).
Continella, A., Polino, M., Pogliani, M. & Zanero, S. There’s a hole in that bucket!: a large-scale analysis of misconfigured S3 buckets. In Proc. 34th Annual Computer Security Applications Conference 702–711 (ACM, 2018).
Cruz, F. & Lombrozo, T. How laypeople evaluate scientific explanations containing jargon. Nat. Hum. Behav. 9, 2038–2053 (2025).
Google Scholar
Measuring the Persuasiveness of Language Models. Anthropic https://www.anthropic.com/news/measuring-model-persuasiveness (2024).
Salvi, F., Horta Ribeiro, M., Gallotti, R. & West, R. On the conversational persuasiveness of GPT-4. Nat. Hum. Behav. 9, 1645–1653 (2025).
Google Scholar
Shekar, S., Pataranutaporn, P., Sarabu, C., Cecchi, G. A. & Maes, P. People overtrust AI-generated medical advice despite low accuracy. NEJM AI 2, AIoa2300015 (2025).
Google Scholar
Li, W. et al. Can a large language model be a gaslighter? In The 13th International Conference on Learning Representations (eds Yue, Y. et al.) https://proceedings.iclr.cc/paper_files/paper/2025/file/0769598fdeb4f23ee86fec1bc0777f44-Paper-Conference.pdf (ICLR, 2024).
Yeung, J. A., Dalmasso, J., Foschini, L., Dobson, R. J. B. & Kraljevic, Z. The psychogenic machine: simulating AI psychosis, delusion reinforcement and harm enablement in large language models. Preprint at https://doi.org/10.48550/arXiv.2509.109702025 (2025).
Turpin, M., Michael, J., Perez, E. & Bowman, S. R. Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting. Preprint at https://doi.org/10.48550/arXiv.2305.04388 (2023).
Baker, B. et al. Monitoring reasoning models for misbehavior and the risks of promoting obfuscation. Preprint at https://doi.org/10.48550/arXiv.2503.11926 (2025).
Perez, E. et al. Discovering language model behaviors with model-written evaluations. Preprint at https://doi.org/10.48550/arXiv.2212.09251 (2022).
Xiao, J. et al. On the algorithmic bias of aligning large language models with RLHF: Preference collapse and matching regularization. Preprint at https://doi.org/10.48550/arXiv.2405.16455 (2024).
Chen, S. et al. When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior. NPJ Digit. Med. 8, 605 (2025). This study identifies sycophantic behaviour in LLMs as a potential source of false medical information in healthcare settings.
Google Scholar
Kalai, A. T., Nachum, O., Vempala, S. S. & Zhang, E. Why language models hallucinate. Preprint at https://doi.org/10.48550/arXiv.2509.04664 (2025). This study characterizes hallucinations in LLMs as a consequence of misaligned training objectives that reward plausible guessing over admitting uncertainty.
Sharma, M. et al. Towards understanding sycophancy in language models. In The 12th International Conference on Learning Representations (eds Kim, B. et al.) https://proceedings.iclr.cc/paper_files/paper/2024/file/0105f7972202c1d4fb817da9f21a9663-Paper-Conference.pdf (ICLR, 2023).
Cheng, M. et al. Sycophantic AI decreases prosocial intentions and promotes dependence. Science 391, eaec8352 (2026).
Google Scholar
Miton, H., Claidière, N. & Mercier, H. Universal cognitive mechanisms explain the cultural success of bloodletting. Evol. Hum. Behav. 36, 303–312 (2015).
Google Scholar
Saposnik, G., Redelmeier, D., Ruff, C. C. & Tobler, P. N. Cognitive biases associated with medical decisions: a systematic review. BMC Med. Inform. Decis. Mak. 16, 138 (2016).
Google Scholar
Huang, L. et al. A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. ACM Trans. Inf. Syst. 43, 42 (2025). This study provides a systematic analysis and taxonomy of hallucinations in LLMs.
Google Scholar
Ferber, D. et al. In-context learning enables multimodal large language models to classify cancer pathology images. Nat. Commun. 15, 10104 (2024).
Google Scholar
Gourabathina, A., Gerych, W., Pan, E. & Ghassemi, M. The medium is the message: how non-clinical information shapes clinical decisions in LLMs. In Proc. 2025 ACM Conference on Fairness, Accountability, and Transparency 1805–1828 (ACM, 2025).
Coda-Forno, J. et al. Inducing anxiety in large language models can induce bias. Preprint at https://doi.org/10.48550/arXiv.2304.11111 (2023).
Corbeil, J.-P., Kim, M., Sordoni, A., Beaulieu, F. & Vozila, P. Medical red teaming protocol of language models: on the importance of user perspectives in healthcare settings. Preprint at https://doi.org/10.48550/arXiv.2507.07248 (2025).
Callahan, A. et al. Standing on FURM ground: a framework for evaluating fair, useful, and reliable AI models in health care systems. NEJM Catal. Innov. Care Deliv. https://doi.org/10.1056/CAT.24.0131 (2024).
Google Scholar
Huang, Y. et al. Position: TrustLLM: trustworthiness in large language models. In Proc. 41st International Conference on Machine Learning 20166–20270 (PMLR, 2024).
Gabriel, I. Artificial intelligence, values, and alignment. Minds Mach. 30, 411–437 (2020).
Google Scholar
Mökander, J., Schuett, J., Kirk, H. R. & Floridi, L. Auditing large language models: a three-layered approach. AI Ethics 4, 1085–1115 (2023).
Google Scholar
Wu, D. et al. First, do NOHARM: towards clinically safe large language models. Preprint at https://doi.org/10.48550/arXiv.2512.01241 (2025).
Bedi, S. et al. Holistic evaluation of large language models for medical tasks with MedHELM. Nat. Med. 32, 943–951 (2026). This work proposes MedHELM, a comprehensive holistic evaluation suite for LLMs for medicine.
Google Scholar
Han, T., Kumar, A., Agarwal, C. & Lakkaraju, H. in ICML 2024 Workshop on Models of Human Feedback for AI Alignment https://icml.cc/virtual/2024/39424 (2024).
Shihab, I. F., Akter, S. & Sharma, A. Detecting and mitigating reward hacking in Reinforcement Learning systems: A comprehensive empirical study. Preprint at https://doi.org/10.48550/arXiv.2507.05619 (2025).
Wu, S. et al. A comparative study on reasoning patterns of OpenAI’s o1 model. Preprint at https://doi.org/10.48550/arXiv.2410.13639 (2024).
El, B. & Zou, J. Moloch’s bargain: emergent misalignment when LLMs compete for audiences. Preprint at https://doi.org/10.48550/arXiv.2510.06105 (2025).
Betley, J. et al. Training large language models on narrow tasks can lead to broad misalignment. Nature 649, 584–589 (2026). This study characterizes ‘emergent misalignment’ and shows that that fine-tuning LLMs on narrow tasks can induce broad, unintended and harmful behaviours.
Google Scholar
Truhn, D., Reis-Filho, J. S. & Kather, J. N. Large language models should be used as scientific reasoning engines, not knowledge databases. Nat. Med. 29, 2983–2984 (2023).
Google Scholar
Asgari, E. et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. NPJ Digit. Med. 8, 274 (2025).
Google Scholar
Kim, Y. et al. Medical hallucinations in foundation models and their impact on healthcare. Preprint at https://doi.org/10.48550/arXiv.2503.05777 (2025).
Bang, Y. et al. HalluLens: LLM hallucination benchmark. In Proc. 63rd Meeting of the Association for Computational Linguistics Vol. 1, 24128–24156 (ACL, 2025).
Soffer, S., Sorin, V., Nadkarni, G. N. & Klang, E. Pitfalls of large language models in medical ethics reasoning. NPJ Digit. Med. 8, 461 (2025).
Google Scholar
Xu, H. et al. Reducing tool hallucination via reliability alignment. In Proc. 42nd International Conference on Machine Learning 2799 (ACM, 2025).
Chung, P. et al. Verifying facts in patient care documents generated by large language models using electronic health records. NEJM AI 3, AIdbp2500418 (2025).
World Medical Association. WMA Declaration of Helsinki—Ethical Principles for Medical Research Involving Human Participants. WMA https://www.wma.net/policies-post/wma-declaration-of-helsinki (2024).
Yu, K.-H., Healey, E., Leong, T.-Y., Kohane, I. S. & Manrai, A. K. Medical artificial intelligence and human values. N. Engl. J. Med. 390, 1895–1904 (2024).
Google Scholar
Mazeika, M. et al. Utility engineering: Analyzing and controlling emergent value systems in AIs. Preprint at https://doi.org/10.48550/arXiv.2502.08640 (2025).
Greenblatt, R. et al. Alignment faking in large language models. Preprint at https://doi.org/10.48550/arXiv.2412.14093 (2024).
Bedi, S. et al. Testing and evaluation of health care applications of large language models: A systematic review: A systematic review. JAMA 333, 319–328 (2024).
Google Scholar
Zack, T. et al. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. Lancet Digit. Health 6, e12–e22 (2024). This study highlights that LLMs can perpetuate racial and gender biases beyond evidence-based variation.
Google Scholar
Omar, M. et al. Sociodemographic biases in medical decision making by large language models. Nat. Med. 31, 1873–1881 (2025). This study assesses sociodemographic biases in medical LLMs, demonstrating differences in clinical decision-making that extend beyond evidence-based variation.
Google Scholar
Gruber, V.-E. et al. A women’s health benchmark for large language models. Preprint at https://doi.org/10.48550/arXiv.2512.17028 (2025).
Yang, J., Soltan, A. A. S., Eyre, D. W., Yang, Y. & Clifton, D. A. An adversarial training framework for mitigating algorithmic biases in clinical machine learning. NPJ Digit. Med. 6, 55 (2023).
Google Scholar
Gichoya, J. W. et al. AI pitfalls and what not to do: mitigating bias in AI. Br. J. Radiol. 96, 20230023 (2023).
Google Scholar
Ktena, I. et al. Generative models improve fairness of medical classifiers under distribution shifts. Nat. Med. 30, 1166–1173 (2024).
Google Scholar
Xu, Z. Mitigating social bias in large language models: a multi-objective approach within a multi-agent framework. In Proc. 39th Conf. AAAI Artificial Intelligence 25587–25579 (AAAI, 2025).
Binz, M. et al. A foundation model to predict and capture human cognition. Nature 644, 1002–1009 (2025).
Google Scholar
Mendu, S. K., Yenala, H., Gulati, A., Kumar, S. & Agrawal, P. Towards safer pretraining: analyzing and filtering harmful content in webscale datasets for responsible LLMs. Preprint at https://doi.org/10.48550/arXiv.2505.02009 (2025).
Griot, M., Hemptinne, C., Vanderdonckt, J. & Yuksel, D. Large language models lack essential metacognition for reliable medical reasoning. Nat. Commun. 16, 642 (2025). This study reveals a gap between benchmark performance and metacognitive awareness in LLMs.
Google Scholar
van Buchem, M. M. et al. Impact of a digital scribe system on clinical documentation time and quality: Usability study. JMIR AI 3, e60020 (2024).
Google Scholar
Ma, S. P. et al. Ambient artificial intelligence scribes: utilization and impact on documentation time. J. Am. Med. Inform. Assoc. 32, 381–385 (2025).
Google Scholar
You, J. G. et al. Ambient documentation technology in clinician experience of documentation burden and burnout. JAMA Netw. Open 8, e2528056 (2025).
Google Scholar
Blease, C. R., Locher, C., Gaab, J., Hägglund, M. & Mandl, K. D. Generative artificial intelligence in primary care: an online survey of UK general practitioners. BMJ Health Care Inform. 31, e101102 (2024).
Google Scholar
Eppler, M. et al. Awareness and use of ChatGPT and large language models: a prospective cross-sectional global survey in urology. Eur. Urol. 85, 146–153 (2024).
Google Scholar
Poon, E. G., Lemak, C. H., Rojas, J. C., Guptill, J. & Classen, D. Adoption of artificial intelligence in healthcare: survey of health system priorities, successes, and challenges. J. Am. Med. Inform. Assoc. 32, 1093–1100 (2025).
Google Scholar
Egli, S. B., Arpagaus, A., Amacher, S. A., Hunziker, S. & Bassetti, S. Use, knowledge and perception of large language models in clinical practice: a cross-sectional mixed-methods survey among clinicians in Switzerland. BMJ Health Care Inform. 32, e101470 (2025).
Google Scholar
Castiblanco Jimenez, I. A., Gomez Acevedo, J. S., Marcolin, F., Vezzetti, E. & Moos, S. Towards an integrated framework to measure user engagement with interactive or physical products. Int. J. Interact. Des. Manuf. 17, 45–67 (2023).
Google Scholar
Chang, C. T. et al. Red teaming ChatGPT in medicine to yield real-world insights on model behavior. NPJ Digit. Med. 8, 149 (2025).
Google Scholar
Artsi, Y. et al. Large language models in real-world clinical workflows: a systematic review of applications and implementation. Front. Digit. Health 7, 1659134 (2025).
Google Scholar
Bondi-Kelly, E. et al. Taking off with AI: lessons from aviation for healthcare. In Proc. 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization https://doi.org/10.1145/3617694.3623224 (ACM, 2023).
Kolbinger, F. R. & Kather, J. N. Adaptive validation strategies for real-world clinical artificial intelligence. Nat. Comput. Sci. 5, 980–986 (2025).
Google Scholar
Crowe, B. et al. Recommendations for clinicians, technologists, and healthcare organizations on the use of generative artificial intelligence in medicine: a position statement from the Society of General Internal Medicine. J. Gen. Intern. Med. 40, 694–702 (2025).
Google Scholar
Zhou, A. et al. AutoRedTeamer: autonomous red teaming with lifelong attack integration. NeurIPS 2025 Conference https://openreview.net/forum?id=xQH4lDLIC0 (NeurIPS, 2025).
Bastani, H. et al. Generative AI without guardrails can harm learning: evidence from high school mathematics. Proc. Natl Acad. Sci. USA 122, e2422633122 (2025).
Google Scholar
Mollick, E. R. et al. AI Agents and Education: Simulated Practice at Scale. The Wharton School Research Paper https://doi.org/10.2139/ssrn.4871171 (SSRN, 2024).
Kosmyna, N. et al. Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task. Preprint at https://doi.org/10.48550/arXiv.2506.08872 (2025).
Hoffmann, M., Boysel, S., Nagle, F., Peng, S. & Xu, K. Generative AI and the nature of work. Harvard Business School Working Paper 25-021 https://doi.org/10.2139/ssrn.5007084 (SSRN, 2024).
Wekenborg, M. K., Gilbert, S. & Kather, J. N. Examining human–AI interaction in real-world healthcare beyond the laboratory. NPJ Digit. Med. 8, 169 (2025).
Google Scholar
Berzin, T. M. & Topol, E. J. Preserving clinical skills in the age of AI assistance. Lancet 406, 1719 (2025).
Google Scholar
Budzyń, K. et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. Lancet Gastroenterol. Hepatol. 10, 896–903 (2025).
Google Scholar
Huo, W., Li, Q., Liang, B., Wang, Y. & Li, X. When healthcare professionals use AI: exploring work well-being through psychological needs satisfaction and job complexity. Behav. Sci. 15, 88 (2025).
Google Scholar
Rafailov, R. et al. Direct preference optimization: your language model is secretly a reward model. NeurIPS 2023 Conference https://neurips.cc/virtual/2023/poster/72164 (NeurIPS, 2023).
Zou, A. et al. Representation engineering: a top-down approach to AI transparency https://doi.org/10.48550/arXiv.2310.01405 (2023).
Zou, A. et al. Improving alignment and robustness with circuit breakers. Preprint at https://doi.org/10.48550/arXiv.2406.04313 (2024).
Inan, H. et al. Llama Guard: LLM-based input–output safeguard for Human–AI conversations. Preprint at https://doi.org/10.48550/arXiv.2312.06674 (2023).
Rebedea, T., Dinu, R., Sreedhar, M. N., Parisien, C. & Cohen, J. NeMo guardrails: a toolkit for controllable and safe LLM applications with programmable rails. In Proc. 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations 431–445 (Association for Computational Linguistics, 2023).
Alaa, A. et al. Position: medical large language model benchmarks should prioritize construct validity Oral. In Proc. International Conference on Machine Learning 2025 https://icml.cc/virtual/2025/oral/40130 (ICML, 2025).
Gallifant, J. & Bitterman, D. S. Humanity’s next medical exam: preparing to evaluate superhuman systems. NEJM AI 2, AIe2501008 (2025).
Google Scholar
Freyer, O. et al. Consideration of cybersecurity risks in the benefit-risk analysis of medical devices: scoping review. J. Med. Internet Res. 26, e65528 (2024).
Google Scholar
Ostermann, M. et al. Cybersecurity requirements for medical devices in the EU and US—a comparison and gap analysis of the MDCG 2019-16 and FDA premarket cybersecurity guidance. Comput. Struct. Biotechnol. J. 28, 259–266 (2025).
Google Scholar
Moberly, T. Doctors must stop using unregistered AI scribe tools, says NHS England. Brit. Med. J. 389, r1302 (2025).
Google Scholar
Freyer, O., Wiest, I. C., Kather, J. N. & Gilbert, S. A future role for health applications of large language models depends on regulators enforcing safety standards. Lancet Digit. Health 6, e662–e672 (2024).
Google Scholar
Biasin, E., Kamenjašević, E. & Ludvigsen, K. R. in Research Handbook on Health, AI and the Law (eds Solaiman, B. & Cohen, I. G.) 57–74 (Edward Elgar, 2024).
Freyer, O., Jayabalan, S., Kather, J. N. & Gilbert, S. Overcoming regulatory barriers to the implementation of AI agents in healthcare. Nat. Med. 31, 3239–3243 (2025).
Google Scholar
Mathias, R., Schonfelder, A., Welzel, C. & Gilbert, S. Letter to the editor on ‘From concept to clinic: living labs and regulatory sandboxes for health system digitalization and the integration of innovative devices into clinical workflows’. IEEE J. Transl. Eng. Health Med. 13, 214–215 (2025).
Google Scholar
Adler-Milstein, J. et al. Electronic health record adoption in US hospitals: the emergence of a digital ‘advanced use’ divide. J. Am. Med. Inform. Assoc. 24, 1142–1148 (2017).
Google Scholar
Hwang, Y.-M., Ng, M. Y., Pillai, M., Sahai, M. P. & Hernandez-Boussard, T. The landscape of AI implementation in US hospitals. Nat. Health 1, 99–112 (2026).
Google Scholar
Shanahan, M. Talking about large language models. Commun. ACM 67, 68–79 (2024).
Google Scholar
Ferber, D. et al. Towards autonomous medical artificial intelligence agents. Nature 655, 1282–1291 (2026).
Keep following us for the latest insights.

















