top of page

Can Artificial Intelligence Pass Neonatal Resuscitation Exams? 🤖👶

  • 2 days ago
  • 3 min read



Artificial intelligence (AI) is rapidly transforming healthcare, from supporting clinical decision-making to assisting with medical education. One of the most exciting developments has been the emergence of large language models (LLMs)such as ChatGPT and DeepSeek. But how well do these tools actually perform when tested on neonatal resuscitation knowledge?

Our latest publication, "Performance of Large Language Models in Neonatal Resuscitation Assessments versus Healthcare Providers: An Exploratory Study," explores this important question. The findings suggest that AI has tremendous potential as an educational tool—but also highlight why expert oversight remains essential.


Why is this important?

Neonatal resuscitation is a high-acuity, low-frequency event. Healthcare professionals must maintain knowledge and decision-making skills despite rarely encountering critically ill newborns in clinical practice.

Simulation training and structured education are highly effective, but they require experienced instructors, considerable preparation, and significant resources. As AI continues to evolve, an important question emerges:

Could large language models help support neonatal resuscitation education?

Before AI can become part of training programs, we first need to understand how accurately these systems perform on neonatal resuscitation assessments.


How did we study this?

This multicentre collaboration between investigators in Canada and China evaluated two leading large language models:

  • ChatGPT-5

  • DeepSeek-R1

The models completed three different sources of neonatal resuscitation examinations:

  • Neonatal resuscitation workshop examinations (Chinese)

  • Questions from the NRP® 8th Edition textbook

  • Kahoot quizzes used during neonatal resuscitation training in Canada

Their performance was compared with historical results from more than 500 healthcare professionals who had completed neonatal resuscitation training. To assess consistency, each AI model answered every question on three separate occasions over several weeks.


What did we find?

The results were encouraging.

Both ChatGPT and DeepSeek performed at a level comparable to healthcare professionals on selected written neonatal resuscitation examinations.

Some key findings included:

  • ChatGPT and DeepSeek achieved overall accuracy similar to healthcare providers.

  • ChatGPT performed particularly well on scenario-based questions requiring clinical reasoning.

  • Both models showed excellent consistency when the same questions were repeated over time.

  • Performance was strongest on multiple-choice questions and lower on short-answer questions requiring more detailed responses.

These findings suggest that modern AI systems can successfully interpret neonatal resuscitation concepts and apply them to written clinical scenarios.


But AI is not ready to replace educators

While the overall performance was impressive, the study also identified important limitations.

Both models occasionally:

  • made medication calculation errors,

  • produced different answers depending on whether questions were asked in English or Chinese,

  • and struggled with some short-answer questions requiring precise responses.

These findings reinforce an important message:

AI should support neonatal education—not replace experienced instructors.

Any AI-generated educational material should continue to be reviewed by Neonatal Resuscitation Program (NRP®) instructors and clinical experts before being used in teaching.


What does this mean for the future?

This publication is part of our growing research program exploring how artificial intelligence can improve neonatal education.

Earlier this year, we demonstrated that ChatGPT can generate high-quality neonatal resuscitation simulation scenarios when reviewed by expert instructors. This new study extends that work by showing that AI can also perform at a level comparable to healthcare professionals on written neonatal resuscitation assessments.

Together, these studies suggest that AI could become a valuable partner in neonatal education by helping to:

  • generate educational content,

  • create simulation scenarios,

  • develop self-assessment tools,

  • provide structured feedback,

  • and support learners between formal training sessions.

However, neonatal resuscitation remains a team-based clinical skill that depends on communication, technical performance, judgement, and experience—qualities that cannot be fully assessed through written examinations alone.


Looking ahead

Artificial intelligence is advancing at an extraordinary pace, and its role in healthcare education will continue to expand.

Our study demonstrates that large language models have the potential to become valuable educational tools in neonatal resuscitation. The next challenge is determining how best to integrate AI into training programs while ensuring patient safety, educational quality, and adherence to established clinical guidelines.

At Research4Babies, we believe the future is not about replacing clinicians with AI—it is about using AI to help clinicians learn better, teach better, and ultimately improve outcomes for newborn infants around the world.

Because every newborn deserves the very best start to life. 🌍👶


Reference

Xu C, Chen Y, Skelding S, Wang D, Zhang Q, Schmölzer GM, Cheung P-Y, on behalf of the TRAINinG Interest Group. Performance of large language models in neonatal resuscitation assessments versus healthcare providers: an exploratory study. Frontiers in Artificial Intelligence. 2026.


Research4Babies

Advancing neonatal research. Sharing knowledge. Improving newborn lives.





Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
Featured Posts
Recent Posts
Archive
Search By Tags
Follow Us
  • Facebook Basic Square
  • Twitter Basic Square
  • Google+ Basic Square

© 2014-2026 by CSAR

  • Spotify
  • Twitter Social Icon
  • LinkedIn Social Icon
  • YouTube Social  Icon
bottom of page