September 3, 2026

AI Research Webinar

Nikolai Schnittke, MD, PhD, FAEMUS
Ryan Bellinger, MD
Nicole Duggan, MD
Shawn J Sethi, DO, FACEP

Overview

As applications of artificial intelligence (AI) in medicine expand, it becomes essential for clinicians and healthcare stakeholders to understand how AI models used to guide healthcare decisions are created, designed, and tested. Computer vision is a subset of AI algorithms that can be leveraged for image analysis and applied to point-of-care ultrasound (POCUS), with the potential to improve image acquisition, interpretation, and support for administrative burdens. Understanding of how computer vision algorithms are developed and deployed helps clinicians better evaluate where AI tools are useful, how reliable they are, and where their limitations lie in clinical practice. This article is a summary of a webinar organized by the EUS POCUS Research Subcommittee. Our expert panel of investigators focused on AI research and discussed the fundamental concepts of research and development of POCUS AI algorithms. They focused on the challenges associated with developing datasets, clinically validating algorithms, and limitations of regulatory approval.

AI fundamentally consists of computerized or automated performance of tasks like pattern recognition or prediction typically performed by humans.1 AI systems can assist clinicians in obtaining optimal ultrasound images by guiding probe positioning, anatomical labeling, and identifying correct views. This is particularly useful in POCUS, where operator experience varies widely. Such tools may reduce workflow burden and improve consistency in measurements across users. AI models are also being developed to identify and classify pathology across many ultrasound domains including thyroid nodules, breast lesions, liver fibrosis, vascular disease, and myocardial infarction. In the emergency department, applications including calculation of ejection fraction, identification of B-lines or neurovascular anatomy, and automated image archiving show clinical promise.2 AI-assisted systems in low-resource or rural settings, may enable broader use, and automated guidance and interpretation may allow non-expert users to perform basic diagnostic scans with acceptable accuracy.3

We would like to thank Dr. Srikar Adhikari, Dr. Cristiana Baloescu, Dr. Michael Blaivas, Dr. Nicole Duggan, Dr. Joseph Pare, and Dr. Nikolai Schnittke for their participation and expertise, as well as Dr. Lynn Roppolo for her assistance in coordinating this webinar.

Dataset Considerations

At their most basic level, the majority of AI models are predictive - they accept an input (such as an image or video clip) and produce an output (such as a measurement or structure identification). Their performance depends heavily on the quantity and quality of the training data, both of which are difficult to achieve. Most computer vision models rely on the creation of image datasets that are labeled by ultrasound experts to train a neural network to learn relevant features. These networks are then tested to predict those features on images not previously seen by the network, and the results of the prediction are compared to the expert labels.1

The quality of models depends on the quality of a massive amount of training clips or images and the quality of the labeling scheme coded to represent interpretations. Therefore, prospective dataset acquisition and labeling is resource intensive. Acquiring retrospective, externally sourced data allows for a much larger training sample size but standard imaging conventions and protocols (depth, gain, views, etc.) may vary. There is an inherent trade-off between highly controlled, pristine, prospectively collected datasets and developing a model that is large and varied enough to be generalizable to the array of real-world clinical environments. While a carefully curated dataset may produce strong internal performance, POCUS images are often imperfect in practice. Additionally, different diagnostic modalities have variable “reference standard” comparators, which themselves may have diagnostic imperfections. These challenges highlight the importance of a careful approach to defining the image labeling scheme and maintaining consistency of labeling throughout the dataset. Therefore, developers must balance internal consistency with real-world variability to ensure the tool functions effectively in diverse clinical settings.

How can we address these limitations and appraise whether a model was developed using best practices to address these challenges? Our panelists discussed the following concepts and strategies:

  1. Local and national efforts to harmonize datasets: choosing a labeling strategy, standardizing imaging protocols and quality, etc. will improve pooled data quality.
  2. Addressing class imbalance: datasets typically skew toward a larger number of negative studies. This poses a challenge in algorithm development and can affect model training and performance. While techniques such as data augmentation, synthetic image generation, and data harmonization can partially address imbalance, they cannot compensate for fundamentally inadequate datasets.
  3. Addressing Spectrum bias: if a model is trained only on very obvious findings, it may fail to recognize subtle or intermediate cases.
  4. Stepwise escalation: this strategy can help address class imbalance and spectrum bias.  Similar to medical education, this approach begins model training with clear and obvious examples of pathology before progressing to training of more subtle findings.

Model Validation

Our panelists emphasized the importance of rigorous model validation, noting that testing algorithms across diverse and independent datasets may be even more critical than the initial creation of training datasets. In contrast to traditional clinical decision tools used in the emergency department (HEART score, PERC, etc.), AI algorithms require more extensive prospective, external, and ongoing validation to ensure safe clinical implementation. If we are to implement POCUS-based AI models to populations that differ from those immediately used for initial training and testing (including variations in age, gender, and other physiologic features) then we must establish robust external validity.4 As patient demographics and clinical characteristics evolve, maintaining model performance will require sustained institutional collaboration and systematic monitoring.  In addition, advances in POCUS hardware and changes in image acquisition protocols will alter the data encountered by AI systems, further emphasizing the need for continual reassessment and recalibration.

Regulatory Considerations

The final portion of the panel discussion focused on federal regulation and its role in the availability of AI-based POCUS tools.

The regulatory environment for AI in POCUS largely falls within the broader regulatory framework called software as a medical device (SaMD), which is applied to AI-enabled medical devices. In the United States, this framework is overseen by the U.S. Food and Drug Administration (FDA) through its medical device regulatory pathways.5 While there are multiple federal pathways for approval in clinical use, most AI applications have entered the market through the 510(k) pathway, demonstrating safety, efficacy, and substantial equivalence to another legally marketed device (referred to as a “predicate”). In contrast, for new devices or applications without precedent, the “de novo” approval pathway must be utilized for federal approval, generally with higher scrutiny and longer duration of review. The FDA has highlighted many of the previously discussed concerns such as dataset composition, validity, generalizability, and bias.

Regulatory bodies such as the FDA require a validation study (also referred to as a “pivotal study”) where the AI prediction is tested against image interpretation by an expert panel. The methods surrounding the pivotal study are of greater importance for the FDA than the original dataset. FDA approval is dependent on the “intended use” of the application, which is based on the methods used in the pivotal study. This means that tools are only approved for use in patients and by clinicians specified in the pivotal study originally supporting the application.6  Achieving FDA clearance requires substantial effort, however POCUS is used by clinicians with a broad spectrum of experience, and in patients with a broad range of clinical and demographic characteristics. Federal clearance for clinical application does not guarantee a tool's appropriateness for clinical or educational use.

Conclusion

Computer Vision is a subset of AI relating to image analysis that can enhance POCUS quality and clinical utility. The value of AI in POCUS depends on the quality of methods used to develop and validate algorithms. Researchers and clinicians who understand the technical and clinical challenges of POCUS are essential to ensure that validation and implementation of AI goes beyond initial regulatory approval and addresses real clinical challenges.

To quote Dr. Adhikari’s concluding remark from the webinar: “Just because we can, doesn’t mean we should”. It is up to clinicians to decide what we should and should not build to provide benefit to our learners, clinicians, and patients.

References

  1. Baloescu C, Liu R. Making Sense of the AI Alphabet Soup in POCUS. Acep.org. Published 2026. Accessed March 15, 2026. https://www.acep.org/emultrasound/newsroom/february-2026/making-sense-of-the-ai-alphabet-soup-in-pocus
  2. Ultrasound Guidelines: Emergency, Point-of-Care and Clinical Ultrasound Guidelines in Medicine. Ann Emerg Med. 2017;69(5):e27-e54. doi:https://doi.org/10.1016/j.annemergmed.2016.08.457
  3. Kayarian F, Patel D, O’Brien JR, Schraft EK, Gottlieb M. Artificial intelligence and point-of-care ultrasound: Benefits, limitations, and implications for the future. Am J Emerg Med. 2024;80:119-122. doi:https://doi.org/10.1016/j.ajem.2024.03.023
  4. Puticiu M, Cimpoesu D, Pop F, et al. AI-Enhanced POCUS in Emergency Care. Diagnostics. 2026;16(2):353. doi:https://doi.org/10.3390/diagnostics16020353
  5. Artificial Intelligence in Software. U.S. Food and Drug Administration. Published 2025. https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-software-medical-device
  6. Yasudda N. Final Document Good Machine Learning Practice for Medical Device Development: Guiding Principles AUTHORING GROUP Artificial Intelligence/Machine Learning-Enabled Working Group.; 2025. https://www.imdrf.org/sites/default/files/2025-02/IMDRF_AIML%20WG_GMLP_N88%20Final.pdf
[ Feedback → ]