2026-04-11 2026, Volume 3 Issue 2

  • Select all
  • research-article
    Dustin A. Silverman, John S. Howard, Priscilla F. A. Pichardo, Yash J. Patil, Mekibib Altaye, Chad A. Zender, Alice L. Tang
    2026, 3(2): 025240053. https://doi.org/10.36922/AIH025240053

    Artificial intelligence models such as chat generative pre-trained transformer (ChatGPT) are being increasingly used to inform treatment-related decisions. Among otolaryngology subspecialties, there is a paucity of literature examining the role of ChatGPT within head and neck surgical oncology. The utility of ChatGPT in addressing questions related to surgically relevant anatomy and lymphadenectomy procedures remains poorly understood. The primary pilot study objective was to determine the reliability of ChatGPT in answering neck dissection-related inquiries compared to expert head and neck surgical oncologists. Five neck dissection-related questions were presented to ChatGPT v3.5. Three fellowship-trained head and neck surgeons compared AI-generated responses to those of an expert head and neck surgeon. Raters, blinded to the author’s identity, evaluated the responses given based on a Likert scale ranging from 1 (strongly disagree) to 5 (strongly agree). The median level of agreement between raters for the ChatGPT responses was 1.0 (interquartile range [IQR]: 1.0, 2.5; minimum = 1 and maximum = 4), while the median level of agreement between raters for the surgeon responses was 5.0 (IQR: 5.0, 5.0; minimum = 5 and maximum = 5). The Mann-Whitney U test yielded a significance level of p=0.007 when comparing the level of agreement between ChatGPT and surgeon responses. Raters showed minimal consistency when evaluating ChatGPT responses (intraclass correlation coefficient = 0.05; 95% confidence interval: 0.0-0.88), in contrast to perfect agreement observed for the surgeon responses. In summary, ChatGPT is a promising tool in the acquisition of surgical knowledge. For neck dissection-related inquiries, a discrepancy between the reliability of ChatGPT-generated responses and surgeon expertise exists. Further refinement in AI models is needed to strengthen the utility of ChatGPT in head and neck oncologic surgery.

  • research-article
    Malik Sallam, Johan Snygg, Mazin Aljabiri, Edward Cody, Reem Allateef, Chadia Beaini, Mohammed Sallam
    2026, 3(2): 025340067. https://doi.org/10.36922/AIH025340067

    The evolution of artificial intelligence (AI) raises questions about the future roles of physicians. This study aimed to propose an exploratory foresight model for stratifying risk across medical specialties, using board-defined competencies and generative AI (genAI) evaluation as the assessment tool. We developed a heuristic framework, the Machine automat-ability, Diagnostic Ambiguity, Legal/ethical complexity, Interpersonal intensity, Knowledge codifiability, Evidence in data, Difficulty of procedures (MALIKED) score, to capture dimensions of displacement vulnerability for 27 board-recognized specialties. To minimize individual bias, ratings were generated by three genAI models (ChatGPT, DeepSeek, and Gemini). Data-centric fields—Clinical Pathology (30.3/35), Anatomic/Clinical Pathology (29.3/35), and both Anatomic Pathology and Radiology (28.0/35 each)—clustered in the highest-vulnerability tier. In contrast, procedurally intensive or patient-interaction-heavy specialties—including Psychiatry (11.0/35), Neurosurgery (11.7/35), Obstetrics/Gynecology (13.0/35), General Surgery (13.0/35), Pediatrics (14.3/35), Emergency Medicine (14.3/35), and Family Medicine (14.3/35)—formed the lowest-vulnerability tier. Between these extremes, mixed-mode specialties, such as Internal Medicine (17.0/35) and Neurology (17.0/35), along with Ophthalmology (19.3/35) and Anesthesiology (21.3/35), occupied an intermediate zone. Displacement risk was driven by knowledge codifiability and data-centricity, while procedural complexity and interpersonal interaction intensity exerted protective effects. This exploratory foresight framework suggests that the risk of displacement by advanced or potentially superintelligent AI is unevenly distributed across medical specialties. While data-driven fields appear most exposed, no specialty is categorically insulated, as multimodal AI and robotics continue to evolve. The MALIKED framework is not predictive but intended as a structured lens for debate, education, and workforce planning regarding the long-term implications of AI in medicine.

  • research-article
    Aref Zribi, Omar Ayaad
    2026, 3(2): 025350070. https://doi.org/10.36922/AIH025350070

    Artificial intelligence (AI) holds huge potential in improving diagnosis and streamlining workflows in health care. However, several challenges remain, hampering the widespread adoption in clinical settings, such as for assessing data quality, bias, interoperability, and privacy, as well as for use in regulation and clinician training. Potent data channels are vital for assuring the exactness and trustworthiness of diagnostic performance. They boost the transmission of high-quality information, which is essential for expert annotations. Interoperable electronic health record integration and federated or privacy-enhancing training approaches allow real-time analytics while guarding patient data. Regulatory indecision and the comprehensive and continuous supervision of the process require transparent, explainable AI and shared accountability among developers, doctors, and institutions. In addition, prospective clinical validation, physician education, and governance are paramount to building trust and guaranteeing safe AI deployment in health care. This review outlines the difficulties faced when integrating these technological advancements into everyday clinical practice.

  • research-article
    Leor Franco, Milan Toma
    2026, 3(2): 025360073. https://doi.org/10.36922/AIH025360073

    Alzheimer’s disease (AD) is the most prevalent cause of dementia worldwide, yet early and accurate diagnosis remains a significant clinical challenge. This study systematically evaluates machine learning (ML) models for AD classification using a shared magnetic resonance imaging (MRI) dataset, focusing on both binary (e.g., healthy vs. demented) and multiclass (four-stage dementia) tasks. MRI images were preprocessed and deep features were extracted using a pretrained GoogLeNet architecture. Classification was performed using support vector machines for binary tasks and error-correcting output codes (ECOC) for multiclass tasks. Model performance was assessed using both hold-out and k-fold (5- and 10-fold) cross-validation (CV) strategies to ensure robust evaluation. Results indicate that CV yields substantially higher and more reliable accuracy than the hold-out method, with binary classification achieving up to 84% accuracy (10-fold CV) and multiclass classification reaching 87% (5-fold CV). The model demonstrated high specificity and precision, particularly for moderate-stage AD, but lower and less stable performance for early disease (very mild) cases. Learning curve analysis confirmed improved generalization with increased training data and minimal overfitting in well-validated models. However, our findings also reveal that high-performance metrics alone are insufficient to judge clinical utility. For example, a hold-out model distinguishing healthy from moderate Alzheimer’s achieved a seemingly impressive 99% test accuracy, but learning curves revealed severe overfitting and likely data leakage, indicating that such results would not generalize to new patients. In contrast, the healthy vs. mild Alzheimer’s task, with a more modest ~70% test accuracy, demonstrated well-behaved learning dynamics and genuine generalization. These results highlight that high reported accuracy is only clinically meaningful when supported by healthy training dynamics; otherwise, models risk underperforming in real-world clinical settings regardless of their reported metrics. We advocate for rigorous validation and learning curve analysis as prerequisites for any clinically actionable ML tool. These findings underscore the importance of transparent per-class performance reporting and robust validation to ensure that ML models can truly support early detection and staging of AD in clinical practice.

  • research-article
    Wolfgang Lederer, Mary Heaney Margreiter
    2026, 3(2): 025370075. https://doi.org/10.36922/AIH025370075

    The development of robots incorporating artificial general intelligence has achieved rapid advancements, with tangible changes being observed on the labor market, such as in the healthcare industry. With the expansion of research targeted at robot development, it is anticipated that robots demonstrating artificial intelligence-associated cognitive properties that reflect individuality and distinctiveness will soon be entering the care system. These nursing robots take over some of the work traditionally done by humans in the elderly and nursing care system, encountering very vulnerable elderly individuals. Therefore, there is a need to reconsider and redefine the aspects of social values, data integrity, and the dignity of working with and on people when the nursing robots are taken into account. Their introduction presents opportunities for workforce redistribution and potential benefits for employees and underscores the urgent need for clear legal regulations. In this context, employing our knowledge in social science to foster constructive cooperation between humans and intelligent machines has become more important than ever.

  • research-article
    Xue Wu, Shengting Cao, Jiaqi Gong
    2026, 3(2): 025380076. https://doi.org/10.36922/AIH025380076

    Timely and accurate monitoring of opioid overdose risks is critical for public health, particularly in underserved rural regions. Traditional surveillance systems often lack the spatial and temporal resolution needed to support proactive interventions. While socioeconomic indicators, such as the social vulnerability index and housing value, show moderate correlations with opioid-related outcomes, existing methods rarely incorporate high-resolution environmental data. Previous research relies largely on static census data and coarse geographic indicators, limiting its ability to detect localized risk patterns. Moreover, the connection between built-environment features and opioid overdose remains underexplored—especially in rural areas like Alabama’s Black Belt. To address this gap, we propose a multiscale spatio-temporal framework that integrates satellite imagery and machine learning to monitor opioid-related emergency room (ER) visit rates. We collected 201,967 housing images from Black Belt counties and classified them using computer vision models, including ResNet and external attention transformers. To overcome limitations in labeled data, we developed four unsupervised pipelines combining k-means clustering with autoencoders, masked autoencoders, VGG16, and household-image ratios. Our results show that unsupervised embeddings outperform supervised classification in capturing signals associated with ER visits. Descriptive features, such as roof type, road layout, and environmental openness, significantly inform predictions. Although Black Belt counties report lower absolute ER visit rates, they show faster year-over-year growth. Our study demonstrates the potential of combining satellite imagery with multimodal artificial intelligence to improve rural health surveillance and supports the development of scalable, interpretable monitoring tools for early intervention and policy planning.

  • research-article
    Sameh Shamroukh
    2026, 3(2): 025390078. https://doi.org/10.36922/AIH025390078

    This study examines the performance of Medicare providers, with a focus on patient backgrounds, number of chronic illnesses, and efficacy of drug treatments as compared to traditional medical services. Over 1.2 million records from the Centers for Medicare & Medicaid Services and different methods, such as random forest regression and K-means clustering, were utilized to identify top-performing providers, categorize patients, and evaluate the results of treatments. The results reveal that while spending more on Medicare and offering more services can sometimes improve patient outcomes, these improvements are not always steady. This inconsistency points out some ongoing problems in the system, primarily affecting older adults and those in underserved communities, who often struggle with worse health and limited access to care. In addition, the study found that the effectiveness and cost of different treatment methods can vary widely. Drug treatments and direct medical services had varying impacts on resource use and health benefits. Combining large-scale public data with advanced analytic techniques, this research provides a reference for policymakers and healthcare organizations and offers insights into designing targeted interventions, with the ultimate aim to preserve fairness and sustainability in the U.S. healthcare system.

  • research-article
    Rashid Nawaz, Saba Kainat, Muhammad Shoaib
    2026, 3(2): 025400081. https://doi.org/10.36922/AIH025400081

    Diabetes mellitus is a complex metabolic disorder with diverse complications, which motivates the use of advanced computational intelligence methods for modeling its nonlinear dynamics. The goal of the current research is to obtain numerical solutions to a diabetes mellitus model by employing log-sigmoid neural networks together with both local and global search strategies. For this model, the genetic algorithm (GA) serves as a global search method, while the sequential quadratic programming approach is employed as a local optimizer. In this study, the problem is addressed using a hybrid solution strategy, introducing a new aspect to the existing research on diabetes mellitus modeling. The proposed model comprises five groups: susceptible, exposed, infected without treatment, infected with treatment, and recovered individuals. To compare the reliability, precision, and consistency of the proposed technique, the log-sigmoid neural network optimized through GA and sequential quadratic programming is compared with the Adam numerical solver. An absolute error within very small ranges is achieved, demonstrating the solver’s proficiency. In addition, 100 independent trials and a network of 5 neurons were employed to test the validity of the proposed stochastic approach, together with statistical measures including root mean squared error, Theil’s inequality coefficients, and mean absolute deviation.

  • research-article
    Gracy Singh, Nidhi Verma, Sonali Bhatt, Saurabh Mukherjee
    2026, 3(2): 025400087. https://doi.org/10.36922/AIH025400087

    Alzheimer’s disease (AD), a progressive neurodegenerative disease, is a major global public health problem. Early and accurate diagnosis is crucial for timely intervention, especially with the prevalence of the condition expected to triple by 2050. Traditional methods have been enhanced by artificial intelligence (AI)-driven techniques, particularly Convolutional Neural Networks (CNNs) and Support Vector Machines (SVMs). These approaches enhance the early detection and classification of diseases by evaluating complex neuroimaging data. Using Matrix Laboratory, we developed hybrid models integrating CNNs and SVMs to detect AD, focusing on feature extraction, predictive accuracy, and model interpretability for clinical use. While Internet of Things--based wearable devices are reviewed for their potential in large-scale data processing and real-time monitoring, our empirical work emphasizes practical AI solutions. In conclusion, integration of multimodal neuroimaging data and advanced feature selection techniques holds the potential to enhance diagnostic precision of AD.

  • research-article
    Seble Frehywot, Yianna Vovides
    2026, 3(2): 025420090. https://doi.org/10.36922/AIH025420090

    Technologies invented in the five industrial revolutions (IRs) have profoundly transformed Global Health Workforce Education (GHWFE), reshaping teaching methodologies, faculty approaches, and student learning. This article first reflects on the influence of technology on GHWFE from the first to the fourth IRs. Then, it focuses on the present, Fifth IR (5IR), the era of human-artificial intelligence (AI) centric collaboration, and the fact that the global health workforce educators are not trained for being nimble to utilize AI and its related technologies in 5IR. The manuscript envisions new directions for the future with the goal of establishing nimbler educators that acknowledge the benefits of interdisciplinary dialogue as a means of deepening AI knowledge and community. The article expands the AI algorithmic literacy framework and proposes a Human-AI Centric Workshop Series that moves global health workforce educators from awareness to knowledge, to applied innovation, and toward expertise in 5IR.

  • research-article
    Carly Hudson, Adrian Goldsworthy, Thuy Linh Phan, Anu Joy, Oystein Tronstad, Marcus Randall
    2026, 3(2): 025450097. https://doi.org/10.36922/AIH025450097

    Machine learning (ML) and artificial intelligence are increasingly ubiquitous in healthcare data analytics. To date, however, ML has been largely restricted to the analysis of structured data. While natural language processing (NLP) is gaining prominence in healthcare, substantial challenges remain in the generation and analysis of unstructured data. Emergency departments, which are increasingly under-resourced and overburdened, may benefit from the implementation of ML techniques that incorporate NLP to support clinical decision-making and improve patient care. Historically, regional and cultural variations have posed significant challenges to the widespread application of ML algorithms beyond their original training datasets. The rapidly increasing use of NLP within clinical note-taking applications provides avenues to assist in standardizing unstructured data and extracting meaningful insights to improve generalization and clinical translation.

  • research-article
    Tara Mansour
    2026, 3(2): 025450099. https://doi.org/10.36922/AIH025450099

    Effective communication is a core competency in health professions’ education, yet opportunities for deliberate, scalable practice remain limited. This perspective explores how voice-enabled generative artificial intelligence (AI), such as ChatGPT Voice, can enhance communication training by enabling students to rehearse clinical dialogues, receive immediate feedback, and engage in reflection within psychologically safe environments. Drawing on recent evidence, the article illustrates how AI-mediated conversation supports skill development in empathy, adaptability, and professional identity formation across patient, family, and interprofessional contexts. Practical examples demonstrate how educators can integrate voice-based AI into coursework and clinical preparation to complement, not replace, human mentorship. Ethical and pedagogical considerations, including privacy, bias, and authenticity, are also discussed. Used thoughtfully, voice-enabled AI can extend the reach of communication education, preparing students to engage confidently and compassionately in the complex interpersonal dynamics of healthcare practice.

  • research-article
    Nuno Soares Domingues
    2026, 3(2): 025470102. https://doi.org/10.36922/AIH025470102

    Clinical decision support systems (CDSS) are increasingly reliant on purely data-driven machine learning models, leading to significant challenges in clinical adoption due to their “black-box” nature, high risk of algorithmic bias, and inability to enforce hard safety constraints. This lack of transparency and clinical alignment poses major challenges for regulatory compliance and professional trust. This study proposes a novel hybrid artificial intelligence (AI) meta-model for CDSS design, which formally translates established engineering decision support paradigms into the clinical domain to create systems that are inherently safer and more explainable. The framework rigorously integrates: (i) Web Ontology Language 2 for formalizing medical concepts; (ii) a semantic web rule language rule base to serve as “clinical guardrails” for enforcing evidence-based guidelines and safety constraints; and (iii) a modular inference policy that intelligently matches decision problems with specific data-driven or probabilistic methods. A prototype was implemented, combining this explicit knowledge layer with modular inference engines and template-based explanations. Evaluation across two large-scale clinical tasks (oncology using the Surveillance, Epidemiology, and End Results database and intensive care using the Medical Information Mart for Intensive Care-IV database) demonstrated superior performance over data-driven baselines. Specifically, the hybrid system achieved a 78% reduction in guideline-violation errors (reducing the contraindication rate from 18% to 6%) while maintaining high predictive accuracy (area under the receiver operating characteristic curve of 0.84 and 0.87, respectively). Furthermore, a clinician usability study confirmed that the transparent, knowledge-driven explanations resulted in significantly higher decision clarity (Likert rating of 6.2/7) and reduced cognitive load. The findings validate that this hybrid AI architecture represents a robust and transferable approach to designing clinically aligned, trustworthy, and explainable CDSS, directly addressing critical requirements for the responsible deployment of AI in healthcare.

  • research-article
    Sheikh Usman Iqbal, Dmitriy Podolskiy, Mona G. Flores, Victoria L. Chiou, Jennifer Sheng, Roberto Araujo, Naveed Afzal, David A. Hall, Ingrid Vasiliu-Feltes, Dipu Patel, Manolis Kellis
    2026, 3(2): 025470103. https://doi.org/10.36922/AIH025470103

    Artificial intelligence (AI) heralds a transformative shift in drug development, with speed, precision, and predictive power as its core features. Advances in systems-level biology platforms, coupled with substantial investments in generative AI-centric pharma integration, have fostered healthy optimism among stakeholders about identifying new cures through renewed approaches and improved productivity. However, navigating epistemological, ethical, patient safety, and ontological dimensions within research and development (R&D) presents challenges that AI must address to enhance its mainstream adoption and practical utility. Here, multidisciplinary experts discuss key applications of AI across the full continuum of drug development, examine the challenges encountered, and propose solution frameworks. Drug development remains fraught with unknown biology, patient heterogeneity, and perplexing therapeutic risks. Stringent regulatory and compliance guidelines further necessitate that conventional pharma processes, practices, and strategies remain paramount in R&D execution, while guiding the integration of AI in a “value-for-effort,” evidence-based, yet Promethean fashion.

  • research-article
    Jia Weng, Jiacheng Weng, Antony Kam, Shining Loo, Lina Zhou, Rencai Fan, Runwei Guan, Shicheng Li, Kai Chen
    2026, 3(2): 025470105. https://doi.org/10.36922/AIH025470105

    Gamma delta (γδ) T cells exert a pivotal anti-tumor role in the breast cancer (BC) tumor microenvironment, highlighting the importance of investigating their prognostic value for improved patient stratification. We analyzed single-cell ribonucleic acid sequencing data to cluster immune cells and identify marker genes. Prognostic features were selected using Least Absolute Shrinkage and Selection Operator regression in the Cancer Genome Atlas-Breast Invasive Carcinoma cohort and validated across five external Gene Expression Omnibus cohorts. The expression of these prognostic genes was further validated by immunohistochemistry (IHC) in an in-house cohort of BC patients. These selected features were used to construct machine learning models, with the best-performing model undergoing hyperparameter tuning to optimize its performance. γδ T cells were identified as one of the major immune cell populations in BC. Twelve γδ T cell-associated genes were selected based on their prognostic significance in external validation cohorts. The random forest (RF) model achieved the highest accuracy (0.835) after hyperparameter tuning. In external validation, the final model showed the highest performance with area under the curve/accuracy values of 0.81/0.849. IHC analysis confirmed dysregulated expression of key signature proteins in BC tissues. In conclusion, we developed an efficient prognostic RF model based on a 12-gene signature from γδ T cells, which may serve as a clinically valuable tool for risk stratification in BC patients.