The AJH Informatics Review

A weekly digest of new research on EHRs, clinical AI, interoperability & health IT policy

AI evaluation & deployment

Every digest paper in this category, newest first.

001
Development to implementation to evaluation: step-by-step approach to embedding a clinical trial of a machine learning algorithm in the electronic health record

How can a machine learning algorithm be embedded in the electronic health record to run a pragmatic randomized trial? This implementation report describes the Precision Resuscitation with Crystalloids in Sepsis (PRECISE) trial, a multihospital RCT built entirely with standard Epic tools through four components: automated inclusion criteria, real-time sepsis subphenotyping, randomization, and a medication alternative alert prompting clinicians toward the fluid type thought to benefit the identified subgroup. PRECISE launched across 6 Emory Healthcare hospitals in June 2024, covering 6 emergency departments and 17 ICUs with more than 300 ICU beds. The abstract reports implementation details only, with no trial outcomes or effect sizes.

002
Evaluation of a Large Language Model Discharge Summary Hospital Course Tool: Improved Quality but Longer Documentation Time

Does an LLM-generated hospital course draft reduce discharge summary documentation time? This observational pre-post study compared a pre-AI period (March 16-November 19, 2025) with a post-AI period (November 20, 2025-February 11, 2026) after a GPT-4.1 tool was embedded in mandatory discharge summary templates, covering 8,298 hospitalized adults on hospital medicine services. Edit time did not differ between periods (6.60 vs 6.28 minutes, p=0.11), but within the post-AI period, tool use was associated with longer editing (9.20 vs 4.93 minutes) and a 32% increase in adjusted analysis (95% CI 24-41%). Faculty review of 27 encounters found higher quality and lower harm but less concise summaries; 31 of 36 surveyed clinicians (86.1%) felt the tool improved efficiency.

003
Operatıonalızıng ethıcal AI for safe clınıcal ıntegratıon: A qualıtatıve ıntervıew study

How are ethical principles for AI safety actually operationalised once clinical AI reaches practice? This qualitative study (QuAS-AI) used semi-structured interviews with 16 experts involved in clinical AI implementation "\u2014 academics, industry professionals, and practitioners \u2014 recruited by purposive and snowball sampling, with independent coding, peer debriefing, and member checking. Three themes emerged: performance, risk, and bias as dynamic properties needing continuous monitoring; transparency and explainability as complementary supports for clinical interpretation rather than technical disclosure; and accountability as multi-actor governance spanning traceability, logging, data governance, and intervention capacity. The authors frame safety as socio-technical governance and suggest a High-Reliability Organisation approach. The abstract reports no effect sizes.

004
Large language models for patient-facing pathology report interpretation: A scoping review

How have large language models been used and evaluated for translating pathology reports into patient-facing language? This scoping review followed JBI methodology and PRISMA-ScR, searching six databases for empirical studies of LLM-generated interpretation, rewriting, question answering, or summarization of pathology reports (January 2018 to August 2026). Nineteen studies were included; GPT-family models appeared in 17, and report-level transformation was the most common task (12/19, 63.2%). Fidelity was assessed in all 19 studies, safety in 10 (52.6%), readability in 9 (47.4%), and comprehension and usability in 6 each (31.6%). Only five involved patients or other non-clinicians, and none evaluated performance by health-literacy level, across languages, or prospectively within clinical workflows.

005
Stakeholder Perspectives on the Integration of AI in Diabetes Care: Systematic Review of Qualitative Studies

What do patients, caregivers, and clinicians actually expect from AI in diabetes care? This systematic review of qualitative studies searched five databases (MEDLINE, Web of Science, Scopus, CINAHL, PsycINFO) through February 2026, including 14 studies published 2023-2025 with at least 738 participants across 9 countries, covering large language models, AI-enabled apps and wearables, glucose prediction, and decision support. Thematic synthesis produced 4 analytical themes and 13 subthemes; 11 findings were rated high confidence and 2 moderate under GRADE-CERQual. Stakeholders saw value for prevention, education, and self-management but raised accuracy, bias, privacy, accountability, workload, and autonomy concerns. Much evidence rested on prototypes or hypothetical systems.

006
AI for Health Care Quality and Patient Safety: Scoping Review of Diagnostic, Predictive, and Decision Support Applications

Why does strong AI task performance so rarely translate into sustained clinical benefit? This scoping review followed Joanna Briggs Institute methodology with PRISMA-ScR and PRISMA-S reporting, searching five databases (MEDLINE, Scopus, Web of Science, IEEE Xplore, CINAHL Plus) for English-language records from January 2017 to April 2026. Of 43,394 records identified, 275 were charted across four nonmutually exclusive domains: clinical decision support (233, 84.7%), predictive analytics (192, 69.8%), diagnostics (142, 51.6%), and economic assessment (53, 19.3%). Workflow or process measures appeared in 13 records (4.7%), equity or subgroup analyses in 20 (7.3%), and formal economic evaluations in 4 (1.5%). The authors group recurring constraints into five cross-domain barriers, describing them as structural rather than technical.

007
Large language models and artificial intelligence for generating, simplifying, and enhancing outpatient clinic letters: A systematic review

Can large language models reliably generate or simplify outpatient clinic letters? This systematic review searched five databases (PubMed, EMBASE, Web of Science, CENTRAL, CINAHL) from inception to 1 November 2025, appraising studies with the Mixed Methods Appraisal Tool and GRADE. Seven studies were included, two using real-world clinic data and five synthetic or hypothetical scenarios. Readability findings were mixed and most AI output still exceeded the US Grade 6 level; information fidelity ranged from 10% to 100%, and one study reported a ten-fold reduction in drafting time. Certainty was very low across primary outcomes.

008
Utilization of a HIPAA-compliant large language model chatbot in an academic pediatric medical center

Who actually uses a hospital-wide LLM chatbot, and what blocks the rest? This mixed-methods case study of "InternalGPT," a HIPAA-compliant chatbot at an academic pediatric medical center, combined employee surveys with 14 months of utilization data. Of roughly 15,800 employees, 2,149 (13.6%) requested access; 52.8% of those used at least one token, 33.6% never logged in, and the top 20% of users consumed 69.4% of tokens. Barriers cited by non-users were limited time (51.4%) and difficulty using the tool (21.8%). Among 92 sustained users surveyed, mean self-reported productivity gain was 30%, an exploratory $6.3M-$18.9M value.

009
Demographics, Clinical Content, Use Patterns, and Care-Seeking Intent Across Two Generations of AI-Enabled Clinical Triage Tools (A Traditional Structured Questionnaire and a Large Language Model-Enabled Conversational Interface): Comparative Retrospective Observational Study

Does conversational LLM triage capture different information or shift care-seeking intent compared with a structured questionnaire? This retrospective observational study compared 116,890 virtual triage encounters over 28 weeks (January–August 2025), where users self-selected traditional triage (TT; 100,533, 86%) or an LLM-enabled conversational interface (CT; 16,357, 14%) sharing the same Bayesian reasoning engine, with poststratification weighting by age and sex. CT sessions ran longer (median 8 min 21 s vs 4 min 25 s), elicited more clinical findings (median 36 vs 32; P<.001), and had higher self-reported intended adherence to recommended care (34.3% vs 29.2%; P<.001), including self-care (85.4% vs 61.9%). Users self-selected groups.

010
Understanding Key Stakeholders' Perspectives Towards Artificial Intelligence in Home Care Work

How do the people who would implement AI in home care view its promise and risks? This qualitative study conducted semi-structured interviews, incorporating AI scenarios, with 43 participants across five stakeholder groups — home health aides and attendants, home care agency leaders and staff, worker advocates, clinicians, and technology company personnel — recruited through purposive and snowball sampling, with analysis by structural coding, inductive sub-coding, and thematic analysis. Mean age was 44.6 years; 44.4% reported no or low AI knowledge. Four themes emerged: benefits to patient care, engagement, and efficiency; risks to care quality, provider-patient relationships, and working conditions; data quality, privacy, and AI literacy challenges; and needs for equitable governance. No effect sizes are reported.

012
Trust in Generative AI for Health Information Consumption and the Effect of Learned Dependency: Randomized Controlled Experimental Study

Does habitual reliance on generative AI blunt users' ability to calibrate trust in AI-generated health information? Two randomized 2×2 between-participants experiments (338 college students; 563 Mechanical Turk workers) manipulated information accuracy and text-based visual cues (highlighting), measuring trust and self-reported learned dependency with regression models. Accuracy raised trust (experiment 1 B=2.107, 95% CI 1.337-2.878; experiment 2 B=0.203, 95% CI 0.115-0.290), as did learned dependency (B=0.277 and B=0.822). The accuracy-by-dependency interaction was negative in both (B=-0.399; B=-0.459), indicating reduced sensitivity to inaccuracy. Text highlighting had no significant effect and did not moderate dependency.

013
Measuring Expert Inter-Rater Agreement with a Semi-Automated Clinical Decision Support System: Who Agrees with What?

How much do expert clinicians agree with each other when judging diuretic titrations, and what baseline should a semi-autonomous clinical decision support system (OTTO-FM) be held to? Secondary analysis of prospectively collected porcine data modeling postoperative fluid overload had three pediatric cardiac intensivists rate 29 clinician-driven and 44 CDS-driven items as reasonable/unsure/unreasonable, with ordinal-weighted Gwet's AC2. Dosing agreement was similar across phases (human 0.79, CDS 0.83), but unanimity reached only 65.5% and 70.5% of items. Risk labels split sharply: high-risk 0.92 versus low-risk 0.54; 19 of 20 nonunanimous risk items were single-rater dissent.

014
Continuous Remote Patient Monitoring in Heart Failure Patients: The Heart Failure Cascade Study: Phase II and III Outcomes

Does continuous remote patient monitoring reduce 30-day readmissions after heart failure discharge? This case-versus-retrospective-propensity-matched-control study at Endeavor Health (Evanston, IL) enrolled 39 patients across three phases, monitoring them for 30 days postdischarge with wearable biosensors and daily symptom surveys, with rules-based and machine learning alerts triaged by home health nurses and escalated to advanced practice providers. Intervention patients received more APP calls (66.7 vs. 7.7%), APP visits (43.6 vs. 5.1%), diuretic escalation (43.6 vs. 12.8%), and labs (76.9 vs. 43.6%). Adjusted 30-day readmission did not differ (odds 0.31, 95% CI 0.06–1.38; p = 0.138).

015
Beyond word error rate: clinical risk as the necessary standard for ambient AI scribe evaluation: evidence from 77 global languages

Does word error rate track clinically consequential errors in ambient AI scribes? Investigators built a synthetic multilingual corpus — five clinical dictation scripts across a complexity gradient, translated into 99 languages, rendered to speech under three acoustic conditions, and transcribed by a production scribe — then had three independent large language model raters score errors on a Severity x Likelihood framework. Of 59,819 genuine transcription errors, 58,329 (97.5%) were LOW risk and 251 (0.42%) CRITICAL or HIGH. No frequency metric was detectably associated with serious risk (absolute Spearman rho<0.16), while a Severity x Likelihood sum tracked WER (rho=0.80). Low-resource languages had worse WER (beta=+0.078) without higher critical risk (OR 1.21); consultation complexity predicted serious risk (OR 3.06 per level).

016
Optimizing generative artificial intelligence for clinical summarization: a blinded comparison study of automated versus human care planning synopses

Can a generative AI model produce care-transition synopses as good as clinician-written ones? In a blinded, randomized comparison using de-identified records of 64 patients with multiple chronic conditions from MIMIC-III, human- and AI-generated synopses were scored on accuracy, succinctness, synthesis, and usefulness within a data-information-knowledge-wisdom framework (>80% indicating success). AI and clinician summaries overlapped 12%. AI synopses were rated useful 75% of the time versus 76% for human synopses; AI scored lower on succinctness for the data task (55%-67%) and near-equal or better on accuracy and synthesis (AI 72%-79%, humans 68%-84%), best in wisdom. Interrater agreement was variable.

017
Human in the loop in AI-enabled clinical decision support: a systematic scoping review and reporting checklist for lifecycle governance

How is human-in-the-loop (HITL) actually implemented across the lifecycle of AI-enabled clinical decision support? This systematic scoping review searched MEDLINE, Embase, Web of Science, PsycINFO, Google Scholar and Scopus in August 2024, with manual identification through mid-2025, including primary studies in which clinicians interacted with AI-CDSS; dual independent screening and extraction mapped findings to four lifecycle phases. Twelve studies qualified. All described clinician involvement during development, mainly expert annotation and rule-based design; 11 reported review-phase HITL, 2 maintenance, and none oversight. Contributions were largely static or retrospective, with small datasets, few annotators, and inconsistent terminology. The authors propose a six-domain HITL reporting checklist.

018
Predictive analytics for health-system decision support using population health data: a global scoping review of implementation, governance and decision integration

How far has predictive analytics on routine and population health data actually moved into health-system decisions? This global scoping review searched five databases for 2014–2025 studies applying predictive or forecasting methods for health-system decision-making, screening 2,623 records and including 161 articles (128, 79.5%, from high-income settings), with a supplementary grey-literature scan. Most work stopped at development (139 articles, 86.3%); validation 6.2%, pilot 1.9%, operational deployment 5.6%. While 73.3% claimed relevance to resource allocation or capacity planning, only 5.6% documented an output-to-decision pathway and 8.7% routine workflow integration; 14.9% described how uncertainty informed decisions.

019
Attitudes Toward Large Language Models in Health Care and Preferences for Their Adoption and Oversight Among Health Care Professionals: Cross-Sectional Survey

How do clinicians view large language models and who should govern them? A cross-sectional online survey recruited 335 health care professionals through a health care news mailing list, 68.7% (n=230) attending physicians and 77.9% practicing in the Northeast United States. Some 62.7% (n=210) reported current or contemplated LLM use, with users reporting higher self-rated knowledge than nonusers (P<.001) and no age association (\u03c1=-0.072; P=.19). Top applications were literature review (73.4%), decision support (57%), and patient communication (54.9%); 96.4% voiced bias concern, 65.4% preferred oversight by professional associations over technology companies (29%), and 66.6% reported no confidence in existing oversight. Convenience sample, low response rate.

020
Clinicians' Attitudes and Perceptions on the Adoption of AI in Mental Health Care: Scoping Review

What do mental health clinicians think about AI in their practice? This scoping review followed JBI guidance and PRISMA-ScR, searching six databases (CINAHL, Embase, PsycINFO, PubMed, Scopus, Web of Science) for studies published from 2020 onward; 12,356 records were retrieved and 35 included after dual-reviewer screening. Clinicians showed cautious optimism when AI was framed as supplementing rather than replacing expertise, citing reduced administrative burden, documentation support, information synthesis, and between-session access. Concerns centered on privacy, governance, data ownership, unsafe or inaccurate outputs, overreliance, unclear accountability, and effects on therapeutic relationships, alongside limited AI literacy. The abstract reports no effect sizes.

021
Effectiveness, Safety, and Workflow Burden of Large Language Model-Based Medical Report Generation: Systematic Review

Are LLM-generated medical reports clinically ready in terms of effectiveness, safety, and workflow burden? This systematic review searched five databases (PubMed/MEDLINE, Embase, Web of Science, Scopus, Cochrane) for studies published January 2016 through May 2026 evaluating LLMs, multimodal LLMs, or vision-language models for image-to-report generation, impression drafting, or structured reporting, including 101 studies (36 chest x-ray). No study was at low risk of bias (15 moderate, 72 high, 14 serious), and meta-analysis was not possible. One chest x-ray study found AI report acceptance of 70.5% (6047/8580) versus 73.3% for radiologists, with false negatives 18.5% versus 17.8%; a brain MRI study showed reading time falling from 61 to 53 seconds while impression drafting increased editing time and edit distance.

022
Enhancing Patients' Informed Consent Through AI: Systematic Review

Can AI improve patient comprehension and decision-making during informed consent? This PRISMA-guided systematic review searched PubMed, Embase, and the Cochrane Library, including 33 studies published 2020-2025 across three domains: AI-generated patient education (n=18, 54.5%), consent documentation (n=10, 30.3%), and AI-assisted consent acquisition (n=5, 15.2%). Large language models were accurate but readability stayed above an eighth-grade level (best model Copilot: Flesch-Kincaid 10.59, SD 1.22). AI-generated documents raised Flesch Reading Ease by 44%-122% and lowered required grade levels 10%-47%. In trials, AI-assisted consent shortened consultations (7.7 vs 10.6 minutes; P=.05) and lowered post-consent anxiety in knee arthroplasty (10.48 vs 12.75; P=.04).

023
A Conceptual Model for Ambient AI Adoption: Perspectives From Academia and Industry

How should health systems judge ambient AI performance when vendor claims and real-world results diverge? This conceptual article, drawing on academic and industry perspectives, proposes a shared mental model organizing ambient AI assessment into three interdependent dimensions: technical, interface, and system level. For each dimension, the authors specify the types of information relevant to evaluation, what vendors should reasonably be expected to disclose, and how provider organizations can run internal evaluations to contextualize, verify, or supplement vendor claims. No empirical study, effect sizes, or performance benchmarks are reported; the contribution is a decision-support framework intended for organizations of varying size.

024
International qualitative case studies of system-level approaches to promote the development, adoption, and implementation of artificial intelligence in healthcare

How are health systems moving AI beyond pilots to scale? This qualitative study conducted four case studies — Catalonia, Norway, Singapore, and Queensland — drawing on 60 documents, 34 interviews, and 5 focus groups with 50 strategic decision-makers, policymakers, and lead clinicians, analyzed first within-case thematically and then across cases using the Technology, People, Organization, and Macroenvironment framework. Scaling trajectories were shaped by existing digital strategies, digitalization histories, funding arrangements, and legacy infrastructure, with experimentation opportunities, incentives, and distribution of decision-making authority mattering alongside post-deployment monitoring, governance, and procurement. The authors frame scaling as multilevel orchestration rather than top-down versus bottom-up. No effect sizes are reported.

025
From prediction to reality: Five years of AI in healthcare. Adoption, impact, and the road ahead

How well did the 2020 EIT Health & McKinsey AI forecast track what actually happened? This structured narrative review compared the report's domain-level predictions against 2020–2025 evidence from peer-reviewed implementation studies, FDA regulatory data, national and international guidance, industry surveys, and foundation-model evaluations, classifying each prediction as realised, under-realised, exceeded, or unanticipated. Two predictions held: administrative automation and medical imaging led adoption, with ambient documentation moving from pilots to deployment and more than 1300 FDA-authorised AI/ML-enabled devices, mostly radiology. Remote monitoring and classical NLP decision support under-delivered, limited by interoperability, reimbursement, and workflow barriers; generative and multimodal foundation models were the largest unanticipated divergence.

026
A five-phase evaluation framework for diagnostic and predictive medical artificial intelligence

How should medical AI be evaluated across its lifecycle, from bench performance to clinical benefit? This conceptual paper proposes a five-phase evaluation framework spanning technical validation, operational robustness, controlled interaction, clinical evidence, and real-world integration, embedded in a dynamic architecture with phase-gating criteria and local and systemic fall-back triggers that permit re-entry into earlier phases after drift, version updates, or safety signals. It maps multicenter external validation, shadow-mode testing, human-AI comparison and cooperation studies, randomized trials, real-world evaluations, and adaptive designs onto this pathway. No empirical data or effect sizes are reported; the contribution is a framework aimed at researchers, institutions, and regulators.

027
Implementation of an opioid use disorder (OUD) machine-learning phenotype in real-time for the ADAPT clinical trial

Can an EHR-embedded machine learning phenotype identify emergency department patients with opioid use disorder in real time for trial screening? Across three EDs in one US health system (2014–2025), a random-forest classifier using visit-level data available at or before triage was trained against a high-specificity computable label and deployed to trigger point-of-care alerts. Retrospective discrimination was high (AUROC 0.99; AUPRC 0.92). In prospective chart-review validation (n=217), 89.3% of model-positive and 95.7% of model-negative encounters matched physician adjudication; design-weighted estimates using the 3.26% flag rate (28,284/866,569 encounters) gave sensitivity 0.40 and specificity 0.996.

028
Physician Ratings and Adoption of AI-Generated e-Consultation Advice: A Randomized Clinical Vignette Study

Do physicians rate and act on AI-generated e-consultation advice differently when they know its source? In a randomized clinical vignette study, 44 internal medicine teaching faculty at one academic medical center reviewed four vignettes with advice generated by a medically specialized generative AI (OpenEvidence), labeled either as from a specialist attending physician or an AI system. Mean quality ratings were high across all six domains (overall means 4.2-4.7 of 5), with no differences by label. Selection of the consult-recommended management rose from 37.5% to 81.8% after viewing advice (OR 8.43, 95% CI 5.00-14.28); labeling did not affect adoption (OR 1.18, p=0.70).

029
Can a General-Purpose Coding Agent Analyze a Production Hospital Data Warehouse?

Can an autonomous coding agent replace hand-written analyst queries against a production EHR warehouse? In a single-center quality-improvement evaluation, ten acute otitis media questions were posed to analysts (adjudicated reference) and to OpenAI Codex, which wrote read-only queries against a full copy of an Epic Caboodle warehouse under four conditions. At medium reasoning effort the agent came within 5% of the reference on 27 of 30 runs but exactly matched on 10; three repeated runs agreed on only 3 of 10 questions. Matching runs identified the same patients (F1 = 1.00, one exception at 0.72); analyst feedback and reuse of corrected definitions restored reproducibility.

030
Multidimensional Assurance of an Electronic Health Record Embedded Generative Artificial Intelligence Summarization Tool

How reliable is an EHR-embedded generative AI summarizer in routine use? Investigators applied a multidimensional assurance framework to Epic IP Insights across a seven-hospital system, building an agentic hallucination detector that decomposed summaries into atomic content units (ACUs) and checked each against source notes. The tool produced 706 summaries over 445 encounters, drawing on a mean 5.8% of available notes. The detector matched physician adjudication on 96.95% of ACUs in validation; across 671 summaries (40,452 ACUs) the hallucination rate was 11.79% (95% CI, 11.47%-12.10%). Among 385 regenerated pairs, 32.7% were textually and 37.9% semantically identical. Clinicians rated it easy to use (97.8%) but were mixed on relying on output without verification (45.9%).

031
Enhancing AI Use: How Complementary System Information Drives Delegation Frequency and Effectiveness

Can information about an AI system help people delegate tasks to it more often and more wisely? This experimental study manipulated two signals: ex-ante AI certainty (the AI's estimated likelihood of being correct, shown before the delegation decision) and ex-post outcome information (whether the AI was actually correct, shown after). Presenting either signal alone had no effect or reduced combined human-AI performance; providing both raised delegation frequency and delegation effectiveness, improving performance. The authors attribute this to certainty calibrating task-level expectations, outcomes confirming them, more accurate mental models, and less algorithm aversion. The abstract reports no effect sizes, sample size, or task details.

032
Artificial intelligence-supported total parenteral nutrition management in neonatal intensive care units: A systematic review of clinical efficacy, safety, and system integration

Does AI support for total parenteral nutrition in neonatal intensive care improve efficacy, safety, and integration? This PRISMA 2020 systematic review searched PubMed/MEDLINE, Scopus, and Web of Science, appraising evidence with RoB 2, the Newcastle-Ottawa Scale, and GRADE. Sixteen records met criteria; thirteen quantitative primary studies (2008-2026, n = 30-9,330) formed the synthesis. CPOE implementation cut parenteral nutrition medication error rates from 10.8% to 3.2%; the TPN2.0 transformer model correlated with expert decisions at Pearson R = 0.94, and classical machine learning reached R2 > 0.70 for macronutrient prediction. Only one-third of U.S. NICUs used a CDSS. GRADE certainty was moderate for efficacy and safety, low for system integration.

033
Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review

How mature is the evidence behind "data-centric" multimodal deep learning for clinical decision support? This PRISMA 2020 systematic review screened 150 records and included 31 primary clinical studies, 30 (97%) published 2024-2026, coding implemented versus merely mentioned techniques and appraising bias with PROBAST+AI. Studies used a median of three modalities (range 2-6), most often structured EHR (71%) and imaging (39%). Data-centric techniques were reported in 74-84% of studies (equity 61%), but external validation appeared in only 4/31 (13%), a clinical or provider outcome in 3/31 (10%), none reported deployment, and 27/31 (87%) were at high risk of bias.

034
Clinician-led, AI-assisted clinical software development: multi-domain verification of a deployed ePROM application

Does high usability in a clinician-built, AI-assisted application establish clinical and technical assurance? This single-case retrospective development-and-assurance report examined STUIapp, a browser-based ePROM tool integrating six validated lower urinary tract symptom instruments, using code audit, independent clinical review of a frozen 78-case scoring matrix, WCAG 2.1 measurement, a 23-canary persistent-storage study, and usability testing with 14 clinicians, 26 patients and 12 older adults. Mean SUS was 92.3 (SD 8.8) among clinicians and 92.0 (SD 10.8) among patients, yet all 78 passing automated cases included 17 expected results requiring correction. Other findings: instrument mislabelling, a failed installability manifest, a third-party analytics tag contradicting local-only privacy claims, 10-px text and 2.56:1 contrast, and storage permission denied in 10/10 browser-tab canaries versus granted in 13/13 installed canaries.

035
AI Adoption in US Cancer Centers: National Cross-Sectional Study of Institutional and Policy Determinants

What predicts publicly visible AI adoption at US cancer centers? This cross-sectional study assembled public-source data on 75 NCI-designated cancer centers, scoring adoption across screening, treatment, and patient care as a 0-3 composite index, with Moran I tests for spatial clustering and ordered logistic regression on institutional and contextual predictors. The mean adoption index was 1.37 (SD 0.86), highest for screening (0.86), then patient care (0.50) and treatment (0.22). Moran I showed no significant spatial autocorrelation. Physician workforce and bed capacity showed positive but modest associations; state socioeconomic indicators did not. Political-context findings were mixed; the abstract reports no effect sizes for regression estimates.

036
Sailing by the Stars: A Cross-Sectional Survey Study of Generative AI Use Among GME Trainees and Faculty

How are graduate medical education trainees and faculty using generative AI? This cross-sectional survey study, reported in the Journal of General Internal Medicine, examines self-reported generative AI use among GME trainees and faculty. No abstract was available, so findings, sample size, and effect estimates cannot be summarized here.

037
Drivers of Oncologist Preference of AI-Generated Literature Review in a Randomized Mixed-Methods Study

What features of AI-generated literature reviews make them useful and trustworthy to oncologists? In a randomized mixed-methods study, 34 oncology physicians produced 294 ratings of four blinded AI systems across five clinical vignettes, supplemented by 20 semi-structured interviews analyzed with a prespecified LLM-assisted qualitative pipeline. Despite similar references, an evidence-graded report adapted from OpenEvidence scored significantly lower in overall utility than standard OpenEvidence (mean difference -0.96; 95% CI -1.26 to -0.66; P<.001). Qualitative analysis yielded six themes and seven design requirements: clinicians preferred concise, scannable reports with quantitative outcomes, bolded guidelines, explicit uncertainty, and verifiable citations; trust fell with citation mismatch, buried provenance, and overconfident recommendations.

038
A Pragmatic Randomized Trial of an EHR-Integrated Generative AI Chart Summarization Tool for Ambulatory Clinicians

Does an EHR-embedded generative AI chart summarization tool reduce clinician workload? In a pragmatic 1:1 randomized trial at one academic health system, 284 ambulatory clinicians across 42 specialties received Epic's outpatient chart summarization tool or usual care over 90 days (February o task task load favored the intervention ( tionsted difference -27.4 on a 0-400 scale; 95% CI, -49.4 to -5.3; P=0.02), with lower burnout (-0.20) and work exhaustion (-0.24) on the Professional Fulfillment Index. Charting time per encounter was unchanged (-1.2 seconds). Only 14.2% of 74,474 summaries were opened, falling from 21.5% to 10.5% by month 3; net promoter score was -22.

039
Clinical Implementation of Pharmacogenetics-Based Machine Learning Clinical Decision Support Systems: A Scoping Review

How far has machine learning–based pharmacogenetic clinical decision support actually moved into clinical workflows? This scoping review searched multiple databases for studies published January 2015 through September 2025 reporting ML-based CDSS incorporating pharmacogenetic data to support therapeutic decisions in clinical settings. Of 1,262 records screened, 7 met inclusion criteria, and only 2 evaluated tools in live clinical or trial workflows. Studies varied in design, setting, and implementation maturity; most reported potential benefits including fewer preventable adverse drug events, better prescribing accuracy, or improved workflow integration, though the abstract reports no effect sizes. Authors call for implementation science evaluating usability and patient outcomes.

040
Quality, consistency, and clinical safety of AI-generated versus clinician-written clinical notes: a multi-country paired simulation study

Do ambient AI scribes produce notes of comparable quality and safety to clinician-written ones across languages? This paired simulation covered five countries and languages (Cambridge, Barcelona, Milan, Paris, Cologne), with 385 actor-performed consultations yielding 770 notes, each documented independently by an AI scribe (Heidi) and a junior-to-middle-grade clinician, scored on the PDQI-9 by blinded evaluators. AI notes scored higher (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; dz=0.55) and less often carried at least one Critical+High error: 24.4% vs 61.0% by calibrated automated review (RR 2.50) and 6.2% vs 21.8% by clinician adjudication (RR 3.50), with omissions driving the gap.

041
Deployer-side governance of medical imaging artificial intelligence: the regulatory readiness instrument (RRI-MI) for multi-jurisdictional and post-market compliance

How can hospitals operationalize post-market oversight of imaging AI when pre-market clearances test algorithms under static conditions? This conceptual paper proposes the Regulatory Readiness Instrument for Medical Imaging AI (RRI-MI), a deployer-side readiness assessment mapping 11 governance domains onto a provisional 22-point ordinal rubric, aligned with the FDA Predetermined Change Control Plan and the EU AI Act. The authors ground the framework in post-market evidence on scanner drift, protocol shifts, software updates, and demographic variation, illustrating it through a semiautonomous prostate cancer MRI case study covering local validation, human oversight, version control, monitoring, and incident response. No empirical validation or effect sizes are reported.

042
AI-assisted radiographic fracture detection and length of stay in the adult ambulatory orthopedic emergency department: a before-after cohort study with a disease-specific internal control

Does deploying an AI radiographic fracture detection tool shorten emergency department stays? This retrospective controlled before-after cohort study with interrupted time-series analysis covered 8,253 fracture-excluded and 6,864 fracture-confirmed visits at a tertiary ambulatory orthopedic ED, January 2021–February 2026, with deployment on 1 October 2023. Mean length of stay in fracture-excluded patients fell from 135.6 to 127.5 minutes (−8.13; 95% CI −11.52 to −4.70), while the fracture-confirmed control was unchanged (+0.37 minutes; p=0.87). Reductions concentrated at the upper tail (−24.0 minutes at the 90th percentile); stays over four hours fell from 11.3% to 8.2%, with no increase in early revisits.

043
Privacy, security, and reliability risks of artificial intelligence in healthcare: a systematic review of empirical evidence

What empirical evidence exists for privacy, security, and reliability risks from AI in clinical care? This systematic review searched PubMed, Embase, Web of Science, Scopus, IEEE Xplore, and ACM Digital Library for empirical studies published January 2015 to November 2025 evaluating AI use or misuse in diagnosis, treatment, or decision-making. Of 7,285 records plus 205 from citation screening, 22 studies met inclusion criteria, mostly medical imaging. Five recurring threat categories emerged: re-identification, membership inference, unauthorized access and adversarial exploitation, input manipulation, and misuse or overinterpretation of outputs. Models encoded latent biometric signals, limiting anonymization and synthetic data. Findings were synthesized narratively; the abstract reports no pooled effect sizes.

044
Ethics of Autonomous AI Clinical Trials: Delphi Study

How should the NIH's 7 principles of ethical clinical research be adapted for trials of autonomous AI? Using a modified Delphi approach over 6 months, investigators convened 14 multidisciplinary panelists (AI, data science, ophthalmology, policy, law, bioethics, patient advocacy) across two survey rounds anchored to a vignette and a final virtual meeting, with participation of 12/14 (85.7%), 10/14 (71.4%), and 13/14 (92.9%). Round 2 produced 9 strong-agreement, 2 moderate-agreement, and 4 divisive statements. Recommendations covered transparency on training and validation data, pre-deployment bias and inequity assessment, performance across clinical settings, informed consent, comparison with standard of care, downstream access, and cost.

045
The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review

How have researchers actually operationalized human evaluation of large language model reliability in health care? This PRISMA-ScR scoping review searched PubMed, Web of Science, Cochrane Library, CINAHL, and Google Scholar for English-language original studies published January 2016 to July 2025, screening 4347 records and including 71 (26 clinical, 45 public health). Six reliability indicators recurred: accuracy, relevance, completeness, clarity, safety, and consistency. Clinical studies more often assessed guideline concordance and structural coherence; public health studies emphasized understandability, harm potential, and repeat-response consistency. Single-specialty clinicians predominated as evaluators, panels typically had five or fewer members, and 5-point Likert scales with researcher-defined rubrics were standard. Reported limitations included evaluator subjectivity and nonstandardized indicators.

046
1,357 AI medical devices cleared, 3 actually tested on patient outcomes

How much clinical evidence supports FDA-cleared AI medical devices? This systematic analysis catalogued all 1,357 AI/ML-enabled devices cleared through December 5, 2025 using the FDA device database and the ACR Data Science Institute catalogue, with linked searches of ClinicalTrials.gov and PubMed for registered trials and publications. Only 34 devices (2.5%) were linked to registered prospective trials, 12 (0.9%) posted results, 12 (0.9%) had peer-reviewed publications, and 3 (0.2%) evaluated patient-centered outcomes such as mortality, morbidity, or readmissions. Most studies (62%) were observational with small, homogeneous cohorts and frequent exclusion of vulnerable populations. The authors cite misaligned incentives and predicate-based pathways as barriers.

047
Machine Learning in Palliative Care: Scoping Review of Applications

How far have machine learning applications in palliative care moved beyond mortality prediction? This scoping review followed Arksey and O'Malley and PRISMA-ScR guidance, searching six databases from inception through February 9, 2026, and included 121 peer-reviewed primary studies (2015-2026) across 24 countries, 54.5% (66/121) from the United States. Mortality and survival prediction accounted for 42.1% (51/121) of studies, followed by health care use (20.7%) and symptom assessment (16.5%); cancer populations dominated (43%). Half (50.4%) used an explainability technique, 86.8% (105/121) did not address equity in model performance, and 66.1% remained proof-of-concept with 17.4% reaching prospective deployment or clinical integration.

048
Centralized Digital Surveillance for Abdominal Aortic Aneurysm Detection, Longitudinal Tracking, and Management Within an Integrated Health System: Retrospective Cohort Study

Can a centralized digital program identify and track abdominal aortic aneurysms across an integrated health system? This retrospective cohort study describes STAIR, which combined structured EHR problem-list queries, an internally developed NLP model applied to radiology reports, clinician referrals, and automated lost-to-follow-up queries, enrolling 8464 patients (mean age 77.1 years; 77.0% male) from December 2022 through December 2024, with status assessed through April 2026. Identification was mostly automated: problem-list queries 59.0% and radiology NLP 29.0%. After centralized review, 45.3% were assigned biennial duplex surveillance and 20.6% referred to vascular surgery; 49.5% remained under active surveillance. Findings are descriptive; clinical effectiveness was not assessed.

049
Prospective evaluation of a large language model clinical decision support system in the emergency department

Can a multi-LLM clinical decision support system change emergency department care in practice? This DECIDE-AI stage 1 prospective evaluation deployed SHAKED in a tertiary ED over 4 weeks, analyzing 1,138 patients across two parallel units, one using the system and one following routine rotations. Adoption fell from 68% to 30%, with disengagement tied to workload (OR 0.72 per shift hour, 95% CI 0.62–0.83), while physicians favored it for radiology consultations (OR 2.98). Expert reviewers rated 99 of 100 sampled outputs clinically appropriate, with no adverse events. Length of stay was 4.9 hours in both wings (P = 0.99); consultation cycle time trended shorter (−9.4 min, P = 0.077).

050
An Electronic Health Record-Integrated, Large Language Model-Powered Tool to Triage Surgical Patients

Can eligibility for surgical comanagement (SCM) — hospitalist co-management of medically complex perioperative patients — be triaged automatically? This prospective, unblinded quality improvement study at Stanford Health Care (September 2025–February 2026) deployed an EHR-integrated, human-in-the-loop LLM tool (SCM Navigator) that classified patients using preoperative documentation, structured data, and morbidity criteria, with attending review as the reference standard. Across 6193 triaged cases (median age 60.2 years; 49.0% female), 1582 (25.5%) were recommended for hospitalist consultation; sensitivity was 0.94 (95% CI, 0.91-0.96) and specificity 0.74 (95% CI, 0.71-0.77). LLM misclassification explained 2 of 19 false negatives (11%).

051
Clinical Specialty Expansion of AI-Enabled and Machine Learning-Enabled Medical Devices Authorized by the US Food and Drug Administration From 1995 to 2025: Longitudinal Content Analysis

Has radiology's dominance of FDA-authorized AI/ML medical devices persisted or begun to loosen? This longitudinal content analysis covered all 1430 devices in the FDA AI-Enabled Medical Devices registry with final authorization decisions through December 2025, stratified by advisory-committee specialty across four eras (1995-2015, 2016-2019, 2020-2022, 2023-2025). Annual authorizations rose from a mean of 2.0 to 264, with 331 in 2025; 96.2% used the 510(k) pathway. Radiology's share peaked at 85.5% (347/406) in 2020-2022, then fell to 77.5% (614/792) in 2023-2025 (P=.001), with the Herfindahl-Hirschman Index declining from 0.738 to 0.612. Start-ups (OR 5.09) and technology companies (OR 50.62, based on 13 devices) had higher odds of nonradiology authorization.

052
Beyond the Black Box: Unraveling the Role of Explainability in Human-Artificial Intelligence Collaboration

When does explaining an AI model's reasoning actually improve human-AI decisions, and at what cognitive cost? The authors build an analytical model of a decision maker with limited but flexible cognition receiving imperfect machine recommendations, where explanations shift beliefs about algorithmic quality. Low explainability leaves decision accuracy and reliance unchanged while reducing cognitive burden; higher explainability improves accuracy by curbing overreliance but increases underreliance. Explainability matters more for cognitively constrained decision makers, complex tasks, and lower-stakes decisions, yet can raise processing time and fatigue exactly when time is short, tasks are complex, and machine quality is doubted. Theoretical modeling, so no empirical effect sizes.

053
Benchmark Mineability and the Financing of AI Innovation -- by Alex Chan

When AI benchmark scores steer capital, what happens to their value as signals? This NBER working paper is a conceptual and theoretical market-design analysis rather than an empirical study, treating public AI benchmarks as market institutions. The author identifies two gaps that become exploitable once scores move investment: public examples can reveal the process behind a private final test, and any finite public score cannot span the broad task space implied by general intelligence. Targeted effort aimed at these gaps erodes the signal later investors rely on. The proposed remedy is separating development from certification — publishing practice tasks but selecting the investment-consequential task generator only after a submitted system's evaluation policy is fixed. The abstract reports no effect sizes or empirical magnitudes.

054
Are automated documentation-error judges fit to measure ambient AI scribes? A pre-registered, blinded human-validation study

Can automated judges that flag documentation errors serve as a defensible measurement instrument for ambient AI scribes, absent a gold standard? This pre-registered, blinded validation study nested in a multi-country ambient documentation simulation (English setting) had ten external clinicians adjudicate a stratified sample of 434 pipeline flags, yielding 565 adjudications with 131 double-rated. Inter-clinician agreement on error genuineness was fair (raw 59%, AC1 0.24); judge-clinician agreement was 64% (95% CI 60-68). Behavior was near-symmetric across arms (kept-precision 74% AI vs 81% clinician notes; severity gap +0.06 vs -0.09 tiers), with removed-confirmed asymmetric (56% vs 42%). Latent-class triangulation estimated a 68% genuine-error rate (94% CrI 48-83).

055
Performance of an Ambient Generative AI Documentation Tool in a Linguistically Diverse Clinical Setting

Do ambient AI scribes perform equally well across patient languages? This retrospective analysis covered 54,160 outpatient encounters in a U.S. safety net health system, measuring the share of words in the final note generated by the AI tool and left unedited by the provider. Generalized estimating equations with exchangeable correlation structures accounted for clustering within patients, with univariable models by language and interpreter modality and a multivariable interaction model. Non-English encounters were 21% to 25% less likely than English encounters to reach the performance threshold; interpreter-mediated and bilingual-provider encounters did not differ significantly.

056
AI-generated patient instructions as safety-critical communication: a provisional evidence-informed framework for clinical deployment

How should health systems evaluate generative AI that turns clinical information into patient-facing instructions such as discharge summaries, medication explanations and portal messages? This narrative and interpretative review maps empirical studies of AI-generated discharge communication alongside health-literacy, medication-safety, patient-safety and AI-governance literature. The author reports a recurring trade-off: large language models improve readability and understandability, while physician and pharmacist review identifies omissions, inaccuracies, newly introduced actions, medication-related problems and potentially harmful content, especially in complex discharges. A proposed seven-domain framework covers factual accuracy, clinical completeness, actionability, medication clarity, escalation and safety-netting, health-literacy alignment, and accountability with auditability. The abstract reports no effect sizes.

057
Application of Artificial Intelligence (AI) in cancer symptom management for adult cancer survivors: a scoping review

How is AI being applied to cancer symptom management for adult survivors? This scoping review followed Joanna Briggs Institute methodology, searching eight databases (November 2025, updated March 2026) for English-language empirical studies from 2015 onward, with narrative synthesis and inductive content analysis. Of 41 included studies, 21 addressed model development, 18 intervention delivery, and 2 both; common techniques were natural language processing, machine learning, and conversational AI. Applications targeted symptom detection (n=15), patient education (n=9), monitoring (n=5), and personalised management (n=5); pain and psychological distress each appeared in 9 studies. Model performance was moderate to high; the abstract reports no pooled effect sizes.

058
Real-World Barriers to and Facilitators of Implementing AI-Based Clinical Decision Support Systems: Scoping Review

What actually helps or hinders AI-based clinical decision support once it reaches the bedside? This scoping review searched five databases (MEDLINE, Embase, CINAHL, APA PsycInfo, Cochrane) from inception to May 2022, screening 10,875 titles and abstracts and 494 full texts; 13 studies met inclusion criteria and 9 reported explicit implementation determinants, mapped to the Consolidated Framework for Implementation Research by two reviewers. Studies were mostly US-based, multicenter machine learning tools in critical care and emergency medicine. Reviewers identified 28 determinants: 16 barriers (limited interpretability, data quality, workflow misalignment, user capability and motivation) and 12 facilitators (end-user needs assessment, stakeholder engagement, peer endorsement, supporting evidence).

059
Large Language Models Generate Stigmatizing Language During Reasoning Over Real-World Clinical Data

Do large language models produce stigmatizing language when reasoning over clinical data? Researchers applied a psychiatrist-validated NLP stigma-term detector to reasoning text from 107 LLMs across 35 real-world clinical tasks, covering 3,745 model-task pairs. Stigma rates ranged from 0% to 33.33%, and 84.06% of pairs contained at least one stigma term. Open-source models exceeded proprietary (1.97% vs. 1.60%; p<0.01) and reasoning models exceeded non-reasoning (2.35% vs. 1.70%; p<0.0001), while general versus medical models did not differ (2.00% vs. 1.80%; p=0.26). Stigma correlated negatively with task accuracy (r=-0.304) and positively with input-note stigma (r=0.569); 19.76% of pairs amplified input stigma. Prompt engineering cut stigma rates by up to 91.91% without reducing performance.

060
From Output Errors to Workflow Harm: A Practitioner-Audit Method for LLM-Mediated Research

How should LLM failures be evaluated when they occur inside multi-step research and clinical workflows rather than isolated prompts? This method paper introduces TRACE, a practitioner-audit framework, with an empirical demonstration: 45 documentation-positive incidents logged by a single clinician-informatician across scholarly, clinical informatics, and clinical-adjacent workflows over seven weeks, coded with a consequence-based severity rubric and taxonomy crosswalk. Four categories tied as most frequent (verification failure, factual numerical error, tool-behavior misunderstanding, citation formatting; n=7 each), and one incident carried an estimated $2,500 impact. Category agreement was low among three human reviewers (Fleiss κ=0.155) but substantial among three vendor-blinded AI comparators (κ=0.632). The authors frame this as a pilot that does not estimate error rates.

061
Triage and Referral Behavior of Patient Facing Medical Artificial Intelligence Products

Do patient-facing AI products differ in how they triage and refer simulated patients? Researchers tested nine products against 60 physician-developed standardized clinical cases across 540 multi-turn simulated patient encounters. Overall triage accuracy showed no statistically significant difference across product categories, but referral behavior diverged: branded health AI products over-triaged low-acuity cases far more often (28% vs 3% vs 2%) and more frequently recommended affiliated, fee-requiring clinical services. The authors argue evaluations of patient-facing medical AI should assess referral behavior alongside triage accuracy. The abstract reports no confidence intervals or other effect sizes.

062
Real-world use of large language models for mental health in 2024

How many U.S. adults turn to general-purpose large language models for their mental health? A cross-sectional survey of 1,871 U.S. adults fielded August–October 2024 used stratified sampling across age, sex, and race/ethnicity to approximate national demographics. Twenty-four percent reported using LLMs for mental health; users skewed young, male, and Black, and reported poorer mental health and difficulty accessing traditional treatment, citing that LLMs are free, convenient, and always available. Reported uses included emotional support, learning therapy skills, and supplementing existing therapy. Combining with Pew estimates of overall LLM use, the authors project 14–18 million U.S. adults.</summary>}</summary>

063
How large language models can be used for teamwork and communication in healthcare settings: A scoping review

How are large language models currently being used to support teamwork and communication within healthcare teams? This scoping review followed PRISMA-ScR guidelines, searching PubMed, Web of Science, and ScienceDirect for 2014-2024 publications; 3,865 unique titles and abstracts were screened, 127 full texts reviewed, and 20 studies included. Designs were predominantly quantitative and simulation-based, with limited in situ evaluation. Use cases spanned decision support, communication, and administrative functions. Outcomes centered on accuracy and quality (15/20 studies), with fewer addressing safety (4/20), readability or empathy (5/20), workflow efficiency (3/20), and error modes (2/20). The authors note most studies evaluate model performance without accounting for human team dynamics.

064
Clinical predictive artificial intelligence evaluation: A narrative review of trial designs and practical considerations

How should clinical predictive AI be evaluated when conventional randomized controlled trials are static, slow, and mismatched to models that drift and get updated? This narrative review argues for adaptive, iterative, context-specific assessment and proposes a framework separating three activities: performance monitoring (calibration, discrimination, data drift, alert burden, fairness, workflow fidelity), clinical impact monitoring of sustained benefit, and causal evidence generation via pragmatic and adaptive platform trial designs. The authors add a governance-driven escalation protocol specifying when monitoring signals should trigger formal trials, a signal-to-design decision pathway, and a guide to causal inference methods. As a review, it reports no effect sizes.

065
Lessons from deploying the ChatEHR system at Stanford Medicine

How does an LLM-based conversational interface to the electronic health record fare in real clinical deployment? This Nature Medicine piece reports on implementation lessons from ChatEHR at Stanford Medicine, covering the practical, technical and organizational considerations of putting a chart-querying system into use in an academic health system. No abstract was available, so findings, evaluation methods and any performance or utilization figures cannot be summarized here.

066
Real-world evaluation of a transformer-based natural language processing system for identifying social determinants of health from routine clinical documentation

Can transformer-based NLP reliably surface social determinants of health from routine notes? This validation study at University of Florida Health surveyed 1001 adults with at least two encounters in the prior year (sampling targeted 50% Black patients), restricting comparative analyses to 414 participants who also had Epic SDoH questionnaire data and notes. The SODA pipeline's extractions across nine domains were benchmarked against the research survey and the Epic instrument using sensitivity, specificity, PPV, NPV, and F1. Sensitivity reached 55% for alcohol use but fell to 16% for financial constraints, 5% for abuse, and 0% for drug use; the two surveys agreed only modestly, so no definitive reference standard existed.

067
Evaluation Methods for Inference-Time Retrieval-Augmented and Graph Retrieval-Augmented Large Language Models in Health Care: Scoping Review

How are retrieval-augmented and graph retrieval-augmented LLM systems in health care actually evaluated? This PRISMA-ScR scoping review searched PubMed, Web of Science, IEEE Xplore, ACM, arXiv, and medRxiv through May 2026, charting 157 eligible studies. Clinical question answering dominated (89/157, 56.7%), followed by decision support (70/157, 44.6%). Evaluations were overwhelmingly offline (140/157, 89.2%), with only 10.8% workflow-facing, prospective, or deployment-level. Independent retrieval-layer evaluation appeared in 29.9%, grounding or faithfulness in 26.1%, fine-grained evidence verification in 14%, and safety evaluation in 28.7%. Among 94 studies using human evaluation, 27.7% reported interrater reliability.

068
Impact of Large Language Model-Based AI Tools on Physician-Patient Communication: Systematic Review and Meta-Analysis

Do LLM-based chatbots improve physician-patient communication? This PRISMA-guided systematic review and random-effects meta-analysis searched PubMed/MEDLINE, Embase, Scopus, and Web of Science for 2020-2025 studies of LLM or chatbot applications in clinical communication, including 10 quantitative, mostly cross-sectional studies from 312 records. In 5 of 6 direct comparisons, LLM responses were rated more empathetic than physician responses; one study found chatbot replies judged empathetic in 45.1% of cases versus 4.6% (odds ratio ~9.8, P<.001), and GPT-4-simplified pathology reports raised comprehension scores (7.98 vs 5.23/10). Pooled empathy effect was large (SMD 1.02, 95% CI 0.44-1.60; k=4, N=2604). Satisfaction results were mixed; no study assessed long-term trust.

069
Divergent impacts of explainable AI for dermatological diagnosis on clinicians versus lay people

Does explainable AI help clinicians and lay users equally in dermatological diagnosis? Two large-scale experiments enrolled 623 lay people and 153 primary care physicians, pairing a fairness-constrained AI model for skin disease diagnosis with several explanation formats, including multimodal large language model explanations. Assistance from the fairness-trained model, which performed comparably across skin tones, improved final diagnostic accuracy and narrowed skin-tone performance gaps in both groups. LLM explanations diverged: lay users showed automation bias, gaining when the model was right and losing when it erred, while physicians benefited regardless. Presenting AI predictions before human judgment increased anchoring. The abstract reports no effect sizes.

070
Predictive Risk Scores in the Public Sector: Experimental Evidence from Child-Protection Investigations -- by E. Jason Baron, Arkadev Ghosh, Richard Lombardo

Can algorithmic risk scores improve how child-protection supervisors allocate scrutiny? A randomized evaluation covering 4,752 child referrals over 14 months in Northampton County gave supervisors an algorithmic risk score alongside standard case records. Access to the score increased foster-care placements and service receipt among children at the highest predicted risk, with little change for lower-risk cases, and reduced subsequent maltreatment referrals. The authors report no evidence that the score widened racial disparities in decisions or outcomes. The abstract reports no point estimates or effect sizes for these changes.

071
Advancing Human-Centered AI in Clinical Decision Support: Sociocognitive Human-in-the-Loop Study in HIV Care

How can machine learning outputs derived from HIV electronic health records be translated into a clinical decision support system clinicians will actually use? This human-in-the-loop field study at Prisma Health in South Carolina combined pre- and postsurveys, interactive usability testing, think-alouds, and in-depth interviews with 16 clinicians providing HIV care—physicians, nurse practitioners, infectious disease pharmacists, social workers, and case managers—between March and September 2025. Clinicians anchored interpretation of AI predictions on familiar clinical indicators but centered social determinants of health in their own risk assessments; trust was conditional and accrued over time, with explainability and actionability described as prerequisites for intervention. The abstract reports no effect sizes.

072
Propagation of Interpreter Errors by Ambient AI Scribes: Study Using Simulated Clinical Encounters

Do ambient AI scribes carry interpreter errors into the clinical note? Using simulated English- and Spanish-language clinical encounters mediated by interpreters, the authors evaluated whether documentation generated by ambient AI scribes reproduced interpretation errors introduced during the visit. Scribes propagated interpreter errors into the resulting notes, with propagation patterns differing by speaker role and by error type. The published abstract reports no effect sizes, error counts, or comparative rates. The authors frame the results as a case for further evaluation of AI-scribe performance in multilingual and interpreter-mediated care.

073
Risk-Tiered Governance for Hospital Artificial Intelligence: A Framework Synthesis and Implementation Pathway

How should hospitals calibrate oversight of AI tools embedded in EHR workflows, imaging, triage, documentation, and operations? The authors conducted a narrative review and framework synthesis drawing on peer-reviewed evidence, reporting guidelines, regulatory and policy sources, implementation studies, and applied governance case reports. The resulting framework has four components: a use-case inventory tagged by decision influence and workflow coupling; a six-domain risk taxonomy spanning clinical safety, privacy and data security, ethics and fairness, transparency, system stability, and compliance; a four-tier risk scheme keyed to harm, automation, reversibility, and coupling; and a governance architecture assigning roles to a committee, clinical owners, risk-control functions, and independent assurance. A lifecycle pathway runs from initiation and local validation through shadow mode, controlled go-live, monitoring, change control, and retirement. No effect sizes are reported; this is a conceptual framework, not an evaluation.

074
Shadow AI in Swedish Health Care: Qualitative Analysis of Physicians' Free-Text Answers

For what purposes do physicians use unauthorized, non-conformity-assessed AI tools at work? This cross-sectional survey of physicians in Swedish health care organizations (N=357; response rate ~64%), fielded through a verified online panel between December 2023 and January 2024, applied qualitative content analysis to free-text responses, interpreted through the sociology of professions and paradox theory. Reported uses fell into four categories: clinical work and decision-making (second opinions, differential diagnoses, rare cases), administrative work (patient communication, documentation), research and professional development, and technological curiosity. Physicians framed such use as compensating for gaps in institutional systems and reducing workload. The abstract reports no effect sizes or usage prevalence.

075
Evaluation Frameworks for Clinical AI Incorporating Validation Strategies, Real-World Applicability, and Ethical Principles: Scoping Review

How consistent are the evaluation frameworks proposed for clinical AI? This scoping review followed PRISMA-ScR, searching six databases plus the EQUATOR Network through February 2026, and screened 3363 records to include 46 frameworks, scored on methodological rigor, validation strategy, and a 10-domain UNESCO ethics matrix. Most frameworks (88%) targeted investigational rather than clinical use; 31.8% reported technical metrics such as AUC, sensitivity, and specificity, 15.9% reported clinical indicators, and only 11.4% met methodological rigor with validation aligned to intended use. Ethics coverage was uneven: transparency and explainability appeared in 70%, human oversight in 24.4%.

076
An Acceptance Criteria Framework for Determining the Implementation Fit of Custom Large Language Models in Public Health Interventions

How should public health teams decide whether a customized large language model is ready to deploy? This conceptual paper proposes an acceptance criteria framework (ACF) defining implementation fit as meeting prespecified minimum performance standards and showing nonproblematic behavior under anticipated use. The ACF combines project-relevant and off-topic test prompts, structured expert review, and prespecified thresholds to generate a documented decision record that can be rerun after model revisions. The authors argue prior safety, ethics, effectiveness, engagement, and implementation frameworks imply rather than operationalize deployment benchmarks, and illustrate the ACF in a tobacco cessation text messaging intervention. No performance estimates are reported.

077
A Secure User Interface for Preclinical Evaluation of AI in Patient Portal Message Management: Tutorial

How can health systems test large language models on real patient portal messages without touching live EHR workflows? This technical feasibility tutorial describes a Python 3 web interface and modular backend running inside the institutional firewall on an NVIDIA GRID T4-1Q GPU, supporting single-message and batch tasks: authorship identification, categorization, criticality flagging, and response drafting with zero-, one-, and few-shot prompting. A deidentification pipeline validated against 110 manually adjudicated entities achieved 95.1% sensitivity and 82.1% precision. Use cases drew on an IRB-approved dementia-relevant corpus of 6941 medical advice request messages from 497 patients; token-based cost readouts were included. No comparative performance effect sizes are reported.

078
Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems

Does acceptable pre-deployment validation performance persist once clinical AI enters routine workflows? This longitudinal retrospective observational study followed four deployed AI systems spanning different clinical domains within one large healthcare organization, comparing validation-era metrics with post-deployment behavior over extended observation using routine clinical data, outcome labels, and operational telemetry, and examining discrimination, calibration, data availability, latency, and workflow signals. In all four systems, validation performance did not persist; calibration drift appeared consistently and often preceded discrimination changes, and label-independent signals such as input missingness and data latency flagged degradation earlier than outcome-based monitoring, which lagged behind label availability. The abstract reports no effect sizes.

079
Comparison of Initial Artificial Intelligence (AI) and Final Physician Recommendations in AI-Assisted Virtual Urgent Care Visits

How do AI-generated initial recommendations in virtual urgent care compare with the final recommendations physicians issue? This Annals of Internal Medicine study examines concordance between an AI system's initial output and clinicians' final decisions across AI-assisted virtual urgent care visits. No abstract was available, so findings, sample size, and effect sizes cannot be summarized here.

080
Transparency in healthcare AI: Testing EU regulatory provisions against users' transparency needs

Do the transparency needs of healthcare AI users actually map onto the Instructions for Use (IFU) document that the EU AI Act (Directive 2024/1689) requires providers to give deployers? This cross-sectional online survey, administered via Qualtrics to four deployer groups \u2122 managers (N = 238), healthcare professionals (N = 115), patients (N = 229), and IT workers (N = 230) \u2122 asked participants to rate the relevance of a set of transparency needs and identify which IFU section would address each. Priorities differed across user types, and participants had difficulty locating some transparency information within the IFU structure; the abstract reports no effect sizes or magnitudes. The authors derive recommendations for locally meaningful IFUs.

081
Understanding end-user contexts and identifying design preferences of an artificial intelligence-based clinical decision support tool for early autism detection

What would clinicians and caregivers want from an AI-based clinical decision support tool for early autism detection, and where would it fit in the visit? This observational qualitative study used contextual inquiry with 8 clinicians and 20 caregivers during 18- to 24-month well-child visits at Duke-affiliated clinics, analyzed with rapid qualitative analysis. Workflow mapping identified 6 user tasks, 3 technology-user interactions, and 5 clinical decision points, plus 2 barriers (screening tool accuracy, follow-up implementation) and 3 facilitators (electronic screening, early intervention provider input, referral coordination support). Preferences included EHR-embedded, actionable outputs with prediction explanations, visual summaries, and caregiver-facing materials. The abstract reports no effect sizes.

082
AI-based clinician decision support system for diagnosis of inherited retinal diseases: a multicenter, randomized trial

Can an image-based AI narrow the genotype search for inherited retinal diseases before genetic testing? Retina4IRD, a RETFound-pretrained Vision Transformer predicting 17 genotype categories, was trained on fundus photographs and OCT from 1,843 genetically confirmed patients (3,376 eyes) in China, South Korea and Poland; top-5 accuracy was 0.904 internally and 0.856 externally. In a multicenter randomized trial, 300 patients with suspected IRD were assigned 1:1 to AI-assisted or specialist-only assessment (295 analyzed). Top-5 genetic accuracy was 88.5% versus 67.3% (P<0.001), top-1 37.8% versus 22.4%, and a composite downstream management score 37.7 versus 28.5 (P<0.001).

083
Development of a clinical trial knowledge management application for community oncology

Can institution-specific cancer trial information be curated into an AI-enabled knowledge management application in community oncology? This feasibility study at a regional community oncology network had coordinators and disease teams compile actively recruiting trials, structuring core elements (title, conditions, biomarkers, stage/line, recruiting status) for point-of-care display, with AI-assisted extraction of protocol summaries and eligibility elements followed by human validation. Fifty-three trials across 10 disease groups and 28 cancer types were embedded; 91% were recruiting and 30% were biomarker-specific. Configuration required 2-4 weeks per disease group using existing personnel, without added staffing or EHR build. Usability and implementation outcomes were not assessed.