How can a machine learning algorithm be embedded in the electronic health record to run a pragmatic randomized trial? This implementation report describes the Precision Resuscitation with Crystalloids in Sepsis (PRECISE) trial, a multihospital RCT built entirely with standard Epic tools through four components: automated inclusion criteria, real-time sepsis subphenotyping, randomization, and a medication alternative alert prompting clinicians toward the fluid type thought to benefit the identified subgroup. PRECISE launched across 6 Emory Healthcare hospitals in June 2024, covering 6 emergency departments and 17 ICUs with more than 300 ICU beds. The abstract reports implementation details only, with no trial outcomes or effect sizes.
AI evaluation & deployment
Does an LLM-generated hospital course draft reduce discharge summary documentation time? This observational pre-post study compared a pre-AI period (March 16-November 19, 2025) with a post-AI period (November 20, 2025-February 11, 2026) after a GPT-4.1 tool was embedded in mandatory discharge summary templates, covering 8,298 hospitalized adults on hospital medicine services. Edit time did not differ between periods (6.60 vs 6.28 minutes, p=0.11), but within the post-AI period, tool use was associated with longer editing (9.20 vs 4.93 minutes) and a 32% increase in adjusted analysis (95% CI 24-41%). Faculty review of 27 encounters found higher quality and lower harm but less concise summaries; 31 of 36 surveyed clinicians (86.1%) felt the tool improved efficiency.
How are ethical principles for AI safety actually operationalised once clinical AI reaches practice? This qualitative study (QuAS-AI) used semi-structured interviews with 16 experts involved in clinical AI implementation "\u2014 academics, industry professionals, and practitioners \u2014 recruited by purposive and snowball sampling, with independent coding, peer debriefing, and member checking. Three themes emerged: performance, risk, and bias as dynamic properties needing continuous monitoring; transparency and explainability as complementary supports for clinical interpretation rather than technical disclosure; and accountability as multi-actor governance spanning traceability, logging, data governance, and intervention capacity. The authors frame safety as socio-technical governance and suggest a High-Reliability Organisation approach. The abstract reports no effect sizes.
How have large language models been used and evaluated for translating pathology reports into patient-facing language? This scoping review followed JBI methodology and PRISMA-ScR, searching six databases for empirical studies of LLM-generated interpretation, rewriting, question answering, or summarization of pathology reports (January 2018 to August 2026). Nineteen studies were included; GPT-family models appeared in 17, and report-level transformation was the most common task (12/19, 63.2%). Fidelity was assessed in all 19 studies, safety in 10 (52.6%), readability in 9 (47.4%), and comprehension and usability in 6 each (31.6%). Only five involved patients or other non-clinicians, and none evaluated performance by health-literacy level, across languages, or prospectively within clinical workflows.
What do patients, caregivers, and clinicians actually expect from AI in diabetes care? This systematic review of qualitative studies searched five databases (MEDLINE, Web of Science, Scopus, CINAHL, PsycINFO) through February 2026, including 14 studies published 2023-2025 with at least 738 participants across 9 countries, covering large language models, AI-enabled apps and wearables, glucose prediction, and decision support. Thematic synthesis produced 4 analytical themes and 13 subthemes; 11 findings were rated high confidence and 2 moderate under GRADE-CERQual. Stakeholders saw value for prevention, education, and self-management but raised accuracy, bias, privacy, accountability, workload, and autonomy concerns. Much evidence rested on prototypes or hypothetical systems.
Why does strong AI task performance so rarely translate into sustained clinical benefit? This scoping review followed Joanna Briggs Institute methodology with PRISMA-ScR and PRISMA-S reporting, searching five databases (MEDLINE, Scopus, Web of Science, IEEE Xplore, CINAHL Plus) for English-language records from January 2017 to April 2026. Of 43,394 records identified, 275 were charted across four nonmutually exclusive domains: clinical decision support (233, 84.7%), predictive analytics (192, 69.8%), diagnostics (142, 51.6%), and economic assessment (53, 19.3%). Workflow or process measures appeared in 13 records (4.7%), equity or subgroup analyses in 20 (7.3%), and formal economic evaluations in 4 (1.5%). The authors group recurring constraints into five cross-domain barriers, describing them as structural rather than technical.
Can large language models reliably generate or simplify outpatient clinic letters? This systematic review searched five databases (PubMed, EMBASE, Web of Science, CENTRAL, CINAHL) from inception to 1 November 2025, appraising studies with the Mixed Methods Appraisal Tool and GRADE. Seven studies were included, two using real-world clinic data and five synthetic or hypothetical scenarios. Readability findings were mixed and most AI output still exceeded the US Grade 6 level; information fidelity ranged from 10% to 100%, and one study reported a ten-fold reduction in drafting time. Certainty was very low across primary outcomes.
Who actually uses a hospital-wide LLM chatbot, and what blocks the rest? This mixed-methods case study of "InternalGPT," a HIPAA-compliant chatbot at an academic pediatric medical center, combined employee surveys with 14 months of utilization data. Of roughly 15,800 employees, 2,149 (13.6%) requested access; 52.8% of those used at least one token, 33.6% never logged in, and the top 20% of users consumed 69.4% of tokens. Barriers cited by non-users were limited time (51.4%) and difficulty using the tool (21.8%). Among 92 sustained users surveyed, mean self-reported productivity gain was 30%, an exploratory $6.3M-$18.9M value.
Does conversational LLM triage capture different information or shift care-seeking intent compared with a structured questionnaire? This retrospective observational study compared 116,890 virtual triage encounters over 28 weeks (January–August 2025), where users self-selected traditional triage (TT; 100,533, 86%) or an LLM-enabled conversational interface (CT; 16,357, 14%) sharing the same Bayesian reasoning engine, with poststratification weighting by age and sex. CT sessions ran longer (median 8 min 21 s vs 4 min 25 s), elicited more clinical findings (median 36 vs 32; P<.001), and had higher self-reported intended adherence to recommended care (34.3% vs 29.2%; P<.001), including self-care (85.4% vs 61.9%). Users self-selected groups.
How do the people who would implement AI in home care view its promise and risks? This qualitative study conducted semi-structured interviews, incorporating AI scenarios, with 43 participants across five stakeholder groups — home health aides and attendants, home care agency leaders and staff, worker advocates, clinicians, and technology company personnel — recruited through purposive and snowball sampling, with analysis by structural coding, inductive sub-coding, and thematic analysis. Mean age was 44.6 years; 44.4% reported no or low AI knowledge. Four themes emerged: benefits to patient care, engagement, and efficiency; risks to care quality, provider-patient relationships, and working conditions; data quality, privacy, and AI literacy challenges; and needs for equitable governance. No effect sizes are reported.
How much artificial intelligence training is offered in US internal medicine residency programs? This national survey reports on the state of AI education across internal medicine training programs. No abstract was available, so findings, sample size, and effect estimates cannot be summarized here.
Does habitual reliance on generative AI blunt users' ability to calibrate trust in AI-generated health information? Two randomized 2×2 between-participants experiments (338 college students; 563 Mechanical Turk workers) manipulated information accuracy and text-based visual cues (highlighting), measuring trust and self-reported learned dependency with regression models. Accuracy raised trust (experiment 1 B=2.107, 95% CI 1.337-2.878; experiment 2 B=0.203, 95% CI 0.115-0.290), as did learned dependency (B=0.277 and B=0.822). The accuracy-by-dependency interaction was negative in both (B=-0.399; B=-0.459), indicating reduced sensitivity to inaccuracy. Text highlighting had no significant effect and did not moderate dependency.
How much do expert clinicians agree with each other when judging diuretic titrations, and what baseline should a semi-autonomous clinical decision support system (OTTO-FM) be held to? Secondary analysis of prospectively collected porcine data modeling postoperative fluid overload had three pediatric cardiac intensivists rate 29 clinician-driven and 44 CDS-driven items as reasonable/unsure/unreasonable, with ordinal-weighted Gwet's AC2. Dosing agreement was similar across phases (human 0.79, CDS 0.83), but unanimity reached only 65.5% and 70.5% of items. Risk labels split sharply: high-risk 0.92 versus low-risk 0.54; 19 of 20 nonunanimous risk items were single-rater dissent.
Does continuous remote patient monitoring reduce 30-day readmissions after heart failure discharge? This case-versus-retrospective-propensity-matched-control study at Endeavor Health (Evanston, IL) enrolled 39 patients across three phases, monitoring them for 30 days postdischarge with wearable biosensors and daily symptom surveys, with rules-based and machine learning alerts triaged by home health nurses and escalated to advanced practice providers. Intervention patients received more APP calls (66.7 vs. 7.7%), APP visits (43.6 vs. 5.1%), diuretic escalation (43.6 vs. 12.8%), and labs (76.9 vs. 43.6%). Adjusted 30-day readmission did not differ (odds 0.31, 95% CI 0.06–1.38; p = 0.138).
Does word error rate track clinically consequential errors in ambient AI scribes? Investigators built a synthetic multilingual corpus — five clinical dictation scripts across a complexity gradient, translated into 99 languages, rendered to speech under three acoustic conditions, and transcribed by a production scribe — then had three independent large language model raters score errors on a Severity x Likelihood framework. Of 59,819 genuine transcription errors, 58,329 (97.5%) were LOW risk and 251 (0.42%) CRITICAL or HIGH. No frequency metric was detectably associated with serious risk (absolute Spearman rho<0.16), while a Severity x Likelihood sum tracked WER (rho=0.80). Low-resource languages had worse WER (beta=+0.078) without higher critical risk (OR 1.21); consultation complexity predicted serious risk (OR 3.06 per level).
Can a generative AI model produce care-transition synopses as good as clinician-written ones? In a blinded, randomized comparison using de-identified records of 64 patients with multiple chronic conditions from MIMIC-III, human- and AI-generated synopses were scored on accuracy, succinctness, synthesis, and usefulness within a data-information-knowledge-wisdom framework (>80% indicating success). AI and clinician summaries overlapped 12%. AI synopses were rated useful 75% of the time versus 76% for human synopses; AI scored lower on succinctness for the data task (55%-67%) and near-equal or better on accuracy and synthesis (AI 72%-79%, humans 68%-84%), best in wisdom. Interrater agreement was variable.
How is human-in-the-loop (HITL) actually implemented across the lifecycle of AI-enabled clinical decision support? This systematic scoping review searched MEDLINE, Embase, Web of Science, PsycINFO, Google Scholar and Scopus in August 2024, with manual identification through mid-2025, including primary studies in which clinicians interacted with AI-CDSS; dual independent screening and extraction mapped findings to four lifecycle phases. Twelve studies qualified. All described clinician involvement during development, mainly expert annotation and rule-based design; 11 reported review-phase HITL, 2 maintenance, and none oversight. Contributions were largely static or retrospective, with small datasets, few annotators, and inconsistent terminology. The authors propose a six-domain HITL reporting checklist.
How far has predictive analytics on routine and population health data actually moved into health-system decisions? This global scoping review searched five databases for 2014–2025 studies applying predictive or forecasting methods for health-system decision-making, screening 2,623 records and including 161 articles (128, 79.5%, from high-income settings), with a supplementary grey-literature scan. Most work stopped at development (139 articles, 86.3%); validation 6.2%, pilot 1.9%, operational deployment 5.6%. While 73.3% claimed relevance to resource allocation or capacity planning, only 5.6% documented an output-to-decision pathway and 8.7% routine workflow integration; 14.9% described how uncertainty informed decisions.
How do clinicians view large language models and who should govern them? A cross-sectional online survey recruited 335 health care professionals through a health care news mailing list, 68.7% (n=230) attending physicians and 77.9% practicing in the Northeast United States. Some 62.7% (n=210) reported current or contemplated LLM use, with users reporting higher self-rated knowledge than nonusers (P<.001) and no age association (\u03c1=-0.072; P=.19). Top applications were literature review (73.4%), decision support (57%), and patient communication (54.9%); 96.4% voiced bias concern, 65.4% preferred oversight by professional associations over technology companies (29%), and 66.6% reported no confidence in existing oversight. Convenience sample, low response rate.
What do mental health clinicians think about AI in their practice? This scoping review followed JBI guidance and PRISMA-ScR, searching six databases (CINAHL, Embase, PsycINFO, PubMed, Scopus, Web of Science) for studies published from 2020 onward; 12,356 records were retrieved and 35 included after dual-reviewer screening. Clinicians showed cautious optimism when AI was framed as supplementing rather than replacing expertise, citing reduced administrative burden, documentation support, information synthesis, and between-session access. Concerns centered on privacy, governance, data ownership, unsafe or inaccurate outputs, overreliance, unclear accountability, and effects on therapeutic relationships, alongside limited AI literacy. The abstract reports no effect sizes.
Are LLM-generated medical reports clinically ready in terms of effectiveness, safety, and workflow burden? This systematic review searched five databases (PubMed/MEDLINE, Embase, Web of Science, Scopus, Cochrane) for studies published January 2016 through May 2026 evaluating LLMs, multimodal LLMs, or vision-language models for image-to-report generation, impression drafting, or structured reporting, including 101 studies (36 chest x-ray). No study was at low risk of bias (15 moderate, 72 high, 14 serious), and meta-analysis was not possible. One chest x-ray study found AI report acceptance of 70.5% (6047/8580) versus 73.3% for radiologists, with false negatives 18.5% versus 17.8%; a brain MRI study showed reading time falling from 61 to 53 seconds while impression drafting increased editing time and edit distance.
Can AI improve patient comprehension and decision-making during informed consent? This PRISMA-guided systematic review searched PubMed, Embase, and the Cochrane Library, including 33 studies published 2020-2025 across three domains: AI-generated patient education (n=18, 54.5%), consent documentation (n=10, 30.3%), and AI-assisted consent acquisition (n=5, 15.2%). Large language models were accurate but readability stayed above an eighth-grade level (best model Copilot: Flesch-Kincaid 10.59, SD 1.22). AI-generated documents raised Flesch Reading Ease by 44%-122% and lowered required grade levels 10%-47%. In trials, AI-assisted consent shortened consultations (7.7 vs 10.6 minutes; P=.05) and lowered post-consent anxiety in knee arthroplasty (10.48 vs 12.75; P=.04).
How should health systems judge ambient AI performance when vendor claims and real-world results diverge? This conceptual article, drawing on academic and industry perspectives, proposes a shared mental model organizing ambient AI assessment into three interdependent dimensions: technical, interface, and system level. For each dimension, the authors specify the types of information relevant to evaluation, what vendors should reasonably be expected to disclose, and how provider organizations can run internal evaluations to contextualize, verify, or supplement vendor claims. No empirical study, effect sizes, or performance benchmarks are reported; the contribution is a decision-support framework intended for organizations of varying size.
How are health systems moving AI beyond pilots to scale? This qualitative study conducted four case studies — Catalonia, Norway, Singapore, and Queensland — drawing on 60 documents, 34 interviews, and 5 focus groups with 50 strategic decision-makers, policymakers, and lead clinicians, analyzed first within-case thematically and then across cases using the Technology, People, Organization, and Macroenvironment framework. Scaling trajectories were shaped by existing digital strategies, digitalization histories, funding arrangements, and legacy infrastructure, with experimentation opportunities, incentives, and distribution of decision-making authority mattering alongside post-deployment monitoring, governance, and procurement. The authors frame scaling as multilevel orchestration rather than top-down versus bottom-up. No effect sizes are reported.
How well did the 2020 EIT Health & McKinsey AI forecast track what actually happened? This structured narrative review compared the report's domain-level predictions against 2020–2025 evidence from peer-reviewed implementation studies, FDA regulatory data, national and international guidance, industry surveys, and foundation-model evaluations, classifying each prediction as realised, under-realised, exceeded, or unanticipated. Two predictions held: administrative automation and medical imaging led adoption, with ambient documentation moving from pilots to deployment and more than 1300 FDA-authorised AI/ML-enabled devices, mostly radiology. Remote monitoring and classical NLP decision support under-delivered, limited by interoperability, reimbursement, and workflow barriers; generative and multimodal foundation models were the largest unanticipated divergence.
How should medical AI be evaluated across its lifecycle, from bench performance to clinical benefit? This conceptual paper proposes a five-phase evaluation framework spanning technical validation, operational robustness, controlled interaction, clinical evidence, and real-world integration, embedded in a dynamic architecture with phase-gating criteria and local and systemic fall-back triggers that permit re-entry into earlier phases after drift, version updates, or safety signals. It maps multicenter external validation, shadow-mode testing, human-AI comparison and cooperation studies, randomized trials, real-world evaluations, and adaptive designs onto this pathway. No empirical data or effect sizes are reported; the contribution is a framework aimed at researchers, institutions, and regulators.
Can an EHR-embedded machine learning phenotype identify emergency department patients with opioid use disorder in real time for trial screening? Across three EDs in one US health system (2014–2025), a random-forest classifier using visit-level data available at or before triage was trained against a high-specificity computable label and deployed to trigger point-of-care alerts. Retrospective discrimination was high (AUROC 0.99; AUPRC 0.92). In prospective chart-review validation (n=217), 89.3% of model-positive and 95.7% of model-negative encounters matched physician adjudication; design-weighted estimates using the 3.26% flag rate (28,284/866,569 encounters) gave sensitivity 0.40 and specificity 0.996.
Do physicians rate and act on AI-generated e-consultation advice differently when they know its source? In a randomized clinical vignette study, 44 internal medicine teaching faculty at one academic medical center reviewed four vignettes with advice generated by a medically specialized generative AI (OpenEvidence), labeled either as from a specialist attending physician or an AI system. Mean quality ratings were high across all six domains (overall means 4.2-4.7 of 5), with no differences by label. Selection of the consult-recommended management rose from 37.5% to 81.8% after viewing advice (OR 8.43, 95% CI 5.00-14.28); labeling did not affect adoption (OR 1.18, p=0.70).
Can an autonomous coding agent replace hand-written analyst queries against a production EHR warehouse? In a single-center quality-improvement evaluation, ten acute otitis media questions were posed to analysts (adjudicated reference) and to OpenAI Codex, which wrote read-only queries against a full copy of an Epic Caboodle warehouse under four conditions. At medium reasoning effort the agent came within 5% of the reference on 27 of 30 runs but exactly matched on 10; three repeated runs agreed on only 3 of 10 questions. Matching runs identified the same patients (F1 = 1.00, one exception at 0.72); analyst feedback and reuse of corrected definitions restored reproducibility.
How reliable is an EHR-embedded generative AI summarizer in routine use? Investigators applied a multidimensional assurance framework to Epic IP Insights across a seven-hospital system, building an agentic hallucination detector that decomposed summaries into atomic content units (ACUs) and checked each against source notes. The tool produced 706 summaries over 445 encounters, drawing on a mean 5.8% of available notes. The detector matched physician adjudication on 96.95% of ACUs in validation; across 671 summaries (40,452 ACUs) the hallucination rate was 11.79% (95% CI, 11.47%-12.10%). Among 385 regenerated pairs, 32.7% were textually and 37.9% semantically identical. Clinicians rated it easy to use (97.8%) but were mixed on relying on output without verification (45.9%).
Can information about an AI system help people delegate tasks to it more often and more wisely? This experimental study manipulated two signals: ex-ante AI certainty (the AI's estimated likelihood of being correct, shown before the delegation decision) and ex-post outcome information (whether the AI was actually correct, shown after). Presenting either signal alone had no effect or reduced combined human-AI performance; providing both raised delegation frequency and delegation effectiveness, improving performance. The authors attribute this to certainty calibrating task-level expectations, outcomes confirming them, more accurate mental models, and less algorithm aversion. The abstract reports no effect sizes, sample size, or task details.
Does AI support for total parenteral nutrition in neonatal intensive care improve efficacy, safety, and integration? This PRISMA 2020 systematic review searched PubMed/MEDLINE, Scopus, and Web of Science, appraising evidence with RoB 2, the Newcastle-Ottawa Scale, and GRADE. Sixteen records met criteria; thirteen quantitative primary studies (2008-2026, n = 30-9,330) formed the synthesis. CPOE implementation cut parenteral nutrition medication error rates from 10.8% to 3.2%; the TPN2.0 transformer model correlated with expert decisions at Pearson R = 0.94, and classical machine learning reached R2 > 0.70 for macronutrient prediction. Only one-third of U.S. NICUs used a CDSS. GRADE certainty was moderate for efficacy and safety, low for system integration.
How mature is the evidence behind "data-centric" multimodal deep learning for clinical decision support? This PRISMA 2020 systematic review screened 150 records and included 31 primary clinical studies, 30 (97%) published 2024-2026, coding implemented versus merely mentioned techniques and appraising bias with PROBAST+AI. Studies used a median of three modalities (range 2-6), most often structured EHR (71%) and imaging (39%). Data-centric techniques were reported in 74-84% of studies (equity 61%), but external validation appeared in only 4/31 (13%), a clinical or provider outcome in 3/31 (10%), none reported deployment, and 27/31 (87%) were at high risk of bias.
Does high usability in a clinician-built, AI-assisted application establish clinical and technical assurance? This single-case retrospective development-and-assurance report examined STUIapp, a browser-based ePROM tool integrating six validated lower urinary tract symptom instruments, using code audit, independent clinical review of a frozen 78-case scoring matrix, WCAG 2.1 measurement, a 23-canary persistent-storage study, and usability testing with 14 clinicians, 26 patients and 12 older adults. Mean SUS was 92.3 (SD 8.8) among clinicians and 92.0 (SD 10.8) among patients, yet all 78 passing automated cases included 17 expected results requiring correction. Other findings: instrument mislabelling, a failed installability manifest, a third-party analytics tag contradicting local-only privacy claims, 10-px text and 2.56:1 contrast, and storage permission denied in 10/10 browser-tab canaries versus granted in 13/13 installed canaries.
What predicts publicly visible AI adoption at US cancer centers? This cross-sectional study assembled public-source data on 75 NCI-designated cancer centers, scoring adoption across screening, treatment, and patient care as a 0-3 composite index, with Moran I tests for spatial clustering and ordered logistic regression on institutional and contextual predictors. The mean adoption index was 1.37 (SD 0.86), highest for screening (0.86), then patient care (0.50) and treatment (0.22). Moran I showed no significant spatial autocorrelation. Physician workforce and bed capacity showed positive but modest associations; state socioeconomic indicators did not. Political-context findings were mixed; the abstract reports no effect sizes for regression estimates.
How are graduate medical education trainees and faculty using generative AI? This cross-sectional survey study, reported in the Journal of General Internal Medicine, examines self-reported generative AI use among GME trainees and faculty. No abstract was available, so findings, sample size, and effect estimates cannot be summarized here.
What features of AI-generated literature reviews make them useful and trustworthy to oncologists? In a randomized mixed-methods study, 34 oncology physicians produced 294 ratings of four blinded AI systems across five clinical vignettes, supplemented by 20 semi-structured interviews analyzed with a prespecified LLM-assisted qualitative pipeline. Despite similar references, an evidence-graded report adapted from OpenEvidence scored significantly lower in overall utility than standard OpenEvidence (mean difference -0.96; 95% CI -1.26 to -0.66; P<.001). Qualitative analysis yielded six themes and seven design requirements: clinicians preferred concise, scannable reports with quantitative outcomes, bolded guidelines, explicit uncertainty, and verifiable citations; trust fell with citation mismatch, buried provenance, and overconfident recommendations.
Does an EHR-embedded generative AI chart summarization tool reduce clinician workload? In a pragmatic 1:1 randomized trial at one academic health system, 284 ambulatory clinicians across 42 specialties received Epic's outpatient chart summarization tool or usual care over 90 days (February o task task load favored the intervention ( tionsted difference -27.4 on a 0-400 scale; 95% CI, -49.4 to -5.3; P=0.02), with lower burnout (-0.20) and work exhaustion (-0.24) on the Professional Fulfillment Index. Charting time per encounter was unchanged (-1.2 seconds). Only 14.2% of 74,474 summaries were opened, falling from 21.5% to 10.5% by month 3; net promoter score was -22.
How far has machine learning–based pharmacogenetic clinical decision support actually moved into clinical workflows? This scoping review searched multiple databases for studies published January 2015 through September 2025 reporting ML-based CDSS incorporating pharmacogenetic data to support therapeutic decisions in clinical settings. Of 1,262 records screened, 7 met inclusion criteria, and only 2 evaluated tools in live clinical or trial workflows. Studies varied in design, setting, and implementation maturity; most reported potential benefits including fewer preventable adverse drug events, better prescribing accuracy, or improved workflow integration, though the abstract reports no effect sizes. Authors call for implementation science evaluating usability and patient outcomes.
Do ambient AI scribes produce notes of comparable quality and safety to clinician-written ones across languages? This paired simulation covered five countries and languages (Cambridge, Barcelona, Milan, Paris, Cologne), with 385 actor-performed consultations yielding 770 notes, each documented independently by an AI scribe (Heidi) and a junior-to-middle-grade clinician, scored on the PDQI-9 by blinded evaluators. AI notes scored higher (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; dz=0.55) and less often carried at least one Critical+High error: 24.4% vs 61.0% by calibrated automated review (RR 2.50) and 6.2% vs 21.8% by clinician adjudication (RR 3.50), with omissions driving the gap.
How can hospitals operationalize post-market oversight of imaging AI when pre-market clearances test algorithms under static conditions? This conceptual paper proposes the Regulatory Readiness Instrument for Medical Imaging AI (RRI-MI), a deployer-side readiness assessment mapping 11 governance domains onto a provisional 22-point ordinal rubric, aligned with the FDA Predetermined Change Control Plan and the EU AI Act. The authors ground the framework in post-market evidence on scanner drift, protocol shifts, software updates, and demographic variation, illustrating it through a semiautonomous prostate cancer MRI case study covering local validation, human oversight, version control, monitoring, and incident response. No empirical validation or effect sizes are reported.
Does deploying an AI radiographic fracture detection tool shorten emergency department stays? This retrospective controlled before-after cohort study with interrupted time-series analysis covered 8,253 fracture-excluded and 6,864 fracture-confirmed visits at a tertiary ambulatory orthopedic ED, January 2021–February 2026, with deployment on 1 October 2023. Mean length of stay in fracture-excluded patients fell from 135.6 to 127.5 minutes (−8.13; 95% CI −11.52 to −4.70), while the fracture-confirmed control was unchanged (+0.37 minutes; p=0.87). Reductions concentrated at the upper tail (−24.0 minutes at the 90th percentile); stays over four hours fell from 11.3% to 8.2%, with no increase in early revisits.
What empirical evidence exists for privacy, security, and reliability risks from AI in clinical care? This systematic review searched PubMed, Embase, Web of Science, Scopus, IEEE Xplore, and ACM Digital Library for empirical studies published January 2015 to November 2025 evaluating AI use or misuse in diagnosis, treatment, or decision-making. Of 7,285 records plus 205 from citation screening, 22 studies met inclusion criteria, mostly medical imaging. Five recurring threat categories emerged: re-identification, membership inference, unauthorized access and adversarial exploitation, input manipulation, and misuse or overinterpretation of outputs. Models encoded latent biometric signals, limiting anonymization and synthetic data. Findings were synthesized narratively; the abstract reports no pooled effect sizes.
How should the NIH's 7 principles of ethical clinical research be adapted for trials of autonomous AI? Using a modified Delphi approach over 6 months, investigators convened 14 multidisciplinary panelists (AI, data science, ophthalmology, policy, law, bioethics, patient advocacy) across two survey rounds anchored to a vignette and a final virtual meeting, with participation of 12/14 (85.7%), 10/14 (71.4%), and 13/14 (92.9%). Round 2 produced 9 strong-agreement, 2 moderate-agreement, and 4 divisive statements. Recommendations covered transparency on training and validation data, pre-deployment bias and inequity assessment, performance across clinical settings, informed consent, comparison with standard of care, downstream access, and cost.
How have researchers actually operationalized human evaluation of large language model reliability in health care? This PRISMA-ScR scoping review searched PubMed, Web of Science, Cochrane Library, CINAHL, and Google Scholar for English-language original studies published January 2016 to July 2025, screening 4347 records and including 71 (26 clinical, 45 public health). Six reliability indicators recurred: accuracy, relevance, completeness, clarity, safety, and consistency. Clinical studies more often assessed guideline concordance and structural coherence; public health studies emphasized understandability, harm potential, and repeat-response consistency. Single-specialty clinicians predominated as evaluators, panels typically had five or fewer members, and 5-point Likert scales with researcher-defined rubrics were standard. Reported limitations included evaluator subjectivity and nonstandardized indicators.
How much clinical evidence supports FDA-cleared AI medical devices? This systematic analysis catalogued all 1,357 AI/ML-enabled devices cleared through December 5, 2025 using the FDA device database and the ACR Data Science Institute catalogue, with linked searches of ClinicalTrials.gov and PubMed for registered trials and publications. Only 34 devices (2.5%) were linked to registered prospective trials, 12 (0.9%) posted results, 12 (0.9%) had peer-reviewed publications, and 3 (0.2%) evaluated patient-centered outcomes such as mortality, morbidity, or readmissions. Most studies (62%) were observational with small, homogeneous cohorts and frequent exclusion of vulnerable populations. The authors cite misaligned incentives and predicate-based pathways as barriers.
How far have machine learning applications in palliative care moved beyond mortality prediction? This scoping review followed Arksey and O'Malley and PRISMA-ScR guidance, searching six databases from inception through February 9, 2026, and included 121 peer-reviewed primary studies (2015-2026) across 24 countries, 54.5% (66/121) from the United States. Mortality and survival prediction accounted for 42.1% (51/121) of studies, followed by health care use (20.7%) and symptom assessment (16.5%); cancer populations dominated (43%). Half (50.4%) used an explainability technique, 86.8% (105/121) did not address equity in model performance, and 66.1% remained proof-of-concept with 17.4% reaching prospective deployment or clinical integration.
Can a centralized digital program identify and track abdominal aortic aneurysms across an integrated health system? This retrospective cohort study describes STAIR, which combined structured EHR problem-list queries, an internally developed NLP model applied to radiology reports, clinician referrals, and automated lost-to-follow-up queries, enrolling 8464 patients (mean age 77.1 years; 77.0% male) from December 2022 through December 2024, with status assessed through April 2026. Identification was mostly automated: problem-list queries 59.0% and radiology NLP 29.0%. After centralized review, 45.3% were assigned biennial duplex surveillance and 20.6% referred to vascular surgery; 49.5% remained under active surveillance. Findings are descriptive; clinical effectiveness was not assessed.
Can a multi-LLM clinical decision support system change emergency department care in practice? This DECIDE-AI stage 1 prospective evaluation deployed SHAKED in a tertiary ED over 4 weeks, analyzing 1,138 patients across two parallel units, one using the system and one following routine rotations. Adoption fell from 68% to 30%, with disengagement tied to workload (OR 0.72 per shift hour, 95% CI 0.62–0.83), while physicians favored it for radiology consultations (OR 2.98). Expert reviewers rated 99 of 100 sampled outputs clinically appropriate, with no adverse events. Length of stay was 4.9 hours in both wings (P = 0.99); consultation cycle time trended shorter (−9.4 min, P = 0.077).
Can eligibility for surgical comanagement (SCM) — hospitalist co-management of medically complex perioperative patients — be triaged automatically? This prospective, unblinded quality improvement study at Stanford Health Care (September 2025–February 2026) deployed an EHR-integrated, human-in-the-loop LLM tool (SCM Navigator) that classified patients using preoperative documentation, structured data, and morbidity criteria, with attending review as the reference standard. Across 6193 triaged cases (median age 60.2 years; 49.0% female), 1582 (25.5%) were recommended for hospitalist consultation; sensitivity was 0.94 (95% CI, 0.91-0.96) and specificity 0.74 (95% CI, 0.71-0.77). LLM misclassification explained 2 of 19 false negatives (11%).
Has radiology's dominance of FDA-authorized AI/ML medical devices persisted or begun to loosen? This longitudinal content analysis covered all 1430 devices in the FDA AI-Enabled Medical Devices registry with final authorization decisions through December 2025, stratified by advisory-committee specialty across four eras (1995-2015, 2016-2019, 2020-2022, 2023-2025). Annual authorizations rose from a mean of 2.0 to 264, with 331 in 2025; 96.2% used the 510(k) pathway. Radiology's share peaked at 85.5% (347/406) in 2020-2022, then fell to 77.5% (614/792) in 2023-2025 (P=.001), with the Herfindahl-Hirschman Index declining from 0.738 to 0.612. Start-ups (OR 5.09) and technology companies (OR 50.62, based on 13 devices) had higher odds of nonradiology authorization.
When does explaining an AI model's reasoning actually improve human-AI decisions, and at what cognitive cost? The authors build an analytical model of a decision maker with limited but flexible cognition receiving imperfect machine recommendations, where explanations shift beliefs about algorithmic quality. Low explainability leaves decision accuracy and reliance unchanged while reducing cognitive burden; higher explainability improves accuracy by curbing overreliance but increases underreliance. Explainability matters more for cognitively constrained decision makers, complex tasks, and lower-stakes decisions, yet can raise processing time and fatigue exactly when time is short, tasks are complex, and machine quality is doubted. Theoretical modeling, so no empirical effect sizes.
When AI benchmark scores steer capital, what happens to their value as signals? This NBER working paper is a conceptual and theoretical market-design analysis rather than an empirical study, treating public AI benchmarks as market institutions. The author identifies two gaps that become exploitable once scores move investment: public examples can reveal the process behind a private final test, and any finite public score cannot span the broad task space implied by general intelligence. Targeted effort aimed at these gaps erodes the signal later investors rely on. The proposed remedy is separating development from certification — publishing practice tasks but selecting the investment-consequential task generator only after a submitted system's evaluation policy is fixed. The abstract reports no effect sizes or empirical magnitudes.
Can automated judges that flag documentation errors serve as a defensible measurement instrument for ambient AI scribes, absent a gold standard? This pre-registered, blinded validation study nested in a multi-country ambient documentation simulation (English setting) had ten external clinicians adjudicate a stratified sample of 434 pipeline flags, yielding 565 adjudications with 131 double-rated. Inter-clinician agreement on error genuineness was fair (raw 59%, AC1 0.24); judge-clinician agreement was 64% (95% CI 60-68). Behavior was near-symmetric across arms (kept-precision 74% AI vs 81% clinician notes; severity gap +0.06 vs -0.09 tiers), with removed-confirmed asymmetric (56% vs 42%). Latent-class triangulation estimated a 68% genuine-error rate (94% CrI 48-83).
Do ambient AI scribes perform equally well across patient languages? This retrospective analysis covered 54,160 outpatient encounters in a U.S. safety net health system, measuring the share of words in the final note generated by the AI tool and left unedited by the provider. Generalized estimating equations with exchangeable correlation structures accounted for clustering within patients, with univariable models by language and interpreter modality and a multivariable interaction model. Non-English encounters were 21% to 25% less likely than English encounters to reach the performance threshold; interpreter-mediated and bilingual-provider encounters did not differ significantly.
How should health systems evaluate generative AI that turns clinical information into patient-facing instructions such as discharge summaries, medication explanations and portal messages? This narrative and interpretative review maps empirical studies of AI-generated discharge communication alongside health-literacy, medication-safety, patient-safety and AI-governance literature. The author reports a recurring trade-off: large language models improve readability and understandability, while physician and pharmacist review identifies omissions, inaccuracies, newly introduced actions, medication-related problems and potentially harmful content, especially in complex discharges. A proposed seven-domain framework covers factual accuracy, clinical completeness, actionability, medication clarity, escalation and safety-netting, health-literacy alignment, and accountability with auditability. The abstract reports no effect sizes.
How is AI being applied to cancer symptom management for adult survivors? This scoping review followed Joanna Briggs Institute methodology, searching eight databases (November 2025, updated March 2026) for English-language empirical studies from 2015 onward, with narrative synthesis and inductive content analysis. Of 41 included studies, 21 addressed model development, 18 intervention delivery, and 2 both; common techniques were natural language processing, machine learning, and conversational AI. Applications targeted symptom detection (n=15), patient education (n=9), monitoring (n=5), and personalised management (n=5); pain and psychological distress each appeared in 9 studies. Model performance was moderate to high; the abstract reports no pooled effect sizes.
What actually helps or hinders AI-based clinical decision support once it reaches the bedside? This scoping review searched five databases (MEDLINE, Embase, CINAHL, APA PsycInfo, Cochrane) from inception to May 2022, screening 10,875 titles and abstracts and 494 full texts; 13 studies met inclusion criteria and 9 reported explicit implementation determinants, mapped to the Consolidated Framework for Implementation Research by two reviewers. Studies were mostly US-based, multicenter machine learning tools in critical care and emergency medicine. Reviewers identified 28 determinants: 16 barriers (limited interpretability, data quality, workflow misalignment, user capability and motivation) and 12 facilitators (end-user needs assessment, stakeholder engagement, peer endorsement, supporting evidence).
Do large language models produce stigmatizing language when reasoning over clinical data? Researchers applied a psychiatrist-validated NLP stigma-term detector to reasoning text from 107 LLMs across 35 real-world clinical tasks, covering 3,745 model-task pairs. Stigma rates ranged from 0% to 33.33%, and 84.06% of pairs contained at least one stigma term. Open-source models exceeded proprietary (1.97% vs. 1.60%; p<0.01) and reasoning models exceeded non-reasoning (2.35% vs. 1.70%; p<0.0001), while general versus medical models did not differ (2.00% vs. 1.80%; p=0.26). Stigma correlated negatively with task accuracy (r=-0.304) and positively with input-note stigma (r=0.569); 19.76% of pairs amplified input stigma. Prompt engineering cut stigma rates by up to 91.91% without reducing performance.
How should LLM failures be evaluated when they occur inside multi-step research and clinical workflows rather than isolated prompts? This method paper introduces TRACE, a practitioner-audit framework, with an empirical demonstration: 45 documentation-positive incidents logged by a single clinician-informatician across scholarly, clinical informatics, and clinical-adjacent workflows over seven weeks, coded with a consequence-based severity rubric and taxonomy crosswalk. Four categories tied as most frequent (verification failure, factual numerical error, tool-behavior misunderstanding, citation formatting; n=7 each), and one incident carried an estimated $2,500 impact. Category agreement was low among three human reviewers (Fleiss κ=0.155) but substantial among three vendor-blinded AI comparators (κ=0.632). The authors frame this as a pilot that does not estimate error rates.
Do patient-facing AI products differ in how they triage and refer simulated patients? Researchers tested nine products against 60 physician-developed standardized clinical cases across 540 multi-turn simulated patient encounters. Overall triage accuracy showed no statistically significant difference across product categories, but referral behavior diverged: branded health AI products over-triaged low-acuity cases far more often (28% vs 3% vs 2%) and more frequently recommended affiliated, fee-requiring clinical services. The authors argue evaluations of patient-facing medical AI should assess referral behavior alongside triage accuracy. The abstract reports no confidence intervals or other effect sizes.
How many U.S. adults turn to general-purpose large language models for their mental health? A cross-sectional survey of 1,871 U.S. adults fielded August–October 2024 used stratified sampling across age, sex, and race/ethnicity to approximate national demographics. Twenty-four percent reported using LLMs for mental health; users skewed young, male, and Black, and reported poorer mental health and difficulty accessing traditional treatment, citing that LLMs are free, convenient, and always available. Reported uses included emotional support, learning therapy skills, and supplementing existing therapy. Combining with Pew estimates of overall LLM use, the authors project 14–18 million U.S. adults.</summary>}</summary>
How are large language models currently being used to support teamwork and communication within healthcare teams? This scoping review followed PRISMA-ScR guidelines, searching PubMed, Web of Science, and ScienceDirect for 2014-2024 publications; 3,865 unique titles and abstracts were screened, 127 full texts reviewed, and 20 studies included. Designs were predominantly quantitative and simulation-based, with limited in situ evaluation. Use cases spanned decision support, communication, and administrative functions. Outcomes centered on accuracy and quality (15/20 studies), with fewer addressing safety (4/20), readability or empathy (5/20), workflow efficiency (3/20), and error modes (2/20). The authors note most studies evaluate model performance without accounting for human team dynamics.
How should clinical predictive AI be evaluated when conventional randomized controlled trials are static, slow, and mismatched to models that drift and get updated? This narrative review argues for adaptive, iterative, context-specific assessment and proposes a framework separating three activities: performance monitoring (calibration, discrimination, data drift, alert burden, fairness, workflow fidelity), clinical impact monitoring of sustained benefit, and causal evidence generation via pragmatic and adaptive platform trial designs. The authors add a governance-driven escalation protocol specifying when monitoring signals should trigger formal trials, a signal-to-design decision pathway, and a guide to causal inference methods. As a review, it reports no effect sizes.
How does an LLM-based conversational interface to the electronic health record fare in real clinical deployment? This Nature Medicine piece reports on implementation lessons from ChatEHR at Stanford Medicine, covering the practical, technical and organizational considerations of putting a chart-querying system into use in an academic health system. No abstract was available, so findings, evaluation methods and any performance or utilization figures cannot be summarized here.
Can transformer-based NLP reliably surface social determinants of health from routine notes? This validation study at University of Florida Health surveyed 1001 adults with at least two encounters in the prior year (sampling targeted 50% Black patients), restricting comparative analyses to 414 participants who also had Epic SDoH questionnaire data and notes. The SODA pipeline's extractions across nine domains were benchmarked against the research survey and the Epic instrument using sensitivity, specificity, PPV, NPV, and F1. Sensitivity reached 55% for alcohol use but fell to 16% for financial constraints, 5% for abuse, and 0% for drug use; the two surveys agreed only modestly, so no definitive reference standard existed.
How are retrieval-augmented and graph retrieval-augmented LLM systems in health care actually evaluated? This PRISMA-ScR scoping review searched PubMed, Web of Science, IEEE Xplore, ACM, arXiv, and medRxiv through May 2026, charting 157 eligible studies. Clinical question answering dominated (89/157, 56.7%), followed by decision support (70/157, 44.6%). Evaluations were overwhelmingly offline (140/157, 89.2%), with only 10.8% workflow-facing, prospective, or deployment-level. Independent retrieval-layer evaluation appeared in 29.9%, grounding or faithfulness in 26.1%, fine-grained evidence verification in 14%, and safety evaluation in 28.7%. Among 94 studies using human evaluation, 27.7% reported interrater reliability.
Do LLM-based chatbots improve physician-patient communication? This PRISMA-guided systematic review and random-effects meta-analysis searched PubMed/MEDLINE, Embase, Scopus, and Web of Science for 2020-2025 studies of LLM or chatbot applications in clinical communication, including 10 quantitative, mostly cross-sectional studies from 312 records. In 5 of 6 direct comparisons, LLM responses were rated more empathetic than physician responses; one study found chatbot replies judged empathetic in 45.1% of cases versus 4.6% (odds ratio ~9.8, P<.001), and GPT-4-simplified pathology reports raised comprehension scores (7.98 vs 5.23/10). Pooled empathy effect was large (SMD 1.02, 95% CI 0.44-1.60; k=4, N=2604). Satisfaction results were mixed; no study assessed long-term trust.
Does explainable AI help clinicians and lay users equally in dermatological diagnosis? Two large-scale experiments enrolled 623 lay people and 153 primary care physicians, pairing a fairness-constrained AI model for skin disease diagnosis with several explanation formats, including multimodal large language model explanations. Assistance from the fairness-trained model, which performed comparably across skin tones, improved final diagnostic accuracy and narrowed skin-tone performance gaps in both groups. LLM explanations diverged: lay users showed automation bias, gaining when the model was right and losing when it erred, while physicians benefited regardless. Presenting AI predictions before human judgment increased anchoring. The abstract reports no effect sizes.
Can algorithmic risk scores improve how child-protection supervisors allocate scrutiny? A randomized evaluation covering 4,752 child referrals over 14 months in Northampton County gave supervisors an algorithmic risk score alongside standard case records. Access to the score increased foster-care placements and service receipt among children at the highest predicted risk, with little change for lower-risk cases, and reduced subsequent maltreatment referrals. The authors report no evidence that the score widened racial disparities in decisions or outcomes. The abstract reports no point estimates or effect sizes for these changes.
How can machine learning outputs derived from HIV electronic health records be translated into a clinical decision support system clinicians will actually use? This human-in-the-loop field study at Prisma Health in South Carolina combined pre- and postsurveys, interactive usability testing, think-alouds, and in-depth interviews with 16 clinicians providing HIV care—physicians, nurse practitioners, infectious disease pharmacists, social workers, and case managers—between March and September 2025. Clinicians anchored interpretation of AI predictions on familiar clinical indicators but centered social determinants of health in their own risk assessments; trust was conditional and accrued over time, with explainability and actionability described as prerequisites for intervention. The abstract reports no effect sizes.
Do ambient AI scribes carry interpreter errors into the clinical note? Using simulated English- and Spanish-language clinical encounters mediated by interpreters, the authors evaluated whether documentation generated by ambient AI scribes reproduced interpretation errors introduced during the visit. Scribes propagated interpreter errors into the resulting notes, with propagation patterns differing by speaker role and by error type. The published abstract reports no effect sizes, error counts, or comparative rates. The authors frame the results as a case for further evaluation of AI-scribe performance in multilingual and interpreter-mediated care.
How should hospitals calibrate oversight of AI tools embedded in EHR workflows, imaging, triage, documentation, and operations? The authors conducted a narrative review and framework synthesis drawing on peer-reviewed evidence, reporting guidelines, regulatory and policy sources, implementation studies, and applied governance case reports. The resulting framework has four components: a use-case inventory tagged by decision influence and workflow coupling; a six-domain risk taxonomy spanning clinical safety, privacy and data security, ethics and fairness, transparency, system stability, and compliance; a four-tier risk scheme keyed to harm, automation, reversibility, and coupling; and a governance architecture assigning roles to a committee, clinical owners, risk-control functions, and independent assurance. A lifecycle pathway runs from initiation and local validation through shadow mode, controlled go-live, monitoring, change control, and retirement. No effect sizes are reported; this is a conceptual framework, not an evaluation.
For what purposes do physicians use unauthorized, non-conformity-assessed AI tools at work? This cross-sectional survey of physicians in Swedish health care organizations (N=357; response rate ~64%), fielded through a verified online panel between December 2023 and January 2024, applied qualitative content analysis to free-text responses, interpreted through the sociology of professions and paradox theory. Reported uses fell into four categories: clinical work and decision-making (second opinions, differential diagnoses, rare cases), administrative work (patient communication, documentation), research and professional development, and technological curiosity. Physicians framed such use as compensating for gaps in institutional systems and reducing workload. The abstract reports no effect sizes or usage prevalence.
How consistent are the evaluation frameworks proposed for clinical AI? This scoping review followed PRISMA-ScR, searching six databases plus the EQUATOR Network through February 2026, and screened 3363 records to include 46 frameworks, scored on methodological rigor, validation strategy, and a 10-domain UNESCO ethics matrix. Most frameworks (88%) targeted investigational rather than clinical use; 31.8% reported technical metrics such as AUC, sensitivity, and specificity, 15.9% reported clinical indicators, and only 11.4% met methodological rigor with validation aligned to intended use. Ethics coverage was uneven: transparency and explainability appeared in 70%, human oversight in 24.4%.
How should public health teams decide whether a customized large language model is ready to deploy? This conceptual paper proposes an acceptance criteria framework (ACF) defining implementation fit as meeting prespecified minimum performance standards and showing nonproblematic behavior under anticipated use. The ACF combines project-relevant and off-topic test prompts, structured expert review, and prespecified thresholds to generate a documented decision record that can be rerun after model revisions. The authors argue prior safety, ethics, effectiveness, engagement, and implementation frameworks imply rather than operationalize deployment benchmarks, and illustrate the ACF in a tobacco cessation text messaging intervention. No performance estimates are reported.
How can health systems test large language models on real patient portal messages without touching live EHR workflows? This technical feasibility tutorial describes a Python 3 web interface and modular backend running inside the institutional firewall on an NVIDIA GRID T4-1Q GPU, supporting single-message and batch tasks: authorship identification, categorization, criticality flagging, and response drafting with zero-, one-, and few-shot prompting. A deidentification pipeline validated against 110 manually adjudicated entities achieved 95.1% sensitivity and 82.1% precision. Use cases drew on an IRB-approved dementia-relevant corpus of 6941 medical advice request messages from 497 patients; token-based cost readouts were included. No comparative performance effect sizes are reported.
Does acceptable pre-deployment validation performance persist once clinical AI enters routine workflows? This longitudinal retrospective observational study followed four deployed AI systems spanning different clinical domains within one large healthcare organization, comparing validation-era metrics with post-deployment behavior over extended observation using routine clinical data, outcome labels, and operational telemetry, and examining discrimination, calibration, data availability, latency, and workflow signals. In all four systems, validation performance did not persist; calibration drift appeared consistently and often preceded discrimination changes, and label-independent signals such as input missingness and data latency flagged degradation earlier than outcome-based monitoring, which lagged behind label availability. The abstract reports no effect sizes.
How do AI-generated initial recommendations in virtual urgent care compare with the final recommendations physicians issue? This Annals of Internal Medicine study examines concordance between an AI system's initial output and clinicians' final decisions across AI-assisted virtual urgent care visits. No abstract was available, so findings, sample size, and effect sizes cannot be summarized here.
Do the transparency needs of healthcare AI users actually map onto the Instructions for Use (IFU) document that the EU AI Act (Directive 2024/1689) requires providers to give deployers? This cross-sectional online survey, administered via Qualtrics to four deployer groups \u2122 managers (N = 238), healthcare professionals (N = 115), patients (N = 229), and IT workers (N = 230) \u2122 asked participants to rate the relevance of a set of transparency needs and identify which IFU section would address each. Priorities differed across user types, and participants had difficulty locating some transparency information within the IFU structure; the abstract reports no effect sizes or magnitudes. The authors derive recommendations for locally meaningful IFUs.
What would clinicians and caregivers want from an AI-based clinical decision support tool for early autism detection, and where would it fit in the visit? This observational qualitative study used contextual inquiry with 8 clinicians and 20 caregivers during 18- to 24-month well-child visits at Duke-affiliated clinics, analyzed with rapid qualitative analysis. Workflow mapping identified 6 user tasks, 3 technology-user interactions, and 5 clinical decision points, plus 2 barriers (screening tool accuracy, follow-up implementation) and 3 facilitators (electronic screening, early intervention provider input, referral coordination support). Preferences included EHR-embedded, actionable outputs with prediction explanations, visual summaries, and caregiver-facing materials. The abstract reports no effect sizes.
Can an image-based AI narrow the genotype search for inherited retinal diseases before genetic testing? Retina4IRD, a RETFound-pretrained Vision Transformer predicting 17 genotype categories, was trained on fundus photographs and OCT from 1,843 genetically confirmed patients (3,376 eyes) in China, South Korea and Poland; top-5 accuracy was 0.904 internally and 0.856 externally. In a multicenter randomized trial, 300 patients with suspected IRD were assigned 1:1 to AI-assisted or specialist-only assessment (295 analyzed). Top-5 genetic accuracy was 88.5% versus 67.3% (P<0.001), top-1 37.8% versus 22.4%, and a composite downstream management score 37.7 versus 28.5 (P<0.001).
Can institution-specific cancer trial information be curated into an AI-enabled knowledge management application in community oncology? This feasibility study at a regional community oncology network had coordinators and disease teams compile actively recruiting trials, structuring core elements (title, conditions, biomarkers, stage/line, recruiting status) for point-of-care display, with AI-assisted extraction of protocol summaries and eligibility elements followed by human validation. Fifty-three trials across 10 disease groups and 28 cancer types were embedded; 91% were recruiting and 30% were biomarker-specific. Configuration required 2-4 weeks per disease group using existing personnel, without added staffing or EHR build. Usability and implementation outcomes were not assessed.