Do ambient AI scribes carry interpreter errors into the clinical note? Using simulated English- and Spanish-language clinical encounters mediated by interpreters, the authors evaluated whether documentation generated by ambient AI scribes reproduced interpretation errors introduced during the visit. Scribes propagated interpreter errors into the resulting notes, with propagation patterns differing by speaker role and by error type. The published abstract reports no effect sizes, error counts, or comparative rates. The authors frame the results as a case for further evaluation of AI-scribe performance in multilingual and interpreter-mediated care.
AI evaluation & deployment
How should hospitals calibrate oversight of AI tools embedded in EHR workflows, imaging, triage, documentation, and operations? The authors conducted a narrative review and framework synthesis drawing on peer-reviewed evidence, reporting guidelines, regulatory and policy sources, implementation studies, and applied governance case reports. The resulting framework has four components: a use-case inventory tagged by decision influence and workflow coupling; a six-domain risk taxonomy spanning clinical safety, privacy and data security, ethics and fairness, transparency, system stability, and compliance; a four-tier risk scheme keyed to harm, automation, reversibility, and coupling; and a governance architecture assigning roles to a committee, clinical owners, risk-control functions, and independent assurance. A lifecycle pathway runs from initiation and local validation through shadow mode, controlled go-live, monitoring, change control, and retirement. No effect sizes are reported; this is a conceptual framework, not an evaluation.
For what purposes do physicians use unauthorized, non-conformity-assessed AI tools at work? This cross-sectional survey of physicians in Swedish health care organizations (N=357; response rate ~64%), fielded through a verified online panel between December 2023 and January 2024, applied qualitative content analysis to free-text responses, interpreted through the sociology of professions and paradox theory. Reported uses fell into four categories: clinical work and decision-making (second opinions, differential diagnoses, rare cases), administrative work (patient communication, documentation), research and professional development, and technological curiosity. Physicians framed such use as compensating for gaps in institutional systems and reducing workload. The abstract reports no effect sizes or usage prevalence.
How consistent are the evaluation frameworks proposed for clinical AI? This scoping review followed PRISMA-ScR, searching six databases plus the EQUATOR Network through February 2026, and screened 3363 records to include 46 frameworks, scored on methodological rigor, validation strategy, and a 10-domain UNESCO ethics matrix. Most frameworks (88%) targeted investigational rather than clinical use; 31.8% reported technical metrics such as AUC, sensitivity, and specificity, 15.9% reported clinical indicators, and only 11.4% met methodological rigor with validation aligned to intended use. Ethics coverage was uneven: transparency and explainability appeared in 70%, human oversight in 24.4%.
How should public health teams decide whether a customized large language model is ready to deploy? This conceptual paper proposes an acceptance criteria framework (ACF) defining implementation fit as meeting prespecified minimum performance standards and showing nonproblematic behavior under anticipated use. The ACF combines project-relevant and off-topic test prompts, structured expert review, and prespecified thresholds to generate a documented decision record that can be rerun after model revisions. The authors argue prior safety, ethics, effectiveness, engagement, and implementation frameworks imply rather than operationalize deployment benchmarks, and illustrate the ACF in a tobacco cessation text messaging intervention. No performance estimates are reported.
How can health systems test large language models on real patient portal messages without touching live EHR workflows? This technical feasibility tutorial describes a Python 3 web interface and modular backend running inside the institutional firewall on an NVIDIA GRID T4-1Q GPU, supporting single-message and batch tasks: authorship identification, categorization, criticality flagging, and response drafting with zero-, one-, and few-shot prompting. A deidentification pipeline validated against 110 manually adjudicated entities achieved 95.1% sensitivity and 82.1% precision. Use cases drew on an IRB-approved dementia-relevant corpus of 6941 medical advice request messages from 497 patients; token-based cost readouts were included. No comparative performance effect sizes are reported.
Does acceptable pre-deployment validation performance persist once clinical AI enters routine workflows? This longitudinal retrospective observational study followed four deployed AI systems spanning different clinical domains within one large healthcare organization, comparing validation-era metrics with post-deployment behavior over extended observation using routine clinical data, outcome labels, and operational telemetry, and examining discrimination, calibration, data availability, latency, and workflow signals. In all four systems, validation performance did not persist; calibration drift appeared consistently and often preceded discrimination changes, and label-independent signals such as input missingness and data latency flagged degradation earlier than outcome-based monitoring, which lagged behind label availability. The abstract reports no effect sizes.
How do AI-generated initial recommendations in virtual urgent care compare with the final recommendations physicians issue? This Annals of Internal Medicine study examines concordance between an AI system's initial output and clinicians' final decisions across AI-assisted virtual urgent care visits. No abstract was available, so findings, sample size, and effect sizes cannot be summarized here.
Do the transparency needs of healthcare AI users actually map onto the Instructions for Use (IFU) document that the EU AI Act (Directive 2024/1689) requires providers to give deployers? This cross-sectional online survey, administered via Qualtrics to four deployer groups \u2122 managers (N = 238), healthcare professionals (N = 115), patients (N = 229), and IT workers (N = 230) \u2122 asked participants to rate the relevance of a set of transparency needs and identify which IFU section would address each. Priorities differed across user types, and participants had difficulty locating some transparency information within the IFU structure; the abstract reports no effect sizes or magnitudes. The authors derive recommendations for locally meaningful IFUs.
What would clinicians and caregivers want from an AI-based clinical decision support tool for early autism detection, and where would it fit in the visit? This observational qualitative study used contextual inquiry with 8 clinicians and 20 caregivers during 18- to 24-month well-child visits at Duke-affiliated clinics, analyzed with rapid qualitative analysis. Workflow mapping identified 6 user tasks, 3 technology-user interactions, and 5 clinical decision points, plus 2 barriers (screening tool accuracy, follow-up implementation) and 3 facilitators (electronic screening, early intervention provider input, referral coordination support). Preferences included EHR-embedded, actionable outputs with prediction explanations, visual summaries, and caregiver-facing materials. The abstract reports no effect sizes.
Can an image-based AI narrow the genotype search for inherited retinal diseases before genetic testing? Retina4IRD, a RETFound-pretrained Vision Transformer predicting 17 genotype categories, was trained on fundus photographs and OCT from 1,843 genetically confirmed patients (3,376 eyes) in China, South Korea and Poland; top-5 accuracy was 0.904 internally and 0.856 externally. In a multicenter randomized trial, 300 patients with suspected IRD were assigned 1:1 to AI-assisted or specialist-only assessment (295 analyzed). Top-5 genetic accuracy was 88.5% versus 67.3% (P<0.001), top-1 37.8% versus 22.4%, and a composite downstream management score 37.7 versus 28.5 (P<0.001).
Can institution-specific cancer trial information be curated into an AI-enabled knowledge management application in community oncology? This feasibility study at a regional community oncology network had coordinators and disease teams compile actively recruiting trials, structuring core elements (title, conditions, biomarkers, stage/line, recruiting status) for point-of-care display, with AI-assisted extraction of protocol summaries and eligibility elements followed by human validation. Fifty-three trials across 10 disease groups and 28 cancer types were embedded; 91% were recruiting and 30% were biomarker-specific. Configuration required 2-4 weeks per disease group using existing personnel, without added staffing or EHR build. Usability and implementation outcomes were not assessed.