The AJH Informatics Review

A weekly digest of new research on EHRs, clinical AI, interoperability & health IT policy

AI evaluation & deployment

Every digest paper in this category, newest first.

001
Propagation of Interpreter Errors by Ambient AI Scribes: Study Using Simulated Clinical Encounters

Do ambient AI scribes carry interpreter errors into the clinical note? Using simulated English- and Spanish-language clinical encounters mediated by interpreters, the authors evaluated whether documentation generated by ambient AI scribes reproduced interpretation errors introduced during the visit. Scribes propagated interpreter errors into the resulting notes, with propagation patterns differing by speaker role and by error type. The published abstract reports no effect sizes, error counts, or comparative rates. The authors frame the results as a case for further evaluation of AI-scribe performance in multilingual and interpreter-mediated care.

002
Risk-Tiered Governance for Hospital Artificial Intelligence: A Framework Synthesis and Implementation Pathway

How should hospitals calibrate oversight of AI tools embedded in EHR workflows, imaging, triage, documentation, and operations? The authors conducted a narrative review and framework synthesis drawing on peer-reviewed evidence, reporting guidelines, regulatory and policy sources, implementation studies, and applied governance case reports. The resulting framework has four components: a use-case inventory tagged by decision influence and workflow coupling; a six-domain risk taxonomy spanning clinical safety, privacy and data security, ethics and fairness, transparency, system stability, and compliance; a four-tier risk scheme keyed to harm, automation, reversibility, and coupling; and a governance architecture assigning roles to a committee, clinical owners, risk-control functions, and independent assurance. A lifecycle pathway runs from initiation and local validation through shadow mode, controlled go-live, monitoring, change control, and retirement. No effect sizes are reported; this is a conceptual framework, not an evaluation.

003
Shadow AI in Swedish Health Care: Qualitative Analysis of Physicians' Free-Text Answers

For what purposes do physicians use unauthorized, non-conformity-assessed AI tools at work? This cross-sectional survey of physicians in Swedish health care organizations (N=357; response rate ~64%), fielded through a verified online panel between December 2023 and January 2024, applied qualitative content analysis to free-text responses, interpreted through the sociology of professions and paradox theory. Reported uses fell into four categories: clinical work and decision-making (second opinions, differential diagnoses, rare cases), administrative work (patient communication, documentation), research and professional development, and technological curiosity. Physicians framed such use as compensating for gaps in institutional systems and reducing workload. The abstract reports no effect sizes or usage prevalence.

004
Evaluation Frameworks for Clinical AI Incorporating Validation Strategies, Real-World Applicability, and Ethical Principles: Scoping Review

How consistent are the evaluation frameworks proposed for clinical AI? This scoping review followed PRISMA-ScR, searching six databases plus the EQUATOR Network through February 2026, and screened 3363 records to include 46 frameworks, scored on methodological rigor, validation strategy, and a 10-domain UNESCO ethics matrix. Most frameworks (88%) targeted investigational rather than clinical use; 31.8% reported technical metrics such as AUC, sensitivity, and specificity, 15.9% reported clinical indicators, and only 11.4% met methodological rigor with validation aligned to intended use. Ethics coverage was uneven: transparency and explainability appeared in 70%, human oversight in 24.4%.

005
An Acceptance Criteria Framework for Determining the Implementation Fit of Custom Large Language Models in Public Health Interventions

How should public health teams decide whether a customized large language model is ready to deploy? This conceptual paper proposes an acceptance criteria framework (ACF) defining implementation fit as meeting prespecified minimum performance standards and showing nonproblematic behavior under anticipated use. The ACF combines project-relevant and off-topic test prompts, structured expert review, and prespecified thresholds to generate a documented decision record that can be rerun after model revisions. The authors argue prior safety, ethics, effectiveness, engagement, and implementation frameworks imply rather than operationalize deployment benchmarks, and illustrate the ACF in a tobacco cessation text messaging intervention. No performance estimates are reported.

006
A Secure User Interface for Preclinical Evaluation of AI in Patient Portal Message Management: Tutorial

How can health systems test large language models on real patient portal messages without touching live EHR workflows? This technical feasibility tutorial describes a Python 3 web interface and modular backend running inside the institutional firewall on an NVIDIA GRID T4-1Q GPU, supporting single-message and batch tasks: authorship identification, categorization, criticality flagging, and response drafting with zero-, one-, and few-shot prompting. A deidentification pipeline validated against 110 manually adjudicated entities achieved 95.1% sensitivity and 82.1% precision. Use cases drew on an IRB-approved dementia-relevant corpus of 6941 medical advice request messages from 497 patients; token-based cost readouts were included. No comparative performance effect sizes are reported.

007
Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems

Does acceptable pre-deployment validation performance persist once clinical AI enters routine workflows? This longitudinal retrospective observational study followed four deployed AI systems spanning different clinical domains within one large healthcare organization, comparing validation-era metrics with post-deployment behavior over extended observation using routine clinical data, outcome labels, and operational telemetry, and examining discrimination, calibration, data availability, latency, and workflow signals. In all four systems, validation performance did not persist; calibration drift appeared consistently and often preceded discrimination changes, and label-independent signals such as input missingness and data latency flagged degradation earlier than outcome-based monitoring, which lagged behind label availability. The abstract reports no effect sizes.

008
Comparison of Initial Artificial Intelligence (AI) and Final Physician Recommendations in AI-Assisted Virtual Urgent Care Visits

How do AI-generated initial recommendations in virtual urgent care compare with the final recommendations physicians issue? This Annals of Internal Medicine study examines concordance between an AI system's initial output and clinicians' final decisions across AI-assisted virtual urgent care visits. No abstract was available, so findings, sample size, and effect sizes cannot be summarized here.

009
Transparency in healthcare AI: Testing EU regulatory provisions against users' transparency needs

Do the transparency needs of healthcare AI users actually map onto the Instructions for Use (IFU) document that the EU AI Act (Directive 2024/1689) requires providers to give deployers? This cross-sectional online survey, administered via Qualtrics to four deployer groups \u2122 managers (N = 238), healthcare professionals (N = 115), patients (N = 229), and IT workers (N = 230) \u2122 asked participants to rate the relevance of a set of transparency needs and identify which IFU section would address each. Priorities differed across user types, and participants had difficulty locating some transparency information within the IFU structure; the abstract reports no effect sizes or magnitudes. The authors derive recommendations for locally meaningful IFUs.

010
Understanding end-user contexts and identifying design preferences of an artificial intelligence-based clinical decision support tool for early autism detection

What would clinicians and caregivers want from an AI-based clinical decision support tool for early autism detection, and where would it fit in the visit? This observational qualitative study used contextual inquiry with 8 clinicians and 20 caregivers during 18- to 24-month well-child visits at Duke-affiliated clinics, analyzed with rapid qualitative analysis. Workflow mapping identified 6 user tasks, 3 technology-user interactions, and 5 clinical decision points, plus 2 barriers (screening tool accuracy, follow-up implementation) and 3 facilitators (electronic screening, early intervention provider input, referral coordination support). Preferences included EHR-embedded, actionable outputs with prediction explanations, visual summaries, and caregiver-facing materials. The abstract reports no effect sizes.

011
AI-based clinician decision support system for diagnosis of inherited retinal diseases: a multicenter, randomized trial

Can an image-based AI narrow the genotype search for inherited retinal diseases before genetic testing? Retina4IRD, a RETFound-pretrained Vision Transformer predicting 17 genotype categories, was trained on fundus photographs and OCT from 1,843 genetically confirmed patients (3,376 eyes) in China, South Korea and Poland; top-5 accuracy was 0.904 internally and 0.856 externally. In a multicenter randomized trial, 300 patients with suspected IRD were assigned 1:1 to AI-assisted or specialist-only assessment (295 analyzed). Top-5 genetic accuracy was 88.5% versus 67.3% (P<0.001), top-1 37.8% versus 22.4%, and a composite downstream management score 37.7 versus 28.5 (P<0.001).

012
Development of a clinical trial knowledge management application for community oncology

Can institution-specific cancer trial information be curated into an AI-enabled knowledge management application in community oncology? This feasibility study at a regional community oncology network had coordinators and disease teams compile actively recruiting trials, structuring core elements (title, conditions, biomarkers, stage/line, recruiting status) for point-of-care display, with AI-assisted extraction of protocol summaries and eligibility elements followed by human validation. Fifty-three trials across 10 disease groups and 28 cancer types were embedded; 91% were recruiting and 30% were biomarker-specific. Configuration required 2-4 weeks per disease group using existing personnel, without added staffing or EHR build. Usability and implementation outcomes were not assessed.