The AJH Informatics Review

A weekly digest of new research on EHRs, clinical AI, interoperability & health IT policy

AI scribes & ambient AI

Every digest paper in this category, newest first.

001
Beyond word error rate: clinical risk as the necessary standard for ambient AI scribe evaluation: evidence from 77 global languages

Does word error rate track clinically consequential errors in ambient AI scribes? Investigators built a synthetic multilingual corpus — five clinical dictation scripts across a complexity gradient, translated into 99 languages, rendered to speech under three acoustic conditions, and transcribed by a production scribe — then had three independent large language model raters score errors on a Severity x Likelihood framework. Of 59,819 genuine transcription errors, 58,329 (97.5%) were LOW risk and 251 (0.42%) CRITICAL or HIGH. No frequency metric was detectably associated with serious risk (absolute Spearman rho<0.16), while a Severity x Likelihood sum tracked WER (rho=0.80). Low-resource languages had worse WER (beta=+0.078) without higher critical risk (OR 1.21); consultation complexity predicted serious risk (OR 3.06 per level).

002
A Conceptual Model for Ambient AI Adoption: Perspectives From Academia and Industry

How should health systems judge ambient AI performance when vendor claims and real-world results diverge? This conceptual article, drawing on academic and industry perspectives, proposes a shared mental model organizing ambient AI assessment into three interdependent dimensions: technical, interface, and system level. For each dimension, the authors specify the types of information relevant to evaluation, what vendors should reasonably be expected to disclose, and how provider organizations can run internal evaluations to contextualize, verify, or supplement vendor claims. No empirical study, effect sizes, or performance benchmarks are reported; the contribution is a decision-support framework intended for organizations of varying size.

003
Early insights into nurses' perceptions, opportunities, and concerns regarding Ambient artificial intelligence in nursing practice

How widely has ambient AI reached inpatient nursing, and what do nurses expect from it? This cross-sectional online survey recruited practicing U.S. inpatient nurses through nursing-focused social media communities and used descriptive analyses of adoption, experience, and perceptions. Among 61 eligible respondents, nearly 60% (n=37) said ambient AI had been piloted or implemented at their institution, and 49.2% (n=30) had used it clinically; 73.3% of those (n=22) reported positive experiences. Cited benefits included reduced documentation time, improved workflow, and better data quality; concerns included patient acceptance, job displacement, and skill loss from overreliance.

004
Ambient Artificial Intelligence Scribes in Ambulatory Care: Patient Acceptability and Trust

What makes patients willing to accept ambient AI scribes during ambulatory visits, and does baseline trust in AI shape that acceptance? This secondary, convergent mixed-methods analysis re-analyzed survey and interview data from 20 patients seen after ambient AI scribe implementation, summarizing trust items descriptively and coding interviews deductively against the Theoretical Framework of Acceptability plus inductively. Trust ranged from low to high, with 60% reporting moderate trust. Acceptability tracked with minimal ethicality concerns, supportive affective attitudes, low patient burden, perceived benefits, and strong intervention coherence; themes differed minimally by trust level. Patients urged patient education and advance notice. The abstract reports no effect sizes.

005
The Effect of Ambient AI Documentation on Clinician Workload, Efficiency, and Patient Experience in a Multisite Emergency Department Network

Does ambient AI documentation change clinician workload, efficiency, and patient experience in emergency care? A retrospective observational cohort study examined voluntary ambient AI use across 14 EDs from May 2024 to June 2025, covering 315,242 notes, with within-clinician paired comparisons and pre/post NASA-TLX surveys. AI was used in 8.6% of notes. Top-box ratings for clinician listening rose (82.0% vs. 76.6%; OR 1.39), while likelihood to recommend did not differ. Active editing time was unchanged (median +0.17 minutes), though clinicians typed 722 fewer characters; copied content halved (9.4% vs. 18.3%) and NASA-TLX workload fell 40.2 points.

006
Large-Scale Implementation of Ambient AI Documentation and Its Effects on EHR Efficiency and Clinician Well-Being

Does ambient AI documentation improve efficiency and clinician well-being when deployed across a large multi-site system? A retrospective index-date-aligned pre-post analysis of Epic Signal data covered 210 ambulatory physicians and APPs using DAX Copilot at a multi-state health system, June 2024 to November 2025, with 8 months before and after activation, plus a 45-day post-activation survey (n=233, 26.4% response). Active note time fell from 5.54 to 4.09 minutes per appointment (26.2%; p<0.001), while note length rose 28.8% and typing fell 51.7%, copy/paste 42.7%, and conventional voice recognition 67.8%. Survey respondents reported reduced burnout and higher satisfaction; no effect sizes given.

007
Readability of AI-Generated Patient Visit Summaries in Orthopedic Surgery: Retrospective Analysis

Do AI scribe–generated patient visit summaries meet patient literacy standards? This retrospective analysis scored 982 consecutive summaries from a commercial AI scribe platform at an academic orthopedic surgery outpatient clinic (December 2023–May 2024, 25 incomplete summaries excluded) using five validated readability indices. Mean Flesch-Kincaid Grade Level was 9.3 (SD 1.2) and mean Flesch Reading Ease was 57.6 ("fairly difficult"); only 0.4% (4/982) met the sixth-grade benchmark and 14.2% (139/982) the eighth-grade threshold. Word count was uncorrelated with grade level (ρ=0.012, P=.70), and indices agreed strongly (W=0.882). No comparison group was included.

008
Quality, consistency, and clinical safety of AI-generated versus clinician-written clinical notes: a multi-country paired simulation study

Do ambient AI scribes produce notes of comparable quality and safety to clinician-written ones across languages? This paired simulation covered five countries and languages (Cambridge, Barcelona, Milan, Paris, Cologne), with 385 actor-performed consultations yielding 770 notes, each documented independently by an AI scribe (Heidi) and a junior-to-middle-grade clinician, scored on the PDQI-9 by blinded evaluators. AI notes scored higher (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; dz=0.55) and less often carried at least one Critical+High error: 24.4% vs 61.0% by calibrated automated review (RR 2.50) and 6.2% vs 21.8% by clinician adjudication (RR 3.50), with omissions driving the gap.

009
Comparative effectiveness of ambient documentation tools in primary care tool-specific variations in efficiency, documentation burden, and productivity over time

Do ambient documentation tools differ from one another in real-world impact? This comparative-effectiveness study followed 163 primary care providers in a large integrated health system (January 2024–June 2025), analyzing 59 130 provider-days across a tablet-based virtual human-assisted tool (A, n=65), an EHR-integrated tool (B, n=68), and a standalone tool (C, n=17), using intention-to-treat and per-protocol models with provider-clustered SEs and month fixed effects. Versus Tool B, Tool A showed more after-hours EHR time (+0.022 hours/provider-day), more manual note composition (+0.046), and lower 48-hour visit closure (−0.120; 95% CI −0.126 to −0.115). Tool C cut after-hours work by 0.055 hours/provider-day (~30 hours annually) but also reduced timely closure (−0.028).

010
Are automated documentation-error judges fit to measure ambient AI scribes? A pre-registered, blinded human-validation study

Can automated judges that flag documentation errors serve as a defensible measurement instrument for ambient AI scribes, absent a gold standard? This pre-registered, blinded validation study nested in a multi-country ambient documentation simulation (English setting) had ten external clinicians adjudicate a stratified sample of 434 pipeline flags, yielding 565 adjudications with 131 double-rated. Inter-clinician agreement on error genuineness was fair (raw 59%, AC1 0.24); judge-clinician agreement was 64% (95% CI 60-68). Behavior was near-symmetric across arms (kept-precision 74% AI vs 81% clinician notes; severity gap +0.06 vs -0.09 tiers), with removed-confirmed asymmetric (56% vs 42%). Latent-class triangulation estimated a 68% genuine-error rate (94% CrI 48-83).

011
Performance of an Ambient Generative AI Documentation Tool in a Linguistically Diverse Clinical Setting

Do ambient AI scribes perform equally well across patient languages? This retrospective analysis covered 54,160 outpatient encounters in a U.S. safety net health system, measuring the share of words in the final note generated by the AI tool and left unedited by the provider. Generalized estimating equations with exchangeable correlation structures accounted for clustering within patients, with univariable models by language and interpreter modality and a multivariable interaction model. Non-English encounters were 21% to 25% less likely than English encounters to reach the performance threshold; interpreter-mediated and bilingual-provider encounters did not differ significantly.

012
Cognitive Workload and Mental Burden in Health Care Professionals Interacting With AI: Systematic Review and Meta-Analysis

Does clinical AI reduce or add to clinicians' cognitive workload? This systematic review and meta-analysis searched MEDLINE, Embase, Web of Science, and CENTRAL (January 2015-2026), including 21 studies of 2885 health care professionals in 7 countries that used validated instruments (NASA-TLX, Professional Fulfillment Index). Ambient documentation AI was associated with lower temporal demand (SMD -1.46, 95% CI -2.81 to -0.11; k=2), lower effort (SMD -1.29, -2.16 to -0.42), reduced work exhaustion (MD -0.35, -0.58 to -0.12), and lower burnout prevalence (OR 0.47, 0.25-0.86). Diagnostic imaging AI and decision support showed mixed or increased workload. GRADE certainty was moderate at best; prediction intervals crossed the null.

013
Changes in Clinician Time Expenditure and Visit Quantity With Artificial Intelligence-Powered Scribes

No abstract was available for this JAMA report on artificial intelligence-powered ambient scribes. The title indicates the study examines whether AI scribe adoption changes how clinicians allocate their time and whether visit volume shifts as a result — outcomes central to arguments that ambient documentation tools reduce administrative burden or, alternatively, free capacity that gets absorbed by added throughput. Design, setting, population, and effect sizes cannot be characterized without the full text; readers should consult the article directly for the magnitude and direction of any measured changes.

014
Adoption and utility of digital scribes in clinical practice - A scoping review

What shapes adoption of digital scribes, and what do they change for patients, clinicians, and organisations? This PRISMA-ScR scoping review searched MEDLINE, CINAHL, Web of Science, SCOPUS, and EMBASE for original studies or case reports evaluating digital scribe implementation in real-world care, mapping themes to the updated Consolidated Framework for Implementation Research and its Outcomes Addendum. Of 4772 studies screened, 29 were included. Scribes were generally acceptable (n=11) and usable (n=8), though nine reported accuracy concerns; reported impacts included reduced documentation burden (n=18), improved clinician wellbeing (n=12), and better patient-clinician interaction (n=10). Only three examined cost or productivity. The abstract reports no pooled effect sizes.

015
Bridging the trust-adoption gap for AI scribes in rural communities: A machine learning approach using the 2024 Canadian digital health survey

How do rural patients view ambient AI scribes, and which characteristics predict acceptance? A cross-sectional analysis of 1,050 rural respondents in the 2024 Canadian Digital Health Survey dichotomized three outcomes — trust in documentation accuracy, perceived interaction benefit, and future-use preference — and fit XGBoost classifiers using sex, age, race/ethnicity, education, employment, income, chronic disease, self-reported health, and high-speed internet access, summarizing subgroup differences as marginally standardized predicted probabilities. Endorsement declined across the three domains. Future-use preference was 0.388 for males versus 0.313 for females, 0.466 for graduate degrees versus 0.308 for less than high school, and 0.408 with chronic disease versus 0.298 without. Internet access showed similar probabilities throughout.

016
Deployment of an ambient AI scribe in emergency care: A 12-month evaluation in a large Spanish hospital network

Can an ambient AI scribe scale across emergency departments without degrading documentation or patient experience? This 12-month multicenter retrospective observational study covered five emergency specialties at 48 Spanish hospitals (February 2025–January 2026), including all level 4 and 5 consultations among roughly 2.27 million ED visits. The scribe was used in 1,032,558 consultations (45.3%), with monthly adoption rising from 7.7% to 57.8% and 2,097 physicians using it at least once. Scribe-assisted consultations were shorter (mean relative time savings 21.8%, p<0.001), transcription accuracy averaged 93.9%, and audited report quality and patient Net Promoter Scores were higher; the abstract reports no effect sizes for the quality and experience comparisons.

017
Propagation of Interpreter Errors by Ambient AI Scribes: Study Using Simulated Clinical Encounters

Do ambient AI scribes carry interpreter errors into the clinical note? Using simulated English- and Spanish-language clinical encounters mediated by interpreters, the authors evaluated whether documentation generated by ambient AI scribes reproduced interpretation errors introduced during the visit. Scribes propagated interpreter errors into the resulting notes, with propagation patterns differing by speaker role and by error type. The published abstract reports no effect sizes, error counts, or comparative rates. The authors frame the results as a case for further evaluation of AI-scribe performance in multilingual and interpreter-mediated care.