The AJH Informatics Review

A weekly digest of new research on EHRs, clinical AI, interoperability & health IT policy

Issue No. 005

August 26, 2026 · 23 papers

Three papers should reset expectations about clinical AI evidence. A systematic analysis of all 1,357 FDA-cleared AI/ML devices found just three evaluated patient-centered outcomes, a stark backdrop for anyone writing deployment policy. In the first prospective DECIDE-AI stage 1 evaluation of a multi-LLM decision support system in an emergency department, outputs were judged clinically appropriate but adoption fell from 68% to 30% and length of stay was identical, a reminder that workload, not accuracy, governs use. A five-country paired simulation found ambient scribe notes scored higher on the PDQI-9 and carried fewer critical errors than clinician-written ones, mostly through fewer omissions. Also notable: portal message writing style explained roughly half the response-rate gap for Black patients, and Bayesian reconstruction recovered 76.6% of suppressed cells in aggregated count sharing.

All EHR use & audit-log metadataDocumentation burden & workloadAI scribes & ambient AIPatient-facing techAI evaluation & deploymentHealth IT policy & regulationOther applied informaticsClinical decision supportEconomics of health ITInteroperability & HIE
001
Leveraging Individualized Electronic Health Record Learner Analytics to Improve Resident Inbasket Management

Can individualized EHR analytics paired with structured coaching improve resident inbasket performance? Investigators at a large Mid-Atlantic academic center conducted a stratified longitudinal analysis of EHR metrics with pre- and post-intervention surveys among PGY-2 and PGY-3 internal medicine residents and continuity clinic attendings, delivering personalized efficiency and quality reports in one-on-one R2C2-model feedback sessions across two training periods. Turnaround time for patient calls fell by 2.2 days (p<0.001), and mean time in patient calls declined from 0.99±0.53 to 0.74±0.38 minutes/day (p=0.03). Self-reported burnout (p=0.09), confidence (p=0.18), and perceived efficiency (p=0.56) were unchanged; 6 of 7 faculty found the analytics at least moderately useful.

002
The Effect of Ambient AI Documentation on Clinician Workload, Efficiency, and Patient Experience in a Multisite Emergency Department Network

Does ambient AI documentation change clinician workload, efficiency, and patient experience in emergency care? A retrospective observational cohort study examined voluntary ambient AI use across 14 EDs from May 2024 to June 2025, covering 315,242 notes, with within-clinician paired comparisons and pre/post NASA-TLX surveys. AI was used in 8.6% of notes. Top-box ratings for clinician listening rose (82.0% vs. 76.6%; OR 1.39), while likelihood to recommend did not differ. Active editing time was unchanged (median +0.17 minutes), though clinicians typed 722 fewer characters; copied content halved (9.4% vs. 18.3%) and NASA-TLX workload fell 40.2 points.

003
Large-Scale Implementation of Ambient AI Documentation and Its Effects on EHR Efficiency and Clinician Well-Being

Does ambient AI documentation improve efficiency and clinician well-being when deployed across a large multi-site system? A retrospective index-date-aligned pre-post analysis of Epic Signal data covered 210 ambulatory physicians and APPs using DAX Copilot at a multi-state health system, June 2024 to November 2025, with 8 months before and after activation, plus a 45-day post-activation survey (n=233, 26.4% response). Active note time fell from 5.54 to 4.09 minutes per appointment (26.2%; p<0.001), while note length rose 28.8% and typing fell 51.7%, copy/paste 42.7%, and conventional voice recognition 67.8%. Survey respondents reported reduced burnout and higher satisfaction; no effect sizes given.

004
Readability of AI-Generated Patient Visit Summaries in Orthopedic Surgery: Retrospective Analysis

Do AI scribe–generated patient visit summaries meet patient literacy standards? This retrospective analysis scored 982 consecutive summaries from a commercial AI scribe platform at an academic orthopedic surgery outpatient clinic (December 2023–May 2024, 25 incomplete summaries excluded) using five validated readability indices. Mean Flesch-Kincaid Grade Level was 9.3 (SD 1.2) and mean Flesch Reading Ease was 57.6 ("fairly difficult"); only 0.4% (4/982) met the sixth-grade benchmark and 14.2% (139/982) the eighth-grade threshold. Word count was uncorrelated with grade level (ρ=0.012, P=.70), and indices agreed strongly (W=0.882). No comparison group was included.

005
★ Quality, consistency, and clinical safety of AI-generated versus clinician-written clinical notes: a multi-country paired simulation study

Do ambient AI scribes produce notes of comparable quality and safety to clinician-written ones across languages? This paired simulation covered five countries and languages (Cambridge, Barcelona, Milan, Paris, Cologne), with 385 actor-performed consultations yielding 770 notes, each documented independently by an AI scribe (Heidi) and a junior-to-middle-grade clinician, scored on the PDQI-9 by blinded evaluators. AI notes scored higher (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; dz=0.55) and less often carried at least one Critical+High error: 24.4% vs 61.0% by calibrated automated review (RR 2.50) and 6.2% vs 21.8% by clinician adjudication (RR 3.50), with omissions driving the gap.

006
Deployer-side governance of medical imaging artificial intelligence: the regulatory readiness instrument (RRI-MI) for multi-jurisdictional and post-market compliance

How can hospitals operationalize post-market oversight of imaging AI when pre-market clearances test algorithms under static conditions? This conceptual paper proposes the Regulatory Readiness Instrument for Medical Imaging AI (RRI-MI), a deployer-side readiness assessment mapping 11 governance domains onto a provisional 22-point ordinal rubric, aligned with the FDA Predetermined Change Control Plan and the EU AI Act. The authors ground the framework in post-market evidence on scanner drift, protocol shifts, software updates, and demographic variation, illustrating it through a semiautonomous prostate cancer MRI case study covering local validation, human oversight, version control, monitoring, and incident response. No empirical validation or effect sizes are reported.

007
AI-assisted radiographic fracture detection and length of stay in the adult ambulatory orthopedic emergency department: a before-after cohort study with a disease-specific internal control

Does deploying an AI radiographic fracture detection tool shorten emergency department stays? This retrospective controlled before-after cohort study with interrupted time-series analysis covered 8,253 fracture-excluded and 6,864 fracture-confirmed visits at a tertiary ambulatory orthopedic ED, January 2021–February 2026, with deployment on 1 October 2023. Mean length of stay in fracture-excluded patients fell from 135.6 to 127.5 minutes (−8.13; 95% CI −11.52 to −4.70), while the fracture-confirmed control was unchanged (+0.37 minutes; p=0.87). Reductions concentrated at the upper tail (−24.0 minutes at the 90th percentile); stays over four hours fell from 11.3% to 8.2%, with no increase in early revisits.

008
Privacy, security, and reliability risks of artificial intelligence in healthcare: a systematic review of empirical evidence

What empirical evidence exists for privacy, security, and reliability risks from AI in clinical care? This systematic review searched PubMed, Embase, Web of Science, Scopus, IEEE Xplore, and ACM Digital Library for empirical studies published January 2015 to November 2025 evaluating AI use or misuse in diagnosis, treatment, or decision-making. Of 7,285 records plus 205 from citation screening, 22 studies met inclusion criteria, mostly medical imaging. Five recurring threat categories emerged: re-identification, membership inference, unauthorized access and adversarial exploitation, input manipulation, and misuse or overinterpretation of outputs. Models encoded latent biometric signals, limiting anonymization and synthetic data. Findings were synthesized narratively; the abstract reports no pooled effect sizes.

009
Ethics of Autonomous AI Clinical Trials: Delphi Study

How should the NIH's 7 principles of ethical clinical research be adapted for trials of autonomous AI? Using a modified Delphi approach over 6 months, investigators convened 14 multidisciplinary panelists (AI, data science, ophthalmology, policy, law, bioethics, patient advocacy) across two survey rounds anchored to a vignette and a final virtual meeting, with participation of 12/14 (85.7%), 10/14 (71.4%), and 13/14 (92.9%). Round 2 produced 9 strong-agreement, 2 moderate-agreement, and 4 divisive statements. Recommendations covered transparency on training and validation data, pre-deployment bias and inequity assessment, performance across clinical settings, informed consent, comparison with standard of care, downstream access, and cost.

010
The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review

How have researchers actually operationalized human evaluation of large language model reliability in health care? This PRISMA-ScR scoping review searched PubMed, Web of Science, Cochrane Library, CINAHL, and Google Scholar for English-language original studies published January 2016 to July 2025, screening 4347 records and including 71 (26 clinical, 45 public health). Six reliability indicators recurred: accuracy, relevance, completeness, clarity, safety, and consistency. Clinical studies more often assessed guideline concordance and structural coherence; public health studies emphasized understandability, harm potential, and repeat-response consistency. Single-specialty clinicians predominated as evaluators, panels typically had five or fewer members, and 5-point Likert scales with researcher-defined rubrics were standard. Reported limitations included evaluator subjectivity and nonstandardized indicators.

011
★ 1,357 AI medical devices cleared, 3 actually tested on patient outcomes

How much clinical evidence supports FDA-cleared AI medical devices? This systematic analysis catalogued all 1,357 AI/ML-enabled devices cleared through December 5, 2025 using the FDA device database and the ACR Data Science Institute catalogue, with linked searches of ClinicalTrials.gov and PubMed for registered trials and publications. Only 34 devices (2.5%) were linked to registered prospective trials, 12 (0.9%) posted results, 12 (0.9%) had peer-reviewed publications, and 3 (0.2%) evaluated patient-centered outcomes such as mortality, morbidity, or readmissions. Most studies (62%) were observational with small, homogeneous cohorts and frequent exclusion of vulnerable populations. The authors cite misaligned incentives and predicate-based pathways as barriers.

012
Machine Learning in Palliative Care: Scoping Review of Applications

How far have machine learning applications in palliative care moved beyond mortality prediction? This scoping review followed Arksey and O'Malley and PRISMA-ScR guidance, searching six databases from inception through February 9, 2026, and included 121 peer-reviewed primary studies (2015-2026) across 24 countries, 54.5% (66/121) from the United States. Mortality and survival prediction accounted for 42.1% (51/121) of studies, followed by health care use (20.7%) and symptom assessment (16.5%); cancer populations dominated (43%). Half (50.4%) used an explainability technique, 86.8% (105/121) did not address equity in model performance, and 66.1% remained proof-of-concept with 17.4% reaching prospective deployment or clinical integration.

013
Centralized Digital Surveillance for Abdominal Aortic Aneurysm Detection, Longitudinal Tracking, and Management Within an Integrated Health System: Retrospective Cohort Study

Can a centralized digital program identify and track abdominal aortic aneurysms across an integrated health system? This retrospective cohort study describes STAIR, which combined structured EHR problem-list queries, an internally developed NLP model applied to radiology reports, clinician referrals, and automated lost-to-follow-up queries, enrolling 8464 patients (mean age 77.1 years; 77.0% male) from December 2022 through December 2024, with status assessed through April 2026. Identification was mostly automated: problem-list queries 59.0% and radiology NLP 29.0%. After centralized review, 45.3% were assigned biennial duplex surveillance and 20.6% referred to vascular surgery; 49.5% remained under active surveillance. Findings are descriptive; clinical effectiveness was not assessed.

014
★ Prospective evaluation of a large language model clinical decision support system in the emergency department

Can a multi-LLM clinical decision support system change emergency department care in practice? This DECIDE-AI stage 1 prospective evaluation deployed SHAKED in a tertiary ED over 4 weeks, analyzing 1,138 patients across two parallel units, one using the system and one following routine rotations. Adoption fell from 68% to 30%, with disengagement tied to workload (OR 0.72 per shift hour, 95% CI 0.62–0.83), while physicians favored it for radiology consultations (OR 2.98). Expert reviewers rated 99 of 100 sampled outputs clinically appropriate, with no adverse events. Length of stay was 4.9 hours in both wings (P = 0.99); consultation cycle time trended shorter (−9.4 min, P = 0.077).

015
An Electronic Health Record-Integrated, Large Language Model-Powered Tool to Triage Surgical Patients

Can eligibility for surgical comanagement (SCM) — hospitalist co-management of medically complex perioperative patients — be triaged automatically? This prospective, unblinded quality improvement study at Stanford Health Care (September 2025–February 2026) deployed an EHR-integrated, human-in-the-loop LLM tool (SCM Navigator) that classified patients using preoperative documentation, structured data, and morbidity criteria, with attending review as the reference standard. Across 6193 triaged cases (median age 60.2 years; 49.0% female), 1582 (25.5%) were recommended for hospitalist consultation; sensitivity was 0.94 (95% CI, 0.91-0.96) and specificity 0.74 (95% CI, 0.71-0.77). LLM misclassification explained 2 of 19 false negatives (11%).

016
Clinical Specialty Expansion of AI-Enabled and Machine Learning-Enabled Medical Devices Authorized by the US Food and Drug Administration From 1995 to 2025: Longitudinal Content Analysis

Has radiology's dominance of FDA-authorized AI/ML medical devices persisted or begun to loosen? This longitudinal content analysis covered all 1430 devices in the FDA AI-Enabled Medical Devices registry with final authorization decisions through December 2025, stratified by advisory-committee specialty across four eras (1995-2015, 2016-2019, 2020-2022, 2023-2025). Annual authorizations rose from a mean of 2.0 to 264, with 331 in 2025; 96.2% used the 510(k) pathway. Radiology's share peaked at 85.5% (347/406) in 2020-2022, then fell to 77.5% (614/792) in 2023-2025 (P=.001), with the Herfindahl-Hirschman Index declining from 0.738 to 0.612. Start-ups (OR 5.09) and technology companies (OR 50.62, based on 13 devices) had higher odds of nonradiology authorization.

017
Public Reporting Systems in Health Care and the Underconceptualized Technical Substrate, a Core Information Systems Dimension: Scoping Review

How much does the literature on public reporting systems in health care address their technical foundations? This scoping review followed Joanna Briggs Institute guidance and PRISMA-ScR, searching seven databases (PubMed, Web of Science, Scopus, IEEE, ACM, AIS eLibrary, Cochrane) for articles published 2000-2026. Of 1882 records identified and 1127 screened after deduplication, 233 studies were included; 157 (67.4%) came from the United States and 43 (18.4%) from Europe, with 60.5% (n=141) quantitative, 24.0% qualitative, and 15.5% mixed methods. Coding across six information systems dimensions—architecture, interoperability, data governance, APIs, usability, technical performance—found the technical substrate underexplored; the abstract reports no dimension-level counts.

018
Beyond the Black Box: Unraveling the Role of Explainability in Human-Artificial Intelligence Collaboration

When does explaining an AI model's reasoning actually improve human-AI decisions, and at what cognitive cost? The authors build an analytical model of a decision maker with limited but flexible cognition receiving imperfect machine recommendations, where explanations shift beliefs about algorithmic quality. Low explainability leaves decision accuracy and reliance unchanged while reducing cognitive burden; higher explainability improves accuracy by curbing overreliance but increases underreliance. Explainability matters more for cognitively constrained decision makers, complex tasks, and lower-stakes decisions, yet can raise processing time and fatigue exactly when time is short, tasks are complex, and machine quality is doubted. Theoretical modeling, so no empirical effect sizes.

019
The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot

Does an AI pair programmer help or hinder distributed, voluntary software collaboration? Using GitHub's proprietary Copilot usage data linked to public OSS project data, the authors estimate Copilot's effects on contribution and coordination in open-source projects; the abstract does not name the identification strategy, sample size, or study years. Copilot use raised project-level code contributions 5.9%, with a 3.4% increase in developer coding participation and a 2.1% increase in individual contributions, but an 8% increase in coordination time and more code discussion. Net timely merges rose. Peripheral developers showed smaller contribution gains and larger coordination increases than core developers.

020
Benchmark Mineability and the Financing of AI Innovation -- by Alex Chan

When AI benchmark scores steer capital, what happens to their value as signals? This NBER working paper is a conceptual and theoretical market-design analysis rather than an empirical study, treating public AI benchmarks as market institutions. The author identifies two gaps that become exploitable once scores move investment: public examples can reveal the process behind a private final test, and any finite public score cannot span the broad task space implied by general intelligence. Targeted effort aimed at these gaps erodes the signal later investors rely on. The proposed remedy is separating development from certification — publishing practice tasks but selecting the investment-consequential task generator only after a submitted system's evaluation policy is fixed. The abstract reports no effect sizes or empirical magnitudes.

021
★ Care Team Response to Patient Portal Message Content and Writing Style

Why do portal messages from historically marginalized patients get fewer responses? This cross-sectional study applied natural language processing to extract message content and writing style features from 3,619,390 medical advice request threads sent by 511,020 adults to non-trainee primary care clinicians between 2021 and 2023, then used regression to decompose response gaps. Black patients had a 3.7-percentage-point lower response rate from the intended target clinician than White patients (95% CI, -4.1 to -3.3), an 11.6% relative reduction. Message content did not explain the gap, but writing style accounted for 48.0% of it for Black patients, 34.9% for Hispanic patients, 60.5% for patients with only a high school education, and 42.8% for Medicaid beneficiaries.

022
Default Hospice Consultation in Comfort Measures Orders and Hospice Engagement

Does making hospice consultation the default option within comfort measures order sets increase hospice engagement? This JAMA Network Open study examines default hospice consultation embedded in comfort measures orders and its relationship to hospice use among hospitalized patients. No abstract was available, so the design, population, and findings cannot be summarized here; readers should consult the full article for methods and effect estimates.

023
★ Sharing Aggregated Patient Counts in Place of Line-Level EHR Data: Analytic Fidelity and the Limits of Count Suppression for Privacy

Can aggregated count "cubes" with small-cell suppression substitute for line-level EHR sharing without leaking the cells they hide? The authors built a Bayesian count-inference pipeline that both reconstructs suppressed counts and functions as a reconstruction attack, applying it to 285 pediatric kidney-transplant patients at Boston Children's Hospital and comparing against CTGAN synthetic data. Cube analyses reproduced line-level results: across 106 demographic-by-medication subgroups, a bootstrap mean of 3.5 showed significant graft-rejection associations, and cube odds-ratio sign changes reversed no significant associations versus 2.3 for CTGAN. However, 76.6% of suppressed cells (14,554 of 18,994) were recovered exactly, including 85.5% of single-patient cells.