The AJH Informatics Review

A weekly digest of new research on EHRs, clinical AI, interoperability & health IT policy

Issue No. 008

September 16, 2026 · 21 papers

Evaluation and deployment gaps dominate. A synthetic multilingual study across 77-99 languages finds word error rate essentially uncorrelated with clinically serious ambient scribe errors — consultation complexity, not low-resource language, predicted critical risk — a direct challenge to how scribes are benchmarked. A systematic review of 101 LLM report-generation studies found none at low risk of bias, with early evidence of slightly lower acceptance than radiologists and longer editing time for drafted impressions. Two reviews quantify the translation gap: of 161 predictive analytics studies, 86% stopped at development and only 5.6% reached deployment, while just 12 studies documented human-in-the-loop AI-CDSS, none at the oversight phase. For interoperability watchers, a FHIR review finds 12 of 20 implementations still proof-of-concept. Also useful: audit-log scoring identifies ICU nurse care teams at 91% accuracy.

All EHR use & audit-log metadataOther applied informaticsAI scribes & ambient AIAI evaluation & deploymentClinical decision supportHealth IT policy & regulationDocumentation burden & workloadPatient-facing techInteroperability & HIEEconomics of health IT
001
Using electronic health record audit logs to identify pediatric intensive care unit teams

Can EHR audit logs reliably identify which clinicians make up a patient's care team? Researchers observed rounds for 1931 patient days (678 development, 1253 validation) across pediatric, neonatal, and cardiovascular ICUs at a quaternary children's hospital, documenting daytime teams, then compared two audit-log algorithms: clinically informed heuristics and a Longitudinal Contribution Score (LCS). In the PICU development cohort, LCS outperformed heuristics for nurses (92.9% vs 67.6% accuracy), frontline clinicians (83.3% vs 77.9%), and attendings (66.5% vs 61.9%). In validation, only the nurse advantage replicated (91.2% vs 74.2%); frontline and attending accuracy were comparable (roughly 83% and 69-70%).

002
A framework for developing, validating, and utilizing automated measures of order errors using the retract-and-reorder methodology

How can EHR audit log data be turned into automated measures of ordering errors? This methods paper lays out a framework for conceptualizing, implementing, validating, and optimizing measures based on the Retract-and-Reorder (RAR) approach, which flags orders that are placed, retracted, and then reordered as signals of potential error. Using illustrative examples, the authors describe the theoretical model underlying RAR detection, validation steps, and applications to studying error epidemiology, root causes, and intervention evaluation in near-real time. The abstract reports no empirical performance estimates or effect sizes.

003
★ Beyond word error rate: clinical risk as the necessary standard for ambient AI scribe evaluation: evidence from 77 global languages

Does word error rate track clinically consequential errors in ambient AI scribes? Investigators built a synthetic multilingual corpus — five clinical dictation scripts across a complexity gradient, translated into 99 languages, rendered to speech under three acoustic conditions, and transcribed by a production scribe — then had three independent large language model raters score errors on a Severity x Likelihood framework. Of 59,819 genuine transcription errors, 58,329 (97.5%) were LOW risk and 251 (0.42%) CRITICAL or HIGH. No frequency metric was detectably associated with serious risk (absolute Spearman rho<0.16), while a Severity x Likelihood sum tracked WER (rho=0.80). Low-resource languages had worse WER (beta=+0.078) without higher critical risk (OR 1.21); consultation complexity predicted serious risk (OR 3.06 per level).

004
Optimizing generative artificial intelligence for clinical summarization: a blinded comparison study of automated versus human care planning synopses

Can a generative AI model produce care-transition synopses as good as clinician-written ones? In a blinded, randomized comparison using de-identified records of 64 patients with multiple chronic conditions from MIMIC-III, human- and AI-generated synopses were scored on accuracy, succinctness, synthesis, and usefulness within a data-information-knowledge-wisdom framework (>80% indicating success). AI and clinician summaries overlapped 12%. AI synopses were rated useful 75% of the time versus 76% for human synopses; AI scored lower on succinctness for the data task (55%-67%) and near-equal or better on accuracy and synthesis (AI 72%-79%, humans 68%-84%), best in wisdom. Interrater agreement was variable.

005
★ Human in the loop in AI-enabled clinical decision support: a systematic scoping review and reporting checklist for lifecycle governance

How is human-in-the-loop (HITL) actually implemented across the lifecycle of AI-enabled clinical decision support? This systematic scoping review searched MEDLINE, Embase, Web of Science, PsycINFO, Google Scholar and Scopus in August 2024, with manual identification through mid-2025, including primary studies in which clinicians interacted with AI-CDSS; dual independent screening and extraction mapped findings to four lifecycle phases. Twelve studies qualified. All described clinician involvement during development, mainly expert annotation and rule-based design; 11 reported review-phase HITL, 2 maintenance, and none oversight. Contributions were largely static or retrospective, with small datasets, few annotators, and inconsistent terminology. The authors propose a six-domain HITL reporting checklist.

006
★ Predictive analytics for health-system decision support using population health data: a global scoping review of implementation, governance and decision integration

How far has predictive analytics on routine and population health data actually moved into health-system decisions? This global scoping review searched five databases for 2014–2025 studies applying predictive or forecasting methods for health-system decision-making, screening 2,623 records and including 161 articles (128, 79.5%, from high-income settings), with a supplementary grey-literature scan. Most work stopped at development (139 articles, 86.3%); validation 6.2%, pilot 1.9%, operational deployment 5.6%. While 73.3% claimed relevance to resource allocation or capacity planning, only 5.6% documented an output-to-decision pathway and 8.7% routine workflow integration; 14.9% described how uncertainty informed decisions.

007
Attitudes Toward Large Language Models in Health Care and Preferences for Their Adoption and Oversight Among Health Care Professionals: Cross-Sectional Survey

How do clinicians view large language models and who should govern them? A cross-sectional online survey recruited 335 health care professionals through a health care news mailing list, 68.7% (n=230) attending physicians and 77.9% practicing in the Northeast United States. Some 62.7% (n=210) reported current or contemplated LLM use, with users reporting higher self-rated knowledge than nonusers (P<.001) and no age association (\u03c1=-0.072; P=.19). Top applications were literature review (73.4%), decision support (57%), and patient communication (54.9%); 96.4% voiced bias concern, 65.4% preferred oversight by professional associations over technology companies (29%), and 66.6% reported no confidence in existing oversight. Convenience sample, low response rate.

008
Clinicians' Attitudes and Perceptions on the Adoption of AI in Mental Health Care: Scoping Review

What do mental health clinicians think about AI in their practice? This scoping review followed JBI guidance and PRISMA-ScR, searching six databases (CINAHL, Embase, PsycINFO, PubMed, Scopus, Web of Science) for studies published from 2020 onward; 12,356 records were retrieved and 35 included after dual-reviewer screening. Clinicians showed cautious optimism when AI was framed as supplementing rather than replacing expertise, citing reduced administrative burden, documentation support, information synthesis, and between-session access. Concerns centered on privacy, governance, data ownership, unsafe or inaccurate outputs, overreliance, unclear accountability, and effects on therapeutic relationships, alongside limited AI literacy. The abstract reports no effect sizes.

009
★ Effectiveness, Safety, and Workflow Burden of Large Language Model-Based Medical Report Generation: Systematic Review

Are LLM-generated medical reports clinically ready in terms of effectiveness, safety, and workflow burden? This systematic review searched five databases (PubMed/MEDLINE, Embase, Web of Science, Scopus, Cochrane) for studies published January 2016 through May 2026 evaluating LLMs, multimodal LLMs, or vision-language models for image-to-report generation, impression drafting, or structured reporting, including 101 studies (36 chest x-ray). No study was at low risk of bias (15 moderate, 72 high, 14 serious), and meta-analysis was not possible. One chest x-ray study found AI report acceptance of 70.5% (6047/8580) versus 73.3% for radiologists, with false negatives 18.5% versus 17.8%; a brain MRI study showed reading time falling from 61 to 53 seconds while impression drafting increased editing time and edit distance.

010
Enhancing Patients' Informed Consent Through AI: Systematic Review

Can AI improve patient comprehension and decision-making during informed consent? This PRISMA-guided systematic review searched PubMed, Embase, and the Cochrane Library, including 33 studies published 2020-2025 across three domains: AI-generated patient education (n=18, 54.5%), consent documentation (n=10, 30.3%), and AI-assisted consent acquisition (n=5, 15.2%). Large language models were accurate but readability stayed above an eighth-grade level (best model Copilot: Flesch-Kincaid 10.59, SD 1.22). AI-generated documents raised Flesch Reading Ease by 44%-122% and lowered required grade levels 10%-47%. In trials, AI-assisted consent shortened consultations (7.7 vs 10.6 minutes; P=.05) and lowered post-consent anxiety in knee arthroplasty (10.48 vs 12.75; P=.04).

011
★ FHIR as an interoperability enabler for clinical applications: a narrative review of implementations, impacts, and future challenges

How far have FHIR-based clinical applications progressed from prototype to routine use? This narrative review searched Scopus and PubMed for English-language articles published January 2019 to January 2026 describing concrete FHIR-based tools with empirical findings, and applied thematic synthesis to build a functional taxonomy and maturity framework. Across 20 included studies, the authors identify six functional paradigms, including AI and decision support, large-scale surveillance and research, and patient empowerment. Most work remained early-stage: 12 of 20 studies were proof-of-concept, with 4 clinical pilots and 4 institutional integrations. The review reports no effect sizes.

012
A tailored maturity model for REDCap-based patient-facing technology: revealing development stages and guiding future evolution

How mature are patient-facing technologies built on REDCap? This integrative review followed PRISMA guidelines, searching PubMed, Web of Science, and Embase, and synthesized 14 studies (2016-2026) using aggregate and thematic synthesis, then combined literature evidence with pilot case insights to build a four-tier maturity model spanning Data Collection, System Integration, Personalized Insights, and a conceptual Intelligence Hub. Publications clustered after 2021 (n=12, 86%), most were US-based (n=11, 79%), and 10 implementations (71%) remained at Tiers 1-2, with only 4 (29%) reaching Personalized Insights. The authors call for longitudinal evaluation and deeper integration.

013
Cost-Effectiveness of Electronic Patient-Reported Outcome Measure Interventions in Cancer: Systematic Review and Parameter Extraction for Economic Modeling

Are electronic patient-reported outcome measure (ePROM) programs in cancer care cost-effective? This systematic review searched Ovid (MEDLINE, Embase), Scopus, and the INAHTA database for English-language papers through March 2025, including 34 publications from 27 studies covering 26 ePROM-integrated interventions for adult cancer populations, alongside parameter extraction for economic modeling. Most interventions (23/26) included alert handling or automated decision support. Only 5 publications reported full cost-effectiveness analyses; 3 were highly uncertain, while 2 showed cost-effectiveness driven by quality-of-life gains and fewer hospitalizations. Five reported partial results (4 favoring ePROMs). Twelve studies had qualitative components, but only 2 addressed economic themes.

014
The development of a technologic approach to improve access to individualized clinical documentation in caregivers' preferred language

Can an EHR template semi-automatically translate clinical documentation into caregivers' preferred language? This pilot development study applied human-centered design and a Discover, Design/Build, Test framework at a pediatric setting, with an interprofessional team of a speech-language pathologist and a certified translation specialist building an English-to-Spanish template in the electronic health record, then surveying speech-language pathologists and caregivers on acceptability and feasibility. The Discover phase documented barriers to providing written documentation in patients' primary language; clinicians endorsed the template's importance but raised feasibility and usability concerns, while caregivers valued receiving information in their primary language. The abstract reports no sample sizes or effect sizes.

015
Inclusive Digital Health Strategies for Persons With Disabilities: A Scoping Review

How well do digital health technologies, including AI, include persons with disabilities? This scoping review followed PRISMA-ScR guidance, searching MEDLINE and Web of Science plus gray literature for English-language sources on digital health, disability, and health equity published 2019 to 2025 (searches last run December 2024). Of 925 records identified and 836 screened, 137 underwent full-text review, yielding 40 peer-reviewed articles plus 40 gray literature sources (80 documents). Findings were organized into five themes spanning access enablers, stakeholders, government initiatives, contextual factors, and emerging innovations; participatory codesign and accessible design recurred as enablers. The abstract reports no effect sizes.

016
Telehealth use among US adults with type 2 diabetes: differences in medication fill patterns and care processes

Does telehealth for type 2 diabetes accompany different medication fill patterns and care processes than in-person care alone? This observational cross-sectional study used the nationally representative 2021-2023 Medical Expenditure Panel Survey, comparing unadjusted descriptive measures among 4348 adults with at least one type 2 diabetes visit. Telehealth use was uncommon: 90.7% had only in-person visits and 9.3% had at least one telehealth visit. Insulin prescriptions were more frequent among telehealth users (47.5%, 95% CI 41.0%-54.1%) than nonusers (34.4%, 95% CI 32.1%-36.7%). The authors interpret this as possible greater clinical complexity; comparisons were unadjusted, with no adjusted effect sizes reported.

017
Systemic Challenges in EHR Alert Design: A Thematic Analysis of Physician Discourse on X (Formerly Twitter)

Can physician social media posts surface usable signals about clinical decision support failures? This qualitative study searched X in August 2025 using Grok 4 via an internal API with keywords such as "EHR alert fatigue" and "override EHR," curating 117 posts from 106 self-identified physicians (2016–2025) and analyzing one post per author with Braun and Clarke reflexive thematic analysis plus AI-assisted coding verification. Six usability failure domains emerged, led by burnout and fatigue (43.4%) and alert overwhelm (38.7%); lack of customization was least common (7.5%). Sentiment was 76.4% negative, and half of patient-safety posts co-occurred with alert overwhelm.

018
The human factor in hospital cybersecurity: a high-reliability organization-based maturity model

Do cybersecurity incidents involving human factors recur within the same healthcare organizations, and can recurrence serve as a marker of persistent vulnerability? The authors analyzed 5,752 cyber incidents reported by U.S. healthcare organizations, comprising 3,740 human-factor and 2,012 non-human-factor incidents, classifying human involvement with a deterministic rule-based approach applied to incident metadata and defining recurrence as a later incident in the same category and organization within three years. Human-factor incidents recurred significantly more often, a pattern holding across alternative windows, stricter classification, and organization-level and paired analyses; the abstract reports no effect sizes. The authors propose a High Reliability Organization-informed maturity model.

019
Survivor Treatment Selection Bias in Evaluations of Virtual Transition of Care

How does survivor treatment selection bias affect evaluations of virtual transition-of-care programs? This JMIR Medical Informatics paper addresses a methodological concern in which patients must survive, or remain enrolled, long enough to receive a virtual follow-up intervention, potentially inflating apparent benefit in observational analyses. No abstract was available, so the study design, data source, population, and any quantitative findings or proposed analytic corrections cannot be described here.

020
Community Health Center Adoption of Enabling Technologies to Address Contextual Drivers of Health for Care-Managed Patients: Formative Evaluation

What keeps community health center care managers from using EHR tools to screen for and act on contextual drivers of health such as food, housing, and transportation? This formative evaluation, guided by human-centered design, used semi-structured interviews with 11 care management staff and 5 subject matter experts from a multi-state health center network, analyzed with a rapid qualitative approach. Staff used screening flowsheets and alerts for documentation, but cited tool limitations — including no discrete fields for referral documentation — and insufficient training on referrals and follow-up. Participants requested leadership support, hands-on training, and customizable care plans. The abstract reports no effect sizes; findings will inform implementation strategies for a planned randomized trial.

021
Beyond Healthcare Utilization: A Whole-Child, Multi-Sector Approach to Pediatric Risk Stratification

Does multi-sector data identify different high-need children than claims-based risk scores? This cross-sectional study applied the North Carolina Integrated Care for Kids (NC InCK) risk stratification algorithm — integrating health, education, social, and juvenile justice data — to 99,564 Medicaid participants aged 0-20 in a five-county central North Carolina region in October 2022, comparing assigned Service Integration Levels (SILs) with Medicaid managed care organization risk levels. SIL 1 covered 90,151 children (90.5%), SIL 2 5,482 (5.5%), and SIL 3 3,931 (4%). Roughly 24% of SIL 2 and 20% of SIL 3 children were classified low risk by their MCO.