The AJH Informatics Review

A weekly digest of new research on EHRs, clinical AI, interoperability & health IT policy

Issue No. 009

September 23, 2026 · 24 papers

Two deployment studies anchor the week: a GPT-4.1 hospital course tool improved discharge summary quality but increased editing time by 32%, even as 86% of clinicians believed it saved them time, and the PRECISE trial shows a machine learning sepsis subphenotyping algorithm, randomization, and alerting built entirely in standard Epic tools across six hospitals. Both should reframe how we judge "efficiency" and how pragmatic AI trials get built. A scoping review of 275 AI quality and safety records found workflow measures in under 5% and formal economic evaluation in 1.5%, naming the barriers structural rather than technical. On the policy side, 11.9 million Americans live in broadband deserts, overwhelmingly rural and sometimes overlapping ambulance and care deserts, while inpatient portal use shows sharp racial and insurance gaps.

All AI evaluation & deploymentClinical decision supportDocumentation burden & workloadPatient-facing techEconomics of health ITInteroperability & HIEOther applied informaticsHealth IT policy & regulation
001
★ Development to implementation to evaluation: step-by-step approach to embedding a clinical trial of a machine learning algorithm in the electronic health record

How can a machine learning algorithm be embedded in the electronic health record to run a pragmatic randomized trial? This implementation report describes the Precision Resuscitation with Crystalloids in Sepsis (PRECISE) trial, a multihospital RCT built entirely with standard Epic tools through four components: automated inclusion criteria, real-time sepsis subphenotyping, randomization, and a medication alternative alert prompting clinicians toward the fluid type thought to benefit the identified subgroup. PRECISE launched across 6 Emory Healthcare hospitals in June 2024, covering 6 emergency departments and 17 ICUs with more than 300 ICU beds. The abstract reports implementation details only, with no trial outcomes or effect sizes.

002
★ Evaluation of a Large Language Model Discharge Summary Hospital Course Tool: Improved Quality but Longer Documentation Time

Does an LLM-generated hospital course draft reduce discharge summary documentation time? This observational pre-post study compared a pre-AI period (March 16-November 19, 2025) with a post-AI period (November 20, 2025-February 11, 2026) after a GPT-4.1 tool was embedded in mandatory discharge summary templates, covering 8,298 hospitalized adults on hospital medicine services. Edit time did not differ between periods (6.60 vs 6.28 minutes, p=0.11), but within the post-AI period, tool use was associated with longer editing (9.20 vs 4.93 minutes) and a 32% increase in adjusted analysis (95% CI 24-41%). Faculty review of 27 encounters found higher quality and lower harm but less concise summaries; 31 of 36 surveyed clinicians (86.1%) felt the tool improved efficiency.

003
Operatıonalızıng ethıcal AI for safe clınıcal ıntegratıon: A qualıtatıve ıntervıew study

How are ethical principles for AI safety actually operationalised once clinical AI reaches practice? This qualitative study (QuAS-AI) used semi-structured interviews with 16 experts involved in clinical AI implementation "\u2014 academics, industry professionals, and practitioners \u2014 recruited by purposive and snowball sampling, with independent coding, peer debriefing, and member checking. Three themes emerged: performance, risk, and bias as dynamic properties needing continuous monitoring; transparency and explainability as complementary supports for clinical interpretation rather than technical disclosure; and accountability as multi-actor governance spanning traceability, logging, data governance, and intervention capacity. The authors frame safety as socio-technical governance and suggest a High-Reliability Organisation approach. The abstract reports no effect sizes.

004
Large language models for patient-facing pathology report interpretation: A scoping review

How have large language models been used and evaluated for translating pathology reports into patient-facing language? This scoping review followed JBI methodology and PRISMA-ScR, searching six databases for empirical studies of LLM-generated interpretation, rewriting, question answering, or summarization of pathology reports (January 2018 to August 2026). Nineteen studies were included; GPT-family models appeared in 17, and report-level transformation was the most common task (12/19, 63.2%). Fidelity was assessed in all 19 studies, safety in 10 (52.6%), readability in 9 (47.4%), and comprehension and usability in 6 each (31.6%). Only five involved patients or other non-clinicians, and none evaluated performance by health-literacy level, across languages, or prospectively within clinical workflows.

005
Stakeholder Perspectives on the Integration of AI in Diabetes Care: Systematic Review of Qualitative Studies

What do patients, caregivers, and clinicians actually expect from AI in diabetes care? This systematic review of qualitative studies searched five databases (MEDLINE, Web of Science, Scopus, CINAHL, PsycINFO) through February 2026, including 14 studies published 2023-2025 with at least 738 participants across 9 countries, covering large language models, AI-enabled apps and wearables, glucose prediction, and decision support. Thematic synthesis produced 4 analytical themes and 13 subthemes; 11 findings were rated high confidence and 2 moderate under GRADE-CERQual. Stakeholders saw value for prevention, education, and self-management but raised accuracy, bias, privacy, accountability, workload, and autonomy concerns. Much evidence rested on prototypes or hypothetical systems.

006
★ AI for Health Care Quality and Patient Safety: Scoping Review of Diagnostic, Predictive, and Decision Support Applications

Why does strong AI task performance so rarely translate into sustained clinical benefit? This scoping review followed Joanna Briggs Institute methodology with PRISMA-ScR and PRISMA-S reporting, searching five databases (MEDLINE, Scopus, Web of Science, IEEE Xplore, CINAHL Plus) for English-language records from January 2017 to April 2026. Of 43,394 records identified, 275 were charted across four nonmutually exclusive domains: clinical decision support (233, 84.7%), predictive analytics (192, 69.8%), diagnostics (142, 51.6%), and economic assessment (53, 19.3%). Workflow or process measures appeared in 13 records (4.7%), equity or subgroup analyses in 20 (7.3%), and formal economic evaluations in 4 (1.5%). The authors group recurring constraints into five cross-domain barriers, describing them as structural rather than technical.

007
Large language models and artificial intelligence for generating, simplifying, and enhancing outpatient clinic letters: A systematic review

Can large language models reliably generate or simplify outpatient clinic letters? This systematic review searched five databases (PubMed, EMBASE, Web of Science, CENTRAL, CINAHL) from inception to 1 November 2025, appraising studies with the Mixed Methods Appraisal Tool and GRADE. Seven studies were included, two using real-world clinic data and five synthetic or hypothetical scenarios. Readability findings were mixed and most AI output still exceeded the US Grade 6 level; information fidelity ranged from 10% to 100%, and one study reported a ten-fold reduction in drafting time. Certainty was very low across primary outcomes.

008
Utilization of a HIPAA-compliant large language model chatbot in an academic pediatric medical center

Who actually uses a hospital-wide LLM chatbot, and what blocks the rest? This mixed-methods case study of "InternalGPT," a HIPAA-compliant chatbot at an academic pediatric medical center, combined employee surveys with 14 months of utilization data. Of roughly 15,800 employees, 2,149 (13.6%) requested access; 52.8% of those used at least one token, 33.6% never logged in, and the top 20% of users consumed 69.4% of tokens. Barriers cited by non-users were limited time (51.4%) and difficulty using the tool (21.8%). Among 92 sustained users surveyed, mean self-reported productivity gain was 30%, an exploratory $6.3M-$18.9M value.

009
Demographics, Clinical Content, Use Patterns, and Care-Seeking Intent Across Two Generations of AI-Enabled Clinical Triage Tools (A Traditional Structured Questionnaire and a Large Language Model-Enabled Conversational Interface): Comparative Retrospective Observational Study

Does conversational LLM triage capture different information or shift care-seeking intent compared with a structured questionnaire? This retrospective observational study compared 116,890 virtual triage encounters over 28 weeks (January–August 2025), where users self-selected traditional triage (TT; 100,533, 86%) or an LLM-enabled conversational interface (CT; 16,357, 14%) sharing the same Bayesian reasoning engine, with poststratification weighting by age and sex. CT sessions ran longer (median 8 min 21 s vs 4 min 25 s), elicited more clinical findings (median 36 vs 32; P<.001), and had higher self-reported intended adherence to recommended care (34.3% vs 29.2%; P<.001), including self-care (85.4% vs 61.9%). Users self-selected groups.

010
Understanding Key Stakeholders' Perspectives Towards Artificial Intelligence in Home Care Work

How do the people who would implement AI in home care view its promise and risks? This qualitative study conducted semi-structured interviews, incorporating AI scenarios, with 43 participants across five stakeholder groups — home health aides and attendants, home care agency leaders and staff, worker advocates, clinicians, and technology company personnel — recruited through purposive and snowball sampling, with analysis by structural coding, inductive sub-coding, and thematic analysis. Mean age was 44.6 years; 44.4% reported no or low AI knowledge. Four themes emerged: benefits to patient care, engagement, and efficiency; risks to care quality, provider-patient relationships, and working conditions; data quality, privacy, and AI literacy challenges; and needs for equitable governance. No effect sizes are reported.

011
Deploying the REACHnet Multi-State EHR-Based Network for Disease Surveillance (MENDS) node: lessons learned and tools utilized

What does it take to convert a research-grade EHR network into a public health surveillance asset? This descriptive implementation report documents how REACHnet transformed its curated clinical data from health systems in Louisiana and Texas into the Multi-State EHR-Based Network for Disease Surveillance (MENDS) dataset for chronic disease surveillance, with access extended to the Louisiana and Texas health departments and the National Association of Chronic Disease Directors. The authors describe governance, regulatory, infrastructure, and partnership processes, along with challenges encountered in producing a relatively low-latency dataset. The abstract reports no quantitative results or effect sizes.

012
Education on Artificial Intelligence in US Internal Medicine Residencies: Results of a National Survey

How much artificial intelligence training is offered in US internal medicine residency programs? This national survey reports on the state of AI education across internal medicine training programs. No abstract was available, so findings, sample size, and effect estimates cannot be summarized here.

013
Digital Health Readiness and Medicare Primary Care Spending: Nationwide County-Level Observational Analysis

Is community-level digital health readiness associated with lower Medicare primary care spending in safety net settings? This county-level observational study covered 2993 US counties from 2017 to 2023, using the Digital Health Index digitalization subindex against geographically adjusted per capita Medicare spending on federally qualified health center and rural health clinic services, with a hybrid within-between panel model. Between-county differences drove the association (β=−67.62, 95% CI −73.51 to −61.74), while within-county change was null (β=−2.72, P=.40). Highest versus lowest digitalization tertile differed by β=−165.05. In 2023 cross-sectional models, health care access showed the strongest inverse association (β=−71.13).

014
Data, Privacy Laws and Firm Production: Evidence from the GDPR

How do data privacy regulations affect firm production and performance? This Journal of Political Economy paper examines the European Union's General Data Protection Regulation as a natural experiment in restricting firms' access to and use of consumer data. No abstract was available, so findings, methods, data sources, and effect sizes cannot be summarized here; readers interested in the estimated magnitudes of GDPR's effects on firm data holdings, computation, and output should consult the full text.

015
★ Disparities in Cancer Patients' Portal Usage During Hospitalizations

Who uses the inpatient portal during cancer hospitalizations? This retrospective analysis of 28,386 patients at a high-volume cancer hospital in 2022–2023 used multivariable logistic and Poisson regression to model any portal login and login rates during admissions, including ICU stays. Median age was 65, median length of stay 4 days; 65% (18,588) logged in at least once. Black patients were 32% less likely to log in than White patients (OR 0.68, 95% CI 0.62–0.74), as were single (OR 0.80) and Medicaid-insured patients (OR 0.79). Among 1,782 ICU patients, 55% logged in, with Black patients 41% less likely (OR 0.59).

016
Perceptions of Technology and Digital Health Tools Among Recently Incarcerated Adults With Opioid Use Disorder: Semistructured Interview Study

How do adults with opioid use disorder leaving incarceration perceive digital health tools during reentry? This qualitative study conducted semistructured interviews with a purposive sample of 39 adults recently released from New Hampshire prisons and jails, recruited from the EXIT-CJS randomized trial of extended-release buprenorphine, naltrexone, and enhanced treatment as usual, with deductive-inductive content analysis and COREQ reporting. Most participants preferred digital over in-person care for its convenience and its ability to bypass transportation shortages, limited provider availability, and work scheduling conflicts, while still valuing in-person therapeutic connection. Use depended on devices, affordable connectivity, and digital literacy; the abstract reports no effect sizes.

017
Trust in Generative AI for Health Information Consumption and the Effect of Learned Dependency: Randomized Controlled Experimental Study

Does habitual reliance on generative AI blunt users' ability to calibrate trust in AI-generated health information? Two randomized 2×2 between-participants experiments (338 college students; 563 Mechanical Turk workers) manipulated information accuracy and text-based visual cues (highlighting), measuring trust and self-reported learned dependency with regression models. Accuracy raised trust (experiment 1 B=2.107, 95% CI 1.337-2.878; experiment 2 B=0.203, 95% CI 0.115-0.290), as did learned dependency (B=0.277 and B=0.822). The accuracy-by-dependency interaction was negative in both (B=-0.399; B=-0.459), indicating reduced sensitivity to inaccuracy. Text highlighting had no significant effect and did not moderate dependency.

018
Leveraging a Locally Anchored Learning Health Care System Toward Equitable Telehealth: Assessing Broadband Access and Supporting Use of Federal Internet Subsidies in a Safety-Net Setting

What are the telehealth and broadband barriers facing patients in an urban safety-net clinic, and how aware are they of federal internet subsidies? Roots Community Health, operating as a community-anchored learning health system with a Telehealth Patient Advisory Council, screened 109 adult patients for Affordable Connectivity Program (ACP) eligibility and administered a 66-item cross-sectional survey to 99. Two-thirds (65/99) had used telehealth and 53% (52/99) wanted future telehealth visits. Barriers included slow internet (46/98), no internet access (40/99), and mobile data plan problems (31/98). Most (65/109) had not heard of ACP, though 60/109 were interested in applying.

019
Use, Modality, and Reimbursement Patterns of Outpatient Psychotherapy Among Children With Commercial Insurance

How has outpatient psychotherapy for commercially insured children been used, delivered, and paid for? This JAMA Network Open study examines use, modality (including in-person versus telehealth delivery), and reimbursement patterns for pediatric outpatient psychotherapy in a commercially insured population. No abstract was available at the time of this digest, so the study's design details, sample size, years covered, and findings are not summarized here; readers should consult the full article for effect sizes and payment estimates.

020
★ Geographic Disparities in Access to Broadband, Ambulance Services, and Health Care and Implications for Telehealth

Can telehealth compensate for rural gaps in primary and emergency care where broadband is inadequate? This population-based cross-sectional study mapped broadband, ambulance, and health care deserts across 41 states (249.1 million people), combining FCC 2024 Broadband Data Collection data, ambulance and health care desert data from September 2021 to February 2022, and the 2020 Census. An estimated 11.9 million people (4.8%) lived in broadband deserts, 88.1% of them rural; 649 225 rural residents lived where all three deserts overlapped. Rural broadband subscription was 88.5% versus 92.6% urban; Western states had 31.9% of rural residents in broadband deserts.

021
Measuring Expert Inter-Rater Agreement with a Semi-Automated Clinical Decision Support System: Who Agrees with What?

How much do expert clinicians agree with each other when judging diuretic titrations, and what baseline should a semi-autonomous clinical decision support system (OTTO-FM) be held to? Secondary analysis of prospectively collected porcine data modeling postoperative fluid overload had three pediatric cardiac intensivists rate 29 clinician-driven and 44 CDS-driven items as reasonable/unsure/unreasonable, with ordinal-weighted Gwet's AC2. Dosing agreement was similar across phases (human 0.79, CDS 0.83), but unanimity reached only 65.5% and 70.5% of items. Risk labels split sharply: high-risk 0.92 versus low-risk 0.54; 19 of 20 nonunanimous risk items were single-rater dissent.

022
Continuous Remote Patient Monitoring in Heart Failure Patients: The Heart Failure Cascade Study: Phase II and III Outcomes

Does continuous remote patient monitoring reduce 30-day readmissions after heart failure discharge? This case-versus-retrospective-propensity-matched-control study at Endeavor Health (Evanston, IL) enrolled 39 patients across three phases, monitoring them for 30 days postdischarge with wearable biosensors and daily symptom surveys, with rules-based and machine learning alerts triaged by home health nurses and escalated to advanced practice providers. Intervention patients received more APP calls (66.7 vs. 7.7%), APP visits (43.6 vs. 5.1%), diuretic escalation (43.6 vs. 12.8%), and labs (76.9 vs. 43.6%). Adjusted 30-day readmission did not differ (odds 0.31, 95% CI 0.06–1.38; p = 0.138).

023
The role and impact of hospital dashboards aligned to externally-driven quality care frameworks: A scoping review

What does the evidence say about hospital dashboards built around external regulatory and accreditation quality frameworks? This scoping review searched MEDLINE, Embase, Emcare, Global Health, and Web of Science from inception to October 2025 for studies implementing inpatient dashboards with measures tied to an externally determined framework and reported outcomes. Twenty-one studies were included, 81% from the United States and 76% implemented hospital-wide. Most reported process gains — data timeliness, performance-gap identification, governance, workflow integration, transparency — and 57% reported change in at least one clinical measure over time. The abstract reports no pooled effect sizes, and the authors note limited, heterogeneous studies.

024
Digital competences for the health workforce: a systematic review

What digital competences does the health workforce actually need? This systematic review followed PRISMA, searching MEDLINE, PubMed, CINAHL, PsycInfo, Cochrane Library, Scopus, ProQuest, BASE and the first 10 pages of Google and Google Scholar for English-language literature from 2014 to 2024, with dual independent screening and quality appraisal using Joanna Briggs Institute tools and the Mixed Methods Appraisal Tool. From 9414 records, 108 peer-reviewed studies and competence frameworks were included. Thematic synthesis grouped competences into three domains: leadership (governance, strategic thinking), procedural (digital professional development, quality improvement), and enabling (patient engagement, mindset). The authors report fragmented competence lists rather than integrated frameworks; no effect sizes are reported.