Can pairing a curriculum with a structural inbox-coverage system improve inter-visit care in residency training? This multimethod pre-post evaluation combined a retrospective survey of educational outcomes with chart review of matched resident cohorts among 32 internal medicine residents in an academic continuity clinic during the 2023-2024 academic year. The intervention added firm-based inbox coverage with faculty supervision plus a longitudinal curriculum on EHR efficiency, inter-visit clinical reasoning, and professional responsibility. Clinic-wide test results communicated rose from 57% to 76% (p<0.01), and mean time to communication fell from 18.7 to 8.5 days (p<0.01). Resident-reported confidence improved, without reported effect sizes; resident-level changes varied by baseline tertile and training level.
Issue No. 007
Leading this week: a seven-hospital assurance evaluation of Epic's EHR-embedded generative summarizer found an 11.79% hallucination rate across 40,452 atomic content units and poor regeneration stability, even as clinicians rated it easy to use — a sobering benchmark for anyone deploying summarization. Pair it with a randomized vignette study in which faculty adopted AI-generated e-consult advice at the same high rate whether it was labeled AI or specialist (37.5% to 81.8%), raising real questions about disclosure as a safeguard. Four international case studies (Catalonia, Norway, Singapore, Queensland) reframe AI scaling as multilevel orchestration rather than top-down mandate, while a five-year audit of the 2020 EIT Health/McKinsey forecast shows imaging and administrative automation delivered but remote monitoring and classical NLP stalled on interoperability and reimbursement. Also notable: linking marketplace and HIE data cut unknown race/ethnicity in Maryland Medicaid from 23.0% to 1.0%.
How should health systems judge ambient AI performance when vendor claims and real-world results diverge? This conceptual article, drawing on academic and industry perspectives, proposes a shared mental model organizing ambient AI assessment into three interdependent dimensions: technical, interface, and system level. For each dimension, the authors specify the types of information relevant to evaluation, what vendors should reasonably be expected to disclose, and how provider organizations can run internal evaluations to contextualize, verify, or supplement vendor claims. No empirical study, effect sizes, or performance benchmarks are reported; the contribution is a decision-support framework intended for organizations of varying size.
How widely has ambient AI reached inpatient nursing, and what do nurses expect from it? This cross-sectional online survey recruited practicing U.S. inpatient nurses through nursing-focused social media communities and used descriptive analyses of adoption, experience, and perceptions. Among 61 eligible respondents, nearly 60% (n=37) said ambient AI had been piloted or implemented at their institution, and 49.2% (n=30) had used it clinically; 73.3% of those (n=22) reported positive experiences. Cited benefits included reduced documentation time, improved workflow, and better data quality; concerns included patient acceptance, job displacement, and skill loss from overreliance.
How are health systems moving AI beyond pilots to scale? This qualitative study conducted four case studies — Catalonia, Norway, Singapore, and Queensland — drawing on 60 documents, 34 interviews, and 5 focus groups with 50 strategic decision-makers, policymakers, and lead clinicians, analyzed first within-case thematically and then across cases using the Technology, People, Organization, and Macroenvironment framework. Scaling trajectories were shaped by existing digital strategies, digitalization histories, funding arrangements, and legacy infrastructure, with experimentation opportunities, incentives, and distribution of decision-making authority mattering alongside post-deployment monitoring, governance, and procurement. The authors frame scaling as multilevel orchestration rather than top-down versus bottom-up. No effect sizes are reported.
How well did the 2020 EIT Health & McKinsey AI forecast track what actually happened? This structured narrative review compared the report's domain-level predictions against 2020–2025 evidence from peer-reviewed implementation studies, FDA regulatory data, national and international guidance, industry surveys, and foundation-model evaluations, classifying each prediction as realised, under-realised, exceeded, or unanticipated. Two predictions held: administrative automation and medical imaging led adoption, with ambient documentation moving from pilots to deployment and more than 1300 FDA-authorised AI/ML-enabled devices, mostly radiology. Remote monitoring and classical NLP decision support under-delivered, limited by interoperability, reimbursement, and workflow barriers; generative and multimodal foundation models were the largest unanticipated divergence.
How should medical AI be evaluated across its lifecycle, from bench performance to clinical benefit? This conceptual paper proposes a five-phase evaluation framework spanning technical validation, operational robustness, controlled interaction, clinical evidence, and real-world integration, embedded in a dynamic architecture with phase-gating criteria and local and systemic fall-back triggers that permit re-entry into earlier phases after drift, version updates, or safety signals. It maps multicenter external validation, shadow-mode testing, human-AI comparison and cooperation studies, randomized trials, real-world evaluations, and adaptive designs onto this pathway. No empirical data or effect sizes are reported; the contribution is a framework aimed at researchers, institutions, and regulators.
Can an EHR-embedded machine learning phenotype identify emergency department patients with opioid use disorder in real time for trial screening? Across three EDs in one US health system (2014–2025), a random-forest classifier using visit-level data available at or before triage was trained against a high-specificity computable label and deployed to trigger point-of-care alerts. Retrospective discrimination was high (AUROC 0.99; AUPRC 0.92). In prospective chart-review validation (n=217), 89.3% of model-positive and 95.7% of model-negative encounters matched physician adjudication; design-weighted estimates using the 3.26% flag rate (28,284/866,569 encounters) gave sensitivity 0.40 and specificity 0.996.
Do physicians rate and act on AI-generated e-consultation advice differently when they know its source? In a randomized clinical vignette study, 44 internal medicine teaching faculty at one academic medical center reviewed four vignettes with advice generated by a medically specialized generative AI (OpenEvidence), labeled either as from a specialist attending physician or an AI system. Mean quality ratings were high across all six domains (overall means 4.2-4.7 of 5), with no differences by label. Selection of the consult-recommended management rose from 37.5% to 81.8% after viewing advice (OR 8.43, 95% CI 5.00-14.28); labeling did not affect adoption (OR 1.18, p=0.70).
Can an autonomous coding agent replace hand-written analyst queries against a production EHR warehouse? In a single-center quality-improvement evaluation, ten acute otitis media questions were posed to analysts (adjudicated reference) and to OpenAI Codex, which wrote read-only queries against a full copy of an Epic Caboodle warehouse under four conditions. At medium reasoning effort the agent came within 5% of the reference on 27 of 30 runs but exactly matched on 10; three repeated runs agreed on only 3 of 10 questions. Matching runs identified the same patients (F1 = 1.00, one exception at 0.72); analyst feedback and reuse of corrected definitions restored reproducibility.
How reliable is an EHR-embedded generative AI summarizer in routine use? Investigators applied a multidimensional assurance framework to Epic IP Insights across a seven-hospital system, building an agentic hallucination detector that decomposed summaries into atomic content units (ACUs) and checked each against source notes. The tool produced 706 summaries over 445 encounters, drawing on a mean 5.8% of available notes. The detector matched physician adjudication on 96.95% of ACUs in validation; across 671 summaries (40,452 ACUs) the hallucination rate was 11.79% (95% CI, 11.47%-12.10%). Among 385 regenerated pairs, 32.7% were textually and 37.9% semantically identical. Clinicians rated it easy to use (97.8%) but were mixed on relying on output without verification (45.9%).
Can multi-agency administrative data be linked to flag Medicaid-enrolled children at risk of out-of-home placement? This descriptive development-and-implementation report covers an expert-derived, three-category risk stratification algorithm built by a government-academic team for roughly 27,000 children in two Ohio counties from 2022 to 2024 under Ohio's Integrated Care for Kids Model, drawing on Medicaid claims, child welfare information systems, area-level social determinants, and patient-reported health risk assessments. The authors describe the legal and administrative work required for data linkage and lessons on data use agreements, partner relationships, and lookback periods for retroactively updated data. The abstract reports no predictive performance metrics or effect sizes.
Can linking external data sources fill gaps in Medicaid race and ethnicity records? This data-linkage study matched all Maryland Medicaid enrollees during calendar year 2023 (N=1,898,041) to records from the state health insurance marketplace and the designated health information exchange, assigning each enrollee a single value using a fixed source hierarchy (marketplace, exchange, historic Medicaid, current Medicaid). Most enrollees (97.8%) appeared in at least one external source, and the share with unknown race and ethnicity fell from 23.0% to 1.0%, with the resulting distribution more closely matching American Community Survey benchmarks and permitting greater disaggregation.
Can standards-based, interoperable electronic care planning tools support shared care for people with multiple chronic conditions? Using participatory and agile design, the team built data standards plus clinician-facing (eCarePlanner) and patient/caregiver-facing (MyCarePlanner) apps, then evaluated them with mixed methods informed by CFIR Process Redesign and SEIPS across formative, iterative, and summative stages. The apps connected to 4 electronic health records at 17 institutions; 57 patients/caregivers and 15 clinicians participated, predominantly aged 65+, White, and well educated. Most were comfortable using the app (97%) and found loading timely (90%), but only 63% felt it would support complex care coordination, 48% that it improved care team communication, and 37% cited cross-section inconsistencies. Interviews highlighted barriers to moving information across settings.
Can information about an AI system help people delegate tasks to it more often and more wisely? This experimental study manipulated two signals: ex-ante AI certainty (the AI's estimated likelihood of being correct, shown before the delegation decision) and ex-post outcome information (whether the AI was actually correct, shown after). Presenting either signal alone had no effect or reduced combined human-AI performance; providing both raised delegation frequency and delegation effectiveness, improving performance. The authors attribute this to certainty calibrating task-level expectations, outcomes confirming them, more accurate mental models, and less algorithm aversion. The abstract reports no effect sizes, sample size, or task details.
Can large language models audit whether published research adheres to its pre-analysis plan? This methodological working paper applies an LLM to the authors' own study, prompting it to identify precommitted design choices, evaluate deviations from the plan, and diagnose gaps in pre-specification. The authors report the approach substantially reduces the human labor required for adherence checks, but that audit output varies across different LLMs, leaving human judgment necessary. The abstract reports no quantitative effect sizes, accuracy rates, or time savings. The paper concludes with proposed best practices for AI-assisted PAP auditing by authors and reviewers.
Does AI assistance build or erode professional expertise? This pre-registered three-month randomized controlled trial gave 133 practicing patent lawyers at eleven U.S. intellectual property firms access to a custom AI drafting assistant, with all work scored by blinded expert patent attorneys. AI access raised benchmark drafting quality by 0.34 SD at 10 days (p=0.03) and 0.38 SD at 90 days (p=0.01). On an unassisted redlining task after three months, treated lawyers outperformed controls by 0.32 SD (p=0.04), but gains were concentrated among senior lawyers (0.45 SD, p=0.02); junior lawyers showed no average gain and bifurcated scores.
How did US family caregivers' digital health engagement change around the COVID-19 pandemic? This cross-sectional trend analysis pooled Health Information National Trends Survey data (HINTS 5 Cycles 3 and 4, HINTS 6) from 2019 to 2022, covering 1676 family caregivers, with weighted multivariable logistic regression. Access to caregivers' own online medical records rose from 48.7% to 72.6% (P<.001) and access to care recipients' records from 30.8% to 44.5% (P<.001); sharing health information on social media grew from 22.5% to 39.1%. High-speed internet was strongly associated with engagement (sharing health information: OR 3.98, 95% CI 2.15-7.35).
What helps or hinders consumers using generative AI tools to seek health information? This scoping review followed JBI guidance and PRISMA-ScR/PRISMA-S, searching 10 databases for English-language empirical studies published from 2022 through a final search on January 8, 2026, and included 27 studies covering symptom appraisal, condition understanding, treatment options, and care navigation. Facilitators centered on comprehensibility and presentation quality (11 studies, 40.7%) and efficiency and access (8, 29.6%); barriers were dominated by credibility and trust concerns (13, 48.1%), especially absent or unclear citations, followed by perceived unsuitability for complex or urgent situations and privacy concerns (4, 14.8%).
What multidisciplinary factors shape adoption of digital health technologies such as patient portals, mobile apps and EHRs? This PROSPERO-registered systematic review (CRD420251056883) followed PRISMA 2020, searching multidisciplinary databases via EBSCO Discovery Service for English-language primary studies published between 2015 and June 2025, retaining 82 studies published from 2020 to 2025. Two reviewers screened and extracted using the SPIDER framework mapped to PICOS, appraised quality with the MMAT, and synthesised findings thematically in ATLAS.ti. Five themes emerged: access, equity and affordability; usability, engagement and empowerment; trust, privacy and governance; integration, workforce and sustainability; and clinical effectiveness and quality of care. The abstract reports no effect sizes.
Does the modality of outpatient mental health care — video, phone, or in-person — affect clinical outcomes? This retrospective comparative effectiveness study used VA administrative data for 813,699 patients completing at least 3 outpatient mental health visits from July 2021 to October 2022, with one-year follow-up and inverse probability-weighted regression adjustment. Mental health hospitalization occurred in 0.9% of the video group, 1.6% of the phone group, and 2.1% of the in-person group; average treatment effects favored video by 0.005 points versus both comparators. Appointment completion was 4.1 percentage points higher than phone and 3.6 higher than in-person. Authors note small effect magnitudes and possible residual confounding.
Does AI support for total parenteral nutrition in neonatal intensive care improve efficacy, safety, and integration? This PRISMA 2020 systematic review searched PubMed/MEDLINE, Scopus, and Web of Science, appraising evidence with RoB 2, the Newcastle-Ottawa Scale, and GRADE. Sixteen records met criteria; thirteen quantitative primary studies (2008-2026, n = 30-9,330) formed the synthesis. CPOE implementation cut parenteral nutrition medication error rates from 10.8% to 3.2%; the TPN2.0 transformer model correlated with expert decisions at Pearson R = 0.94, and classical machine learning reached R2 > 0.70 for macronutrient prediction. Only one-third of U.S. NICUs used a CDSS. GRADE certainty was moderate for efficacy and safety, low for system integration.
How reliable are the data produced by mandated inpatient screening for health-related social needs? This observational study used EHR data from 197,305 adult inpatient encounters across 18 Atrium Health hospitals in the southeastern US, May–December 2024, applying a data quality framework plus mixed-effects logistic regression. Screens were completed for 172,519 encounters (87.4%), though completion ranged from 92.1% to 72.4% by market; 12.7% (21,900) screened positive for at least one need. Positive patients more often had Medicaid (23.1% vs 13.7%) and higher 30-day readmission (14.3% vs 11%, p<0.001). Intradomain correlation was moderate to strong (food insecurity Cramer V=0.87), interdomain weaker (0.44).
Can chart review be automated with a large language model without exposing protected health information? This paper describes development of a no-code application built on an in-house ChatGPT 4o-mini deployment within Emplify Health's firewalled Azure cloud, which imports tabular data from Excel or text files, guides users through prompt engineering, submits records row-by-row, and tabulates extracted parameters or summaries. Initially built to identify cancer cases in pathology reports, it was generalized to cancer registries, cardiology CT reports, imaging narratives, and clinical notes. The authors report accelerated data retrieval but the abstract provides no accuracy metrics, time savings, or comparison group.