Can machine learning identify point-of-care ultrasound (POCUS) in free-text notes and gauge whether standardized documentation templates improve charge capture? This retrospective operational cohort study analyzed 559,029 encounters from 109,776 patients across 11 OBGYN clinic sites at one academic medical center (January 2018–August 2024), training LightGBM and BioClinBERT classifiers against manual CPT assignments and comparing periods before and after a February 2023 ProcDoc smart form. BioClinBERT reached 0.97 accuracy (F1 0.55–0.63). ProcDoc adoption hit 75.1% at 12 months; billing recapture fell from 10.0% to 2.4% (OR 0.22, 95% CI 0.17–0.30), with overall POCUS billing up 0.6%.
Issue No. 003
Two papers deserve immediate attention. A quasi-experimental study in Management Science finds that an EHR vendor's covertly biased decision support function raised extended-release opioid prescribing and left prescribers durably changed after the nudge was removed, with about 54% of the effect attributed to decision distortion rather than learning — the strongest case yet for auditing CDS logic as a safety and governance matter. Alongside it, Stanford's account of deploying ChatEHR offers rare organizational detail on putting an LLM chart-querying interface into live clinical use. For evaluation, a narrative review proposes separating performance monitoring, clinical impact monitoring, and causal evidence generation, with escalation triggers for when signals warrant trials. Also notable: interviews with 35 primary care physicians expose accuracy and completeness gaps in federally mandated real-time benefit tools, and a pragmatic trial shows pended orders plus outreach lifted lung cancer rescreening 27.7 points.
Can an online peer coaching program reduce physicians' after-hours administrative work? This voluntary longitudinal survey study followed 280 physicians who completed the Charting Champions Program between 2020 and 2023, using a 14-item Likert survey at program entry and 30-90 days after completion. Respondents reported significant decreases in hours spent charting (P<0.0001) and completing paperwork outside clinical hours (P<0.006), along with less work-related dread, burnout, and thoughts of quitting (all P<0.001), and greater focus, control, and mental energy. Patients seen per clinical day was unchanged (P>0.918). The abstract reports p-values only, with no point estimates or effect sizes, and includes no control group.
Can EHR analyses adjust for how much documentation a patient accumulates, separate from how sick they are? The authors built the EHR Density Index (EDI) using records from 24,987 adult patients at UNC Health (2018-2024), characterizing quantity, depth, and breadth of data per patient-year by assigning a utilization cluster via a Gaussian Mixture Model and computing within-cluster residuals across four clinical domains. Four interpretable clusters emerged; the Charlson Comorbidity Index predicted cluster membership but was only weakly associated with within-cluster residuals, suggesting the EDI captures record dimensions distinct from disease burden. The abstract reports no effect sizes.</summary>}Wait — I must not include stray text. Let me re-emit cleanly.{
How are large language models currently being used to support teamwork and communication within healthcare teams? This scoping review followed PRISMA-ScR guidelines, searching PubMed, Web of Science, and ScienceDirect for 2014-2024 publications; 3,865 unique titles and abstracts were screened, 127 full texts reviewed, and 20 studies included. Designs were predominantly quantitative and simulation-based, with limited in situ evaluation. Use cases spanned decision support, communication, and administrative functions. Outcomes centered on accuracy and quality (15/20 studies), with fewer addressing safety (4/20), readability or empathy (5/20), workflow efficiency (3/20), and error modes (2/20). The authors note most studies evaluate model performance without accounting for human team dynamics.
How should clinical predictive AI be evaluated when conventional randomized controlled trials are static, slow, and mismatched to models that drift and get updated? This narrative review argues for adaptive, iterative, context-specific assessment and proposes a framework separating three activities: performance monitoring (calibration, discrimination, data drift, alert burden, fairness, workflow fidelity), clinical impact monitoring of sustained benefit, and causal evidence generation via pragmatic and adaptive platform trial designs. The authors add a governance-driven escalation protocol specifying when monitoring signals should trigger formal trials, a signal-to-design decision pathway, and a guide to causal inference methods. As a review, it reports no effect sizes.
How does an LLM-based conversational interface to the electronic health record fare in real clinical deployment? This Nature Medicine piece reports on implementation lessons from ChatEHR at Stanford Medicine, covering the practical, technical and organizational considerations of putting a chart-querying system into use in an academic health system. No abstract was available, so findings, evaluation methods and any performance or utilization figures cannot be summarized here.
How automated are electronic early warning/track-and-trigger systems (EW/TTS) in practice? This PRISMA-guided systematic review searched PubMed, Web of Science, and Scopus for real-world clinical implementations published January 2010 to December 2025, screening 1181 records and including 43 studies, with quality appraised using the Joanna Briggs Institute checklist. Clinical deterioration was the primary objective in 54.5% (24/44) of stated aims; vital signs and assessment scores made up 62.7% (42/67) of clinical indexes. Automation was measured in 41.9% (18/43), predictive algorithms in 25.6% (11/43), and interoperable connectivity in 69.8% (30/43). Reported outcomes included earlier warning (20%) and lower specificity (12.9%).
Can patient-centered outreach improve adherence to annual lung cancer screening? A pragmatic 2×2 factorial randomized trial at Kaiser Permanente Washington enrolled 1837 patients with normal low-dose CT findings (November 2022–April 2024; follow-up through July 2025), assigning usual care, health communication (print/video messaging), stepped reminders (pended LDCT orders for primary care physicians plus patient scheduling outreach), or both. Stepped reminders raised 9-to-15-month rescreening 27.7 percentage points (75.5% vs 47.4%; RR 1.59, 95% CI 1.47-1.72), with larger gains among current tobacco users (risk difference 32.3 vs 24.1 points). Health communication was 4.7 points lower (59.2% vs 63.3%; RR 0.93).
How do primary care providers experience federally mandated real-time benefit tools (RTBTs) that display medication out-of-pocket costs in the EHR? Researchers conducted a qualitative descriptive study with semi-structured interviews of 35 PCPs at primary care clinics affiliated with two academic health systems sharing one EHR, using thematic analysis. Most participants were physicians (25/35), female (23/35), and had at least 10 years' experience (19/35). Three themes emerged: minimal RTBT training with openness to more; perceived potential to support cost conversations and reduce administrative burden; and pitfalls including incomplete information, inaccurate cost estimates, and clinically inappropriate lower-cost suggestions. The abstract reports no effect sizes.
Can a manipulated clinical decision support tool durably change prescribing? This quasi-experimental study compares physicians using an EHR whose vendor secretly embedded a biased CDS function promoting extended-release opioids between 2016 and spring 2019 against a control group of physicians who adopted other federally certified vendors in 2011. Affected physicians increased opioid claims during the treatment window and sustained a higher propensity to prescribe after the function was removed, persisting through relocation, affiliation changes, and stricter state opioid regulations; greater physician awareness attenuated the effect. Machine-learning estimates attribute roughly 54% of the treatment effect to decision-making distortion rather than learning. The abstract reports no other effect sizes.
How do clinicians judge the appropriateness, risks, and benefits of emoji in clinical messaging? This qualitative study combined focus groups and a survey from August to October 2025 with 29 clinicians at a large academic health system, spanning four specialties and including physicians, advanced practice providers, genetic counselors, medical students, and other healthcare workers, analyzed using rapid qualitative analysis. Participants described emoji as clarifying tone, building rapport, softening directives, and reducing notification fatigue, but raised concerns about ambiguity, informality, and medicolegal risk from message discoverability. Appropriateness hinged on hierarchy, familiarity, clinical gravity, generation, and platform. They preferred onboarding conversations and curated emoji sets over prescriptive guidelines. The abstract reports no effect sizes.
How should computable phenotypes be built and reported for embedded pragmatic clinical trials that rely on EHR and claims data? The EHR Core Working Group of the NIH Pragmatic Trials Collaboratory compiled investigator experiences developing, adapting, and applying phenotype definitions, presented as four case studies of different approaches. The authors recommend building phenotypes with multidisciplinary teams, validating them appropriately, disseminating details of phenotype construction, and capturing and reporting modifications made during trial conduct. They emphasize teams that understand why the data were collected and potential sources of bias. This is a methods and experience report; the abstract provides no quantitative results or effect sizes.
What do virtual hospital services (VHS) for working-age adults actually look like in the published literature? This scoping review followed JBI methodology and PRISMA-ScR, searching four databases for articles published March 2021 to July 2025 on VHS and hybrid hospital-in-the-home models for adults aged 18-65. Of 1624 records, 28 studies met eligibility: 16 described fully virtual models and 12 hybrid HITH. Most combined synchronous and asynchronous communication; mobile apps and wearables were uncommon. Respiratory conditions predominated (17/28), heart failure exacerbation was the most common specific condition (6/28), and patient satisfaction or experience was the most reported outcome (17/28). The authors note heterogeneous terminology, inconsistent reporting, and dominance of pilot and single-site studies.
What determines successful implementation of an in-house dosimetry quality assurance checklist in radiation oncology? This qualitative implementation study at an academic medical center used semi-structured interviews, field observations, and surveys with dosimetrists, physicists, trainees, and developers across pre-implementation, implementation, and post-implementation phases, coded abductively using an adapted CFIR mapped to UTAUT and the ERIC strategy compilation. Four CFIR constructs and 12 sub-constructs emerged as barriers (structural characteristics, planning most negative); five constructs and 19 sub-constructs as facilitators (relative advantage, culture, leadership engagement). Suggestions mapped to 19 ERIC strategies; CFIR-ERIC matching identified 14. Acceptability, appropriateness, and feasibility improved significantly (p<0.05), but adoption reached 100% only in week six.