Can standardized care-team protocols shift patient portal messages away from primary care clinicians? This quasi-experimental pre/post study evaluated a cost-neutral quality improvement initiative embedding EHR-based protocols in medical assistant and nurse workflows for patient medical advice requests across 11 primary care clinics in a large academic health system, using multivariable mixed-effects regression adjusting for clustering by clinician. The share of requests routed to a PCP fell from 61.6% to 57.6% (p<0.001; adjusted OR 0.83, 95% CI 0.82-0.84), a 4.6% reduction in adjusted mean percentage. PCPs still received more than half of all patient messages.
Issue No. 006
The headline result is a pragmatic randomized trial of Epic's ambulatory chart summarization tool: task load, burnout, and exhaustion improved modestly, yet charting time was unchanged, only 14.2% of 74,474 summaries were ever opened, and net promoter score was -22 — a sober benchmark for anyone forecasting generative AI value. Two papers puncture assurance assumptions: a clinician-built, AI-assisted ePROM app scored ~92 on SUS while harboring scoring errors, contrast failures, and an analytics tag contradicting its privacy claims, and a systematic review of multimodal deep learning found external validation in just 13% of studies and 87% at high risk of bias. On the policy side, machine learning applied to 31,000-plus 510(k) submissions promises large recall and workload reductions, while NLP on Medicaid care coordination notes quantifies administrative burden in patient-hours and dollars.
Can a vendor-derived "order friction" metric guide targeted redesign of pediatric medication ordering? This quasi-experimental pre-post analysis used Epic-supplied order friction data on all inpatient medication orders across a three-hospital pediatric system during 2024. A multidisciplinary workgroup built a hospital medicine preference list prepopulating dosing, frequency, and as-needed indications for 15 medications (45 variants, 25 orderables). Among 22 paired orderables, median changes per order fell from 3.20 (IQR 2.41-3.68) to 0.39 (IQR 0.18-1.04; p < 0.001). System-level volume-weighted changes per order fell from 3.06 to 2.72 as preference-list adoption rose from 9.6% to 15.8%. The design was uncontrolled.
What makes patients willing to accept ambient AI scribes during ambulatory visits, and does baseline trust in AI shape that acceptance? This secondary, convergent mixed-methods analysis re-analyzed survey and interview data from 20 patients seen after ambient AI scribe implementation, summarizing trust items descriptively and coding interviews deductively against the Theoretical Framework of Acceptability plus inductively. Trust ranged from low to high, with 60% reporting moderate trust. Acceptability tracked with minimal ethicality concerns, supportive affective attitudes, low patient burden, perceived benefits, and strong intervention coherence; themes differed minimally by trust level. Patients urged patient education and advance notice. The abstract reports no effect sizes.
How mature is the evidence behind "data-centric" multimodal deep learning for clinical decision support? This PRISMA 2020 systematic review screened 150 records and included 31 primary clinical studies, 30 (97%) published 2024-2026, coding implemented versus merely mentioned techniques and appraising bias with PROBAST+AI. Studies used a median of three modalities (range 2-6), most often structured EHR (71%) and imaging (39%). Data-centric techniques were reported in 74-84% of studies (equity 61%), but external validation appeared in only 4/31 (13%), a clinical or provider outcome in 3/31 (10%), none reported deployment, and 27/31 (87%) were at high risk of bias.
Does high usability in a clinician-built, AI-assisted application establish clinical and technical assurance? This single-case retrospective development-and-assurance report examined STUIapp, a browser-based ePROM tool integrating six validated lower urinary tract symptom instruments, using code audit, independent clinical review of a frozen 78-case scoring matrix, WCAG 2.1 measurement, a 23-canary persistent-storage study, and usability testing with 14 clinicians, 26 patients and 12 older adults. Mean SUS was 92.3 (SD 8.8) among clinicians and 92.0 (SD 10.8) among patients, yet all 78 passing automated cases included 17 expected results requiring correction. Other findings: instrument mislabelling, a failed installability manifest, a third-party analytics tag contradicting local-only privacy claims, 10-px text and 2.56:1 contrast, and storage permission denied in 10/10 browser-tab canaries versus granted in 13/13 installed canaries.
What predicts publicly visible AI adoption at US cancer centers? This cross-sectional study assembled public-source data on 75 NCI-designated cancer centers, scoring adoption across screening, treatment, and patient care as a 0-3 composite index, with Moran I tests for spatial clustering and ordered logistic regression on institutional and contextual predictors. The mean adoption index was 1.37 (SD 0.86), highest for screening (0.86), then patient care (0.50) and treatment (0.22). Moran I showed no significant spatial autocorrelation. Physician workforce and bed capacity showed positive but modest associations; state socioeconomic indicators did not. Political-context findings were mixed; the abstract reports no effect sizes for regression estimates.
How are graduate medical education trainees and faculty using generative AI? This cross-sectional survey study, reported in the Journal of General Internal Medicine, examines self-reported generative AI use among GME trainees and faculty. No abstract was available, so findings, sample size, and effect estimates cannot be summarized here.
What features of AI-generated literature reviews make them useful and trustworthy to oncologists? In a randomized mixed-methods study, 34 oncology physicians produced 294 ratings of four blinded AI systems across five clinical vignettes, supplemented by 20 semi-structured interviews analyzed with a prespecified LLM-assisted qualitative pipeline. Despite similar references, an evidence-graded report adapted from OpenEvidence scored significantly lower in overall utility than standard OpenEvidence (mean difference -0.96; 95% CI -1.26 to -0.66; P<.001). Qualitative analysis yielded six themes and seven design requirements: clinicians preferred concise, scannable reports with quantitative outcomes, bolded guidelines, explicit uncertainty, and verifiable citations; trust fell with citation mismatch, buried provenance, and overconfident recommendations.
Does an EHR-embedded generative AI chart summarization tool reduce clinician workload? In a pragmatic 1:1 randomized trial at one academic health system, 284 ambulatory clinicians across 42 specialties received Epic's outpatient chart summarization tool or usual care over 90 days (February o task task load favored the intervention ( tionsted difference -27.4 on a 0-400 scale; 95% CI, -49.4 to -5.3; P=0.02), with lower burnout (-0.20) and work exhaustion (-0.24) on the Professional Fulfillment Index. Charting time per encounter was unchanged (-1.2 seconds). Only 14.2% of 74,474 summaries were opened, falling from 21.5% to 10.5% by month 3; net promoter score was -22.
Can machine learning help the FDA cut recalls and review workload in the 510(k) substantial-equivalence pathway? The authors trained recall-risk models on submission-time information and embedded them in a data-driven policy recommending acceptance, rejection, or deferral to FDA committees for in-depth review, using an assembled data set of more than 31,000 submissions drawn from FDA and CMS sources. Against current practice (10.3% recall rate, workload normalized to 100%), a conservative evaluation showed a 32.9% improvement in recall rate and a 40.5% workload reduction, with estimated annual savings of roughly $1.7 billion from avoided replacement costs, about 1.1% of US medical device spending.
Can digital health platforms improve patient-physician matching once geography no longer binds? Using nationwide Swedish online care with time-conditional random assignment of patients to physicians, the author estimates reallocation gains from aligning provider heterogeneity with patient needs. Matching high-risk patients to doctors effective at averting emergency room use lowers ER visits by 4.4 percent (SE 1.3), and reallocation reduces counter-guideline antibiotic prescribing by 3.1 percent (SE 1.4). Trade-offs across outcomes were limited, as horizontal differentiation among doctors and varied patient needs permitted simultaneous improvement; efficiency-enhancing reallocations also carried equity consequences.
Why does rating inflation in online reputation systems rise and then fall as reviewers gain tenure? This exploratory mixed-method study of reviewers in a digital platform community combines qualitative and quantitative analysis to trace how inflation behavior evolves with socialization. The authors identify three archetypical phases: Newcomers, not yet socialized, are less likely to inflate; Inflators, following direct reviewer reciprocity, inflate; and Veterans, motivated by generalized reciprocity and commitment to the community, rate more candidly. Rating inflation thus follows an inverted U-shape over the reviewer lifecycle. The abstract reports no effect sizes, sample size, or study years.
Does a model's ability to automate a task predict its ability to help a weaker agent do it? The authors build a benchmark spanning seven economically grounded real-world tasks in which an assistant model writes guidance for a standardized lower-capacity worker model that produces the deliverable, compared against automation mode where the assistant works directly; outputs are scored by blind pairwise comparisons from an LLM judge panel with task-specific rubrics across ten replications. Rankings under the two regimes correlate only modestly, the automation winner loses on augmentation in five of seven tasks, the unaided worker outranks every assisted condition on three tasks, and only one model's guidance beats no guidance on average.
Which jobs and tasks actually use generative AI at work? The authors field a nationally representative survey linking genAI adoption to detailed occupations and tasks, producing the first task-level adoption indexes. Occupational exposure scores explain some but not all variation in adoption across occupations and tasks, and the survey-based indexes differ conceptually from platform chat-log measures, which the authors argue over-classify chats into generic activities spanning many occupations. Adoption is widespread but shallow: within most occupations and tasks, fewer than half of workers adopt. The abstract reports no other effect sizes.
Does firm-level AI investment translate into productivity growth, and through what channel? Using a new firm-level measure of AI investment built from AI-skilled employment—spanning machine learning through generative and agentic AI—the authors relate AI investment to productivity across firms, and construct a measure of organization capital derived from workers' job descriptions. AI investment is associated with productivity growth in recent years but not over the prior decade, with gains concentrated in AI-skilled jobs that build organization capital, that is, durable firm-specific knowledge from learning-by-doing. The abstract reports no effect sizes or sample details.
Can adaptive digital outreach improve statin refills among patients with recent nonadherence? This pragmatic, health system–embedded sequential multiple assignment randomized trial (SMART, PROBE design) at Kaiser Permanente Northern California randomized 20,604 adults (mean age 54.8 years; 40.9% female; 15.3% with established ASCVD) with ASCVD or high risk to portal messaging, SMS, nonsecure email, or usual communication, with second-stage randomization of 14-day nonresponders. Initial outreach raised 14-day refill from 12.0% to 14.4% (adjusted risk difference, 2.3 percentage points; risk ratio, 1.20). A second outreach added gains among nonresponders (13.0% vs 10.4%), while switching modality did not (risk ratio, 1.06).
Has post-pandemic telehealth use declined, and have sociodemographic gaps in access closed? This repeated cross-sectional analysis used the nationally representative 2022 and 2024 Health Information National Trends Survey (11,386 US adults; mean age 48.6 years, 51.1% female), with logistic regression adjusting for sociodemographic, clinical, and access variables and year interactions. Unadjusted telehealth use fell from 39.0% (95% CI 36.9–41.3%) to 34.9% (32.1–37.9%), while the video share of visits held at 71.1% (68.8–73.4%). Use tracked age, internet use, income, and large-metro residence; gender and insurance differences narrowed. About 20% reported technical problems, and over 75% rated telehealth comparable to in-person care.
Why do users of eHealth behavioral interventions taper off over time? The authors extend Expectation-Confirmation Theory by treating engagement as a dynamic learning process, estimating a hierarchical Bayesian structural learning model of how users update perceptions of intervention effectiveness from ongoing experience and how those beliefs drive continued participation. Learning performance was lower for interventions with ambiguous instructions and those targeting short-term health outcomes, which generate noisier feedback and less accurate effectiveness perceptions, associated with reduced sustained engagement. Several denoising design strategies are evaluated in counterfactual simulations. The abstract reports no effect sizes, sample size, or study period.
Can a short informational video sent before a clinical visit raise lung cancer screening uptake? In a randomized feasibility trial at Kaiser Permanente Colorado (March–October 2025), 1,093 screening-eligible patients with upcoming primary care or pulmonology appointments were assigned by birth month to a text-delivered video nudge (with or without a rooming QR code; n=549) or usual care (n=544). Intervention patients more often received a screening order within one day (22.6% vs 16.4%; p=.010) and during follow-up (32.6% vs 24.1%; p=.002). Baseline LDCT completion was 8.6% vs 5.7% (p=.078). Seventeen percent viewed the video, watching 79% on average.
Can a consolidated, human factors-informed guideline improve how clinical decision support is designed in practice? Using a two-phase explanatory sequential mixed methods design, the authors surveyed 25 guideline users on usefulness, ease of use, satisfaction, and influence on decision-making, then conducted 10 semi-structured interviews analyzed thematically with a general inductive approach. Respondents described the guideline as easy to use, reported greater confidence selecting and designing CDS interventions, and valued its consolidated format and step-by-step structure; requested additions included practical tools, worked examples, and case studies. The abstract reports no quantitative effect sizes or survey score magnitudes.
How far has machine learning–based pharmacogenetic clinical decision support actually moved into clinical workflows? This scoping review searched multiple databases for studies published January 2015 through September 2025 reporting ML-based CDSS incorporating pharmacogenetic data to support therapeutic decisions in clinical settings. Of 1,262 records screened, 7 met inclusion criteria, and only 2 evaluated tools in live clinical or trial workflows. Studies varied in design, setting, and implementation maturity; most reported potential benefits including fewer preventable adverse drug events, better prescribing accuracy, or improved workflow integration, though the abstract reports no effect sizes. Authors call for implementation science evaluating usability and patient outcomes.
No abstract was available for this report, which describes an electronic clinical pathway intended to improve diagnosis and management of obesity hypoventilation syndrome in hospitalized patients. Published in the Journal of General Internal Medicine, the title indicates a clinical decision support or pathway implementation focus on an underrecognized condition in inpatient care; no methods, population, or outcome data can be summarized here.
How often do administrative burdens surface in Medicaid care coordination, and what do they cost patients in time? A retrospective cohort study applied natural language processing classifiers to encounter notes from a community-based care coordination program in Washington, Virginia, and Ohio (January 2023-November 2025), covering 142,473 beneficiaries, 49,282 (34.6%) with at least one encounter. Paperwork was most prevalent (25.3%), followed by scheduling (16.2%), prior authorization (9.8%), and transportation (6.1%). Mean per-patient time cost at the clinician-equivalent rate was highest for transportation ($47.58) and lowest for scheduling ($12.35); documented burdens totaled 18,822 patient-hours ($628,665). African American beneficiaries had higher unadjusted burden prevalence than White beneficiaries (rate ratio, 1.22), attributable to plan enrollment rather than within-plan differences.