How do European general practitioners experience digital health technologies in daily practice? This systematic review (PRISMA 2020, Synthesis Without Meta-Analysis) searched eight databases plus Elicit AI for studies published January 2014 to October 2025, yielding 65 studies covering 24,994 GPs in 14 countries, 93.8% (61/65) from Western and Northern Europe. Perceived usefulness was positive in 86.2% (56/65) of studies, while ease of use was positive in none (0/65) and negative in 66.2% (43/65); actual use was near universal (63/65, 96.9%). Barriers included poor interoperability, limited training, and unreimbursed digital workload, alongside concerns about digital exclusion.Voosh.aiVoosh.aiVoosh.aiVoosh.aiVoosh.ai Normalization process theory constructs showed only partial integration.Voosh.ai Ratings.Voosh.ai bias risk used MMAT and JBI tools.
Issue No. 004
Ambient documentation dominates, and the news is that the tools are not interchangeable: across 163 primary care providers and 59,130 provider-days, one product cut after-hours EHR time while another increased it, and both slowed timely visit closure. A companion safety-net analysis of 54,160 encounters found non-English visits 21%-25% less likely to meet the AI performance threshold regardless of interpreter modality, a concrete equity signal for anyone buying scribes. On the model-behavior side, a sweep of 107 LLMs across 35 clinical tasks found stigmatizing language in 84% of model-task pairs, correlated with lower accuracy but reducible up to 92% by prompt engineering. Also worth your time: nine patient-facing AI products matched on triage accuracy but branded ones over-triaged and steered patients to affiliated paid services, and real-time prescription benefit data lifted fills 14.5 points only for the costliest drugs.
What drives whether a patient shows up (informative presence) versus whether a lab is actually ordered at that visit (informative observation)? The authors model the two stages separately using a stochastic recurrent-event model for outpatient visits and a visit-weighted GEE for per-visit biomarker recording, across three EHR-linked cohorts: All of Us (n=599,423), Yale New Haven Health System (n=319,666), and Michigan Genomics Initiative (n=82,372), covering 68 biomarkers with detailed analysis of ten. Median outpatient visits ranged from 1.7 to 6.1 per year over 4.4-7.2 years of follow-up, while the median within-person share of visits containing a given biomarker ranged from 0.4% to 19.5%. Chronic disease burden and recent visits predicted higher visit rates consistently; race, ethnicity, and neighborhood income associations varied by cohort, and covariate effects on measurement varied by biomarker (prior cancer predicted more blood counts, fewer lipids).
Do ambient documentation tools differ from one another in real-world impact? This comparative-effectiveness study followed 163 primary care providers in a large integrated health system (January 2024–June 2025), analyzing 59 130 provider-days across a tablet-based virtual human-assisted tool (A, n=65), an EHR-integrated tool (B, n=68), and a standalone tool (C, n=17), using intention-to-treat and per-protocol models with provider-clustered SEs and month fixed effects. Versus Tool B, Tool A showed more after-hours EHR time (+0.022 hours/provider-day), more manual note composition (+0.046), and lower 48-hour visit closure (−0.120; 95% CI −0.126 to −0.115). Tool C cut after-hours work by 0.055 hours/provider-day (~30 hours annually) but also reduced timely closure (−0.028).
Can automated judges that flag documentation errors serve as a defensible measurement instrument for ambient AI scribes, absent a gold standard? This pre-registered, blinded validation study nested in a multi-country ambient documentation simulation (English setting) had ten external clinicians adjudicate a stratified sample of 434 pipeline flags, yielding 565 adjudications with 131 double-rated. Inter-clinician agreement on error genuineness was fair (raw 59%, AC1 0.24); judge-clinician agreement was 64% (95% CI 60-68). Behavior was near-symmetric across arms (kept-precision 74% AI vs 81% clinician notes; severity gap +0.06 vs -0.09 tiers), with removed-confirmed asymmetric (56% vs 42%). Latent-class triangulation estimated a 68% genuine-error rate (94% CrI 48-83).
Do ambient AI scribes perform equally well across patient languages? This retrospective analysis covered 54,160 outpatient encounters in a U.S. safety net health system, measuring the share of words in the final note generated by the AI tool and left unedited by the provider. Generalized estimating equations with exchangeable correlation structures accounted for clustering within patients, with univariable models by language and interpreter modality and a multivariable interaction model. Non-English encounters were 21% to 25% less likely than English encounters to reach the performance threshold; interpreter-mediated and bilingual-provider encounters did not differ significantly.
How should health systems evaluate generative AI that turns clinical information into patient-facing instructions such as discharge summaries, medication explanations and portal messages? This narrative and interpretative review maps empirical studies of AI-generated discharge communication alongside health-literacy, medication-safety, patient-safety and AI-governance literature. The author reports a recurring trade-off: large language models improve readability and understandability, while physician and pharmacist review identifies omissions, inaccuracies, newly introduced actions, medication-related problems and potentially harmful content, especially in complex discharges. A proposed seven-domain framework covers factual accuracy, clinical completeness, actionability, medication clarity, escalation and safety-netting, health-literacy alignment, and accountability with auditability. The abstract reports no effect sizes.
How is AI being applied to cancer symptom management for adult survivors? This scoping review followed Joanna Briggs Institute methodology, searching eight databases (November 2025, updated March 2026) for English-language empirical studies from 2015 onward, with narrative synthesis and inductive content analysis. Of 41 included studies, 21 addressed model development, 18 intervention delivery, and 2 both; common techniques were natural language processing, machine learning, and conversational AI. Applications targeted symptom detection (n=15), patient education (n=9), monitoring (n=5), and personalised management (n=5); pain and psychological distress each appeared in 9 studies. Model performance was moderate to high; the abstract reports no pooled effect sizes.
What actually helps or hinders AI-based clinical decision support once it reaches the bedside? This scoping review searched five databases (MEDLINE, Embase, CINAHL, APA PsycInfo, Cochrane) from inception to May 2022, screening 10,875 titles and abstracts and 494 full texts; 13 studies met inclusion criteria and 9 reported explicit implementation determinants, mapped to the Consolidated Framework for Implementation Research by two reviewers. Studies were mostly US-based, multicenter machine learning tools in critical care and emergency medicine. Reviewers identified 28 determinants: 16 barriers (limited interpretability, data quality, workflow misalignment, user capability and motivation) and 12 facilitators (end-user needs assessment, stakeholder engagement, peer endorsement, supporting evidence).
Do large language models produce stigmatizing language when reasoning over clinical data? Researchers applied a psychiatrist-validated NLP stigma-term detector to reasoning text from 107 LLMs across 35 real-world clinical tasks, covering 3,745 model-task pairs. Stigma rates ranged from 0% to 33.33%, and 84.06% of pairs contained at least one stigma term. Open-source models exceeded proprietary (1.97% vs. 1.60%; p<0.01) and reasoning models exceeded non-reasoning (2.35% vs. 1.70%; p<0.0001), while general versus medical models did not differ (2.00% vs. 1.80%; p=0.26). Stigma correlated negatively with task accuracy (r=-0.304) and positively with input-note stigma (r=0.569); 19.76% of pairs amplified input stigma. Prompt engineering cut stigma rates by up to 91.91% without reducing performance.
How should LLM failures be evaluated when they occur inside multi-step research and clinical workflows rather than isolated prompts? This method paper introduces TRACE, a practitioner-audit framework, with an empirical demonstration: 45 documentation-positive incidents logged by a single clinician-informatician across scholarly, clinical informatics, and clinical-adjacent workflows over seven weeks, coded with a consequence-based severity rubric and taxonomy crosswalk. Four categories tied as most frequent (verification failure, factual numerical error, tool-behavior misunderstanding, citation formatting; n=7 each), and one incident carried an estimated $2,500 impact. Category agreement was low among three human reviewers (Fleiss κ=0.155) but substantial among three vendor-blinded AI comparators (κ=0.632). The authors frame this as a pilot that does not estimate error rates.
Do patient-facing AI products differ in how they triage and refer simulated patients? Researchers tested nine products against 60 physician-developed standardized clinical cases across 540 multi-turn simulated patient encounters. Overall triage accuracy showed no statistically significant difference across product categories, but referral behavior diverged: branded health AI products over-triaged low-acuity cases far more often (28% vs 3% vs 2%) and more frequently recommended affiliated, fee-requiring clinical services. The authors argue evaluations of patient-facing medical AI should assess referral behavior alongside triage accuracy. The abstract reports no confidence intervals or other effect sizes.
What blocks health information exchange in substance use disorder care, as providers experience it? A qualitative study convened 11 focus groups (n=31) and 5 validation interviews (n=5) with behavioral health providers (52% prescribers) from 4 SUD treatment organizations across 14 US states, using HEDIS-based scenarios analyzed thematically and via Unified Modeling Language workflow diagrams. Incomplete data access at the point of care routinely forced manual exchange by fax, phone, and secure email, and confusion about HIPAA, 42 CFR Part 2, and state release requirements was ubiquitous. UML modeling of 4 care scenarios identified 3 shared data-sharing subprocesses. Providers prioritized interoperable consent management, HIE/PDMP-EHR integration, and harmonized privacy rules. No effect sizes are reported.
Which state factors explain how quickly Medicaid programs adopted telehealth policies? This legal mapping study (50-state survey) identified state laws and policies governing Medicaid telehealth delivery in effect through December 31, 2023, with two researchers independently abstracting policies and multivariable regressions testing state-level correlates of monthly policy counts from 2018 to 2023. States averaged 2.7 telehealth policies in effect (SD 1.4; range 0-6), with most adoption concentrated in 2020-2021; audio-only reimbursement spread fastest, from 0 to 43 states over five years. Adoption was unrelated to rurality, health professional shortage, Medicaid expansion, broadband availability, demographics, income, unemployment, or COVID-19 cases and deaths.
Why do online shoppers concentrate their search and purchases on top-ranked products — lower search costs at prominent positions, or beliefs that higher-ranked items offer better returns to search? The authors build an experimental paradigm and run incentivized experiments to separate the two mechanisms. Both are present, and short-term randomization of rankings alone does not disentangle them; ignoring beliefs yields biased search-cost estimates and incorrect consumer welfare predictions for alternative recommendation systems, including platform self-preferencing. The paper proposes approaches for recovering unbiased search costs in real search settings. The abstract reports no effect sizes or point estimates.
Do complex telehealth tasks themselves contribute to uptake inequities among socioeconomically marginalized patients? Researchers combined a complexity walkthrough inspection by eight researcher-evaluators of six telehealth tasks required at a Federally-Qualified Health Center with remote user testing of 24 FQHC patients, paired with complexity-focused interviews and surveys and mixed-methods analysis. Patients completed only 33.9% of required subtasks without issues, and took twice as long as walkthrough evaluators. Ambiguity (unclear inputs, unfamiliar concepts and words) and relationship dimensions (context switching, deep navigational hierarchies) most affected perceived difficulty; cognitive load was high for two tasks, and many patients abandoned tasks early.
What drives health systems and payers to adopt — or drop — patient-facing digital health tools? This qualitative study used semistructured interviews conducted August through December 2025 with nine senior leaders from a large Midwestern academic health system and affiliated payers, including a provider-owned health plan and a state Medicaid program, using an evidence-based mobile intervention for alcohol use disorder as the use case. Thematic analysis with inductive and deductive coding identified four decision-making mechanisms: prioritization under organizational constraint, risk mitigation, operational fit, and value determination. The authors argue these processes are largely invisible to patients, clinicians, and developers. No effect sizes are reported.
How well do patients stay in buprenorphine treatment when care is delivered entirely by telehealth? This retrospective cohort study followed 22,064 adults initiating opioid use disorder treatment in a multistate, telehealth-only addiction program between May 2019 and April 2025 (mean age 41.2 years, 48.6% female, 33.8% rural, 78.5% Medicaid), using chi-square tests and absolute risk differences. Retention fell from 86.8% at month 1 to 71.0% at month 6, then declined gradually through 24 months. At month 6, retention was 70.4% among patients with prior buprenorphine exposure versus 52.0% without (risk difference 18.5 percentage points, 95% CI 15.5-21.5).
How many U.S. adults turn to general-purpose large language models for their mental health? A cross-sectional survey of 1,871 U.S. adults fielded August–October 2024 used stratified sampling across age, sex, and race/ethnicity to approximate national demographics. Twenty-four percent reported using LLMs for mental health; users skewed young, male, and Black, and reported poorer mental health and difficulty accessing traditional treatment, citing that LLMs are free, convenient, and always available. Reported uses included emotional support, learning therapy skills, and supplementing existing therapy. Combining with Pew estimates of overall LLM use, the authors project 14–18 million U.S. adults.</summary>}</summary>
How do patients view and understand test results released immediately to them under information-blocking rules? This JAMA Network Open study examines patient access to and comprehension of immediately released test results. No abstract was available, so the study's design, population, and findings — including any magnitudes — cannot be summarized here.
Did 2025 federal immigration policy changes shift outpatient and telehealth use among likely undocumented patients? This cross-sectional study of electronic health record data from a public safety-net system compared 184 541 visits from January–June 2024 with 182 573 from the same months in 2025, proxying documentation status by non-English primary language without a Social Security number. Likely undocumented patients had 15% higher in-person visit completion overall (IRR, 1.15; 95% CI, 1.13-1.16), but month-specific declines of 5% to 12% in 2025, significant in March (IRR, 0.88) and June (IRR, 0.93). Their telehealth share rose 12% (IRR, 1.12) while falling 5% among patients likely with legal status; total visits were unchanged.
Can automated post-discharge check-ins surface adverse events in adults with multiple chronic conditions? This mixed-methods, user-centered design study used semi-structured interviews with 37 patients and 23 clinicians to set requirements, then field-tested a prototype in 20 patients for up to 7 days after discharge, drawing on interoperable EHR data to deliver symptom and patient-reported outcome questionnaires with risk-stratified advice and real-time clinician escalation. Patients completed 60% of questionnaires; 7 received Level 2 or 3 advice, 3 triggered Level 3 escalation emails, and 4 of the 7 had chart-confirmed emergency department visits within a week. Clinicians found PRO trends hard to interpret.
Can a large language model reliably categorize the content of patient portal messages at scale? Researchers built a zero-shot GPT-4o-mini pipeline applied to all medical advice request messages sent to ambulatory clinicians at an academic medical center in 2024-2025, using an 11-category taxonomy derived by expert panel via modified Delphi; two annotators labeled 750 messages (Cohen kappa 0.80), with 500 held out for evaluation. Micro- and macro-averaged F1 were 0.89 and 0.86, and labels were identical across runs for 93.6% of messages. Across 2.4 million messages, Problems & Management and Medications & Prescriptions appeared in 67.9%, the top four topics in 93.9%, and 51.7% spanned multiple topics.
How can an EHR-based intervention be sustained after a pragmatic trial ends? This qualitative case study of the NOHARM stepped-wedge cluster-randomized trial, which tested the Healing After Surgery educational bundle promoting non-pharmacological perioperative pain care across multiple surgical practices and hospital sites, reviewed minutes from biweekly sustainment committee meetings and conducted debriefings with committee members guided by the Clinical Sustainability Assessment Tool. Six themes emerged: automating low-touch components, reliance on internal champions, intentional handoff communication, institutional attention, relaxed fidelity, and continuous evaluation. The authors emphasize early structured planning and transition to clinical ownership; the abstract reports no effect sizes.
Does giving prescribers real-time prescription benefit (RTPB) information at the point of ordering improve whether patients actually fill prescriptions? This post hoc analysis of a cluster randomized trial in an urban ambulatory network randomized practices during January–December 2021, analyzing 38 289 RTPB-eligible orders of 1 386 577 outpatient prescriptions (2.8%). Overall fill rates were unchanged (54% control vs 55% RTPB; adjusted difference 1.2 percentage points; 95% CI, −1.3 to 3.7). In the highest out-of-pocket cost quartile (>$120.83 per 30-day fill), fills rose from 33% to 49% (14.5 pp; 95% CI, 8.4-20.6), with the largest gain in lowest-income communities (30.3 pp).