Does an LLM-generated hospital course draft reduce discharge summary documentation time? This observational pre-post study compared a pre-AI period (March 16-November 19, 2025) with a post-AI period (November 20, 2025-February 11, 2026) after a GPT-4.1 tool was embedded in mandatory discharge summary templates, covering 8,298 hospitalized adults on hospital medicine services. Edit time did not differ between periods (6.60 vs 6.28 minutes, p=0.11), but within the post-AI period, tool use was associated with longer editing (9.20 vs 4.93 minutes) and a 32% increase in adjusted analysis (95% CI 24-41%). Faculty review of 27 encounters found higher quality and lower harm but less concise summaries; 31 of 36 surveyed clinicians (86.1%) felt the tool improved efficiency.
Documentation burden & workload
Can large language models reliably generate or simplify outpatient clinic letters? This systematic review searched five databases (PubMed, EMBASE, Web of Science, CENTRAL, CINAHL) from inception to 1 November 2025, appraising studies with the Mixed Methods Appraisal Tool and GRADE. Seven studies were included, two using real-world clinic data and five synthetic or hypothetical scenarios. Readability findings were mixed and most AI output still exceeded the US Grade 6 level; information fidelity ranged from 10% to 100%, and one study reported a ten-fold reduction in drafting time. Certainty was very low across primary outcomes.
Are LLM-generated medical reports clinically ready in terms of effectiveness, safety, and workflow burden? This systematic review searched five databases (PubMed/MEDLINE, Embase, Web of Science, Scopus, Cochrane) for studies published January 2016 through May 2026 evaluating LLMs, multimodal LLMs, or vision-language models for image-to-report generation, impression drafting, or structured reporting, including 101 studies (36 chest x-ray). No study was at low risk of bias (15 moderate, 72 high, 14 serious), and meta-analysis was not possible. One chest x-ray study found AI report acceptance of 70.5% (6047/8580) versus 73.3% for radiologists, with false negatives 18.5% versus 17.8%; a brain MRI study showed reading time falling from 61 to 53 seconds while impression drafting increased editing time and edit distance.
Can physician social media posts surface usable signals about clinical decision support failures? This qualitative study searched X in August 2025 using Grok 4 via an internal API with keywords such as "EHR alert fatigue" and "override EHR," curating 117 posts from 106 self-identified physicians (2016–2025) and analyzing one post per author with Braun and Clarke reflexive thematic analysis plus AI-assisted coding verification. Six usability failure domains emerged, led by burnout and fatigue (43.4%) and alert overwhelm (38.7%); lack of customization was least common (7.5%). Sentiment was 76.4% negative, and half of patient-safety posts co-occurred with alert overwhelm.
Can pairing a curriculum with a structural inbox-coverage system improve inter-visit care in residency training? This multimethod pre-post evaluation combined a retrospective survey of educational outcomes with chart review of matched resident cohorts among 32 internal medicine residents in an academic continuity clinic during the 2023-2024 academic year. The intervention added firm-based inbox coverage with faculty supervision plus a longitudinal curriculum on EHR efficiency, inter-visit clinical reasoning, and professional responsibility. Clinic-wide test results communicated rose from 57% to 76% (p<0.01), and mean time to communication fell from 18.7 to 8.5 days (p<0.01). Resident-reported confidence improved, without reported effect sizes; resident-level changes varied by baseline tertile and training level.
How widely has ambient AI reached inpatient nursing, and what do nurses expect from it? This cross-sectional online survey recruited practicing U.S. inpatient nurses through nursing-focused social media communities and used descriptive analyses of adoption, experience, and perceptions. Among 61 eligible respondents, nearly 60% (n=37) said ambient AI had been piloted or implemented at their institution, and 49.2% (n=30) had used it clinically; 73.3% of those (n=22) reported positive experiences. Cited benefits included reduced documentation time, improved workflow, and better data quality; concerns included patient acceptance, job displacement, and skill loss from overreliance.
Can standardized care-team protocols shift patient portal messages away from primary care clinicians? This quasi-experimental pre/post study evaluated a cost-neutral quality improvement initiative embedding EHR-based protocols in medical assistant and nurse workflows for patient medical advice requests across 11 primary care clinics in a large academic health system, using multivariable mixed-effects regression adjusting for clustering by clinician. The share of requests routed to a PCP fell from 61.6% to 57.6% (p<0.001; adjusted OR 0.83, 95% CI 0.82-0.84), a 4.6% reduction in adjusted mean percentage. PCPs still received more than half of all patient messages.
Can a vendor-derived "order friction" metric guide targeted redesign of pediatric medication ordering? This quasi-experimental pre-post analysis used Epic-supplied order friction data on all inpatient medication orders across a three-hospital pediatric system during 2024. A multidisciplinary workgroup built a hospital medicine preference list prepopulating dosing, frequency, and as-needed indications for 15 medications (45 variants, 25 orderables). Among 22 paired orderables, median changes per order fell from 3.20 (IQR 2.41-3.68) to 0.39 (IQR 0.18-1.04; p < 0.001). System-level volume-weighted changes per order fell from 3.06 to 2.72 as preference-list adoption rose from 9.6% to 15.8%. The design was uncontrolled.
Does an EHR-embedded generative AI chart summarization tool reduce clinician workload? In a pragmatic 1:1 randomized trial at one academic health system, 284 ambulatory clinicians across 42 specialties received Epic's outpatient chart summarization tool or usual care over 90 days (February o task task load favored the intervention ( tionsted difference -27.4 on a 0-400 scale; 95% CI, -49.4 to -5.3; P=0.02), with lower burnout (-0.20) and work exhaustion (-0.24) on the Professional Fulfillment Index. Charting time per encounter was unchanged (-1.2 seconds). Only 14.2% of 74,474 summaries were opened, falling from 21.5% to 10.5% by month 3; net promoter score was -22.
Can individualized EHR analytics paired with structured coaching improve resident inbasket performance? Investigators at a large Mid-Atlantic academic center conducted a stratified longitudinal analysis of EHR metrics with pre- and post-intervention surveys among PGY-2 and PGY-3 internal medicine residents and continuity clinic attendings, delivering personalized efficiency and quality reports in one-on-one R2C2-model feedback sessions across two training periods. Turnaround time for patient calls fell by 2.2 days (p<0.001), and mean time in patient calls declined from 0.99±0.53 to 0.74±0.38 minutes/day (p=0.03). Self-reported burnout (p=0.09), confidence (p=0.18), and perceived efficiency (p=0.56) were unchanged; 6 of 7 faculty found the analytics at least moderately useful.
Does ambient AI documentation change clinician workload, efficiency, and patient experience in emergency care? A retrospective observational cohort study examined voluntary ambient AI use across 14 EDs from May 2024 to June 2025, covering 315,242 notes, with within-clinician paired comparisons and pre/post NASA-TLX surveys. AI was used in 8.6% of notes. Top-box ratings for clinician listening rose (82.0% vs. 76.6%; OR 1.39), while likelihood to recommend did not differ. Active editing time was unchanged (median +0.17 minutes), though clinicians typed 722 fewer characters; copied content halved (9.4% vs. 18.3%) and NASA-TLX workload fell 40.2 points.
Does ambient AI documentation improve efficiency and clinician well-being when deployed across a large multi-site system? A retrospective index-date-aligned pre-post analysis of Epic Signal data covered 210 ambulatory physicians and APPs using DAX Copilot at a multi-state health system, June 2024 to November 2025, with 8 months before and after activation, plus a 45-day post-activation survey (n=233, 26.4% response). Active note time fell from 5.54 to 4.09 minutes per appointment (26.2%; p<0.001), while note length rose 28.8% and typing fell 51.7%, copy/paste 42.7%, and conventional voice recognition 67.8%. Survey respondents reported reduced burnout and higher satisfaction; no effect sizes given.
Why do portal messages from historically marginalized patients get fewer responses? This cross-sectional study applied natural language processing to extract message content and writing style features from 3,619,390 medical advice request threads sent by 511,020 adults to non-trainee primary care clinicians between 2021 and 2023, then used regression to decompose response gaps. Black patients had a 3.7-percentage-point lower response rate from the intended target clinician than White patients (95% CI, -4.1 to -3.3), an 11.6% relative reduction. Message content did not explain the gap, but writing style accounted for 48.0% of it for Black patients, 34.9% for Hispanic patients, 60.5% for patients with only a high school education, and 42.8% for Medicaid beneficiaries.
How do European general practitioners experience digital health technologies in daily practice? This systematic review (PRISMA 2020, Synthesis Without Meta-Analysis) searched eight databases plus Elicit AI for studies published January 2014 to October 2025, yielding 65 studies covering 24,994 GPs in 14 countries, 93.8% (61/65) from Western and Northern Europe. Perceived usefulness was positive in 86.2% (56/65) of studies, while ease of use was positive in none (0/65) and negative in 66.2% (43/65); actual use was near universal (63/65, 96.9%). Barriers included poor interoperability, limited training, and unreimbursed digital workload, alongside concerns about digital exclusion.Voosh.aiVoosh.aiVoosh.aiVoosh.aiVoosh.ai Normalization process theory constructs showed only partial integration.Voosh.ai Ratings.Voosh.ai bias risk used MMAT and JBI tools.
Can a large language model reliably categorize the content of patient portal messages at scale? Researchers built a zero-shot GPT-4o-mini pipeline applied to all medical advice request messages sent to ambulatory clinicians at an academic medical center in 2024-2025, using an 11-category taxonomy derived by expert panel via modified Delphi; two annotators labeled 750 messages (Cohen kappa 0.80), with 500 held out for evaluation. Micro- and macro-averaged F1 were 0.89 and 0.86, and labels were identical across runs for 93.6% of messages. Across 2.4 million messages, Problems & Management and Medications & Prescriptions appeared in 67.9%, the top four topics in 93.9%, and 51.7% spanned multiple topics.
Can machine learning identify point-of-care ultrasound (POCUS) in free-text notes and gauge whether standardized documentation templates improve charge capture? This retrospective operational cohort study analyzed 559,029 encounters from 109,776 patients across 11 OBGYN clinic sites at one academic medical center (January 2018–August 2024), training LightGBM and BioClinBERT classifiers against manual CPT assignments and comparing periods before and after a February 2023 ProcDoc smart form. BioClinBERT reached 0.97 accuracy (F1 0.55–0.63). ProcDoc adoption hit 75.1% at 12 months; billing recapture fell from 10.0% to 2.4% (OR 0.22, 95% CI 0.17–0.30), with overall POCUS billing up 0.6%.
Can an online peer coaching program reduce physicians' after-hours administrative work? This voluntary longitudinal survey study followed 280 physicians who completed the Charting Champions Program between 2020 and 2023, using a 14-item Likert survey at program entry and 30-90 days after completion. Respondents reported significant decreases in hours spent charting (P<0.0001) and completing paperwork outside clinical hours (P<0.006), along with less work-related dread, burnout, and thoughts of quitting (all P<0.001), and greater focus, control, and mental energy. Patients seen per clinical day was unchanged (P>0.918). The abstract reports p-values only, with no point estimates or effect sizes, and includes no control group.
How common are overlapping administrative burdens in family medicine, and do health IT and staffing supports help? This cross-sectional study surveyed 8419 US family physicians completing American Board of Family Medicine certification requirements in 2024, measuring effort tracking down external health information, completing prior authorizations, and after-hours documentation. More than three-quarters reported at least one substantial burden and 15% reported all three. Satisfaction with EHR support for external information was associated with less effort on that task (OR 0.47), while ability to complete prior authorizations in the EHR was not associated with lower burden. Helpful EHR templates were associated with less after-hours documentation (OR 0.70) and less triple burden (OR 0.63).
Does clinical AI reduce or add to clinicians' cognitive workload? This systematic review and meta-analysis searched MEDLINE, Embase, Web of Science, and CENTRAL (January 2015-2026), including 21 studies of 2885 health care professionals in 7 countries that used validated instruments (NASA-TLX, Professional Fulfillment Index). Ambient documentation AI was associated with lower temporal demand (SMD -1.46, 95% CI -2.81 to -0.11; k=2), lower effort (SMD -1.29, -2.16 to -0.42), reduced work exhaustion (MD -0.35, -0.58 to -0.12), and lower burnout prevalence (OR 0.47, 0.25-0.86). Diagnostic imaging AI and decision support showed mixed or increased workload. GRADE certainty was moderate at best; prediction intervals crossed the null.
No abstract was available for this JAMA report on artificial intelligence-powered ambient scribes. The title indicates the study examines whether AI scribe adoption changes how clinicians allocate their time and whether visit volume shifts as a result — outcomes central to arguments that ambient documentation tools reduce administrative burden or, alternatively, free capacity that gets absorbed by added throughput. Design, setting, population, and effect sizes cannot be characterized without the full text; readers should consult the article directly for the magnitude and direction of any measured changes.
Can an EHR interface mirroring the physical ICU room layout speed documentation and lower cognitive load? In a randomized two-period crossover simulation, 36 ICU nurses completed two standardized documentation scenarios using a Spatial Awareness Integrated EHR prototype versus a traditional linear flowsheet interface, analyzed with linear mixed-effects models. The spatial prototype cut documentation time by 177 seconds per task (27% faster; d = 0.88), reduced NASA-TLX workload 51.3% (18.44 vs. 37.82), and improved accuracy 6.3% (98.70% vs. 92.86%). Behavioral intention to adopt rose 25.5%, but System Usability Scale (77.99 vs. 71.60) and perceived usefulness differences were not significant.
How do emergency physicians actually allocate shift time, and how much of it goes to the computer? This cross-sectional observational time-motion study in a high-volume urban ED used the validated TimeCaT application to track 20 physicians across one 8- to 9-hour shift each, totaling more than 150 hours of real-time observation, supplemented by EHR event logs for after-shift work. Physicians spent a median 34.1% of shift minutes on the computer (156.5 min) versus 26.9% with patients (115.2 min), plus 15.9% on verbal communication with staff. EHR logs showed an additional median 1.3 hours of post-shift computer use, or 29.8 combined computer minutes per scheduled hour. Visualizations showed frequent task switching and variable fragmentation.
What shapes adoption of digital scribes, and what do they change for patients, clinicians, and organisations? This PRISMA-ScR scoping review searched MEDLINE, CINAHL, Web of Science, SCOPUS, and EMBASE for original studies or case reports evaluating digital scribe implementation in real-world care, mapping themes to the updated Consolidated Framework for Implementation Research and its Outcomes Addendum. Of 4772 studies screened, 29 were included. Scribes were generally acceptable (n=11) and usable (n=8), though nine reported accuracy concerns; reported impacts included reduced documentation burden (n=18), improved clinician wellbeing (n=12), and better patient-clinician interaction (n=10). Only three examined cost or productivity. The abstract reports no pooled effect sizes.
Can an ambient AI scribe scale across emergency departments without degrading documentation or patient experience? This 12-month multicenter retrospective observational study covered five emergency specialties at 48 Spanish hospitals (February 2025–January 2026), including all level 4 and 5 consultations among roughly 2.27 million ED visits. The scribe was used in 1,032,558 consultations (45.3%), with monthly adoption rising from 7.7% to 57.8% and 2,097 physicians using it at least once. Scribe-assisted consultations were shorter (mean relative time savings 21.8%, p<0.001), transcription accuracy averaged 93.9%, and audited report quality and patient Net Promoter Scores were higher; the abstract reports no effect sizes for the quality and experience comparisons.