Can an EHR interface mirroring the physical ICU room layout speed documentation and lower cognitive load? In a randomized two-period crossover simulation, 36 ICU nurses completed two standardized documentation scenarios using a Spatial Awareness Integrated EHR prototype versus a traditional linear flowsheet interface, analyzed with linear mixed-effects models. The spatial prototype cut documentation time by 177 seconds per task (27% faster; d = 0.88), reduced NASA-TLX workload 51.3% (18.44 vs. 37.82), and improved accuracy 6.3% (98.70% vs. 92.86%). Behavioral intention to adopt rose 25.5%, but System Usability Scale (77.99 vs. 71.60) and perceived usefulness differences were not significant.
Issue No. 001
The strongest signal comes from monitoring: a longitudinal study of four deployed clinical AI systems found validation-era performance did not persist, with calibration drift and label-independent signals like input missingness flagging decay before outcome-based surveillance could. Read it alongside a scoping review of 46 clinical AI evaluation frameworks, 88% aimed at investigational rather than deployed use. Ambient documentation is scaling fast — an AI scribe covered over a million emergency consultations across 48 Spanish hospitals with 93.9% transcription accuracy — but simulated interpreter-mediated visits show scribes propagate interpretation errors straight into the note, a safety gap for multilingual care. On interoperability, VA data link higher HIE volume to fewer community-care readmissions and deaths, yet more in direct care. And a multicenter randomized trial shows image-based AI meaningfully narrowing genotype searches for inherited retinal disease.
How do emergency physicians actually allocate shift time, and how much of it goes to the computer? This cross-sectional observational time-motion study in a high-volume urban ED used the validated TimeCaT application to track 20 physicians across one 8- to 9-hour shift each, totaling more than 150 hours of real-time observation, supplemented by EHR event logs for after-shift work. Physicians spent a median 34.1% of shift minutes on the computer (156.5 min) versus 26.9% with patients (115.2 min), plus 15.9% on verbal communication with staff. EHR logs showed an additional median 1.3 hours of post-shift computer use, or 29.8 combined computer minutes per scheduled hour. Visualizations showed frequent task switching and variable fragmentation.
What shapes adoption of digital scribes, and what do they change for patients, clinicians, and organisations? This PRISMA-ScR scoping review searched MEDLINE, CINAHL, Web of Science, SCOPUS, and EMBASE for original studies or case reports evaluating digital scribe implementation in real-world care, mapping themes to the updated Consolidated Framework for Implementation Research and its Outcomes Addendum. Of 4772 studies screened, 29 were included. Scribes were generally acceptable (n=11) and usable (n=8), though nine reported accuracy concerns; reported impacts included reduced documentation burden (n=18), improved clinician wellbeing (n=12), and better patient-clinician interaction (n=10). Only three examined cost or productivity. The abstract reports no pooled effect sizes.
How do rural patients view ambient AI scribes, and which characteristics predict acceptance? A cross-sectional analysis of 1,050 rural respondents in the 2024 Canadian Digital Health Survey dichotomized three outcomes — trust in documentation accuracy, perceived interaction benefit, and future-use preference — and fit XGBoost classifiers using sex, age, race/ethnicity, education, employment, income, chronic disease, self-reported health, and high-speed internet access, summarizing subgroup differences as marginally standardized predicted probabilities. Endorsement declined across the three domains. Future-use preference was 0.388 for males versus 0.313 for females, 0.466 for graduate degrees versus 0.308 for less than high school, and 0.408 with chronic disease versus 0.298 without. Internet access showed similar probabilities throughout.
Can an ambient AI scribe scale across emergency departments without degrading documentation or patient experience? This 12-month multicenter retrospective observational study covered five emergency specialties at 48 Spanish hospitals (February 2025–January 2026), including all level 4 and 5 consultations among roughly 2.27 million ED visits. The scribe was used in 1,032,558 consultations (45.3%), with monthly adoption rising from 7.7% to 57.8% and 2,097 physicians using it at least once. Scribe-assisted consultations were shorter (mean relative time savings 21.8%, p<0.001), transcription accuracy averaged 93.9%, and audited report quality and patient Net Promoter Scores were higher; the abstract reports no effect sizes for the quality and experience comparisons.
Do ambient AI scribes carry interpreter errors into the clinical note? Using simulated English- and Spanish-language clinical encounters mediated by interpreters, the authors evaluated whether documentation generated by ambient AI scribes reproduced interpretation errors introduced during the visit. Scribes propagated interpreter errors into the resulting notes, with propagation patterns differing by speaker role and by error type. The published abstract reports no effect sizes, error counts, or comparative rates. The authors frame the results as a case for further evaluation of AI-scribe performance in multilingual and interpreter-mediated care.
How should hospitals calibrate oversight of AI tools embedded in EHR workflows, imaging, triage, documentation, and operations? The authors conducted a narrative review and framework synthesis drawing on peer-reviewed evidence, reporting guidelines, regulatory and policy sources, implementation studies, and applied governance case reports. The resulting framework has four components: a use-case inventory tagged by decision influence and workflow coupling; a six-domain risk taxonomy spanning clinical safety, privacy and data security, ethics and fairness, transparency, system stability, and compliance; a four-tier risk scheme keyed to harm, automation, reversibility, and coupling; and a governance architecture assigning roles to a committee, clinical owners, risk-control functions, and independent assurance. A lifecycle pathway runs from initiation and local validation through shadow mode, controlled go-live, monitoring, change control, and retirement. No effect sizes are reported; this is a conceptual framework, not an evaluation.
For what purposes do physicians use unauthorized, non-conformity-assessed AI tools at work? This cross-sectional survey of physicians in Swedish health care organizations (N=357; response rate ~64%), fielded through a verified online panel between December 2023 and January 2024, applied qualitative content analysis to free-text responses, interpreted through the sociology of professions and paradox theory. Reported uses fell into four categories: clinical work and decision-making (second opinions, differential diagnoses, rare cases), administrative work (patient communication, documentation), research and professional development, and technological curiosity. Physicians framed such use as compensating for gaps in institutional systems and reducing workload. The abstract reports no effect sizes or usage prevalence.
How consistent are the evaluation frameworks proposed for clinical AI? This scoping review followed PRISMA-ScR, searching six databases plus the EQUATOR Network through February 2026, and screened 3363 records to include 46 frameworks, scored on methodological rigor, validation strategy, and a 10-domain UNESCO ethics matrix. Most frameworks (88%) targeted investigational rather than clinical use; 31.8% reported technical metrics such as AUC, sensitivity, and specificity, 15.9% reported clinical indicators, and only 11.4% met methodological rigor with validation aligned to intended use. Ethics coverage was uneven: transparency and explainability appeared in 70%, human oversight in 24.4%.
How should public health teams decide whether a customized large language model is ready to deploy? This conceptual paper proposes an acceptance criteria framework (ACF) defining implementation fit as meeting prespecified minimum performance standards and showing nonproblematic behavior under anticipated use. The ACF combines project-relevant and off-topic test prompts, structured expert review, and prespecified thresholds to generate a documented decision record that can be rerun after model revisions. The authors argue prior safety, ethics, effectiveness, engagement, and implementation frameworks imply rather than operationalize deployment benchmarks, and illustrate the ACF in a tobacco cessation text messaging intervention. No performance estimates are reported.
How can health systems test large language models on real patient portal messages without touching live EHR workflows? This technical feasibility tutorial describes a Python 3 web interface and modular backend running inside the institutional firewall on an NVIDIA GRID T4-1Q GPU, supporting single-message and batch tasks: authorship identification, categorization, criticality flagging, and response drafting with zero-, one-, and few-shot prompting. A deidentification pipeline validated against 110 manually adjudicated entities achieved 95.1% sensitivity and 82.1% precision. Use cases drew on an IRB-approved dementia-relevant corpus of 6941 medical advice request messages from 497 patients; token-based cost readouts were included. No comparative performance effect sizes are reported.
Does acceptable pre-deployment validation performance persist once clinical AI enters routine workflows? This longitudinal retrospective observational study followed four deployed AI systems spanning different clinical domains within one large healthcare organization, comparing validation-era metrics with post-deployment behavior over extended observation using routine clinical data, outcome labels, and operational telemetry, and examining discrimination, calibration, data availability, latency, and workflow signals. In all four systems, validation performance did not persist; calibration drift appeared consistently and often preceded discrimination changes, and label-independent signals such as input missingness and data latency flagged degradation earlier than outcome-based monitoring, which lagged behind label availability. The abstract reports no effect sizes.
How do AI-generated initial recommendations in virtual urgent care compare with the final recommendations physicians issue? This Annals of Internal Medicine study examines concordance between an AI system's initial output and clinicians' final decisions across AI-assisted virtual urgent care visits. No abstract was available, so findings, sample size, and effect sizes cannot be summarized here.
Does a higher volume of health information exchange (HIE) improve outcomes, and do effects differ between community and VA direct care? Using VHA electronic health record data from January 2022 to December 2023 (3144 medical center-months; ~2.4 million patients monthly), the authors instrumented HIE volume with each center's count of organizational exchange partners, with center and month fixed effects. In community care, a 1-SD increase in HIE volume was associated with 4.07 fewer 30-day readmissions, 1.45 fewer avoidable hospitalizations, and 0.25 fewer inpatient deaths per center-month; in VHA direct care, 7.07 additional readmissions and 0.19 additional inpatient deaths, with no significant change in avoidable hospitalizations.
Can a health system-governed data platform overcome the fragmentation, latency, and quality limits of real-world data? This descriptive platform paper reports on Truveta's partnership model, architecture, and applications, covering de-identified electronic health record data on 130 million US patients — roughly 1 in 3 Americans — aggregated from participating health systems and linked to closed claims, mortality, and social determinants data. Records are ingested daily, normalized to standard ontologies, and processed with NLP to extract concepts from notes, imaging, and pathology text. The data have supported over 100 publications spanning treatment effectiveness, device surveillance, COVID-19 vaccine safety, and health equity. No comparative effect estimates are reported.
Can national health information networks be repurposed to acquire EHR data for research? In a pilot, the All of Us Center for Linkage and Acquisition of Data worked with eHealth Exchange, the largest US health information network, to route participant-authorized queries to one health information exchange and one hospital system; returned FHIR and C-CDA records were mapped to the OMOP common data model and compared with existing All of Us EHR data. Retrieved records added complementary information and improved completeness, though the abstract reports no effect sizes. Barriers included inconsistent capacity to transact authorizations and variable data quality.
Do the transparency needs of healthcare AI users actually map onto the Instructions for Use (IFU) document that the EU AI Act (Directive 2024/1689) requires providers to give deployers? This cross-sectional online survey, administered via Qualtrics to four deployer groups \u2122 managers (N = 238), healthcare professionals (N = 115), patients (N = 229), and IT workers (N = 230) \u2122 asked participants to rate the relevance of a set of transparency needs and identify which IFU section would address each. Priorities differed across user types, and participants had difficulty locating some transparency information within the IFU structure; the abstract reports no effect sizes or magnitudes. The authors derive recommendations for locally meaningful IFUs.
How much do members of the public support regulating AI-delivered mental health advice? This Health Affairs Scholar paper takes up that question, but no abstract was available at the time of writing, so the study design, sample, and findings cannot be characterized here. Readers interested in public opinion on guardrails for consumer-facing chatbots and other AI tools offering psychological support should consult the full text for the survey methods, population sampled, and reported levels of support for specific regulatory approaches.
Do venture-backed maternal health startups target the populations and problems driving the US maternal health crisis? This cross-sectional study used financial databases and dual-coder content analysis of company websites to characterize US perinatal startups founded between 2014 and 2022, identifying 439 companies, of which 183 met inclusion criteria and 172 were venture-funded in their last round. These firms raised $977.5 million, with 52% ($508.3 million) concentrated in three companies. Among 133 actively operating firms, virtual or hybrid wraparound pregnancy care was most common (34.6%), while 17.3% mentioned health equity and 18.1% maternal mortality; fewer than half of eligible startups accepted insurance and fewer accepted Medicaid.
Do online star ratings shift where patients go for elective inpatient care? The authors link the universe of hospital Yelp reviews to Florida inpatient claims for elective procedures (2012-2017), exploiting exogenous variation in ratings to identify causal effects on hospital selection. A one-standard-deviation increase in a hospital's within-market rating percentile rank — roughly half a star — was associated with patients traveling 7.9% farther for labor and delivery and 33.5% farther for orthopedic surgery. Falsification tests using emergency admissions were null, consistent with ratings influencing elective rather than urgent choices.
Does generative AI change which human skills firms hire for? Using ChatGPT's release as an exogenous shock in a quasi-experimental design, the authors track job-posting demand across 1,820 publicly listed U.S. companies over a ±12-month window, classifying postings into five organizing skills: task division, task allocation, information provision, reward distribution, and exception management, with a queuing-theory model predicting which fall first. Demand declined significantly for monitoring (reward distribution), operational exceptions, and task division, with information provision also falling; task allocation and conflict resolution were more stable. Declines intensified after GPT-4's release. The abstract reports no effect sizes.
Does mandating disclosure of generative AI use change how creators themselves work with AI, before any audience sees the label? The authors theorize an "indirect disclosure effect" grounded in Goffman's impression management and test it in two nested mixed-methods experiments in which participants collaborated with a text-to-image GenAI tool under varying disclosure conditions. When disclosure was anticipated, the majority of creators withdrew from the creative process and ceded image generation to the tool, a shift attributed to fears that audiences would not recognize their creative agency; resulting artifacts reflected computational rather than human creativity and were evaluated as such regardless of the label. The abstract reports no effect sizes or sample sizes.
What happens to user-generated content when a Q&A platform bans generative AI? Using a difference-in-differences design comparing Stack Overflow with Reddit's AskProgramming subreddit around Stack Overflow's prohibition on ChatGPT-generated posts, the authors applied NLP measures of linguistic characteristics plus voting and posting-frequency data. After the restriction, Stack Overflow answers showed greater language complexity, positivity, and length, and received more upvotes, while questions were unaffected; the volume of questions and answers and the number of first-time contributors declined. Findings were corroborated by Italy's temporary nationwide ChatGPT ban and a scenario-based experiment with 440 participants indicating compensatory knowledge signaling. The abstract reports no effect sizes.
Does mandatory electronic reporting with automated auditing improve regulatory compliance and environmental performance? Using a difference-in-differences design, the authors evaluate the first U.S. program requiring online reporting and automated auditing of wastewater discharge releases, framing it as a test case for AI-based compliance tools with automated feedback. The program was associated with more complete reporting, reduced discharges, and a higher rate of reported violations, with larger effects among minor dischargers and publicly owned facilities. The authors also find evidence that state authorities targeted inspections toward plants with recent noncompliance, a possible mechanism. The abstract reports no effect sizes or point estimates.
How much labor and cost does manual fax routing consume in a specialty division? This two-phase quality improvement study at Duke's Division of Cardiology paired an observational time study at three ambulatory clinics (April 1–July 12, 2024) with a volume count of all inbound faxes to the divisional communication hub (July 1–December 31, 2025). Processing took 4.4 to 9.4 minutes per fax (mean 6.0). The hub received 24,420 faxes, averaging 4,070 faxes and 13,341 pages monthly, implying 407 person-hours and $10,663.40 per month, about 2.5 FTEs. No AI automation was evaluated; the authors frame fax routing as an automation target.
Do US counties with limited physical healthcare capacity also lack the broadband needed for telemedicine to substitute? This cross-sectional ecological analysis linked 3,133 counties across the 2017 National Neighborhood Data Archive (outpatient care centers, diagnostic labs, nursing/residential care), 2022 FCC Mapping Broadband Health in America data (split at the median 9.8% of households without broadband), and 2022 American Community Survey covariates, using t-tests and multivariable linear regression. Low-broadband counties had fewer outpatient care centers (10.46 vs. 11.91 per 100,000) and diagnostic labs (1.91 vs. 3.95 per 100,000; both P<0.001), plus higher poverty and rurality. Adjusted associations persisted (β = -0.045, -0.024, and -0.089).
Is engagement with clinical digital health tools associated with psychological distress? This cross-sectional analysis pooled Health Information National Trends Survey cycles (HINTS 5, 2017-2020; HINTS 6, 2022; HINTS 7, 2024) covering 23 682 US adults (mean age 55 years; 59% female), with a composite engagement index spanning secure messaging, online test results, portal access, wellness apps, and device data transmission, and distress measured by the PHQ-4. In survey-weighted regression, higher engagement was associated with higher distress (\u03b2 = 0.51; 95% CI, 0.27-0.75; P < .001), against a mean PHQ-4 of 2.0. Associations were strongest for active behaviors—clinician messaging and app use—and persisted among those reporting good or better health.
Does centralizing appointment scheduling and adding same-day virtual clinician evaluation improve access after nurse triage? This retrospective quasi-experimental evaluation used difference-in-differences and event-study analyses of the VA Health Connect rollout across 18 regions from October 2018 to September 2024, drawing on 11,118,916 encounters (4,560,677 pre-, 6,558,239 post-modernization) from VA Corporate Data Warehouse, Telecare, CRM, and VSignals survey data. Same-day scheduling rose 14.3 percentage points (95% CI 10.1-18.5) and time from call to scheduled appointment fell 0.37 days, though time to completed appointment rose 2.9 days. Callers with no 7-day follow-up declined 2.3 points; ED visits, admissions, and costs were unchanged.
What drives nonresponse to routinely collected patient-reported outcome measures? This retrospective cohort study used iterative mixed-effects logistic regression on all adults seen at five Mass General Brigham radiation oncology clinics over one year (12,214 patients, 71 providers, five clinics), modeling failure to ever complete the portal-administered PROMIS Global-10. Patient- and appointment-level response rates were 35.4% and 10.9%, with patient-level response varying nearly fivefold across clinics (12.8% to 66.2%). After adding provider- and clinic-level factors, sex, education, and employment became nonsignificant, while recent surgery (aOR 1.97) and time since diagnosis >12 months (aOR 0.46) persisted; later program launch (aOR 0.29) and higher historical collection rate (aOR 0.79) predicted lower nonresponse, and academic versus community setting did not.
Can a text-based treatment guideline be translated into executable logic inside an EHR? This implementation case study describes a multidisciplinary team using agile methods to convert the American Diabetes Association Standards of Care for type 2 diabetes with cardiovascular or renal disease into structured algorithms and ontology groupers for diagnoses, labs, and medications, built on SNOMED CT, LOINC, and RxNorm. The work yielded three tools: a real-time registry identifying patients eligible for guideline-directed therapy, filterable to population and individual-provider treatment gaps, plus two clinical decision support instruments embedded in clinician workflows. The authors describe the process as feasible but complex and labor-intensive, and call for guideline organizations to supply technical frameworks, regular updates, and vendor collaboration. The abstract reports no effect sizes or utilization outcomes.
What would clinicians and caregivers want from an AI-based clinical decision support tool for early autism detection, and where would it fit in the visit? This observational qualitative study used contextual inquiry with 8 clinicians and 20 caregivers during 18- to 24-month well-child visits at Duke-affiliated clinics, analyzed with rapid qualitative analysis. Workflow mapping identified 6 user tasks, 3 technology-user interactions, and 5 clinical decision points, plus 2 barriers (screening tool accuracy, follow-up implementation) and 3 facilitators (electronic screening, early intervention provider input, referral coordination support). Preferences included EHR-embedded, actionable outputs with prediction explanations, visual summaries, and caregiver-facing materials. The abstract reports no effect sizes.
What clinician characteristics moderate use of a chronic pain clinical decision support tool? Using electronic health record data from a pragmatic randomized controlled trial (October 2019–May 2022) covering 69 primary care clinicians with access to the OneSheet CDS, investigators modeled tool access within three days of an encounter, with generalized linear models testing clinician gender and years in practice as moderators. OneSheet was used in 959 of 145,511 encounters (0.7%). Use was lower for new-patient encounters (−0.42 percentage points; 95% CI −0.65 to −0.19) and higher for chronic pain diagnoses (3.48 pp) and long-term opioid therapy (4.70 pp). Associations were strongest among female clinicians with over 16 years in practice; prior use did not predict future use.
Can an image-based AI narrow the genotype search for inherited retinal diseases before genetic testing? Retina4IRD, a RETFound-pretrained Vision Transformer predicting 17 genotype categories, was trained on fundus photographs and OCT from 1,843 genetically confirmed patients (3,376 eyes) in China, South Korea and Poland; top-5 accuracy was 0.904 internally and 0.856 externally. In a multicenter randomized trial, 300 patients with suspected IRD were assigned 1:1 to AI-assisted or specialist-only assessment (295 analyzed). Top-5 genetic accuracy was 88.5% versus 67.3% (P<0.001), top-1 37.8% versus 22.4%, and a composite downstream management score 37.7 versus 28.5 (P<0.001).
Can EHR-based clinical decision support safely curb excessive continuous pulse oximetry (CPO) monitoring? This quality improvement study at a 796-bed tertiary center used interrupted time series analysis across all non-ICU inpatient units, comparing June 2022–May 2023 (18,351 CPO-associated hospitalizations) with June 2023–May 2024 (18,713) after revised CPO orders, order sets, and a best practice advisory informed by semi-structured interviews. Against a pre-intervention monthly average of 120,771 CPO hours, the intervention was associated with an immediate reduction of 22,680 hours (95% CI −33,591 to −11,769), with monthly alarms down 332,819 and duration per admission down 34.2 hours. ICU transfers rose 9.6 per 1000 admissions; rapid responses, code blues, and mortality were unchanged.
Does provider engagement with clinical decision support alerts vary by patient race and sex? This retrospective study used EHR data on alert-based CDS during outpatient primary care at a New York City academic health system, applying logistic regression to model alert engagement by patient race and sex with adjustment for encounter and provider factors, and a generalized structural equation model to test mediation by alert type. Direct effects indicated differential provider response by patient demographics; indirect effects indicated unequal assignment of alert types across groups, so uneven exposure alone could yield inequitable outcomes. The abstract reports no effect sizes.</summary
Can institution-specific cancer trial information be curated into an AI-enabled knowledge management application in community oncology? This feasibility study at a regional community oncology network had coordinators and disease teams compile actively recruiting trials, structuring core elements (title, conditions, biomarkers, stage/line, recruiting status) for point-of-care display, with AI-assisted extraction of protocol summaries and eligibility elements followed by human validation. Fifty-three trials across 10 disease groups and 28 cancer types were embedded; 91% were recruiting and 30% were biomarker-specific. Configuration required 2-4 weeks per disease group using existing personnel, without added staffing or EHR build. Usability and implementation outcomes were not assessed.
What determines whether automated waitlists—tools that notify patients of earlier appointment openings—succeed in improving ambulatory access? A convergent, multisite mixed methods study surveyed 127 US health systems, 90 of which reported automated waitlist usage data, plus qualitative and quantitative data from 10 purposively sampled systems, analyzed using the Consolidated Framework for Implementation Research. High performers filled 38.8% (IQR 36.2%-45.7%) of appointments offered through the waitlist, and missed appointment rates were lower for waitlist-scheduled visits (3.1%, IQR 2.5%-4.8%) than for all appointments (6.6%, IQR 4.1%-9.9%). Flexible configuration, cross-functional governance, and leadership endorsement facilitated sustained use; specialty gatekeeping, clinician capacity, insurance requirements, and digital inequities limited reach.
How have logic models and theory of change been applied to health information technology interventions? This PRISMA-ScR scoping review searched PubMed, Web of Science, Academic Search Elite, APA PsycArticles, and CINAHL, with dual independent screening and extraction, identifying 69 publications from 2012 to 2025 across medical informatics, public health, health services research, and implementation science. Use rose after 2020 and clustered in patient-facing mobile health, telehealth, and remote monitoring. Of 69 studies, 60 (87%) included a model visualization, 50 (72%) cited development guidance (most often UK MRC, realist evaluation, or Kellogg), 28 (41%) drew on behavioral or implementation frameworks, and only 3 (4%) reused an existing model.