The AJH Informatics Review

A weekly digest of new research on EHRs, clinical AI, interoperability & health IT policy

Economics of health IT

Every digest paper in this category, newest first.

001
Utilization of a HIPAA-compliant large language model chatbot in an academic pediatric medical center

Who actually uses a hospital-wide LLM chatbot, and what blocks the rest? This mixed-methods case study of "InternalGPT," a HIPAA-compliant chatbot at an academic pediatric medical center, combined employee surveys with 14 months of utilization data. Of roughly 15,800 employees, 2,149 (13.6%) requested access; 52.8% of those used at least one token, 33.6% never logged in, and the top 20% of users consumed 69.4% of tokens. Barriers cited by non-users were limited time (51.4%) and difficulty using the tool (21.8%). Among 92 sustained users surveyed, mean self-reported productivity gain was 30%, an exploratory $6.3M-$18.9M value.

002
Digital Health Readiness and Medicare Primary Care Spending: Nationwide County-Level Observational Analysis

Is community-level digital health readiness associated with lower Medicare primary care spending in safety net settings? This county-level observational study covered 2993 US counties from 2017 to 2023, using the Digital Health Index digitalization subindex against geographically adjusted per capita Medicare spending on federally qualified health center and rural health clinic services, with a hybrid within-between panel model. Between-county differences drove the association (β=−67.62, 95% CI −73.51 to −61.74), while within-county change was null (β=−2.72, P=.40). Highest versus lowest digitalization tertile differed by β=−165.05. In 2023 cross-sectional models, health care access showed the strongest inverse association (β=−71.13).

003
Data, Privacy Laws and Firm Production: Evidence from the GDPR

How do data privacy regulations affect firm production and performance? This Journal of Political Economy paper examines the European Union's General Data Protection Regulation as a natural experiment in restricting firms' access to and use of consumer data. No abstract was available, so findings, methods, data sources, and effect sizes cannot be summarized here; readers interested in the estimated magnitudes of GDPR's effects on firm data holdings, computation, and output should consult the full text.

004
Cost-Effectiveness of Electronic Patient-Reported Outcome Measure Interventions in Cancer: Systematic Review and Parameter Extraction for Economic Modeling

Are electronic patient-reported outcome measure (ePROM) programs in cancer care cost-effective? This systematic review searched Ovid (MEDLINE, Embase), Scopus, and the INAHTA database for English-language papers through March 2025, including 34 publications from 27 studies covering 26 ePROM-integrated interventions for adult cancer populations, alongside parameter extraction for economic modeling. Most interventions (23/26) included alert handling or automated decision support. Only 5 publications reported full cost-effectiveness analyses; 3 were highly uncertain, while 2 showed cost-effectiveness driven by quality-of-life gains and fewer hospitalizations. Five reported partial results (4 favoring ePROMs). Twelve studies had qualitative components, but only 2 addressed economic themes.

005
Enhancing AI Use: How Complementary System Information Drives Delegation Frequency and Effectiveness

Can information about an AI system help people delegate tasks to it more often and more wisely? This experimental study manipulated two signals: ex-ante AI certainty (the AI's estimated likelihood of being correct, shown before the delegation decision) and ex-post outcome information (whether the AI was actually correct, shown after). Presenting either signal alone had no effect or reduced combined human-AI performance; providing both raised delegation frequency and delegation effectiveness, improving performance. The authors attribute this to certainty calibrating task-level expectations, outcomes confirming them, more accurate mental models, and less algorithm aversion. The abstract reports no effect sizes, sample size, or task details.

006
The Utility of AI Tools in Auditing Adherence to Pre-Analysis Plans -- by Jeffrey Clemens, Anwita Mahajan

Can large language models audit whether published research adheres to its pre-analysis plan? This methodological working paper applies an LLM to the authors' own study, prompting it to identify precommitted design choices, evaluate deviations from the plan, and diagnose gaps in pre-specification. The authors report the approach substantially reduces the human labor required for adherence checks, but that audit output varies across different LLMs, leaving human judgment necessary. The abstract reports no quantitative effect sizes, accuracy rates, or time savings. The paper concludes with proposed best practices for AI-assisted PAP auditing by authors and reviewers.

007
Does AI Assistance Enhance or Erode Expertise? Evidence from a Three-Month Field Experiment in Patent Drafting -- by David Autor, Tanya Rodchenko, Josh Martin, Zanna Iscenko, Scott Strand, David Pearl, Melissa Ferere

Does AI assistance build or erode professional expertise? This pre-registered three-month randomized controlled trial gave 133 practicing patent lawyers at eleven U.S. intellectual property firms access to a custom AI drafting assistant, with all work scored by blinded expert patent attorneys. AI access raised benchmark drafting quality by 0.34 SD at 10 days (p=0.03) and 0.38 SD at 90 days (p=0.01). On an unassisted redlining task after three months, treated lawyers outperformed controls by 0.32 SD (p=0.04), but gains were concentrated among senior lawyers (0.45 SD, p=0.02); junior lawyers showed no average gain and bifurcated scores.

008
Harmonizing Safety and Speed: A Human-Algorithm Approach to Enhance the FDA’s Medical Device Clearance Policy

Can machine learning help the FDA cut recalls and review workload in the 510(k) substantial-equivalence pathway? The authors trained recall-risk models on submission-time information and embedded them in a data-driven policy recommending acceptance, rejection, or deferral to FDA committees for in-depth review, using an assembled data set of more than 31,000 submissions drawn from FDA and CMS sources. Against current practice (10.3% recall rate, workload normalized to 100%), a conservative evaluation showed a 32.9% improvement in recall rate and a 40.5% workload reduction, with estimated annual savings of roughly $1.7 billion from avoided replacement costs, about 1.1% of US medical device spending.

009
Defying Distance? The Provision of Medical Services in the Digital Age

Can digital health platforms improve patient-physician matching once geography no longer binds? Using nationwide Swedish online care with time-conditional random assignment of patients to physicians, the author estimates reallocation gains from aligning provider heterogeneity with patient needs. Matching high-risk patients to doctors effective at averting emergency room use lowers ER visits by 4.4 percent (SE 1.3), and reallocation reduces counter-guideline antibiotic prescribing by 3.1 percent (SE 1.4). Trade-offs across outcomes were limited, as horizontal differentiation among doctors and varied patient needs permitted simultaneous improvement; efficiency-enhancing reallocations also carried equity consequences.

010
Inflation in Reputation Systems? Newcomers, Veterans, and Socialization within a Platform Community

Why does rating inflation in online reputation systems rise and then fall as reviewers gain tenure? This exploratory mixed-method study of reviewers in a digital platform community combines qualitative and quantitative analysis to trace how inflation behavior evolves with socialization. The authors identify three archetypical phases: Newcomers, not yet socialized, are less likely to inflate; Inflators, following direct reviewer reciprocity, inflate; and Veterans, motivated by generalized reciprocity and commitment to the community, rate more candidly. Rating inflation thus follows an inverted U-shape over the reviewer lifecycle. The abstract reports no effect sizes, sample size, or study years.

011
CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks -- by Pattaraphon Kenny Wongchamcharoen, Kris Gulati, Min Min Fong, Abhishek Nagaraj

Does a model's ability to automate a task predict its ability to help a weaker agent do it? The authors build a benchmark spanning seven economically grounded real-world tasks in which an assistant model writes guidance for a standardized lower-capacity worker model that produces the deliverable, compared against automation mode where the assistant works directly; outputs are scored by blind pairwise comparisons from an LLM judge panel with task-specific rubrics across ten replications. Rankings under the two regimes correlate only modestly, the automation winner loses on augmentation in five of seven tasks, the unaided worker outranks every assisted condition on three tasks, and only one model's guidance beats no guidance on average.

012
What Work Does Generative AI Do? -- by Alexander Bick, Adam Blandin, David J. Deming, Tyler R. Schumacher

Which jobs and tasks actually use generative AI at work? The authors field a nationally representative survey linking genAI adoption to detailed occupations and tasks, producing the first task-level adoption indexes. Occupational exposure scores explain some but not all variation in adoption across occupations and tasks, and the survey-based indexes differ conceptually from platform chat-log measures, which the authors argue over-classify chats into generic activities spanning many occupations. Adoption is widespread but shallow: within most occupations and tasks, fewer than half of workers adopt. The abstract reports no other effect sizes.

013
Canaries in the Gold Mine: Early Productivity Gains from Artificial Intelligence Creating Organization Capital -- by Tania Babina, Alex X. He, Renhao Jiang

Does firm-level AI investment translate into productivity growth, and through what channel? Using a new firm-level measure of AI investment built from AI-skilled employment—spanning machine learning through generative and agentic AI—the authors relate AI investment to productivity across firms, and construct a measure of organization capital derived from workers' job descriptions. AI investment is associated with productivity growth in recent years but not over the prior decade, with gains concentrated in AI-skilled jobs that build organization capital, that is, durable firm-specific knowledge from learning-by-doing. The abstract reports no effect sizes or sample details.

014
Learning to be Proficient? A Structural Model of User Dynamic Engagement in eHealth Behavioral Interventions

Why do users of eHealth behavioral interventions taper off over time? The authors extend Expectation-Confirmation Theory by treating engagement as a dynamic learning process, estimating a hierarchical Bayesian structural learning model of how users update perceptions of intervention effectiveness from ongoing experience and how those beliefs drive continued participation. Learning performance was lower for interventions with ambiguous instructions and those targeting short-term health outcomes, which generate noisier feedback and less accurate effectiveness perceptions, associated with reduced sustained engagement. Several denoising design strategies are evaluated in counterfactual simulations. The abstract reports no effect sizes, sample size, or study period.

015
Beyond the Black Box: Unraveling the Role of Explainability in Human-Artificial Intelligence Collaboration

When does explaining an AI model's reasoning actually improve human-AI decisions, and at what cognitive cost? The authors build an analytical model of a decision maker with limited but flexible cognition receiving imperfect machine recommendations, where explanations shift beliefs about algorithmic quality. Low explainability leaves decision accuracy and reliance unchanged while reducing cognitive burden; higher explainability improves accuracy by curbing overreliance but increases underreliance. Explainability matters more for cognitively constrained decision makers, complex tasks, and lower-stakes decisions, yet can raise processing time and fatigue exactly when time is short, tasks are complex, and machine quality is doubted. Theoretical modeling, so no empirical effect sizes.

016
The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot

Does an AI pair programmer help or hinder distributed, voluntary software collaboration? Using GitHub's proprietary Copilot usage data linked to public OSS project data, the authors estimate Copilot's effects on contribution and coordination in open-source projects; the abstract does not name the identification strategy, sample size, or study years. Copilot use raised project-level code contributions 5.9%, with a 3.4% increase in developer coding participation and a 2.1% increase in individual contributions, but an 8% increase in coordination time and more code discussion. Net timely merges rose. Peripheral developers showed smaller contribution gains and larger coordination increases than core developers.

017
Benchmark Mineability and the Financing of AI Innovation -- by Alex Chan

When AI benchmark scores steer capital, what happens to their value as signals? This NBER working paper is a conceptual and theoretical market-design analysis rather than an empirical study, treating public AI benchmarks as market institutions. The author identifies two gaps that become exploitable once scores move investment: public examples can reveal the process behind a private final test, and any finite public score cannot span the broad task space implied by general intelligence. Targeted effort aimed at these gaps erodes the signal later investors rely on. The proposed remedy is separating development from certification — publishing practice tasks but selecting the investment-consequential task generator only after a submitted system's evaluation policy is fixed. The abstract reports no effect sizes or empirical magnitudes.

018
Consumer Inferences from Product Rankings: The Role of Beliefs in Search Behavior

Why do online shoppers concentrate their search and purchases on top-ranked products — lower search costs at prominent positions, or beliefs that higher-ranked items offer better returns to search? The authors build an experimental paradigm and run incentivized experiments to separate the two mechanisms. Both are present, and short-term randomization of rankings alone does not disentangle them; ignoring beliefs yields biased search-cost estimates and incorrect consumer welfare predictions for alternative recommendation systems, including platform self-preferencing. The paper proposes approaches for recovering unbiased search costs in real search settings. The abstract reports no effect sizes or point estimates.

019
The Impact of Manipulated Clinical Decision Support Algorithm on Opioid Prescribing Decision

Can a manipulated clinical decision support tool durably change prescribing? This quasi-experimental study compares physicians using an EHR whose vendor secretly embedded a biased CDS function promoting extended-release opioids between 2016 and spring 2019 against a control group of physicians who adopted other federally certified vendors in 2011. Affected physicians increased opioid claims during the treatment window and sustained a higher propensity to prescribe after the function was removed, persisting through relocation, affiliation changes, and stricter state opioid regulations; greater physician awareness attenuated the effect. Machine-learning estimates attribute roughly 54% of the treatment effect to decision-making distortion rather than learning. The abstract reports no other effect sizes.

020
Predictive Risk Scores in the Public Sector: Experimental Evidence from Child-Protection Investigations -- by E. Jason Baron, Arkadev Ghosh, Richard Lombardo

Can algorithmic risk scores improve how child-protection supervisors allocate scrutiny? A randomized evaluation covering 4,752 child referrals over 14 months in Northampton County gave supervisors an algorithmic risk score alongside standard case records. Access to the score increased foster-care placements and service receipt among children at the highest predicted risk, with little change for lower-risk cases, and reduced subsequent maltreatment referrals. The authors report no evidence that the score widened racial disparities in decisions or outcomes. The abstract reports no point estimates or effect sizes for these changes.

021
Harvesting Ratings

Can firms manipulate online ratings through pricing, and does that degrade ratings as quality signals? This is an analytical modeling paper: a two-period model of price competition between an incumbent and an entrant of either high or low quality, in which consumers rate on value-for-money and cannot separate genuine quality from discounting. Low-quality entrants either discount to "harvest" favorable ratings or mimic high prices to signal quality; harvesting inflates positive ratings, reduces their informativeness, worsens the cold-start problem, and deters high-quality entry. Lowering the effort cost of rating yields more but less informative ratings. Suggested remedies include limiting new-seller discounts and displaying price paid. The abstract reports no effect sizes.

022
Replaceable but Employed: Automation and the Meaning of Work -- by Joshua S. Gans

Can automation harm workers without displacing them? This theoretical paper models jobs in which workers derive utility both from producing useful output and from knowing that output depends on their own contribution, so a credible machine alternative erodes the second source of meaning even when the firm keeps the worker. The model implies compensation rises when wages adjust fully but workers absorb part of the loss under partial adjustment, and that automation becomes more likely. An external developer may profit by publicly demonstrating a machine before licensing it, since salience alone devalues human work \u2014 a "meaning externality" under which profitable development can be socially harmful. No empirical estimates are reported.

023
Venture-Backed Maternal Health Startups and the Maternal Health Crisis

Do venture-backed maternal health startups target the populations and problems driving the US maternal health crisis? This cross-sectional study used financial databases and dual-coder content analysis of company websites to characterize US perinatal startups founded between 2014 and 2022, identifying 439 companies, of which 183 met inclusion criteria and 172 were venture-funded in their last round. These firms raised $977.5 million, with 52% ($508.3 million) concentrated in three companies. Among 133 actively operating firms, virtual or hybrid wraparound pregnancy care was most common (34.6%), while 17.3% mentioned health equity and 18.1% maternal mortality; fewer than half of eligible startups accepted insurance and fewer accepted Medicaid.

024
Star Ratings and Willingness to Travel for Hospital Care

Do online star ratings shift where patients go for elective inpatient care? The authors link the universe of hospital Yelp reviews to Florida inpatient claims for elective procedures (2012-2017), exploiting exogenous variation in ratings to identify causal effects on hospital selection. A one-standard-deviation increase in a hospital's within-market rating percentile rank — roughly half a star — was associated with patients traveling 7.9% farther for labor and delivery and 33.5% farther for orthopedic surgery. Falsification tests using emergency admissions were null, consistent with ratings influencing elective rather than urgent choices.

025
Skill Deprioritization: Reorganizing in the Age of Generative Artificial Intelligence

Does generative AI change which human skills firms hire for? Using ChatGPT's release as an exogenous shock in a quasi-experimental design, the authors track job-posting demand across 1,820 publicly listed U.S. companies over a ±12-month window, classifying postings into five organizing skills: task division, task allocation, information provision, reward distribution, and exception management, with a queuing-theory model predicting which fall first. Demand declined significantly for monitoring (reward distribution), operational exceptions, and task division, with information provision also falling; task allocation and conflict resolution were more stable. Declines intensified after GPT-4's release. The abstract reports no effect sizes.

026
The Indirect Disclosure Effect: How Disclosing Generative AI Use Impacts Human Creative Collaboration with AI

Does mandating disclosure of generative AI use change how creators themselves work with AI, before any audience sees the label? The authors theorize an "indirect disclosure effect" grounded in Goffman's impression management and test it in two nested mixed-methods experiments in which participants collaborated with a text-to-image GenAI tool under varying disclosure conditions. When disclosure was anticipated, the majority of creators withdrew from the creative process and ceded image generation to the tool, a shift attributed to fears that audiences would not recognize their creative agency; resulting artifacts reflected computational rather than human creativity and were evaluated as such regardless of the label. The abstract reports no effect sizes or sample sizes.

027
Unraveling the Impact: An Empirical Investigation of ChatGPT’s Exclusion from Stack Overflow

What happens to user-generated content when a Q&A platform bans generative AI? Using a difference-in-differences design comparing Stack Overflow with Reddit's AskProgramming subreddit around Stack Overflow's prohibition on ChatGPT-generated posts, the authors applied NLP measures of linguistic characteristics plus voting and posting-frequency data. After the restriction, Stack Overflow answers showed greater language complexity, positivity, and length, and received more upvotes, while questions were unaffected; the volume of questions and answers and the number of first-time contributors declined. Findings were corroborated by Italy's temporary nationwide ChatGPT ban and a scenario-based experiment with 440 participants indicating compensatory knowledge signaling. The abstract reports no effect sizes.

028
Evidence of How Electronic Reporting and Automated Auditing Affects Regulatory Compliance and Environmental Performance -- by Wayne B. Gray, Ronald Shadbegian, Ann Wolverton

Does mandatory electronic reporting with automated auditing improve regulatory compliance and environmental performance? Using a difference-in-differences design, the authors evaluate the first U.S. program requiring online reporting and automated auditing of wastewater discharge releases, framing it as a test case for AI-based compliance tools with automated feedback. The program was associated with more complete reporting, reduced discharges, and a higher rate of reported violations, with larger effects among minor dischargers and publicly owned facilities. The authors also find evidence that state authorities targeted inspections toward plants with recent noncompliance, a possible mechanism. The abstract reports no effect sizes or point estimates.

029
Agentic Artificial Intelligence as a Catalyst for Administrative Modernization: The Beginning of the End for Traditional Fax Workflows in Healthcare

How much labor and cost does manual fax routing consume in a specialty division? This two-phase quality improvement study at Duke's Division of Cardiology paired an observational time study at three ambulatory clinics (April 1–July 12, 2024) with a volume count of all inbound faxes to the divisional communication hub (July 1–December 31, 2025). Processing took 4.4 to 9.4 minutes per fax (mean 6.0). The hub received 24,420 faxes, averaging 4,070 faxes and 13,341 pages monthly, implying 407 person-hours and $10,663.40 per month, about 2.5 FTEs. No AI automation was evaluated; the authors frame fax routing as an automation target.