Wednesday, May 20, 2026
S&P 500 · NVDA · BTC
Health · Analysis

Geisinger publishes 18-month outcome data on diagnostic agents.

NEJM AI publishes the field's first prospectively-designed multispecialty diagnostic-agent outcome study. 47,318 encounters, four specialties, four-arm comparison. The effect sizes are modest, consistent, and large enough to matter.

Editorial cover: Geisinger publishes 18-month outcome data on diagnostic agents

INTELAR · Editorial cover · Editorial visual for the Health desk.

Geisinger Health System published eighteen months of operational outcome data on its diagnostic-agent deployment in NEJM AI on 19 May 2026 under the title "Multispecialty Operational Outcomes of Diagnostic Agent Deployment in an Integrated Delivery System: A Prospective Cohort Study," lead-authored by Dr. Karen Murphy, Geisinger's Chief Medical Information Officer and the principal investigator on the institution's clinical AI evaluation programme since its launch in late 2023. The publication is the first peer-reviewed, prospectively-designed, multispecialty outcome study at this duration to clear the editorial process at a high-impact venue, and the data it contains substantively advances the field's empirical evidence base on what diagnostic agent deployment actually produces in operational clinical practice. The cohort covers four clinical environments — cardiology, oncology, primary care, and emergency department — across Geisinger's Danville, Wilkes-Barre, Scranton, and State College campuses. The primary endpoints address diagnostic accuracy, time-to-diagnosis, and inpatient length-of-stay. The secondary endpoints include physician burnout score change and patient satisfaction measures. The methodology is documented in extensive detail and the limitations are stated with the kind of analytical precision that distinguishes serious clinical AI research from the marketing-grade studies that have dominated the press cycle since 2023. This report covers what the publication actually shows, where the methodology is strongest, where it is constrained, and how peer institutions have begun to respond.

The cohort: four specialties, four campuses, 18 months

Geisinger's prospective cohort study enrolled 47,318 patient encounters across the four specialty environments during the operational window of 1 November 2024 through 30 April 2026. The cohort is structured as four parallel comparison arms within each specialty, distinguishing between encounters in which the diagnostic agent was deployed and operating under ACIS-aligned governance (the intervention arm), encounters in which the agent was technically available but the attending clinician opted not to invoke it (the within-deployment comparison arm), encounters at Geisinger campuses where the agent had not yet been deployed (the contemporaneous control arm), and historical Geisinger encounters from the eighteen-month window immediately preceding deployment (the historical control arm). The four-arm structure is methodologically more robust than the simple before-and-after comparisons that have dominated the published clinical AI outcome literature since 2023. The within-deployment comparison arm in particular addresses what Murphy's methodology section calls "the selection-effect confounder" — the possibility that clinicians who choose to invoke the agent are systematically different from those who do not, and that any apparent agent-related outcome difference may reflect those clinician differences rather than the agent's contribution.

Cardiology comprised 8,917 encounters in the intervention arm at the Danville campus's Heart Institute and the Wilkes-Barre and Scranton campus cardiology departments. The deployed capability is an ambient documentation and pre-procedure risk-stratification agent built on the Suki AI platform, operating under ACIS Tier 1 governance and integrated through Geisinger's epic EHR. Oncology comprised 6,234 encounters at the Geisinger Cancer Institute, principally at the Danville campus's tumour board operations and the State College campus's medical oncology clinics. The deployed capability is a tumour board summarisation and clinical trial eligibility screening agent built on Suki AI's oncology pipeline, also operating under ACIS Tier 1 governance. Primary care comprised 22,841 encounters across primary care practices at all four campuses. The deployed capability is an ambient documentation agent built on the Abridge platform, operating under ACIS Tier 0 governance (pure summarisation with no clinical content generation). Emergency department comprised 9,326 encounters at the Danville campus's emergency department only. The deployed capability is a structured triage assistance agent built on Aidoc's BriefCase platform, operating under ACIS Tier 1 governance and integrated with the Epic ED workflow. The differential allocation of capabilities across the four specialties reflects Geisinger's deployment sequencing decisions, which in turn reflect the institution's governance review of which capabilities cleared the ACIS validation requirements first.

The cohort enrolment was prospective from each specialty's deployment date, with the four specialties going live at staggered intervals between November 2024 and February 2025. The study design accommodates the staggered onset through a stratified-by-specialty analysis architecture that does not pool the four specialties into a single combined endpoint. This is an important methodological commitment. The published clinical AI outcome literature has been substantially weakened by studies that pool heterogeneous clinical environments into single endpoints, producing effect estimates that obscure the within-specialty variation that is actually clinically relevant. Murphy's study explicitly resists this temptation. The headline outcome data is reported by specialty. The combined-specialty estimates that appear in the discussion are framed as exploratory rather than primary. This is correct methodology and it is the kind of analytical discipline that distinguishes Murphy's study from the bulk of the published clinical AI outcome literature.

The primary endpoints: diagnostic accuracy, time-to-diagnosis, length-of-stay

The diagnostic accuracy endpoint is operationalised as concordance between the attending clinician's documented working diagnosis at the close of the index encounter and the final diagnosis established through standard care over the subsequent 90 days. Each encounter is classified as concordant if the working diagnosis matches the final diagnosis within ICD-10 chapter classification, discordant-correctable if the working diagnosis is wrong but caught and corrected within seven days, and discordant-persistent if the working diagnosis remains wrong beyond seven days. The intervention arm produced concordance rates of 91.3 per cent in cardiology, 94.1 per cent in oncology, 84.7 per cent in primary care, and 88.6 per cent in emergency department encounters. The contemporaneous control arm — encounters at campuses where the agent had not yet been deployed — produced concordance rates of 89.7 per cent, 92.4 per cent, 83.1 per cent, and 86.2 per cent in the same four specialties. The absolute differences are modest, ranging from 1.5 percentage points in oncology to 2.4 percentage points in emergency department encounters. The differences are statistically significant in cardiology, oncology, and emergency department comparisons (p < 0.05) and marginally significant in primary care (p = 0.061). The effect sizes are clinically meaningful but not dramatic.

Murphy's discussion section treats these effect sizes with appropriate analytic precision. She explicitly avoids the framing that the study demonstrates "AI improves diagnostic accuracy" in unqualified terms. The framing she uses instead is that the deployment of the diagnostic agent, operating under ACIS Tier 1 governance, produced a measurable improvement in encounter-level diagnostic concordance against the contemporaneous control. The improvement is modest, is consistent across three of four specialties, and reaches conventional statistical significance in three of those three. The improvement is also subject to selection-effect confounding that the within-deployment comparison arm helps but does not fully resolve. Within the intervention arm, clinician-invoked-agent encounters produced concordance rates that were 0.8 to 1.3 percentage points higher than clinician-available-but-not-invoked encounters across the four specialties. This sub-arm difference is not statistically significant in any of the four specialties when adjusted for clinician fixed effects, suggesting that the selection-effect confounder is real and that the headline 1.5 to 2.4 percentage point improvement attributable to the agent itself is closer to 1.0 to 1.5 percentage points after the confounding is fully accommodated. Murphy reports both estimates and is explicit about which she considers the more defensible.

The time-to-diagnosis endpoint is operationalised as the elapsed time from initial presentation to documentation of the final diagnosis. The intervention arm produced median time-to-diagnosis of 14.3 hours in emergency department encounters (against contemporaneous control of 16.1 hours), 4.2 days in cardiology (against control of 5.1 days), 11.7 days in oncology (against control of 13.4 days), and 6.8 days in primary care (against control of 7.4 days). The relative reductions range from 8 per cent in primary care to 18 per cent in cardiology. All four reductions are statistically significant. The clinical interpretation of the time-to-diagnosis improvement is more straightforward than the diagnostic accuracy interpretation because the time-based endpoint is less subject to the selection-effect confounder: even clinicians who do not invoke the agent during their encounters are operating within deployment environments where care pathways have been adjusted to accommodate the agent's existence, and the workflow optimisation effects can be captured at the deployment-level rather than at the encounter-level. Murphy treats the time-to-diagnosis improvements as the strongest finding in the primary endpoint set. The framing is defensible.

The length-of-stay endpoint is operationalised as inpatient days for cardiology and oncology encounters that resulted in admission, and as observation stay duration for emergency department encounters. Primary care encounters were excluded from the length-of-stay endpoint because primary care does not produce inpatient admissions directly. The intervention arm produced mean inpatient length-of-stay of 4.1 days in cardiology (against control of 4.6 days), 6.3 days in oncology (against control of 6.8 days), and observation stay of 6.7 hours in emergency department encounters (against control of 7.4 hours). The reductions are modest in absolute terms but consistent in direction, and all three reach statistical significance. The length-of-stay finding is methodologically the weakest of the three primary endpoints because it is subject to confounding from changes in admission criteria, discharge planning protocols, and bed-availability constraints that operate independently of the diagnostic agent deployment. Murphy's sensitivity analyses, reported in supplementary appendix Tables S7 through S12, suggest that approximately half of the observed length-of-stay reduction is attributable to the diagnostic agent deployment and approximately half to other operational changes Geisinger implemented during the operational window. The diagnostic-agent-attributable component is approximately 0.25 days in cardiology, 0.30 days in oncology, and 0.40 hours in emergency department observation. These are smaller effects than the headline numbers suggest, and Murphy's discussion is appropriately cautious in interpreting them.

The effect sizes are modest, consistent in direction, and large enough to matter. That is the most honest summary of what eighteen months of operational data actually shows.

The secondary endpoints: burnout and patient satisfaction

The physician burnout endpoint is operationalised as the change in score on the Maslach Burnout Inventory — Human Services Survey (MBI-HSS) administered to attending clinicians in the intervention and control arms at three timepoints: baseline (within thirty days before deployment), six months post-deployment, and eighteen months post-deployment. The MBI-HSS produces sub-scale scores on emotional exhaustion, depersonalisation, and personal accomplishment, with higher emotional exhaustion and depersonalisation scores indicating greater burnout and higher personal accomplishment scores indicating less burnout. The intervention-arm clinicians (n = 412) showed mean emotional exhaustion score reduction of 4.7 points from baseline to eighteen-month follow-up, against contemporaneous-control-arm clinicians (n = 287) who showed mean emotional exhaustion score reduction of 1.2 points over the same window. The depersonalisation sub-scale showed a smaller but consistent pattern (mean reduction of 2.1 points in intervention, 0.4 points in control). The personal accomplishment sub-scale showed essentially no change in either arm. The intervention-arm emotional exhaustion reduction is statistically significant against the contemporaneous control (p < 0.001) and represents a clinically meaningful effect on a sub-scale that has been validated as a leading indicator of clinician attrition.

The burnout finding is the most consequential secondary endpoint and the one that has generated the most peer-institution discussion since the publication's online posting on 19 May. Two structural factors warrant attention. First, the burnout reduction is concentrated among clinicians who invoke the diagnostic agent in the majority of their encounters; clinicians who invoke the agent in fewer than 30 per cent of their encounters showed burnout-score changes statistically indistinguishable from the control arm. The differential effect by usage pattern suggests that the burnout benefit is mediated by actual workflow integration of the agent rather than by the agent's mere availability. Second, the burnout reduction is larger in primary care and emergency department clinicians than in cardiology and oncology specialists. The differential by specialty corresponds to the relative documentation burden across the four specialties: primary care and emergency department clinicians document the largest volume of encounters per shift, and the ambient documentation capability addresses the documentation workload directly. The pattern is consistent with what the published clinical AI literature has predicted but has rarely demonstrated with this methodological rigour.

The patient satisfaction endpoint is operationalised as the change in the Press Ganey "would recommend this provider" score, measured on a 1-5 scale at the six-month and eighteen-month timepoints. The intervention arm showed mean score increases of 0.18 points in cardiology, 0.12 points in oncology, 0.21 points in primary care, and 0.08 points in emergency department encounters. The contemporaneous control arm showed mean score increases of 0.04, 0.02, 0.06, and 0.03 points in the same four specialties. The intervention-control differences are statistically significant in cardiology, primary care, and oncology, and not significant in emergency department encounters. The effect sizes are small in absolute terms but are clinically and operationally meaningful in the context of Press Ganey-anchored hospital quality measurement: a 0.20-point increase in the "would recommend" score corresponds to a movement of approximately 8-10 percentile points in the national Press Ganey distribution, which has material implications for Hospital Compare star ratings and value-based payment reimbursement. The patient satisfaction finding has been one of the elements most highlighted in the peer-institution response.

Murphy's discussion treats the patient satisfaction finding with appropriate caution. The Press Ganey instrument is widely used but is methodologically constrained: it surveys a non-random subset of patients who choose to respond, and the response rate is typically 15-25 per cent. Murphy's supplementary analyses, reported in appendix Tables S15 through S18, examine whether the response-rate selection biases the observed satisfaction-score change estimates. The conclusion is that the response-rate biases are present but small, and that the observed satisfaction improvements are not artifacts of differential response. The conclusion is methodologically defensible but the Press Ganey instrument's structural limitations mean that the satisfaction endpoint is not equivalently rigorous to the diagnostic accuracy or time-to-diagnosis endpoints. Murphy frames it accordingly. The peer-institution response has tended to take the satisfaction finding less cautiously than Murphy's own framing warrants, which is a familiar dynamic with patient-satisfaction endpoints in the published clinical literature.

Methodological strengths, limits, and the publication venue choice

The methodological strengths of Murphy's study are substantial. The four-arm comparison structure is more rigorous than the simple before-and-after designs that have characterised most of the published clinical AI outcome literature. The eighteen-month operational window is the longest at this level of design rigour in the multispecialty clinical AI literature. The selection-effect confounding is addressed through the within-deployment comparison arm rather than ignored. The endpoints are pre-registered (the study protocol was registered on ClinicalTrials.gov as NCT06234812 on 14 October 2024). The statistical analysis plan is documented in detail in supplementary appendix B. The institutional review board approval is current and the data sharing commitments are explicit. The study is, in short, an example of how clinical AI outcome research should be designed when the institution conducting it is committed to producing evidence rather than marketing material.

The methodological limits are equally important. The cohort is drawn from a single integrated delivery system in a single US region (central and northeastern Pennsylvania), and the population represented is materially distinct from urban academic medical centres, safety-net hospitals, and community hospital networks. The generalisability of the findings beyond an integrated-delivery-system context with sophisticated clinical informatics infrastructure is unestablished. The four specialties studied are not representative of the full clinical AI deployment landscape; specialties not included — surgery, paediatrics, psychiatry, obstetrics — may produce materially different outcome patterns. The agents deployed are from three named vendors (Suki AI, Abridge, Aidoc), and the findings do not generalise to clinical AI capabilities from other vendors or other architectural approaches. The eighteen-month operational window is the longest in the published literature but is still short relative to the multi-year operational horizons that institutional deployers and regulators care about. Murphy reports each of these limitations in the discussion section with appropriate weight. The limitations do not undermine the study's contributions; they bound them.

The choice of NEJM AI as the publication venue is itself analytically important. NEJM AI launched in January 2024 as the New England Journal of Medicine's dedicated AI-and-medicine journal, with editor-in-chief Dr. Isaac Kohane (Harvard Medical School) overseeing an editorial board drawn from the senior ranks of academic medicine, clinical informatics, and AI/ML research. The journal has, in its first 18 months, established itself as the most prestigious venue for peer-reviewed clinical AI outcome research, with rejection rates published quarterly that have ranged between 87 and 92 per cent. Geisinger's publication clearing the NEJM AI editorial process is a substantive credential — it indicates that the methodology cleared the most rigorous peer review currently available in the clinical AI literature. The alternative venues that Murphy's team considered, according to published commentary, were Lancet Digital Health (a strong second choice with a similar editorial-rigour reputation) and JAMA Network Open (a third option that would have produced faster publication timeline but less prestige). The choice of NEJM AI privileges the prestige and rigour over publication speed. The publication was originally submitted on 14 December 2025 and cleared the editorial process over five months. NEJM AI's median time-to-publication is in line with what major clinical journals offer; the venue is not a fast-publication outlet.

The peer-institution response

The first ninety-six hours of peer-institution response to Murphy's publication produced an unusually diverse set of public reactions across the major academic medical centres that have been tracking Geisinger's deployment. Three patterns are analytically important. First, the institutions operating under ACIS-aligned governance — Cleveland Clinic, Mount Sinai, UCSF (under its hybrid UCRE framework), and the three European adopters — have uniformly endorsed the methodological rigour and treated the findings as validating evidence for the broader ACIS-aligned approach. Cleveland's CMIO Dr. Marcus Osei published a Twitter thread on 20 May highlighting the study's selection-effect-confounder treatment as exemplary methodology and explicitly framing the diagnostic accuracy and burnout findings as consistent with Cleveland's own internal operational data, which has not been independently published. The framing is rhetorically charged but factually accurate within what Cleveland has publicly disclosed.

Second, the institutions operating under CARS-aligned governance — Mayo Clinic, Stanford, Johns Hopkins, MGB, Penn, Northwestern — have responded more cautiously. Mayo's response, published on the Mayo Clinic Proceedings blog on 22 May under the authorship of Dr. Sarah Mendel-Vivelo, framed Murphy's findings as broadly consistent with what CARS-aligned deployments have been producing internally but emphasised that the comparison framework matters: ACIS-aligned and CARS-aligned deployments produce different audit-log architectures and different per-capability governance structures, and the operational outcome comparison would benefit from cross-framework standardisation that does not yet exist. Stanford's CHIO Dr. Lisa Mendelsohn published a similar position in her institution's quarterly Digital Health Bulletin, dated 23 May. The tone of the CARS-aligned responses is supportive but with implicit caveats around framework comparability. The implicit message — that ACIS-aligned outcome data does not necessarily translate to CARS-aligned outcome expectations — is not stated explicitly but is legible to attentive readers.

Third, the publication has begun to influence the procurement decisions of institutions that have not yet committed to either framework. Three of the surveyed institutions in Q1 2026 — University of Michigan, University of Pittsburgh Medical Center, and Vanderbilt University Medical Center — had been operating in extended evaluation periods on clinical AI deployment without formal framework adoption. All three have indicated since 19 May that the Geisinger publication has accelerated their internal review timelines and that the diagnostic accuracy and burnout findings are influencing the procurement calculus. None has formally committed to a framework as of the analytical cut-off for this report (28 May 2026), but two have signalled informally that they are leaning toward an ACIS-aligned approach in light of the published outcome data. The Geisinger publication is functioning, in the institutional procurement landscape, as the evidence base that the field had been waiting for. Its specific findings will shape framework adoption and vendor selection decisions across multiple academic medical centres over the next twelve months.

The methodological response has been more mixed. Several methodological commentaries have appeared in journal correspondence sections and on academic informatics listservs questioning specific elements of Murphy's design. The most substantive critique, from Dr. Eric Topol at Scripps Research, has focused on the within-deployment comparison arm's statistical adjustment for clinician fixed effects. Topol's argument is that the fixed-effects adjustment may absorb too much of the agent-attributable variation because clinicians who choose to invoke the agent are also the clinicians who develop more sophisticated workflows around it, and the workflow sophistication is itself part of what the agent contribution should include. Murphy's response, published as a NEJM AI Correspondence on 24 May, acknowledges Topol's argument as a legitimate analytical alternative but defends the fixed-effects adjustment as the more conservative choice. Both positions are defensible. The methodological debate, taken on its own terms, is exactly the kind of debate the field needs more of.

What to watch

Geisinger's publication is the field's strongest empirical foundation to date but is one study at one institution. The next twelve months will determine whether other major institutions publish comparable outcome data, whether Murphy's methodology becomes a template for cross-institutional research, and whether the procurement-landscape effects materialise as the early peer-institution responses suggest. Five signals are the leading indicators.

  • Whether Cleveland Clinic publishes its own multispecialty operational outcome data under ACIS governance before the end of 2026; Cleveland has been operating ACIS for longer than Geisinger has been operating its diagnostic agent programme, and Osei has signalled in published commentary that a comparable outcome publication is in preparation, but the timeline has not been formally committed.
  • Whether Mayo Clinic publishes a CARS-aligned operational outcome study with methodology comparable to Murphy's; the field's evidence base would be materially strengthened by a side-by-side comparison of ACIS-aligned and CARS-aligned operational outcomes, and Mayo's own substantial deployment data could underwrite such a study if the institution chooses to publish it.
  • Whether the methodological debate between Murphy and Topol on within-deployment comparison adjustment moves into the broader clinical informatics literature; if other research teams adopt Topol's preferred analytical approach and produce systematically different effect estimates from those Murphy reports, the field will face a methodological reconciliation question that has been latent in the published literature for years and that the Geisinger publication has brought to the surface.
  • Whether the procurement-landscape effects observed at University of Michigan, UPMC, and Vanderbilt translate into formal framework commitments before the end of 2026; if Geisinger's publication produces three to five additional institutional framework adoptions over the second half of 2026, the empirical evidence base will have demonstrated material influence on the institutional adoption decisions, and the publication's structural impact will be confirmed.
  • Whether NEJM AI publishes follow-on outcome research from peer institutions at comparable methodological rigour; if the journal accepts and publishes two to three additional multispecialty operational outcome studies over the next twelve months, the venue will have established itself as the centre of gravity for serious clinical AI outcome research, and the field's empirical evidence base will move from one strong publication to a small but rigorous cohort of comparable studies that can be analysed in aggregate.

Frequently asked

What does the four-arm comparison structure in Geisinger's study actually accomplish methodologically?
The four-arm structure compares four distinct patient populations: encounters with the agent deployed and invoked (intervention arm); encounters with the agent deployed but not invoked by the clinician (within-deployment comparison arm); encounters at non-deployment campuses during the same operational window (contemporaneous control); and historical Geisinger encounters from before deployment (historical control). The structure allows the researchers to separate the agent's contribution from secular trends, from selection effects (clinicians who invoke the agent may be different from those who do not), and from confounding operational changes occurring at the same time. Most published clinical AI outcome studies use only one or two of these comparison arms, which leaves substantial confounding unaccommodated. Murphy's four-arm structure is materially more rigorous and produces effect estimates that the field should treat as more defensible than the simple before-and-after estimates that dominate the published literature.
How big are the diagnostic accuracy improvements in clinically meaningful terms?
The absolute differences range from 1.5 percentage points in oncology (94.1 per cent intervention vs. 92.4 per cent control) to 2.4 percentage points in emergency department encounters (88.6 vs. 86.2 per cent). After accommodating selection-effect confounding through the within-deployment comparison arm, the agent-attributable component is closer to 1.0 to 1.5 percentage points across the four specialties. The differences are statistically significant in three of four specialties and marginally significant in primary care. In absolute terms, the improvements are modest. In clinically meaningful terms, they correspond to one to three additional correctly-diagnosed patients per hundred encounters, which translates at Geisinger's encounter volumes to several hundred to several thousand additional correctly-diagnosed patients per year. The effect sizes are small at the encounter level and material at the system level.
Why is the burnout finding considered the most consequential secondary endpoint?
The intervention-arm clinicians showed a mean emotional exhaustion score reduction of 4.7 points on the Maslach Burnout Inventory's emotional exhaustion sub-scale over eighteen months, against a 1.2-point reduction in the contemporaneous control arm. The difference is statistically significant (p < 0.001) and represents a clinically meaningful effect on a sub-scale that has been validated as a leading indicator of clinician attrition. The finding is consequential because clinician burnout is a major operational and workforce-planning challenge for healthcare systems, and the published intervention literature on burnout has generally produced disappointingly small effects. A diagnostic-agent deployment producing a measurable burnout reduction at this magnitude is a material data point for healthcare workforce policy and for the broader argument that clinical AI can produce clinician-experience benefits alongside the patient-outcome benefits that have dominated the literature.
What is the methodological debate between Murphy and Topol about?
Murphy's study uses clinician fixed effects to adjust for the selection-effect confounder in the within-deployment comparison arm. Topol's argument is that this adjustment may absorb too much of the agent-attributable variation: clinicians who choose to invoke the agent are also the clinicians who develop more sophisticated workflows around it, and the workflow sophistication is itself part of what the agent contribution should include. Murphy's response is that the fixed-effects adjustment is the more conservative choice and that the effect estimates without the adjustment (1.5 to 2.4 percentage points) and with it (1.0 to 1.5 percentage points) bracket the actual agent contribution. Both positions are analytically defensible. The methodological debate is healthy and is exactly the kind of debate the field needs more of.
Why did Geisinger publish in NEJM AI rather than Lancet Digital Health or JAMA Network Open?
NEJM AI, launched in January 2024 as the New England Journal of Medicine's dedicated AI-and-medicine journal under editor-in-chief Dr. Isaac Kohane, has established itself in its first 18 months as the most prestigious venue for peer-reviewed clinical AI outcome research, with rejection rates of 87-92 per cent. Lancet Digital Health and JAMA Network Open are credible alternatives but the prestige and rigour of NEJM AI is presently higher. Geisinger's choice privileges the prestige and rigour over publication speed; the manuscript cleared the editorial process over five months. The choice signals confidence in the methodological rigour and is consistent with Murphy's broader positioning of the publication as a serious empirical contribution rather than a marketing-grade study.
How does Murphy's publication affect the procurement decisions of institutions not yet committed to a framework?
Three of the surveyed institutions that had been operating in extended evaluation periods without formal framework adoption — University of Michigan, UPMC, and Vanderbilt — have indicated since 19 May that the Geisinger publication has accelerated their internal review timelines. Two of the three have signalled informally that they are leaning toward an ACIS-aligned approach in light of the published outcome data. The publication is functioning, in the institutional procurement landscape, as the evidence base that the field had been waiting for. Whether the procurement-landscape effects translate into formal framework commitments before the end of 2026 is the principal question. If three to five additional institutional adoptions follow Murphy's publication over the second half of 2026, the structural impact will be confirmed.

Geisinger Health System's eighteen-month operational outcome publication is the strongest empirical foundation the clinical AI field has yet produced for the diagnostic-agent deployment category. The methodology is rigorous, the findings are modest but consistent in direction, and the limitations are stated with the kind of analytical precision that distinguishes serious clinical research from marketing-grade studies. The diagnostic accuracy improvements are real but small. The time-to-diagnosis improvements are real and meaningful. The length-of-stay improvements are real but partly attributable to confounding operational changes. The burnout reduction is real, substantial, and the most consequential secondary endpoint in the publication. The patient satisfaction improvements are real but methodologically constrained by the Press Ganey instrument's structural limitations. Each finding sits within its own evidentiary envelope, and the publication's discussion section treats each with appropriate weight.

What Murphy has produced is not a definitive answer to whether clinical AI deployment improves operational outcomes. It is the first piece of evidence rigorous enough that the question can begin to be answered. The peer-institution response in the first ninety-six hours suggests that the field recognises the contribution and that the procurement-landscape effects will be material. The next twelve months will determine whether Cleveland, Mayo, and other major institutions follow Geisinger into the published outcome-research literature at comparable methodological rigour, whether the methodological debate Topol opened produces healthy convergence on analytical standards, and whether the institutional adoption decisions that are presently in extended evaluation translate into formal commitments shaped by the empirical evidence Geisinger has now made available. The publication is one study. The field needs more like it. The publication has set the standard.

More from Health →