LINK TO ORIGINAL ARTICLE: https://ai.nejm.org/doi/full/10.1056/AIp2600029
Below; Written by CHAT GPT.
Translating Pathology AI into Clinical Impact
A Review and Assessment of Su et al. (2026)
1. Purpose and Scope
In “Bridging the Gap — Translating AI in Pathology into Clinical Impact,” Su and colleagues address one of the central contradictions in digital pathology: the scientific literature has produced increasingly capable artificial intelligence systems, yet routine clinical deployment remains limited. The paper argues that this gap cannot be explained by algorithm performance alone. Instead, it reflects a layered set of technical, economic, operational, evidentiary, and governance barriers that must be resolved in sequence.
The authors begin with a striking comparison. Among 1,430 artificial intelligence- or machine learning-enabled medical devices in the FDA public inventory, only five are described as modern pathology AI tools based on whole-slide images, compared with 1,094 devices listed under radiology. At the same time, only approximately 4% of U.S. pathology slide volume is reportedly read digitally. These figures are used to establish the paper’s central premise: pathology AI cannot become commonplace until digital pathology itself becomes a reliable production environment. Yet digitization alone is not enough. Radiology, despite being “born digital,” has also experienced slower AI adoption than early expectations suggested. The authors therefore place infrastructure alongside workflow integration, reimbursement, validation, trust, medicolegal accountability, and quality management as coequal determinants of clinical adoption.
The paper defines clinical impact across three domains. The first is workflow performance, including turnaround time, concordance, and reduction in unnecessary ancillary testing. The second is clinical research, including trial recruitment efficiency and consistency of pathology endpoints. The third is patient outcomes, where prospective evidence is available. This is an important framing choice because it avoids equating technical accuracy with clinical value. A model may perform well on a benchmark yet contribute little to care if it slows workflow, fails at new sites, or produces no measurable effect on diagnosis, testing, treatment, or trial operations.
2. The Article’s Central Framework
The paper’s most important conceptual contribution is its hierarchical model of adoption barriers. Rather than presenting a miscellaneous list of implementation problems, the authors organize the field into successive layers. Foundational infrastructure and economic barriers come first. Postdigitization challenges involving sample variability, generalizability, and trust follow. Only after those barriers are addressed do the authors turn to strategic clinical pathways, implementation governance, and global equity.
This sequencing matters. A laboratory cannot derive value from an algorithm if slides are not scanned consistently, images cannot be retrieved quickly, network performance is inadequate, or the result appears in a separate application that disrupts diagnostic workflow. Similarly, a technically integrated system cannot be trusted if the model was trained on narrow data, performs differently across laboratories, or lacks an institutional process for monitoring failure.
The article therefore reframes pathology AI as a system rather than a software product. The relevant system includes tissue preparation, staining, slide production, scanning, image transmission, storage, retrieval, viewing, algorithmic inference, LIS integration, reporting, security, validation, training, monitoring, and governance. The model is clinically useful only when the entire chain functions reliably.
3. Foundational Digital Infrastructure
Su et al. correctly identify digitization as the first barrier. Pathology differs from radiology because the source material is not inherently digital. Glass slides must first be converted into gigapixel whole-slide images. That process introduces capital costs, operational complexity, quality-control requirements, and substantial data-management burdens.
The authors characterize digitization as a costly overlay on existing analog workflows. In many settings, glass slides continue to be prepared, transported, archived, and available for review even after scanning is introduced. Digital pathology may therefore add cost before it replaces any existing expense. This helps explain why apparently favorable technology can encounter institutional resistance even when its long-term value appears plausible.
The paper emphasizes the need for high-throughput scanning, rapid image transmission, scalable storage, and reliable retrieval. These are not merely engineering concerns. If slide loading is slow, overlays lag, or images must be transferred manually, the technology can reduce rather than improve productivity. The authors therefore treat latency, uptime, and workflow continuity as clinical performance variables.
Table 1 is especially useful here. It recommends high-throughput scanning, cloud or tiered storage, edge inference, DICOM adoption, vendor-neutral APIs, and end-to-end HIPAA-compliant data paths. It also proposes measurable endpoints, including per-slide scanner throughput, viewer-overlay latency, system uptime, encryption compliance, and preservation of turnaround time. The insistence on in-viewer overlays rather than application switching is particularly practical. It recognizes that a clinically valuable algorithm must appear where the pathologist is already working.
4. Interoperability and Workflow Integration
The paper appropriately treats interoperability as more than file compatibility. DICOM adoption can improve the exchange of whole-slide images, but the implementation problem extends across case identity, specimen hierarchy, worklist synchronization, user authentication, report generation, and error recovery.
A functioning clinical workflow must preserve the relationship among patient, accession, specimen, part, block, slide, stain, image, algorithm, and result. If those relationships are not maintained accurately, even a technically correct algorithm may be unsafe or unusable. The authors do not explore this full hierarchy in detail, but their emphasis on LIS and viewer integration points in the right direction.
The article’s broader implication is that pathology AI should be evaluated at the level of the case workflow. A scanner may perform well in isolation, an algorithm may achieve a high area under the curve, and a viewer may render images smoothly, yet the complete system may still fail if case routing is unreliable or results cannot be incorporated into the diagnostic report.
This systems perspective is one of the paper’s strengths. It shifts attention from isolated component specifications to operational outcomes such as throughput, latency, uptime, turnaround time, and user burden.
5. Security and Data Governance
Su et al. also recognize that whole-slide images are sensitive clinical data. The image, slide label, associated metadata, and linkage to the laboratory record may all contain protected health information. The paper therefore includes privacy and security across the full data path rather than treating them as late-stage compliance issues.
The authors point to federated and on-device inference as ways to reduce the transfer of sensitive patient data. They also suggest that security controls now associated with controlled genomic data may eventually extend to high-resolution medical imaging. This is a reasonable forward-looking concern. Pathology images may encode not only direct identifiers but also biological information that could become increasingly inferable as models improve.
For infrastructure providers, the implication is that clinical deployment requires more than encrypted storage. It requires identity management, role-based access, audit trails, secure transfer, model isolation, version control, and incident response. The article does not fully specify these requirements, but it correctly places them within the core adoption framework.
6. Economic and Reimbursement Barriers
The economic discussion is central to the paper. Hospitals and pathology departments often bear the cost of scanners, storage, service contracts, software, interfaces, validation, training, and technical support. Yet the financial benefits may accrue elsewhere in the institution.
For example, faster interpretation may benefit surgery or oncology. Reduced molecular testing may benefit the payer or health system. Improved trial recruitment may benefit a research office or pharmaceutical sponsor. Better throughput may support enterprise capacity without generating a direct payment to pathology. This separation between the cost center and the beneficiary creates a persistent barrier to adoption.
The authors note that reimbursement for AI-enabled decision support remains limited. They call for new billing approaches, bundled payments, and value-sharing arrangements. Importantly, they do not assume that fee-for-service reimbursement is the only solution. Their framework allows for clinical-trial sponsors, oncology programs, and health systems to share infrastructure costs when the value appears outside the pathology department.
Table 1 translates this problem into measurable endpoints. It proposes positive return on investment per case, sustainable storage and transmission cost per whole-slide image, documented recovery of digitization costs, increased use of AI-related billing codes, and formal value-sharing arrangements. These are useful starting points, although the article does not provide detailed cost models or distinguish among capital purchase, subscription, per-slide, cloud, and hybrid payment structures.
7. Sample Variability and Domain Shift
One of the paper’s strongest technical sections concerns sample variability. Pathology images differ across institutions because of fixation, processing, section thickness, stain chemistry, scanner optics, focus, compression, and acquisition protocols. These differences can create powerful site-specific signatures.
The authors cite evidence that AI can predict the originating institution of a slide with an area under the curve above 0.9. This is an important warning. A model may appear to recognize disease while actually learning technical features associated with a particular laboratory, scanner, or patient population.
The danger is especially serious when site and outcome are correlated. Suppose one institution treats more advanced disease and also uses a distinctive stain or scanner. A model may learn the institutional signature instead of the biological feature of interest. It may then perform well in internal validation but fail elsewhere.
Su et al. propose color normalization, stain-transfer techniques, multi-institutional training, domain adaptation, personalized federated learning, continual learning, and postdeployment recalibration. They also recommend a particularly insightful endpoint: site-of-origin prediction should approach chance after mitigation. That measure directly tests whether residual institutional information remains embedded in the data.
The article also makes an important distinction between interoperability and biological standardization. DICOM can improve file exchange, but it does not correct fixation, staining, sectioning, or scanner variability. Standardized data formats and standardized tissue preparation are related but different problems.
8. External Validation and the Trust Gap
The authors attribute clinician skepticism largely to concerns about external validity. This is appropriate. Pathologists are unlikely to trust a system that performs well in a curated development cohort but has not been tested under local conditions.
The paper warns that strong aggregate metrics can conceal clinically important failure modes. A model may achieve a high overall area under the curve while performing poorly in certain institutions, demographic groups, tissue types, or scanner environments. It may also fail in the cases where clinical support is most needed.
The authors therefore favor task-specific evaluation. A triage model may require very high sensitivity. A quantitative biomarker may require reproducibility and low interobserver variance. A molecular prescreener may need strong negative predictive value and demonstrable reduction in unnecessary confirmatory testing. A prognostic model may require calibration, discrimination, and evidence that the score adds value beyond standard clinical variables.
This task-oriented approach is far more useful than applying a single evidentiary template to all pathology AI products.
9. Explainability, Human Factors, and Clinical Trust
The paper takes a balanced view of explainability. Heatmaps and visual overlays may help pathologists understand where the model is focusing, but they do not prove that the model is biologically valid. A heatmap can look plausible even when the prediction is partly driven by an artifact.
The authors therefore treat explainability as one component of trust rather than a substitute for external validation. Trust also depends on equity, user training, liability allocation, escalation pathways, and clear definitions of human and algorithmic responsibility.
This is one of the more mature elements of the paper. The authors implicitly recognize that clinical use requires decisions about what happens when the pathologist and algorithm disagree, when the model reports low confidence, when the slide fails quality control, or when the case falls outside the intended-use population.
The paper does not prescribe a universal escalation policy, nor should it. Different applications will require different governance. A triage system, biomarker quantifier, molecular predictor, and autonomous diagnostic tool create different levels of risk and different requirements for human review.
10. AI as a Digital Copilot
The first major translational pathway identified by Su et al. is AI as a digital copilot. This concept positions AI as a second reader, triage assistant, quantitative aid, or diagnostic support tool rather than as a replacement for the pathologist.
The paper divides these applications into three readiness tiers. Near-term uses include second reading, triage, and some forms of quantification. Medium-term uses include tumor-infiltrating lymphocyte assessment, intraoperative evaluation, and prediction of molecular alterations from H&E slides. Longer-term uses include broader treatment stratification and prognostic inference.
This tiered approach is useful because it resists the tendency to describe all pathology AI as equally mature. The readiness of an application depends on the task, tumor type, validation setting, regulatory status, and workflow.
The authors also favor continuous quantitative outputs rather than simple binary labels. Probability scores and continuous measures can communicate uncertainty and allow thresholds to be adapted to different purposes. However, such outputs also require calibration and clear interpretation. A score is not clinically meaningful unless the user understands the population, endpoint, and threshold on which it was validated.
11. Diagnostic and Biomarker Applications
The digital copilot model encompasses several categories of use. AI may help prioritize high-risk cases, quantify biomarkers such as Ki67 or HER2, support differential diagnosis, identify educational cases, or predict molecular alterations from routine H&E images.
The paper is strongest when it presents these as distinct functions rather than as a single category of “AI diagnosis.” Each has different evidence requirements and workflow implications.
Biomarker quantification, for example, may be evaluated against interobserver reproducibility and consistency with adjudicated scoring. Triage may be evaluated by time to review and false-negative rate. Molecular prediction may be evaluated by performance against confirmatory testing and the degree to which testing can be safely reduced.
The authors cite encouraging evidence but appropriately qualify the maturity of these applications. Expert-level performance on narrow benchmarks does not guarantee reliable generalization across independent prospective cohorts.
12. Clinical-Trial Enablement
The second major pathway is clinical-trial enablement. This may represent one of the most practical near-term uses of pathology AI because it can create measurable value even before routine reimbursement is established.
AI may support trial prescreening, reduce unnecessary confirmatory sequencing, accelerate recruitment, predict treatment response, standardize pathology endpoints, and enable large-scale reanalysis of archived specimens.
The authors cite evidence that H&E-based prescreening could spare approximately 40% of confirmatory sequencing in colorectal cancer trial settings. They also describe international deployment that shortened recruitment timelines for specific genomic alterations.
The trial use case has important economic implications. Pharmaceutical sponsors may have a direct incentive to support scanning, infrastructure, storage, algorithm deployment, and site qualification. This creates a possible bridge from research digitization to routine clinical digitization.
The article also notes that AI-assisted pathology review may improve consistency in endpoints such as tumor burden and reticulin fibrosis. That could reduce variability across trial sites and increase statistical efficiency. However, this advantage depends on the algorithm itself being robust across laboratories and sample conditions.
13. Implementation and Governance
Su et al. emphasize that a validated algorithm is not yet a clinical service. Postdevelopment work includes local validation, procurement review, workflow integration, training, monitoring, recalibration, and explicit governance over failure modes.
Local validation is especially important because the receiving site may differ from the development environment in patient population, disease prevalence, tissue preparation, scanners, software versions, case mix, and reporting practices. Validation should therefore test not only analytical performance but also image quality, routing, display, latency, failure handling, and report integration.
The authors recommend that procurement decisions require external validation evidence and model cards. This shifts evidence review earlier in the process. Institutions should understand the model’s development population, intended use, exclusion criteria, scanner compatibility, subgroup performance, known failure modes, and update policy before deployment.
The concept of algorithmic stewardship is particularly important. It implies that responsibility continues after go-live. A deployed model requires an owner, a monitoring plan, a process for reviewing updates, and a protocol for responding to drift or failure.
14. Postdeployment Monitoring and Version Control
The article’s emphasis on monitoring is one of its most consequential implications. AI performance may change after deployment because of new scanners, altered stain protocols, software updates, case-mix changes, or population shifts.
A production platform should therefore monitor performance by site, scanner, stain, specimen type, subgroup, and algorithm version. It should also detect changes in override rates, abstention rates, quality-control failures, and turnaround time.
Version control is equally important. For every result, the system should be able to identify the source image, algorithm version, preprocessing method, threshold, output, and user interaction. Without that information, incident investigation and regulatory review become difficult.
The authors do not provide a detailed monitoring architecture, but their framework clearly implies that observability is a core product requirement. Inference alone is not sufficient.
15. Equity and Global Reach
The paper extends its analysis to low- and middle-income countries, where pathology services and specialist density may be limited. The potential benefit of AI triage and second reading may be substantial, but infrastructure constraints are also greater.
Su et al. recommend systems that tolerate intermittent connectivity, lower-magnification scans, and cloud-light deployment. They also emphasize that slides from low- and middle-income settings must be included in training and validation datasets.
This is essential. A system developed in a small number of well-resourced academic centers may not generalize to laboratories with different tissue processing, staining, scanners, disease prevalence, or case mix.
The paper proposes endpoints such as feasible deployment cost, geographic performance parity, expanded coverage in regions with few pathologists, and reduced time from specimen collection to expert-level interpretation. These measures appropriately connect technical deployment to service access.
16. Overall Assessment and Strategic Significance
Su et al. provide a concise but unusually operational account of why pathology AI has not yet achieved broad clinical impact. Their most important contribution is the recognition that adoption depends on the integration of infrastructure, economics, validation, workflow, governance, and evidence.
The article is not a systematic review, and it does not resolve the central questions of reimbursement, regulatory strategy, return on investment, or multi-vendor accountability. Much of the cited evidence concerns technical performance, concordance, and workflow rather than patient outcomes. The paper therefore should not be read as proof that pathology AI has already achieved widespread clinical utility.
Its value lies instead in the framework it provides. The authors make clear that the next phase of digital pathology will not be determined solely by algorithm accuracy. Success will depend on whether the complete system is reliable, interoperable, secure, locally validated, economically defensible, clinically integrated, and continuously monitored.
The paper also suggests that the most credible near-term pathways are narrower than the broadest claims often made for AI. Digital copilots, biomarker quantification, triage, molecular prescreening, and clinical-trial enablement offer measurable and operationally tractable opportunities. These applications can create value without requiring full autonomous diagnosis.
The central lesson is that pathology AI should be developed and evaluated as part of an end-to-end clinical platform. Scanner performance, storage architecture, viewer integration, LIS connectivity, algorithm output, model governance, and postdeployment monitoring are not separate commercial categories from the standpoint of clinical impact. They are interdependent components of the same operating system.
In that sense, the paper’s title is accurate. The gap between pathology AI research and clinical impact is not primarily a gap in model capability. It is a gap in translation.