The Proof Problem, Part 1: AI Isn’t Replacing Our Job. It’s Adding a New One.
Friday August 28, 2026

by Rajendra Singh MD
Professor of Pathology, University of Pennsylvania
Co-Founder, PathPresenter

Every few months, someone tells me AI is coming for pathology. More recently, the message has become reassuring: AI isn’t coming for our jobs, it’s just coming for the boring parts. I don’t think either framing gets the real change quite right.
A recent talk by Princeton computer scientist Arvind Narayanan gave me a different way to think about it. His argument, in short, is that as AI gets better at doing things, human work doesn’t disappear. It moves from building and executing to evaluating, deciding, and being accountable for the outcome. I think that distinction matters enormously in medicine. Narayanan uses the analogy of a crane operator: a more powerful crane doesn’t eliminate the operator; it changes the job from doing the lifting to directing it and being responsible for what happens. In some ways, that may be what is beginning to happen to us.
Think about what is already happening. A pathologist can experiment with a classifier far more easily than a few years ago. A radiologist can receive AI-generated findings. A cardiologist can get an automated ECG interpretation. Soon, almost every specialty will have models sitting somewhere between clinical data and clinical decisions. Building the model is getting easier. Proving that it deserves to influence patient care is not. That may turn out to be one of the most important responsibilities physicians have in the AI era.
Someone still has to ask: Does this model work on our patients? Does it fail differently in different populations? What happens when the scanner, stain, laboratory, protocol, or patient population changes? Is it learning the biology we think it is learning, or exploiting a shortcut we haven’t noticed? And perhaps most importantly, how much should I trust it when there is a real patient on the other side of the prediction? AI can help us do parts of that evaluation, but the standard for what counts as enough evidence, and the accountability for acting on it cannot simply be delegated back to the system being evaluated. That requires clinical judgment.
For years, much of the conversation has been about whether physicians need to learn to build AI. I increasingly think that is only half the question. We also need physicians who know how to validate AI and to determine when, where, and how much it can be trusted. That doesn’t mean physicians shouldn’t learn to build models, or that model development belongs only to engineers and vendors. Quite the opposite: understanding how these systems are built will make us better users and better evaluators of them.
But building and validating are different responsibilities. Physicians have always had to evaluate evidence before acting on it. AI doesn’t create that responsibility from scratch; it expands it. Determining whether an AI system is safe, generalizable, clinically meaningful, and trustworthy is not only a technical question. It is a clinical one too.
And right now, there is a problem: in most health systems, we have not built the infrastructure to do this well. The vendor has evidence. The paper has an AUC. The regulator may have cleared the product. IT can deploy it. But who determines whether it works on our patients, in our workflow, with our instruments, and under the conditions in which we will actually use it?
That is the proof problem, and I suspect solving it will become one of the defining responsibilities of medicine in the AI era.
Over the next few posts, I want to get practical about what that actually means: what to ask an AI vendor, what makes a validation convincing versus what just sounds convincing, how models quietly fail after they’re deployed, and what infrastructure health systems will need before AI can be trusted as part of routine clinical care. I’ll use pathology as my laboratory, because that’s where I sit. But this problem belongs to every specialty, and solving it will require clinicians, scientists, engineers, health systems, regulators, and industry working together.
Next: why “it worked in the paper” might be the least useful sentence in clinical AI, and what we should be asking instead.
This post was prompted in part by Arvind Narayanan’s talk and essay, “What Will Be Left For Us to Work On?” His framing of how AI shifts human work from execution toward evaluation and accountability is well worth reading..
About the Author
Dr. Rajendra Singh is a Professor of Pathology at the University of Pennsylvania and co-founder of PathPresenter. He serves as a member of the Digital and Computational Pathology Committee of the CAP, Editorial Board of the WHO for Classification of tumors, 5th Edition and the Board of Digital Pathology Association.
More Posts
- The Proof Problem, Part 1: AI Isn’t Replacing Our Job. It’s Adding a New One.
- The Business of Pathology: Unlocking a Lab’s Value
- Expanding Access to Hantavirus Education: Sharing Images Through the PathPresenter Public Library
- The Future of Pathology Part 3: The Delivery Problem and the Infrastructure of Intelligence
- What ASCO Taught Me About the Future of Pathology – Part 2
