What Watson for Clinical Trial Matching taught me about the Generative AI era
The AI project I helped lead in oncology was specifically scoped, measured, and built with quality measures baked in from the start. That is exactly why it worked — and what today’s models risk forgetting.
If you only read one screen
- The project was IBM Watson for Clinical Trial Matching: cognitive computing that read structured and unstructured EMR data to screen cancer patients against trial eligibility criteria and assess whether a site could realistically accrue a given trial.
- We piloted it on both ends of care — Mayo Clinic (academic, with IBM Watson Health and Novartis) and Highlands Oncology Group, a large community cancer center.
- It worked because it was bounded: it did the heavy screening fast and reliably — in our study, cutting screening from 110 minutes to 24 — excluded ineligible patients dependably, and left the final call to a clinician.
- The lesson for the ChatGPT era: value in medicine comes from scoping the problem, keeping a human in the loop, and respecting the workflow and rules — not from a bigger, more general box.
There is a lot of excitement about generative AI in healthcare. Before we treat it as a magic box, it is worth remembering what actually worked the last time the industry put AI to work in oncology.
I helped lead one of those efforts: the IBM Watson Clinical Trial Matching program, as the Novartis partner in an alliance that also included Mayo Clinic and community practices. The problem was real and specific. Fewer than 5% of cancer patients enroll in a clinical trial, and roughly one in five trials closes for poor accrual. Eligible patients were being missed simply because manual screening could not keep up.
What we built was narrow by design. Watson’s cognitive computing read a patient’s structured and unstructured EMR data, extracted the attributes that matter, and compared them against clinician notes training and a corpus of trained trial eligibility criteria - both to match individual patients to trials and to help a site judge whether it could feasibly accrue a given protocol. We deliberately piloted across the spectrum of care: at Mayo Clinic, and at Highlands Oncology Group, a large community cancer center where most patients are actually treated.
And it worked — but read carefully how. In our published study of nearly 1,000 breast cancer patients, the system agreed with expert eligibility determinations 81 to 96 percent of the time, with high specificity and sensitivity, and it reliably excluded the ineligible. Most tellingly, screening that took a person 110 minutes took 24 with the system. The value was speed and reliable triage at scale. The clinician still made the decision.
That is the part I would underline for anyone deploying large language models in medicine today. The Watson project delivered because it was scoped to a bounded workflow with clear criteria and a human decision point — not because it answered any question you could ask it. The hard problems were never the model. They were the messy, unstructured data, the regulatory and clinical context, and fitting into how oncology teams actually work.
Today’s models are the opposite of scoped. They are trained on everything and pointed at everything. The honest questions from the Watson era still apply: what does the system actually know, what can you trust, and where does a human stay accountable? Innovation does not become value in healthcare because the box got bigger. It becomes value when you bound the problem, respect the constraints, and keep people in the loop.
So Watson performed well as a proof of concept - it even outperformed expectations. But yet, the analysis of Watson was that it was widely perceived as a failure. How did that happen? Read about the shortcomings and issues in my companion article.
This draws on work I co-authored in JCO Clinical Cancer Informatics (2020) and presented at ASCO (2018). Want to talk AI in medical affairs and clinical development? Book a conversation.