The technology worked. The program still didn’t.
Six reasons Watson for Clinical Trial Matching never reached scale — and where health-AI programs still fail today.
If you only read one screen
- The network never reached the sites where we had a critical mass of studies — and partial coverage created its own feasibility, training, and IRB burden.
- The pricing exceeded the cost of running the studies themselves.
- The vendor made the technology the story; clinicians cared about endpoints and patients.
- A sidecar outside the system of record does not get used, even by willing teams.
- In regulated research the sponsor owns every risk — and asking a site to onboard another vendor pulls in its entire compliance apparatus. IBM’s GCP auditability was flawless; the barrier was organizational.
- One brand stretched across everything meant every unrelated failure became our problem.
My previous piece covered what genuinely worked about Watson for Clinical Trial Matching: measurable, validated, human-in-the-loop. So it is fair to ask why it never became the standard. Six reasons, from inside the program.
1. The network never reached the sites that mattered
A matching platform is only as valuable as the trials and sites it covers. Our case depended on reaching the sites where Novartis had a critical mass of studies — and it never scaled to them. A tool that performs beautifully where you have two studies does not move a portfolio.
And partial coverage is not neutral — it carries its own cost. The moment a single study runs different technologies at different sites, you inherit feasibility, training, and IRB implications purely because those sites now carry different responsibilities. Uneven adoption does not just fail to help; it adds burden to the study itself.
2. The pricing did not match the value
Watson came to market with premium, platform-level pricing. In our evaluation it could exceed the cost of running the studies themselves — while still not delivering coverage at the sites that mattered most. Health-AI has to price against the value it actually displaces, not the ambition of the platform.
3. The technology became the story
In the vendor’s telling, the technology was the headline. For us — and far more importantly for clinicians — the story was the study endpoints and the impact on patients. When the narrative centers on the vendor’s capability instead of the customer’s mission, you lose the room, and you lose the clinicians whose adoption you need.
4. A sidecar outside the system of record does not get used
This is the one I would most want people to hear. Our site study teams were willing participants; they wanted it to work. But the solution lived outside the EMR — the system of record. In a busy clinic, the ability to simply forget a sidecar tool is real and material. If it is not in the workflow people already live in, enthusiasm will not save you.
5. In regulated research, the sponsor owns the risk — and the site owns the onboarding burden
Technology partners enjoy enormous public recognition, and in consumer contexts they can seem largely unharmed by inaccuracy or hallucination. Clinical research does not work that way. In an IRB-regulated environment, one data breach, one suspicion of harm, one missing GCP audit trail is the sponsor’s responsibility. No amount of hand-waving about novel technology and the promise of the future of clinical trials absolves a sponsor of that.
To IBM’s genuine credit, they met that bar: full GCP auditability, a flawless audit trail of every piece of data consumed and every suggestion ranked to the clinician. The technical compliance was there. The remaining barrier was organizational. Asking a site to take on another technology provider in order to access clinical trials is an enormous ask — it pulls an entire organization’s guardrails and standards into the decision. That is not a software problem, and no amount of model performance solves it.
6. One brand across everything means every failure is your failure
IBM put the Watson name on everything from taxes to trial matching to genomics. A stumble in any of those tainted the name across all of them. By the time we were recruiting new sites, conversations started with two strikes against the brand before we said a word about what the tool actually did.
None of this is an argument against AI in clinical development. The first piece covers what worked, and it worked well. It is an argument that in healthcare the model is rarely what decides the outcome — coverage, pricing, narrative, workflow, institutional risk, and brand decide it. Those are the same six places today’s generative-AI programs are most likely to fail.
Want to talk about AI in medical affairs and clinical development? Book a conversation.