One Med-AI Won’t Fit All

Artur O.'s avatarPosted by
AI is going to transform medicine radically, according to Prof. Isaac Kohane, the Editor-in-Chief of NEJM AI

“AI lacks genuine common sense in the real world. It can’t recognize the subtle cues conveyed by posture, body language, and other nonverbal signals”. Interview with Prof. Isaac “Zak” Kohane, Chair of the Department of Biomedical Informatics at Harvard Medical School and Editor-in-Chief of NEJM AI.

By Artur Olesch, 23-minute read

In 1987, you earned your PhD from Boston University and then completed postdoctoral work at Boston Children’s Hospital, where you still work as a pediatric endocrinologist. How did it happen that a pediatric endocrinologist became Chair of the Department of Biomedical Informatics at Harvard Medical School?

This question brings me back to an important decision in my life.

I grew up in Geneva, Switzerland. Back then, I knew I wanted to do research in biology. But I quickly noticed that people in Switzerland who wanted careers in biology were often doing postdocs in the United States and elsewhere. Opportunities in Switzerland were limited. You were essentially waiting for a professor at the top to retire before positions opened up. So I decided I would have to go elsewhere.

If you wonder where my American accent comes from… I attended the International School of Geneva. Even though television was in French and the kids in my neighborhood were French, the American students gave many of us American accents.

So I decided to go to America for college, and I attended Brown University. Although my major focus was biology, I was able, back in 1977, to take a computer science class.

By the time I graduated from college, I actually had several job offers, mostly from the defense industry, for computer science positions that paid much better than anything I would earn for many years during my clinical training.

Nonetheless, I decided to go to medical school. But within the first year, I was overwhelmed with anxiety – I had a completely different idea of what medicine was about. I thought it was much more like a science, whereas in reality it is much more like an art – and a very important art. I just didn’t know what I was going to do.

Fortunately, I met wonderful mentors. I was referred to my future PhD advisor, Peter Szolovits, a professor of Computer Science and Engineering at MIT. At the time, he was a young professor who had just come from the California Institute of Technology (Caltech). He was working on AI, specifically clinical AI. The technological stack was completely different from today’s AI, but many of the same questions were already being asked. The field was heavily influenced by cognitive science, which, interestingly, has become very relevant again today.

It was a wonderful time to do a PhD.

One day Pete told me, “Zak, you should finish your clinical training.” He probably said it in different words, but the message was clear: “I don’t get much respect from physicians. If you really want to do AI in medicine, you’ll have much more credibility if you complete your clinical training.”

So I did. I trained in pediatrics, and I have to say I loved it. It was probably the happiest period of my life: during my PhD, I was always thinking about research; when I was in the hospital, I was simply focused on taking care of patients.

I also became fascinated by the homeostatic regulatory loops in endocrinology, and the faculty in that field seemed very mathematically oriented. I had already started doing research with a PhD student from my former research group at MIT.

Again, I was fortunate to have great mentors. They recognized that I probably wasn’t going to become a molecular biologist. My future was going to be in computer science. So after finishing my clinical training, I continued practicing medicine, but gradually less and less, while doing more and more research.

Academia is actually much closer to business than people like to acknowledge. When you bring in a lot of grants and funding, people notice. Eventually, someone approached me and asked whether I would like to create a biomedical center at Harvard. Later, they decided to establish the Department of Biomedical Informatics.

At the same time, what had seemed like a rather idiosyncratic decision back in 1983, when I chose to pursue an MD-PhD in computer science, suddenly became very timely. By 2012, AI was experiencing its resurgence. Neural networks were scaling, we finally had enough data, and the computational infrastructure had matured. At least for image analysis, the results were becoming extremely promising.

My own research then shifted much more decisively toward AI.

And then you became an editor of NEJM AI.

About seven years ago, the leadership of the New England Journal of Medicine, where I served on the editorial board, asked me, “Would you like to start an NEJM AI journal?”

I said no.

They asked, “Why not?”

I replied, “AI is great, but the quality of the research right now is going to be interesting primarily to AI researchers, not to practicing physicians. I don’t think the time is right.”

They asked me again about three and a half years ago, and this time I said, “Maybe now is the right time.”

Then we got very lucky. Because of my academic background in computer science, I knew Peter Lee, the head of Microsoft Research. In October 2022, he sent me a mysterious email that said, “Zak, if you answer my phone call, you’ll find it very worthwhile.”

During that conversation, he told me about GPT-4. It was before most people had even heard of ChatGPT.

With Peter, you wrote “The AI Revolution in Medicine: GPT-4 and Beyond” – the first book that explored the potential of AI as the second pair of doctors.

Yes, we decided to write a book, and then the arrival of generative AI into the public consciousness dramatically accelerated the journal we had founded. The first issue came out in January 2023, so the timing was perfect.

At the same time, my department, which had previously focused primarily on machine learning and bioinformatics, began to pivot. Through new faculty recruitment, we increasingly expanded into AI, and it’s been a great journey.

Now you’re at Harvard Medical School, which is one of the places to follow and be for anyone who wants to stay up to date on AI research in medicine. What are the latest developments in AI that have impressed you most recently?

A few things have really surprised me.

First, I expected most of the uptake to happen on the economic and administrative side because I assumed healthcare institutions would be slow to adopt AI, which they have been. They started creating slow, ponderous governance structures: safety committees, ethics committees, and review processes to decide whether AI should be adopted.

But then one company, OpenEvidence, came in from the side. OpenEvidence essentially said, “If you’re a licensed physician with a National Provider Identifier, we’ll give you access to our platform.”

Suddenly, we went from having almost no official use of AI to more than 60–70% of American physicians using it. That was my first big surprise. Today, if you walk into almost any of our hospitals, you’ll see OpenEvidence open on most of the laptops at the nursing stations.

The second surprise has been how quickly we’re moving in robotics. It’s not fully there yet, but if you had asked me two years ago, I would have said robotics would be the last major breakthrough. Now I think you’re going to see a great deal of robotic assistance in specialized procedures: eye surgery, ovarian surgery, gallbladder surgery, and perhaps even some forms of skin surgery.

Part of the reason is that the potential benefit is significant, but surgeons are also natural experimenters; they are constantly refining the techniques. In some ways, they operate in a much less regulated environment than physicians prescribing drugs. There is a long tradition of surgical innovation and experimentation that would make many non-surgical physicians uncomfortable.

I’m less surprised by the fact that drug development is going to accelerate. It certainly will, but probably not as dramatically as some people expect. Every successful drug still has to go through clinical trials, and those still involve real physicians and real patients.

There are many ways AI can accelerate recruitment of both physicians and patients. The preclinical phase of drug development will definitely become much faster. The clinical phase is becoming less expensive because much of the paperwork involved in drug development is being automated using large language models, and patient recruitment is also improving.

But in the end, you still have to administer medications to real patients and observe both their therapeutic responses and their side effects. Biology still moves slowly in that respect.

The other development that’s been particularly interesting is the growing use of AI by patients. Around the world, including in the United States, patients often have to wait a long time to see a physician. So despite all of AI’s imperfections – its mistakes and hallucinations, which are becoming less frequent but still exist – patients are finding it genuinely useful. If they have a health concern, they use AI all the time.

It’s a fascinating situation because, even in Europe, where governments and regulators are very concerned about AI, consumers are quietly using it anyway.

I think this is creating a real change.

On one hand, especially among patients with chronic or serious illnesses, we’re seeing much better-informed patients. On the other hand, physicians, particularly younger physicians, are, in some respects, becoming less expert because they rely on AI in ways that I probably would have done myself if I were in their position.

That means they may not develop some of the skills they otherwise would have acquired. I’m not saying that’s necessarily good or bad, but it does mean that the knowledge gap between physicians and patients is narrowing. I think that’s going to fundamentally change the nature of medicine over the next ten years.

Patients are using AI, and AI companies are giving them new tools. Just a week ago, OpenAI launched ChatGPT Health. Will everyone soon have a personal AI that can access the data in their electronic health record and tailor its answers to health-related questions accordingly? What does it mean for healthcare and physicians?

I wrote an article for the New England Journal of Medicine called Compared with What? Measuring AI against the Health Care We Have.

If the comparison is with not having a doctor at all, which is increasingly the reality – in countries like France, England, and many parts of the United States – people may wait months to see a primary care physician. So what does AI mean in the absence of primary care?

We have a genuine workforce shortage. People are always worried that AI will replace doctors, but the reality is that we simply don’t have enough doctors to do the job. So I’m less concerned about AI agents helping patients.

The real question is whether we’ll develop safe practices.

Ideally, before every doctor’s appointment, AI – whether it’s ChatGPT Health or another system – should help patients prepare by suggesting the questions they should ask. During the consultation, AI should help ensure that the physician covers all the important issues. After the visit, AI should evaluate whether the physician made appropriate decisions.

At the same time, the patient’s AI should also determine whether the patient actually understood what happened during the visit and whether they are following the recommendations, or perhaps even correctly deciding not to follow a particular recommendation.

None of these processes are being formalized. They’re simply evolving in a chaotic, almost anarchic way. As a result, no one is systematically exploring how these systems should work together. Everyone is essentially making it up as they go along.

That’s what should concern us: The implementation is extremely uneven.

Paradoxically, at least in the United States, and I suspect in many other countries as well, people originally feared that AI would increase disparities between those who have access to healthcare and those who don’t.

But take Boston as an example. What do wealthy people do? They hire concierge physicians – doctors who care for only about one-tenth the number of patients seen by a typical primary care physician and are paid directly by their patients. People with money therefore have much greater access to human physicians than people without money.

In some sense, AI is helping reduce that gap.

Your recent research has centered on values embedded in AI systems.

This is something I’m most concerned about, too. When you visit a physician, sometimes you want someone who is conservative, someone who intervenes only when necessary. Other times, you want someone who is more aggressive, someone who says, “Let’s take care of this immediately.” Which approach is appropriate often depends on your own preferences, and patients frequently choose physicians based on these characteristics.

As part of our Human Values Project, we’ve studied different AI models and found that they are not interchangeable. They actually make different kinds of decisions. For example, Gemini tends to be more aggressive than OpenAI models in certain clinical situations. Different Anthropic models also vary, depending on their size and complexity. Some are more aggressive than others.

We’ve also published a preprint showing that these models generally agree with physicians around the world on most ethical questions. One interesting exception involves American physicians. Because of how medicine is practiced in the United States, physicians in the US tend to place greater emphasis on patient autonomy than physicians in countries such as China or Brazil.

These are characteristics of AI systems that most people don’t even realize exist. I think we need to stop treating AI models as though they’re all the same. They have different styles, different alignments, and different value systems, and those differences genuinely matter for patients.

Imagine a young woman with severe anorexia who refuses to eat. Do you wait because she says she’ll start eating tomorrow? Or, because you’re worried about a potentially fatal cardiac arrhythmia, do you insert a feeding tube immediately?

Physicians themselves are divided on that question. AI models are divided as well. Which approach do you want? The answer depends on who you are and what you value. That’s an area I’m spending a great deal of time studying.

The other major area I’m working on relates to systems like ChatGPT Health and the idea of unlocking patients’ healthcare data. The reality is that, in many parts of the world, and certainly in the United States, people are highly mobile. Your medical record is rarely stored in just one place throughout your life.

I believe the right long-term solution is for your data to follow you throughout your lifetime in the cloud, under your control. That would provide a true longitudinal health record. What we’re seeing is that AI models themselves are becoming increasingly interchangeable. They continue to improve, and competition is driving costs down.

What is not interchangeable is data. Data is becoming the most valuable resource. For individuals, personal health data is the most valuable resource because it determines not only whether AI can make the right decisions about your health, but also how your privacy is protected.

I think the next major social revolution will revolve around ownership of health data and the relationships among patients, AI companies, and hospitals.

Historically, hospitals have acted as the custodians of medical data. Various companies have built businesses by repurposing portions of that data. Now that AI is becoming a personal advisor, it is increasingly obvious that individuals should control their own health information. I think this will fundamentally influence both public policy and future business models.

In Europe, on the contrary, data is well protected, but there is a price for it – tools like OpenEvidence and ChatGPT Health are currently available.

I think that’s a problem. Europe has roughly 450 million citizens, and they deserve access to these capabilities. European physicians make mistakes too. They get tired too.

Patients shouldn’t rely on AI as their only doctor, but they should have AI available as a second set of eyes, something that can detect mistakes, omissions, or overlooked information. The challenge is finding a way for patients to control their own medical records without becoming prisoners of either the data itself or bureaucratic systems.

I appreciate the safeguards Europe is putting in place. But when those safeguards also prevent the beneficial uses of health data, we need to ask ourselves how we can do better. I actually think this is more of a technology question than a policy question. The policy only needs to change enough to enable the technology.

Think about how mobile people are within Europe. People move easily between countries, yet it’s still incredibly difficult to share healthcare data across borders. That simply doesn’t make sense. When someone who was born in Berlin lives in the UK, and half of their medical record remains in Berlin, it’s nonsensical that important information relevant to their care in England cannot be accessed and integrated instantly.

I think there is a real and urgent need for a unified, patient-controlled health record.

Hopefully, the European Health Data Space will help make that a reality.

Speaking of human error, some research already shows that AI often makes better diagnostic decisions. But what about AI working together with a human doctor? Should we always keep the human in the loop?

Yes. Those findings are both disturbing and, to some extent, artificial. Let me explain why.

What you said is absolutely correct. There have now been several studies showing that, for certain tasks, though certainly not all tasks, a doctor plus AI can actually perform worse than either the doctor alone or the AI alone. Essentially, there can be a pathological interaction where the physician doesn’t trust the AI, is annoyed by it, or responds to it in some unproductive way.

But part of the problem is that these studies are usually conducted under highly controlled conditions. You present a clinical vignette, everything is organized neatly, and the participant responds to that scenario.

The real world isn’t like that.

When a patient comes in complaining of pain, they don’t tell their story in the orderly way a clinical vignette is written. There’s a tremendous amount of noise. They might say something like, “I dropped the groceries when my son Charles handed them to me.” The challenge is making sense of all that noisy information and deciding what actually matters.

Right now, these systems are trained on data collected in ways that allow expert pattern recognition. They’re very good at recognizing classic textbook presentations. But they’re still not good enough at handling the chaos of the real world and distilling it into what is actually important. My prediction is that, in real clinical trials where an AI simply interacts with patients on its own, it will miss a number of important things. I’ve seen that happen.

The reality is that AI systems used in medicine, including ChatGPT, perform better the more information you provide and the better that information is organized. The more expert you are at collecting and structuring the relevant data, the better the AI performs. It can then take you the rest of the way by recalling details and knowledge that, as human beings, we may have forgotten.

But our common sense reasoning comes from millions of years of evolution. We evolved as animals that had to deal with constant distractions and learn which signals mattered and which did not. We instinctively recognize that a minor pain over here isn’t relevant, that someone looking in one direction doesn’t matter, while another subtle movement or posture might be highly significant.

That kind of real-world reasoning has not yet been properly tested in AI systems.

So I would be very surprised if, today, in actual clinical practice, doctors working with AI were not better than AI working alone. AI, by itself, simply won’t know which questions to ask and won’t recognize the subtle cues conveyed by posture, body language, and other nonverbal signals.

In ten years, AI may very well reach that level. It’s just not there today.

Some experts suggest general medical AI could solve those limitations. Believe it or not, many tech leaders herald AGI within a few years. Do you see any early signs of a single AI system that outperforms specialists in every field of medicine?

I don’t really like the term AGI, frankly, because I see intelligence as part of a spectrum. We’ll continue to see different capabilities improve over time. As these models become better, they’ll become more and more broadly competent.

Let’s even assume that one day we do have an AI that becomes truly expert across every aspect of medicine. If you talk to some of the leaders in the field, they’ll tell you we’ll have AGI tomorrow or next year. We already have AI systems that outperform humans on tasks like the International Mathematical Olympiad. That’s already happened.

Having genuine common sense in the real world, without requiring human interpretation, may still take another couple of years. But even if we get there, I don’t think there is any substitute for human beings acting as moral agents. We’ve lived with AI outperforming humans at chess for decades now, yet that hasn’t stopped us from playing chess with each other.

We’re going to live in a hybrid world.

I recently read about a bestselling novel that was initially self-published on Kindle. It became so successful that a major publisher acquired it for several million dollars. It’s now fairly clear that it was written by AI. So we already have an AI-written bestseller. But here’s the point: people are still going to write books. I don’t know what the balance will ultimately look like.

As far as medicine is concerned, though, over the next ten years I believe anyone undergoing a major operation without first speaking to an experienced physician is making a mistake.

You should talk to a doctor who has performed that operation many times and ask, “What do you think is the right decision for me?” For example, we know there are treatments for throat cancer where patients can choose between removing the voice box or trying to preserve it with other treatments. Removing the voice box may extend life by a few months or perhaps even a year. But many patients who choose that option later regret it deeply. Some ultimately say they would rather have had less time than lose their voice.

An experienced physician understands that. A wise doctor may tell a particular patient, “Don’t have your voice box removed.” That’s the human side of medicine.

Now, perhaps one day AI will become indistinguishable from human beings. But I prefer talking about the world we’re actually going to live in over the next ten years. If you were my friend and you were about to undergo surgery, I would absolutely tell you to speak with a physician who has real experience with that procedure.

You need both the physician’s moral judgment and the perspective that comes from caring for many patients. You don’t want to find yourself on the other side of an operation living with pain, regret, or some kind of existential crisis.

So we can say today – and repeat it again – that AI will never replace doctors in the future.

I think AI is going to transform medicine radically. In fact, I’m not even sure we’ll recognize what being a doctor looks like in the future compared with today. The profession will change dramatically.

At the same time, anyone who claims to know what healthcare will look like more than 20 years from now isn’t being very sober. What I am confident about is this: over the next two decades, doctors will continue to play an essential role. It just won’t be the same role they have today.

What worries you most in the development of AI in medicine?

That the wrong people will end up controlling the alignment of these models. I’m not talking about something obviously evil, like deliberately harming people. I’m talking about much subtler things. Because of cost efficiency, one recommendation may become the default simply because it’s cheaper. That will be very tempting for government bureaucrats or corporate executives when millions, if not billions, of dollars are involved and resources can be shifted from one population to another.

Instead of having an open public conversation about how resources should be allocated, those decisions could simply become embedded in AI systems.

That is going to happen. The real question is: who will be in control? That’s what worries me the most.

What gives you hope?

The incredible acceleration of open-source and open-weight AI models makes it entirely predictable that, if we don’t end up with just a handful of dominant models but instead see many different ones, we’ll eventually see AI systems sponsored by patient organizations and other communities. That would probably be my first choice.

What gives me the greatest optimism is the flourishing of a large ecosystem of AI systems, where different groups can choose the models that best reflect their needs and values.

Human beings are diverse. We have different values, different priorities, and different opinions. I think we should embrace diversity. Some people are naturally risk-averse. Others say, “I’d rather go through all the pain upfront.”

We shouldn’t try to make everyone the same. Instead, we should build AI systems that reflect the needs and values of different populations. Only by having a rich ecosystem of AI models can we achieve that.

So, for me, the combination of data autonomy, giving people control over their own health data, together with the rapid growth of open-weight AI models, is what gives me the greatest optimism.

Thank you!


Leave a comment