Next-Generation Clinical Decision Support Tools
Among the most actively used tools in the modern physician’s arsenal are point-of-care clinical decision support (CDS) tools. For the uninitiated, these tools help bring together the massive amount of medical data that is available, presenting it in an organized and easily referenced way, showing the physician where the current research comes to a consensus, summarizing a particular pathology, outlining the recommended management given a specific diagnosis, etc. The current generation of these tools has been human-curated. There are authors attributed to each page.
These tools have made possible something that would have been impossible a few decades ago: changing management directly at the bedside, or close enough to it, by providing easily referenced, up-to-date evidence-based medicine. This would be impossible if the only option for a physician was to spend time looking up literature, processing toward the most pertinent, reading each article, wading through the statistics, comparing findings, contemplating the data, notating that data in a way that allows for referencing later… you get the picture.
This is not to say that there is not a place for reading literature and running through the aforementioned process. If your only avenue of engaging with pathology, management, and current evidence-based medicine is through point-of-care CDS tools, you are ceding a good deal of medical decision-making and learning, while decreasing your ability to keep important information fresh and correct in your mind, as well as truly surf the bleeding edge of medical research. There is still a large aspect of pattern-recognition and being a good diagnostician that goes into medical decision-making. Reading medicine more widely than simply engaging with these tools opens up a range of topics that come in handy serendipitously for just these aspects of practicing medicine. Alas, I digress. This is, perhaps, a topic for another day.
The Dethroner?
The above point-of-care CDS framework has been held in place with little movement, at least for as long as was in residency and through nearly a decade of clinical practice. It is in this placid setting that OpenEvidence dropped like a boulder into a pond. Previously, OpenEvidence was… honestly, I don’t really know. I know that I ended up there randomly from time to time. I keep reading that it began as an AI company, but I just cannot recall ever thinking it was that. Informative research on the company this is not. It is merely my experiential report, but I would bet it is not far off from many other physicians’ experience. However, seemingly overnight, OpenEvidence has crashed the party, smashing the old regime as it lumbers in, changing the playing field for physicians and providers of point-of-care CDS tools.
The current state and format of OpenEvidence will be extremely familiar to physicians and non-physicians alike, what is surely to be a symbol of the 2020s: a chatbot. There is a dialogue box just waiting for you to send your query to the OpenEvidence large language model (LLM)[1]. A physician can enter a clinical question, ask to learn more about a diagnosis, or even enter their patient’s history of present illness, lab and imaging workup, family history, physical exam findings, etc., and get a thorough differential diagnosis. OpenEvidence will even tell you at times what it “thinks” is most likely[2].
The first time I used OpenEvidence I had three thoughts:
- This is awesome.
- How should I consider fact-checking?
- When do they ask for payment?
Cool, With Caveats
This is awesome
The amount of data available to the modern physician is astronomical. On the one hand, this is great. It means that we have lots of data. On the other hand, translating that into clinical practice is difficult, even for the most well-read physician. Point-of-care CDS tools are arguably a good place to start when you want to look something up. They are quick, ubiquitously available, and well organized.
As touched on a bit above, OpenEvidence allows for something that previous-generation CDS tools did not allow[3]: entering your patient’s symptoms, labs, images, etc. This allows for a much more tailored search experience. You aren’t just looking up rhabdomyolysis, you are looking up elevated CK in the setting of a weeklong febrile illness in a 5-year-old that also presents with anemia and a family history of several autoimmune disorders. This is nearly like looking all of these pieces up, juxtaposing them, pulling out the relevant details, and getting one article with specifically useful information for you and your patient. It is a truly significant development.
And yet…
While asking about a specific diagnosis is not all that different from accessing a page on that diagnosis, navigating the article by heading or search terms, and digging into the specifics you would like to know, conversing on single topics is the simple case. You may have enough knowledge to note inconsistencies and, given the single-topic nature of the query, find it relatively easy to consult more trusted sources to ensure correctness.
However, what of the example above, asking about a patient’s specific findings, tied to the current history of present illness, family history, prior presentations… We are not yet truly certain that LLMs are good at this. Without a doubt, you will receive an answer that sounds plausible, well put together, and confident, but you have introduced so many variables. How will you go about ensuring you are getting the correct information? How do you decide to act on this information? Look up more? Disregard it? Tease out what is really pertinent to your patient? Can you even be sure how all of this information is affecting your clinical decision-making? If you aren’t sure of each detail, piece of information, or conclusion the LLM has offered, how can it be possible to know? These are questions I am grappling with and have yet to come up with concrete answers.
Fact or Fiction (or Hallucination)?
How do I fact-check this?
This thought comes to mind often, variably, depending on the seriousness of what I am using LLMs for. If I am asking it about the proper spacing of plants for my landscaping, I don’t think much of it. If I am using it to care for a patient, it is heavy on my mind. An ongoing, developing heuristic I have been using is, if I am going to use the information in front of me to concretely make a decision for treatment or workup, then I dig in to fact-check. I also consistently remind myself that deciding not to do something is also a decision and one I need to be careful of when consulting LLMs like OpenEvidence, digging into the information all the same. I wish this were foolproof, but I fear it to be relatively brittle. For instance, if I look something up and tell myself it is just out of some small curiosity, perhaps something to potentially consider for my patient, does that fit my developing heuristic? I have no real way of knowing how what I have seen will subconsciously affect my thinking. Maybe I decide not to have a lab drawn and this feels like it is coming only from my mind and thinking, but I have no way of knowing how something I have “just quickly looked up” with the LLM has influenced me. As we as physicians use these tools more and more, we will begin “quickly looking up” a mountain of information, easy to consciously throw away, but subconsciously?
An article on a pathology is static in the moment of reference, but an LLM? It offers a confident answer in a way that a static page cannot be said to do. An LLM may state things like, “the most likely reason” or sycophantically, “Good thinking! That is something that is unlikely in this case.” I do not have a doubt that this sort of interaction is fundamentally different from simply reading an article or web page. I just do not think we have a good handle yet on how that affects our clinical decision-making. On top of that, if you consult the LLM on another occasion with the same question, you may get different answers, or at least answers presented differently.
Societally, we have not truly come to terms with this new burden we have created out of thin air: fact-checking LLMs. When it comes to LLMs used for clinical support, transparent links to references are crucial, but do we follow every single one? Do we read each paper? Just the statistical findings? Use what we know and compare it to the output to see if it makes sense? It kills me a bit to admit, but I do not think we are equipped to handle this burden and may never be. If we take this burden to its extreme, that we are individually responsible for fact-checking every bit of data that comes out of an LLM, we will not be gaining any efficiency, and in fact, given the ease with which we can bring data and information to our fingertips, we will be overloading ourselves instantly and constantly.
Along these lines, are there any ways to both hold an LLM provider completely responsible for any hallucination or incorrect piece of information and continue to push LLM/AI technology forward? If we get the balance incorrect, will medicine be left behind as useful LLM/AI technology is developed? I do not have the answers, but I am in good company with… everyone at this point.
You Are The Payment Model
When do I pay?
None of us have an excuse not to ask this question. As has been true since time immemorial, nothing is free, so if you ain’t paying with money, then you’re paying with something. The nuance is what that something is: time, data, attention (to advertising). Knowing what you are paying is extremely important, especially if you are not told what it is outright.
In the case of OpenEvidence, we are paying with the internet’s favorite model: attention to advertising. Pharmaceutical companies have never stopped attempting to bias physician decision-making in their favor. These tactics have moved from overt to slightly more hidden over time (it is more tasteful), but they remain all the same. So not only do we have questions of subconscious effects of LLMs on medical decision-making and questions of how to fact-check, we must also add concern for how advertising is influencing us (and it is, even if you think you are of too strong a mind for this to happen).
I would much rather pay for a point-of-care CDS tool like OpenEvidence, but that is not as lucrative a model as advertising[4]. Have no doubt, this really matters. Pharmaceutical companies will argue that it doesn’t, or that there are ways to not bias physician thinking (for instance, never advertising directly from the current query), but this is a slippery and unproven slope. Advertising is advertising, and its goal is to make money, not to ensure that clinical decision-making is unaffected. We must be concerned about how it affects our thinking.
In addition to our attention (and potentially our medical decision-making) we are also paying with data. Make no mistake, OpenEvidence knows what you are querying, how your discussions go, what information you dig into, what information causes you to stop interacting (and thus not have eyes on advertising), and how long it has been since you last used the service. They will build a profile of you, anonymized or not, but they will. If they need to crank that money lever and adjust how they advertise, they will know how to reach you. They are gaining so much data that they likely don’t even know what to do with it all yet, but this is the cream the pharmaceutical (and other) companies are coming for. Whether or not OpenEvidence has any nefarious aims (a likely too black-and-white distinction), these are, in fact, the facts. We must keep this in mind. Needing eyes on advertising will affect their decision-making.
In the near term, these concerns must be dealt with on an individual level. Over the long term, they turn societal. It takes time to complete strong studies and to monitor effects. We simply won’t have the data for an extended period of time. We as physicians will need to be individually aware and clear-eyed about all of this, developing strong values and practices around LLM use. This may not be enough, but it is likely all we can control for now. As more information comes out, as best practices develop, we can add our voices, advocating for what is best for our patients and our medical practice. This is uncharted territory. We need not face it with total pessimism, but absolutely with skepticism, as aware as possible, our eyes open, and a mind constantly toward patient care and non-manipulated medical practice.
I will use LLM rather than AI throughout this post in order to put a fine point on what we are actually dealing with and to not get lost in discussions of true intelligence, sentience, or consciousness (none of which I believe OpenEvidence, or any LLM for that matter, has). ↩︎
Speaking of LLMs as thinking is fraught, as LLMs do not think, they simply produce the next best token that fits the prior tokens produced (significantly oversimplified, what word fits best after the word prior). How we engage with this important nuance is a major question of our times. ↩︎
I do realize that the other point-of-care CDS offerings are not simply standing still and merely staring at OpenEvidence, but are adding their own LLM tools. I simply think that OpenEvidence has made a giant splash and offers a good vantage point of what is here and what is coming. ↩︎
Subscription-based offerings exist. Again, focusing on OpenEvidence is a useful perspective as its use (and valuation) skyrocket. Regardless of the existence of subscription models, these considerations matter nonetheless. Advertising dollars are tempting and it can easily become the dominant model. ↩︎