
The Horse Riddle and the Diagnostic Swagger of Machines
By Arthur Lazarus, MD, MBA
Published on 08/23/2026
Solve this riddle: A Horse walks into a room with two other horses. Each Horse has three baby horses next to it. How many horses are in the room?
I answered 12. Seeking validation, I consulted several general-purpose artificial intelligence systems. All gave the same answer: 12. Their reasoning seemed embarrassingly straightforward: three adult horses, each with three babies. Three plus nine equals twelve. Case closed. Next question.
Except the case was not closed.
Two medically oriented AI platforms declined to respond to the riddle. One told me the question was outside its programmed scope. Another politely suggested that, for a nonmedical puzzle, I consult a general-purpose assistant or another trusted resource.
At first glance, those responses seemed almost comically rigid. Here were sophisticated medical AI systems, presumably capable of digesting clinical information far more complex than the reproductive arrangements of three horses, refusing to count livestock.
But after reading the human responses to the riddle on LinkedIn, I began to wonder whether the medical systems had given the wisest answers.
On LinkedIn, people did what we often do when given a simple question and an open comment box: they came up with answers faster than horses can have babies. There were three horses if you count only the adults. There were nine if only two of the three horses had foals. There were six, depending on how “next to” was interpreted. There were infinitely many if each baby also has three babies, creating a barnyard version of exponential growth. One person pointed out that the riddle never says the room started out empty. Another joked about a dozen eggs, as a reminder that any talk of animal arithmetic eventually slides into the absurd.
Although some of the responses were clearly silly, they were, in a way, good mental gymnastics, the type of cognitive flexibility necessary for the practice of medicine. The human commenters spotted something the confident AI did not: words aren’t the same as math. “Each horse” sounds clear until you ask which horses count, whether “next to” means only one neighbor, whether capitalizing “Horse” twice adds secret meaning, or whether baby horses count as horses rather than as foals, colts, or ponies. The riddle has a standard answer, but its wording doesn’t force one. That gap is where AI’s overconfidence lives, and where it can be most dangerous.
Artificial intelligence is exceptionally good at producing something medicine has always found seductive: a coherent story. Give a language model a collection of facts, and it tends to connect them. Symptoms become a diagnosis. Laboratory findings acquire meaning. A sequence of 1 events can be made to look like causation. An ambiguous prompt becomes an arithmetic problem with the answer 12.
The interesting thing about my answer was not that 12 was unreasonable. Under the ordinary interpretation of the riddle, 12 is probably what its author intended. What was interesting was how quickly I turned a plausible interpretation into an answer. That distinction—between an answer and the answer—is where AI introduces risk. AI’s output arrives quickly and sounds polished. The reasoning appears orderly. There are rarely verbal beads of sweat on the page. Such fluency can easily be mistaken for certainty.
Clinical medicine is full of horse riddles. A patient presents with chest discomfort, shortness of breath, an equivocal troponin level, a baseline left bundle branch block, chronic kidney disease, and a history that does not fit neatly into an order set. The question appears to ask for a diagnosis, but it may actually contain several questions: What is the most likely etiology? What is most threatening at present? Which investigations can wait? What does this particular patient value? What additional information should be gathered? What are the consequences of being wrong?
Or, as my wise mentor frequently reminded me during residency, when I was too quick to jump to a reductionist answer, “Sometimes patients can have lice and fleas.” And, of course, the standard retort to his reminder is, “When you hear hoofbeats, don’t think of zebras.” But often the safest answer is: I do not yet have enough information.
That sentence may be among the most important in medicine. Unfortunately, it is not particularly glamorous. “I don’t know” does not look impressive on a dashboard. “Insufficient information” rarely impresses an audience at a technology demonstration. A system that confidently produces a differential diagnosis in three seconds appears more capable than one that halts and says, “Before I answer, I need to know whether the pain is exertional.”
This is why I now have a greater appreciation for the medical AI platforms that refused to answer my horse riddle. They were not demonstrating superior horse knowledge. They were demonstrating something potentially more valuable: boundary recognition. They knew—or at least their designers had instructed them—that this was not their job. Medicine could use more of that epistemic humility.
I worry that AI can produce a perfectly reasonable answer to a subtly unreasonable question. Suppose an AI concludes that a patient has pneumonia. The diagnosis may be plausible, but did the system consider pulmonary embolism, heart failure, drug toxicity, or malignancy? Did it note that the patient’s oxygen saturation was recorded after supplemental oxygen had already been started? Did it recognize that the phrase “doing fine” in yesterday’s progress note may mask a physician’s uncertainty rather than document clinical stability?
Or did it simply count the horses?
The human responses to the riddle reveal another difference worth preserving. Humans did something wonderfully inefficient: they challenged the premises, asking, “What if there were 2 already horses in the room?” “What exactly does ‘next to’ mean?” “Are baby horses… horses?” “Why are horses in a room rather than a stable?”
Questions that might seem out of bounds are often the foundation of sound clinical reasoning. Experienced physicians spend much of their careers realizing that the most dangerous mistake is not choosing the wrong answer but accepting the wrong question.
The patient labeled “noncompliant” may lack transportation. The “frequent flyer” may have nowhere else to seek care. The “treatment failure” may never have received the medication. The “psychiatric symptom” may be neurologic disease.
The algorithm can be exquisitely accurate within the frame it has been given and still fail because the frame itself is flawed. That is why calibration matters more than confidence. A trustworthy medical AI should not merely tell us what it thinks. It should help us understand how much confidence the available information warrants. It should identify missing information, expose assumptions, acknowledge plausible alternatives, and make uncertainty visible rather than smoothing it over.
The ideal clinical AI response may sometimes sound less like an oracle and more like a thoughtful resident presenting a case: “Given these findings, pneumonia is possible, but I cannot reliably distinguish it from several alternatives without additional information.” It’s less dramatic than “Diagnosis: pneumonia,” but it’s considerably more conservative and safer.
There is irony here. We often worry that AI will become too human. In medicine, I sometimes worry about the opposite: that physicians will become too much like AI—accepting the first coherent interpretation, equating speed with accuracy, growing too comfortable with conclusions based on ambiguous information, and feeling compelled to produce an answer.
Perhaps the medical AI systems that refused to puzzle over the horses were onto something. They effectively said: This is outside my competence. Physicians once considered that statement a mark of professional maturity. We should ensure artificial intelligence learns the same lesson— and that we do not unlearn it ourselves.
For the record, I still think there are 12 horses. But now I would answer differently: Probably 12, assuming the three adult horses are distinct, each has three distinct baby horses next to it, and no other horses are already in the room.
It is a longer answer. It is less satisfying. It contains a qualifier. And it sounds considerably less intelligent. Which may be exactly why it is better medicine.
About the Author
Arthur Lazarus, MD, MBA
Physician Executive • Psychiatry
Arthur Lazarus is a physician-author whose work spans narrative medicine, physician leadership, artificial intelligence, healthcare ethics, medical culture, and fiction. He has published numerous books and more than 500 articles and essays across scientific journals, professional publications, and online platforms. His writing explores the forces reshaping modern practice while preserving a central commitment to story and the human relationship at the heart of care.
Sign up for our Newsletter!
Receive the latest articles directly in your inbox




Discussion
Join the conversation! Login if you already have an account, or create an account. We would love to hear your perspective.
Comments
0Loading comments…