In the late 19th century, a high school mathematics teacher in Germany decided to forgo the usual human students in favor of attempting to teach a horse to learn math.
Whilst he initially appeared to have succeeded in this goal, and the horse drew crowds of impressed visitors, the real success of Clever Hans the Math Horse was in the lessons he taught scientists, changing how we conducted experiments forever.
The mathematics teacher in question was Wilhelm von Osten, a man with a wide variety of interests which included animal intelligence as well as the now thoroughly discredited idea of phrenology.
von Osten believed that we had significantly underestimated the intelligence of animals and that he could show this by teaching a cat, a bear, and a horse mathematics.
As you might reasonably expect, the cat didn't seem very interested in anything a human would have to teach it. Meanwhile the bear (Jesus Christ, von Osten, why start with a bear) became aggressive.
But the horse? Why, the horse, a stallion named Hans, appeared to be a verifiable genius on the topic, the Euler of the horse world.
Genius horse?
Let's qualify that statement somewhat – he was still pretty stupid compared with a human, and wasn't sat there doing calculus.
Having seemingly taught the horse how to communicate answers to questions by tapping his hoof on the ground a certain number of times to indicate a letter or a number, von Osten was then apparently able to teach the horse to do some basic counting and arithmetic. As well as this, Hans was able to remember his square roots well enough to keep a crowd engaged.
People did begin to flock to see Clever Hans, as he became known, and appeared to believe that the horse really did possess these skillsets. That must have been quite disconcerting, to learn that the animal you have been riding around on and feeding hay to might actually have the intelligence of a teenager.

One New York Times article covering "Berlin's Wonderful Horse" was notably absent of much in the way of skepticism, telling readers that the horse could, amongst other feats, spell out the name of the Prince of Saxe-Coburg and Gotha.
"Hans can tell the time on a watch and can indicate the exact hour," the article claimed. "At the test yesterday he recognized persons from photographs," it continued, before fully exploring the idea that the horse had a concept of time.
According to that article, Hans demonstrated that he could correctly identify colors, and was able to tap out what day of the week it was, as well as the month, all without having anything like horse networking events to schedule or horse babysitters to arrange.
Other claims include that he was able to guess the composer after listening to a piece of music.
The investigation
Your math horse won't make it onto the pages of the New York Times without some egghead out there saying, "Hold on now, are we absolutely sure this horse is doing math?"
The first to try and rigorously test the horse was the German board of education, setting up a commission to find out what was going on. Over the course of a year and a half, the trials saw Hans's trainer separated from the horse to try and rule out any trickery.
Despite this, the horse continued to be able to perform on demand, getting the answers correct nearly as often as when his own trainer was in the vicinity. The commission eventually concluded in 1904 that there was no trickery involved, which was correct – but only on the human side of things. The horse was playing his own game.
So you're saying the horse could actually do math?
No, the horse wasn't actually doing math. This was very much a case of "he was then transferred to better investigators, who upgraded their condition to fraud". Psychologist Oskar Pfungst took on the investigation, and soon found some telling clues.
At first, he found that the horse was indeed very accurate when asked questions by von Osten under normal conditions, as well as other people now questioning their life decisions as they, in turn, questioned a horse. But when von Osten or the other questioners were asked to move further away from the horse, his accuracy dropped. Far more damning, when the questioner did not know the answer to the question they were asking, the horse's accuracy fell off a cliff.
While von Osten had not meant to present any clues to the horse, Pfungst realized that he had been doing so the whole time unconsciously. Rather than tapping his foot a certain number of times to indicate a letter, the horse was simply tapping his foot and looking for cues in the question-asker as to when it was time to stop tapping his foot, and receive his treat and/or applause.
Pfungst was able to demonstrate that researchers were giving off these subtle clues by playing the role of the horse himself, using Hans's trick of picking up on facial and posture clues even when he hadn't heard the question, and when the question-askers had been informed that they were giving off these signs.
More damningly, the horse was unable to answer any questions when a screen separated him from seeing the human participants' faces.
The horse reminds me of AI somehow?
You're not alone. The story of Clever Hans, a now long-dead horse people thought could do math, has been popping up to a new audience lately: people reading about the evolution of large language model (LLM) artificial intelligence (AI) chatbots.
In AI circles, the "Clever Hans effect" means something entirely different, but is similarly troublesome. AI can, in many instances, provide the correct answer to your questions. But that doesn't necessarily mean that it has done so through the right chain of reasoning, nor, of course, that it knows anything.
"When assessing machine behavior, the general task solving ability must be evaluated (e.g. by measuring the classification accuracy, or the total reward of a reinforcement learning system). At the same time it is important to comprehend the decision-making process itself," a 2019 paper explains.
"In other words, transparency of the what and why in a decision of a nonlinear machine becomes very effective for the essential task of judging whether the learned strategy is valid and generalizable or whether the model has based its decision on a spurious correlation in the training data."
Using AI to assess medical imaging, you would like it to assess patients based on diagnostic criteria. However, it is possible that models sometimes pick up on other spurious correlations instead of clinically relevant features. The diagnosis could be correct, but it is still a demonstration of the "Clever Hans" effect; it got there through looking at other cues.
One team attempting to look for potential biases in medical AIs, for example, found that they were able to identify the race of patients based on heavily pixilated scans. While they were unable to find the exact mechanism behind this, it is possible that the AI had identified a pattern in medical equipment, with areas of lower income using older scanning equipment, a correlation not picked up on ahead of the research.
So how did Clever Hans affect science?
Positively, overall. Researchers now know that as well as all the other variables in experiments, they needed to be aware of the "Clever Hans effect": giving unintentional cues to participants, human or animal, which you are trying to study.
Practices like double blinding – where participants and researchers are unaware of aspects of the experiment until it has been completed – mitigate against the effect.
AI researchers may now have to deal with their own version of the Clever Hans effect, though it may take a little more sorting out than the first.





