← Back to the library

    Simpson's paradox

    holds up

    Researchers have tested this many times, in different places, over many years, and they keep finding the same result. This one is safe to trust.

    In 1973 Berkeley's graduate admissions numbers looked like proof. Around 44 percent of male applicants were admitted and about 35 percent of women, across more than twelve thousand applications. Peter Bickel, a statistician at the university, was asked to look at the data, and he started from a fact about how Berkeley actually works. There is no single admissions process. Each department decides for itself who gets in, and there were more than eighty of them. So instead of one big number he calculated the rate for men and for women inside each individual department, and the pattern came apart. A few departments did show a real bias, but slightly more of them favoured women than men, and when the numbers were pooled properly the small tilt that remained ran gently in women's favour. Same data, opposite conclusion, both readings arithmetically correct. What the aggregate had been measuring was where people applied. Women had applied in far greater numbers to the most competitive departments, the ones rejecting almost every applicant of any gender, and men had gone in larger numbers to departments that admitted most people who applied. The gap was a map of application choices, not of how anyone was judged once their file was open.

    Every dashboard you read is combined data, which means every number on it can be lying in this specific way. Overall conversion can rise while every single segment converts worse, purely because the mix shifted toward an easier segment. Average deal size can fall while every person on the team improves, if the cheaper product grew fastest. Salary averages, churn, response times, and win rates all behave the same way, because in each case the aggregate silently weights groups by size. The habit worth building is a single question asked before you act on any number: what happens to this when we split it? Split by segment, by channel, by team, by period. If the direction survives every split, you have something. If it reverses, the aggregate was measuring composition rather than performance, and acting on it means fixing something that is not broken. It cuts both ways, which is what makes it dangerous. A hidden variable can manufacture a problem that is not there, and it can hide one that is. One caution about this case in particular: showing that an aggregate gap comes from where people applied does not close the question of why the applications were distributed that way. Bickel's own paper made that point, and it is a different question from the one the arithmetic answers. One more thing, because this story carries its own hidden variable. Almost every retelling says Berkeley was sued, and the lawsuit is usually the hook the whole anecdote hangs on. There is no trace of any such case. Bickel has said the university feared being sued and commissioned the analysis itself, and the courtroom got added somewhere in the retelling. Peer-reviewed papers still open with it today. Which is the same lesson in a different form: the version everyone repeats is not always the version that happened, and nobody checked because the story was too good to check.

    Source: Bickel, Hammel and O'Connell, Sex Bias in Graduate Admissions: Data from Berkeley, Science (1975). The formal description: E. H. Simpson, The Interpretation of Interaction in Contingency Tables, Journal of the Royal Statistical Society (1951). Karl Pearson and Udny Yule had noted the reversal decades earlier.

    The book, if you want to go further

    The Art of Statistics

    David Spiegelhalter, 2019

    On reading data honestly, including why aggregates mislead and what to check before believing one.

    Draw your own card. It does not take long, and it rewards taking your time.