Researchers have tested this many times, in different places, over many years, and they keep finding the same result. This one is safe to trust.
Daniel Kahneman spent a career on bias, the way judgement leans predictably in one direction. Late in it he encountered a different kind of error, and the case that convinced him came from an insurance company. He and colleagues gave the same realistic policies to different underwriters and asked each to set a premium independently. Before the results came in, the company's executives were asked how far apart two competent underwriters would typically land. They said about ten percent. The real median gap between any two underwriters on the same policy was fifty-five percent, more than five times what the people running the business had assumed. On one case, one underwriter quoted 9,500 dollars and another 16,700. The claims adjusters, given identical claims, were forty-three percent apart. A senior executive later estimated the annual cost of that scatter, counting business lost to high quotes and losses on low ones, in the hundreds of millions. Nobody at the company had known.
The book that followed, Noise, written with Olivier Sibony and Cass Sunstein in 2021, gives the scatter a name and takes it apart. Bias is the average error across a set of judgements. Noise is their variability: the spread among answers that ought to be the same. The two add to total error in the same way, which is the part people find hardest to accept, because a noisy system can be unbiased on average and still wrong in almost every individual case.
Noise itself divides. Level noise is the simplest and the one people already suspect: some judges are harsher or more generous across the board, so one interviewer gives everyone a seven and another gives everyone a four. Pattern noise is what remains once level differences are removed, and it is the larger component. It is the reaction of a particular judge to a particular kind of case, not harsh in general but harsh toward this type, not lenient in general but taken with that one. A judge who is mild across the board but severe with repeat offenders; an interviewer with a soft spot for people who changed careers; another who writes off anyone who sounds too prepared. Pattern noise has a stable part, which shows up whenever that type of case appears in front of that judge, and an unstable part the authors call occasion noise: the same judge varying with mood, fatigue, the time of day, what came before. Occasion noise is the famous kind, and the smallest.
What makes pattern noise hard to see is that it does not feel like variability from the inside. The judge's reaction is consistent, and consistency is normally what we treat as evidence that a judgement is sound. So each judge experiences their own signature as discernment. And organisations almost never put two judges in front of the same case, so the disagreement that would expose the signature never gets a chance to appear. Each decision arrives singly, with reasons attached, and feels justified.
Two consequences follow, and both go against instinct. First, noise can be measured without knowing the right answer. A bias audit needs ground truth; a noise audit needs only disagreement. If several judges see the same case and land far apart, at most one of them can be right, and the system has been shown unreliable before anyone has found out which. Second, awareness does not reduce it. Because the reaction feels like judgement, trying harder to be objective changes nothing. What reduces noise is procedure: breaking a judgement into separate assessments scored independently before an overall view is allowed to form, using comparative rather than absolute scales, aggregating independent judgements where more than one judge exists, and delaying the holistic intuition to the end rather than banning it.
The authors apply all of this to hiring directly. Different interviewers at the same firm reach different decisions about indistinguishable applicants, and the unstructured interview, a conversation that ends in an impression, is described as an ideal medium for noise. Their verdict on it as a predictor of who will succeed in a job is that it is often useless, which agrees with the decades of pooled validity research behind the structured-process finding elsewhere in this library.
Where the evidence stops. The insurance audit is the authors' own consulting work at one company, reported in the book rather than in a peer-reviewed study, and the numbers are theirs. The decomposition of noise into level, pattern and occasion components is a statistical definition rather than a finding, so it cannot be wrong, but the claim that pattern noise is usually the largest component rests on the authors' data across the settings they examined. The book's headline, that noise is as costly as bias and more neglected, is the part reviewers have argued with; one review criticised its reliance on single studies that are hard to replicate, and a statistician's review called it a four-hundred-page book on variance and meant it both ways. The hiring conclusions draw on the wider interview-validity literature rather than on a study the authors ran.
Two people interview the same candidate and come out having met different people. The instinct is that one of them read the person better, so it gets settled by argument, seniority or a compromise nobody believes in. The research says something less comfortable: both are running a signature, each one consistent, and the disagreement is the finding. It tells you the process is unreliable before anyone knows which of the two was right.
If you hire alone, that finding never arrives, because nothing ever puts a second judge in front of your candidate. Your signature runs unchecked on every hire, and it will feel like experience, because the reactions that make it up are the same every time. That is the specific trap for a founder or an owner doing their own hiring: the consistency of your judgement is not evidence of its accuracy, and there is no one in the building positioned to tell you so.
The fix does not require the second judge you do not have. It requires sequence. Decide the three or four things the job actually demands before you meet anyone, in the order that matters, so the criteria exist before a face does. Score each one separately, in writing, before you allow yourself an overall view, because an overall view formed early will quietly rewrite every score that follows it. Give the candidate a piece of the real work and watch, since a sample of the work predicts performance better than any conversation about it. Then, and only then, take the gut call, because it does add something once the pieces are down and destroys something when it comes first.
Where a second judge does exist, a co-founder, a partner, one colleague, use them properly: both of you score before either of you speaks. Four people who confer before scoring are one opinion wearing four badges, which is the same condition that breaks the wisdom of crowds. The pooled evidence on what actually predicts job performance is validity. The mechanism by which one strong trait colours every other score, a major source of the signature, is the halo effect, and the reason a conversation manufactures its own evidence is confirmation bias. The conditions under which an expert's intuition can be trusted at all, which is the argument this entry pushes against, are set out in Kahneman and Klein.
Kahneman and Klein say expert intuition can be trusted under two conditions: a regular environment and fast, clear feedback. Pattern noise says the felt consistency of a judge's reaction is not evidence that either condition holds, because a signature feels like discernment whether or not it predicts anything. They agree more than they collide. The firefighter and the chess player get thousands of cases with immediate feedback; the interviewer gets a few dozen a year and learns how a hire turned out months later, if at all. Hiring is the environment where intuition's conditions are least likely to be met, which is why the same man who wrote the conditions calls the interview often useless.
Source: Daniel Kahneman, Olivier Sibony and Cass R. Sunstein, Noise: A Flaw in Human Judgment, Little, Brown Spark, 2021. The insurance audit, the executives' ten percent estimate against the fifty-five percent median, and the 9,500 against 16,700 example are the authors' own, reported in the book and in Kahneman's interviews, including Behavioral Scientist, 2021. Olivier Sibony on the three kinds of noise and the noise audit: interview with The Decision Lab, Beyond Bias. The interview-validity finding: Frank L. Schmidt and John E. Hunter, The Validity and Utility of Selection Methods in Personnel Psychology, Psychological Bulletin, 1998. A review for a statistics readership: Chance, 2024.
Daniel Kahneman, Olivier Sibony and Cass R. Sunstein, 2021
The book this entry comes from, and the one that puts a number on a thing every interviewer has felt and nobody had measured. The insurance audit is in it, with the executives guessing ten percent and finding fifty-five, and so is the three-way split of noise that this entry describes. It is long, it argues with itself, and the chapters on hiring and performance reviews are the ones to read first if you have ever walked out of a debrief wondering how two colleagues met such different people.
Draw your own card. It does not take long, and it rewards taking your time.