Researchers have tested this many times, in different places, over many years, and they keep finding the same result. This one is safe to trust.
For decades, people argued about what actually predicts whether someone will be good at a job. Two researchers settled a large part of it by pooling eighty five years of studies into one massive analysis. The ranking that came out was uncomfortable for how most companies hire. The single friendly, unstructured conversation, the classic interview everyone trusts, was one of the weakest predictors of all. Near the top sat things most firms use least: a sample of the actual work, a test of the relevant skill, and a structured interview where every candidate is asked the same questions and scored the same way. In 2022, other researchers re-examined the famous numbers and argued the original analysis had overstated some of them, and structured approaches came out looking even stronger. A rebuttal to that revision exists, and the debate continues, but one thing survives every version: structure beats winging it.
The hiring method almost everyone trusts most, the relaxed conversation where you get a feel for someone, is close to the least predictive thing you can do, and the methods that actually work are the ones that feel colder: see the real work, ask everyone the same questions, score before you discuss. Interviews still have a place, but an unstructured one mostly measures how much you like the person, which is not the same as whether they can do the job. Whatever else is disputed in this research, the core holds across every version of the numbers: in hiring, structure beats instinct, and it is not close. On why instinct feels so reliable regardless, see when expert intuition works.
These estimates were substantially revised in 2022 by Sackett, Zhang, Berry and Lievens, who found the range restriction corrections underneath the original table had been applied where they did not belong. Cognitive ability fell from .51 to .31, work samples by .21, and structured interviews came out with the highest mean validity at .42. That revision is itself still being argued over.
Source: Schmidt and Hunter, Psychological Bulletin, 1998.
Laszlo Bock, 2015
What one of the world's largest hiring machines learned when it finally measured its own methods.
Draw your own card. It does not take long, and it rewards taking your time.