Decide before you look: 10 studies on AI overreliance
In the AI trainings I run, one sequence repeats at every table. Paste the task, read the answer, ship it. The reading lasts about 4 seconds. A meta-analysis of 106 experiments found that a human paired with a model often lands below whichever of the two was stronger alone, and those 4 seconds are where the loss happens.
Two heads, one answer, worse results.
Michelle Vaccaro, Abdullah Almaatouq and Thomas Malone pooled 106 experimental studies and 370 effect sizes for Nature Human Behaviour. Human-AI combinations came out below the best of human alone or AI alone, Hedges' g = -0.23, with a 95% confidence interval from -0.39 to -0.07 1. The average hides a split. Losses concentrated in decision tasks: classify, diagnose, choose. Gains appeared in content creation: draft, generate, produce. The losses ran deepest where the model already outperformed the people it was paired with.
The pattern predates chatbots. In 1999, Linda Skitka, Kathleen Mosier and Mark Burdick ran a flight simulation with an automated monitoring aid that was reliable without being perfect. Participants working without the aid outperformed the ones who had it. Two failure modes showed up: omission, where a person missed an event the aid failed to flag, and commission, where a person followed the aid's directive against other indicators that were fully valid and visible on screen 2.
Artur Klingbeil, Cassandra Grützner and Philipp Schreck reproduced the second failure with nothing but a label. Telling participants that a piece of advice came from an AI system was enough to make them follow it against information sitting in front of them, with no avatar, no voice and no other anthropomorphic cue in the interface 3.
Expertise does not immunise anyone. Susanne Gaube and colleagues showed chest X-rays with diagnostic advice to radiologists and internal medicine physicians. Human experts had written all of the advice, some of it wrong, and part of it carried an AI label. Diagnostic accuracy fell when the advice was wrong, whatever the label said. Radiologists rated AI-labelled advice as lower in quality, while the physicians with less imaging expertise did not 4.
Explanations raise acceptance and leave accuracy flat.
Gagan Bansal and colleagues tested the standard remedy at CHI 2021: show the reasoning. Across three tasks, with an AI whose accuracy sat close to the participants' own, explanations raised the chance that a person accepted the recommendation, and raised it whether the recommendation was right or wrong. Team accuracy stayed flat 5.
Helena Vasconcelos and colleagues found the condition under which explanations earn their place: they help when they cut the cost of checking. People weigh the effort of engaging with an explanation against what they expect to gain, so an explanation that hands over something quick to verify reduces overreliance, while confident prose about the model's reasoning does nothing 6.
Friction works, and people dislike it.
Zana Buçinca, Maja Barbara Malaya and Krzysztof Gajos built three interfaces that force engagement: commit to your own answer before the AI's appears, request the suggestion rather than receive it, wait before it shows. With 199 participants, those designs cut overreliance compared with plain explanation interfaces. Two findings came attached. Participants gave their worst subjective ratings to the designs that helped them most, and the benefit concentrated in people who score high on need for cognition 7.
Charvi Rastogi and co-authors at IBM Research named the mechanism: anchoring. Reading the model's answer first drags your answer toward it. Their de-anchoring approach gives the human more time on the cases where the model is least confident, and it lifted joint performance on exactly the cases that hurt, the ones where the model was unsure and wrong 8.
Pairing is a design choice with a price.
Nikhil Agarwal, Alex Moehring, Pranav Rajpurkar and Tobias Salz ran an information experiment with professional radiologists. Handing them AI predictions did not improve diagnostic performance on average, while handing them clinical context did. The failure sits in belief updating: radiologists underweighted the AI predictions, and they treated their own read and the model's as statistically independent when the two overlap. The authors draw a blunt design conclusion. Until that updating error gets corrected, send a case to a human or to a model, and rarely to an AI-assisted human 9.
That result unsettles the assumption I hear in most rollouts, that putting a person next to a model is the cautious middle option. On decision tasks it is the option that the literature scores worst.
The protocol.
1. Write your answer before you open the model's. One line in a note, plus your confidence out of 10. This is the commit-first condition from Buçinca's interfaces 7 and the de-anchoring result from Rastogi's 8, available without any tooling. If you cannot write a line, you are the novice on this task, and steps 3 and 4 carry the load.
2. Count your overrules. For 2 weeks, tally the times you rejected or rewrote the model's output on something that mattered. Seven days of daily use with zero overrules means you are approving, not deciding 12.
3. Ask for the checkable object rather than the conclusion. The query, the calculation, the two sources, the test that would fail. You check an answer when checking is cheap 6. "Summarise this contract" returns prose to admire. "Quote the 3 clauses that set a deadline, with their article numbers" returns something you can confirm in 40 seconds with the document open.
4. Score the output against an external anchor. Run the code, recompute one number by hand, open 2 of the cited sources. Fluent writing raises your acceptance and leaves accuracy where it was 5.
5. On decision tasks, name one owner per case. You decide it or the model decides it, with a written rule for the routing: model confidence, stakes, your own expertise on that case type. Splitting one judgment across both is where the meta-analysis located its losses 19. Keep the pairing for drafting and generation, where the same meta-analysis found gains.
6. Put the saved minutes into the check. In a survey of 319 knowledge workers describing 936 real uses of generative AI, Hao-Ping Lee and colleagues at Microsoft Research found that confidence in the model went with less critical thinking, while confidence in one's own expertise went with more 10. The verification step is where the time you save is supposed to land.
Take the decision you hand to a model most often. Before the next one, write your own answer and your confidence in a note, open the model's answer, then record which of the two you kept. Ten entries in, count how many times you kept your own.
Sources.
- Vaccaro, M., Almaatouq, A., Malone, T. (2024). When combinations of humans and AI are useful: a systematic review and meta-analysis. Nature Human Behaviour, 8, 2293-2303. doi:10.1038/s41562-024-02024-1
- Skitka, L. J., Mosier, K. L., Burdick, M. (1999). Does automation bias decision-making? International Journal of Human-Computer Studies, 51(5), 991-1006. doi:10.1006/ijhc.1999.0252
- Klingbeil, A., Grützner, C., Schreck, P. (2024). Trust and reliance on AI: an experimental study on the extent and costs of overreliance on AI. Computers in Human Behavior, 160, 108352. doi:10.1016/j.chb.2024.108352
- Gaube, S., Suresh, H., Raue, M., Merritt, A., Berkowitz, S. J., Lermer, E., Coughlin, J. F., Guttag, J. V., Colak, E., Ghassemi, M. (2021). Do as AI say: susceptibility in deployment of clinical decision-aids. npj Digital Medicine, 4, 31. doi:10.1038/s41746-021-00385-9
- Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., Weld, D. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. CHI 2021. doi:10.1145/3411764.3445717
- Vasconcelos, H., Jörke, M., Grunde-McLaughlin, M., Gerstenberg, T., Bernstein, M. S., Krishna, R. (2023). Explanations can reduce overreliance on AI systems during decision-making. Proceedings of the ACM on Human-Computer Interaction, 7 (CSCW1), 129. doi:10.1145/3579605
- Buçinca, Z., Malaya, M. B., Gajos, K. Z. (2021). To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5 (CSCW1), 188. doi:10.1145/3449287
- Rastogi, C., Zhang, Y., Wei, D., Varshney, K. R., Dhurandhar, A., Tomsett, R. (2022). Deciding fast and slow: the role of cognitive biases in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 6 (CSCW1), 83. doi:10.1145/3512930
- Agarwal, N., Moehring, A., Rajpurkar, P., Salz, T. (2023). Combining human expertise with artificial intelligence: experimental evidence from radiology. NBER Working Paper 31422. nber.org/papers/w31422
- Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., Wilson, N. (2025). The impact of generative AI on critical thinking: self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. CHI 2025. doi:10.1145/3706598.3713778