Skip to main content Scroll Top
19th Ave New York, NY 95822, USA
Untitled design

How Fast Thinking and Slow Thinking will shape our interaction with AI, continued

In our previous post, we kicked off our exploration of human-centered AI within the TANGO project by looking at how artificial intelligence has become seamlessly integrated into our daily lives. To build systems we can truly trust, we argued that we must first understand human cognition—specifically the balance between fast, effortless intuition (System 1) and slow, effortful deliberation (System 2).

We highlighted that humans generally prefer deliberative decision-makers and rely on clear explanations to evaluate trustworthiness. However, translating these human expectations directly to machines isn’t straightforward: speed expectations differ radically between humans and computers, and AI inner workings often remain opaque. Ultimately, we concluded that designing effective Explainable AI (XAI) requires a deeper look into how humans produce and evaluate justifications. So today, let’s have a look at our latest research.

Biased against intuition?

Our previous investigations found that people consistently prefer their fellow humans to engage in deliberation rather than rely solely on intuition when solving problems1. In our latest study, we set out to test whether our own tendency toward intuition or deliberation shapes how we evaluate the thinking of others. For instance, more intuition-oriented people might prefer intuition, while more deliberation-oriented people would value it less. However, we found that, while individuals with a more deliberative disposition rate intuitive reasoners less favorably, even those of us who rely most strongly on intuition in our daily lives still maintain an absolute preference for deliberation in others. In other words: no matter how strongly we tend to rely on our own gut feelings, we remain skeptical of intuition in others.

Fast for me, slow for thee

Why do we hold others to a different standard than ourselves? Our findings revealed that our personal reasoning style primarily shifts our attitude toward intuition, rather than changing how much we value deliberation. While intuitive thinkers appreciate intuition in others more than analytical thinkers do, both groups value deliberation equally.

The key reason for this divide lies in cognitive transparency and justification. When we make an intuitive decision ourselves, it is accompanied by an internal “feeling of rightness” that gives us confidence. But because we cannot feel what others feel, someone else’s intuition remains completely opaque to us. Deliberation, by contrast, operates with cognitive transparency—its output comes with an explicit awareness of how a conclusion was derived, enabling people to explain and justify their reasoning, while intuition often cannot provide justifications. When we seek trustworthy advice or solutions from others, we ultimately need a sound justification rather than an opaque feeling.

From human justifications to evaluating arguments

For human-centered AI projects like TANGO, this insight is fundamental. Because we inherently demand transparent justifications when evaluating decision-makers, we bring those same expectations to AI systems. But how do we actually assess whether an explanation or justification holds up?

Judging arguments, fast and slow

Every day, we act as judges: when a colleague explains a decision, a news article backs a claim with a statistic, or a chatbot justifies its answer, we decide in a split second whether the reasoning holds up. In a set of studies2, we asked participants to solve short reasoning problems and then rate the strength of several arguments defending different possible answers. Some participants rated these arguments under time pressure and cognitive load (forcing a fast, intuitive judgment), while others could deliberate freely.

Under time pressure, people leaned heavily on one cue above all others: whether the argument supported the answer they themselves had given, regardless of whether that answer or argument was actually correct. Given time to deliberate, justification validity became the dominant factor, and the belief-consistency (“myside”) bias shrank substantially. Deliberation doesn’t just help us solve problems—it sharpens our ability to tell a strong argument from a weak one.

Enter the machines

In a companion study3 presented at the Cognitive Science Society conference, we ran the identical task on several current large language models—including GPT-4.1, GPT-4o, LLaMA 3.3, Mistral, and Qwen—to compare their ratings to human data.

Most recent models achieved higher overall accuracy than humans on the problems. More importantly, their argument ratings closely mirrored the slow, deliberate evaluations of human participants who correctly solved the problems. Even when a model got a problem wrong, it still tended to rate the objectively correct answer’s argument more favorably—something intuitively-judging humans essentially never did after an incorrect answer.

Why it matters for human-machine decision-making

For a project like TANGO, built around designing decision-support systems that work with human reasoning, this distinction matters far more than a raw accuracy score. Knowing that a model reaches good performance is useful, but the “how” determines whether its judgments generalize and how much we should trust it as a partner in decision-making.

Understanding how AI systems weight different cues during argument evaluation—such as logical validity, explicitness, or alignment with their own answer—helps reveal where their judgments come from, and ultimately what will allow for a fruitful partnership between humans and machines. As AI moves further into roles that involve evaluating, advising, and justifying—not just answering—the process behind a judgment deserves as much scrutiny as the judgment itself. By bridging these insights from human reasoning with AI design, we can create more transparent, effective, and user-centered technologies.

 

Written by: Nicolas Beauvais, Matthieu Raoelison, & Wim De Neys

 

[1] De Neys, W., & Raoelison, M. (2025). Humans and LLMs rate deliberation as superior to intuition on complex reasoning tasks. Communications Psychology, 3, 141.

[2] Beauvais, N., & De Neys, W. (2026). Argument evaluation, fast and slow: Deliberation boosts justification strength discrimination in reasoning problems. Acta Psychologica, 270, 107725.

[3] Beauvais, N., & De Neys, W. (2026). Argument Evaluation Strategies in Human and Machine Reasoning: LLMs Resemble Sound Deliberate (But Not Intuitive) Thinkers. Proceedings of the Annual Meeting of the Cognitive Science Society, 48.