Your team just stopped saying 'I don't know'

A new study found AI access cut people's willingness to admit ignorance from 44% to 3% while their confidence doubled. The real risk isn't the wrong answers.

Your team just stopped saying 'I don't know'

A group of researchers ran a small, clever experiment this month. They asked people questions that a chatbot reliably gets wrong, then gave half of them access to that chatbot. Two things moved. The share of people willing to say “I don’t know” fell from 44% to 3%. Their confidence roughly doubled, from 30% to 76%. Accuracy went the other way, from 27% down to 9%.

The headline everyone repeated was the accuracy one: AI made people three times less accurate. That is the least interesting part, and the least reliable. The researchers deliberately picked a model that got these questions wrong almost every time, so critics were quick to point out that you have really just handed people a confident tool that happens to be wrong. Rig the tool, get bad answers. Fair enough.

But that objection only kills the accuracy claim. The finding that survives it is the one a leadership team should sit with: the collapse of “I don’t know.”

The number that matters is 3%

Forty-four percent of people, unaided, were comfortable admitting they did not know. Give them a chatbot and that comfort nearly vanishes. The tool did not teach them anything. It just made the admission feel unnecessary.

What makes this stick is the incentive result. When the researchers paid people to be accurate, the willingness to admit ignorance crept up only to 8%, and accuracy to 16%, both still well under the no-AI baseline. Money barely moved it. That tells you this is not a motivation problem you can fix with a KPI or a stern reminder to “check the AI’s work.” It is closer to a reflex, and the reflex is being switched off.

“I don’t know” was doing a job

We treat “I’m not sure” as a gap to be closed. In an organization it is closer to a piece of infrastructure. Uncertainty is the trigger that fires the next useful action: check the source, ask the specialist, escalate, sleep on it. A person who says “I don’t know” is routing the decision somewhere safer.

AI does not just supply an answer. It removes the moment where that trigger would have fired. You do not lose accuracy first and confidence second; you lose the small hesitation that used to catch the error before it became a decision. The wrong answer was always possible. What changed is that nobody stops to wonder about it.

The reviewer you were counting on

This quietly inverts the governance conversation. Most AI-safety effort points at the model: hallucination rates, evals, guardrails, red-teaming. The comfort we sell ourselves at the end of every such discussion is “we keep a human in the loop.” I have argued before that with AI writing most of the code, reviewing it is now the actual job.

The catch is that the loop only works if the human brings calibrated doubt to it. This study says the mere presence of the tool erodes exactly that doubt. A reviewer whose willingness to say “I don’t know” has dropped to 3% is not a control. It is a rubber stamp with a pulse. The safeguard and the thing that disables the safeguard are the same system.

A better model makes this worse

The obvious rebuttal is that they used a bad model, and a good one would help. For raw accuracy, partly true. For the confidence effect, it runs the other way. The more reliable your AI, the more trust it earns, and the more completely it retires your habit of checking. So on the rare occasion it is wrong, which is exactly the confidently-wrong output that passes every test, no one in the building is positioned to notice.

High accuracy paired with eroded skepticism is a more dangerous combination than mediocre accuracy paired with healthy suspicion. The strongest teams I have worked with treat a very good tool as more of a threat to their judgment, not less, precisely because it is so easy to stop thinking around it.

What I would actually do

Three things, none of them technical.

Watch your “I don’t know” rate. If confidence in your reviews, forecasts, and status updates is rising while outcomes are flat, that is the warning sign the study describes, not a productivity win. Rising certainty with unchanged results is a symptom.

Put friction back where the thinking used to be. Ask people to show what they verified, not what the tool told them. Effort used to be a free signal that someone had skin in the game; AI broke that signal, so now you have to ask for it on purpose.

Reward calibration over confidence. Make “I don’t know” cheap to say and valued when it is right. The person who flags uncertainty early is worth more to you than the one who is fluently, confidently wrong.

Everyone will soon rent the same handful of models. The edge will not come from which one you pick. It will come from whether your people are still willing to say the three words the tool never volunteers: I don’t know.

If you are trying to work out where AI genuinely improves your decisions and where it just makes them faster and more confident, that is the sort of question I dig into in an AI advisory hour.