Psychological Safety in AI Teams: Safe to Question, Safe to Say No

Turn this article into takeaways for your work.
Each assistant summarizes the article only for you and suggests best practices for your work.
Updated August 2026
Psychological safety in AI teams is Amy Edmondson's original concept, the shared belief that a team is safe for interpersonal risk-taking, extended to a new set of risks that only exist once AI is doing part of the work. It means people feel safe questioning an agent's output, admitting they used AI on a task, flagging an AI error or overriding its recommendation, and saying plainly, "I don't trust this result," without any of that costing them status, credibility, or a mark against their name.
Edmondson built her career studying why people stay quiet in hospitals, cockpits, and boardrooms even when speaking up would save a life or a project. The mechanism she found, that silence is a rational response to a room that punishes risk, transfers almost intact to AI. Nobody used to need permission to distrust a spreadsheet. Now a growing share of the workforce needs exactly that permission, and most teams have never granted it out loud. This article covers what's genuinely new about safety on human-AI teams, why automation bias and over-deference are the failure mode that shows up when it's missing, real research on why people stop asking "why," and what leaders do to build it back.
Extending Edmondson's Model: Four New Things People Need to Feel Safe Saying
Edmondson's original four-stage model, inclusion, learner, contributor, and challenger safety, was built around interactions between people. We cover that full model in psychological safety at work. Once an AI agent is a working part of the team rather than a tool one person quietly uses on the side, the same four stages have to stretch to cover a new set of moments that Edmondson never had to model in 1999, because the thing being questioned used to always be a colleague.

Safe to Question AI Output
Someone looks at what an agent produced and has a nagging sense something is off: a number that doesn't match their intuition, a citation that reads too clean, a step they'd expect a person to flag. Challenger safety, in Edmondson's model, is the stage where people feel free to question the plan or the person in charge. The AI-team version of that stage is questioning the machine, and it carries a different social cost: nobody looks bad for doubting a colleague's idea the same way they might look slow or paranoid for second-guessing a system "everyone else" seems to be accepting without comment.
Safe to Admit You Used AI
This one sits closer to inclusion and learner safety. Using AI on a draft, a summary, or a piece of analysis is not a confession, but it gets treated like one on plenty of teams, because nobody has said out loud that it isn't. A person who hides their AI use isn't lying out of malice; they're reading the room correctly if the room has never made disclosure feel safe. That silence has a direct cost: a manager loses visibility into how work actually got produced, which means they can't catch the exact failure mode this article is about.
Safe to Flag an AI Error or Override Its Recommendation
An agent gets something wrong, or a person overrides what it recommended and does something different. Both of those are contributor-safety moments, using your judgment and having it count, but they now involve overriding a system that was procured, configured, and probably championed by someone senior. Flagging a person's mistake risks an awkward conversation. Flagging or overriding a system's mistake risks looking like you don't trust the investment the company just made in it, which is a different, often heavier, kind of risk to carry alone.
Safe to Say "I Don't Trust This Result"
This is the sentence that matters most, and the one most teams have the least practice saying. It's not a technical claim about accuracy; it's a statement of uncertainty from a person who is supposed to look competent and decisive. Saying it about a colleague's work is uncomfortable but familiar territory. Saying it about a system that produced a fast, fluent, professional-looking answer, with nobody else in the room objecting, is the genuinely new interpersonal risk this article is named for.
Automation Bias and Over-Deference: The New Failure Mode
Here's the part most psychological safety advice misses entirely. When a team is unsafe around AI, the failure doesn't look like conflict. It looks like agreement. Everyone nods, the output ships, and the team looks perfectly aligned right up until the mistake surfaces somewhere expensive.

What Automation Bias Actually Is
Researchers Linda Skitka, Kathleen Mosier, and Malcolm Burdick gave this pattern a name in a 1999 study that is still the anchor citation for the field: automation bias, the tendency to accept an automated system's output as a substitute for the vigilant checking a person would otherwise do themselves. Their research, run in a simulated flight-monitoring task, found something specific enough to matter here: people made two distinct kinds of errors. Errors of omission, failing to notice a problem because the automated system didn't flag it. And errors of commission, actively following an automated recommendation even when other information available to them suggested it was wrong. Source: Skitka, Mosier & Burdick, "Does Automation Bias Decision-Making?"
That second category is the important one for a culture conversation. An error of commission isn't an accident. It's a person choosing the automated answer over their own judgment, in the moment, often while other cues were sitting right in front of them. That's not a training gap. It's a decision shaped by what the room around that person makes comfortable to say and do.
Why Psychological Safety Is the Missing Variable
Most writing about automation bias treats it as an individual cognitive flaw, something to fix with better UI or a training module. That framing misses the team-level mechanism entirely. An error of commission is easiest to make precisely when questioning the system feels socially costly and staying quiet feels free. Flip the incentive and the same person, with the same cognitive wiring, behaves differently. Skitka's lab finding and Edmondson's field research describe the same shape from two directions: automation bias shows people default to acceptance under uncertainty, and psychological safety shows what makes speaking up under uncertainty costly or free. Over-deference to AI isn't a new kind of human failure. It's Edmondson's oldest one, a team that punishes voice and rewards silence, wearing a new interface.
The Loan-Officer Study: When Not Asking "Why" Is the Point
A 2026 study led by Harvard Business School's Alex Chan puts real numbers on how deliberate that silence can be. Researchers had 2,512 online participants act as lending officers reviewing pairs of real $10,000 loan requests, with an AI system providing a risk score for each. Eighty percent of participants wanted to see the AI's risk score, but only 46% went further and asked to see the explanation behind it, even when that explanation was one click away. When researchers randomly assigned some participants a financial bonus tied to how well the loan actually performed rather than a flat fee, explanation-avoidance rose by 20 percentage points, and avoidance jumped by more than 10 additional points when participants had reason to believe the explanation might reveal that race or gender had influenced the AI's score. Source: HBS Working Knowledge, "When AI Gives Advice, Employees Rarely Ask Why"
Chan's own framing of the result is the line worth sitting with: "The biggest risk of AI isn't just bad answers or lack of adoption. It's training people to stop asking why." His participants weren't confused about how to see the AI's reasoning. A meaningful share of them chose not to look, because looking risked surfacing something inconvenient. That's not automation bias as a passive cognitive shortcut. It's automation bias as an active, motivated choice, made easiest by a setup where nobody's incentives rewarded asking the harder question.
Key Facts
- Automation bias, the tendency to accept an automated system's output over a person's own vigilant checking, was defined by Skitka, Mosier, and Burdick's 1999 study, which found people make both errors of omission (missing a problem the system didn't flag) and errors of commission (actively following a bad automated recommendation over contrary evidence). Source: International Journal of Human-Computer Studies, via ACM
- In a 2026 study of 2,512 participants acting as loan officers, 80% wanted to see an AI system's risk score but only 46% asked to see the explanation behind it; explanation-avoidance rose 20 percentage points when a financial bonus was tied to loan performance, and rose more than 10 additional points when the explanation risked revealing racial or gender bias. Source: HBS Working Knowledge
- Amy Edmondson and Jayshree Seth's 2026 Harvard Business Review research found that despite the productivity gains companies expect from AI, team performance often declines instead: people start second-guessing themselves and trust erodes in ways that are hard to pinpoint, a dynamic they describe as "trust ambiguity." Source: Harvard Business Review
- Microsoft's 2026 Work Trend Index found that when managers deliberately created psychological safety around AI experimentation, employees reported up to 20 points higher AI readiness and perceived value, and were 1.4 times more likely to become high-frequency users of agentic AI. Source: Microsoft Work Trend Index 2026
- Deloitte's 2026 Global Human Capital Trends research found 60% of executives already use AI in decision-making, but only 5% say they manage that well, a gap that compounds into what Deloitte calls AI cultural debt when nobody is safe enough to name the problem out loud. Source: Deloitte 2026 Global Human Capital Trends
Why This Kind of Silence Is Easy to Miss
Old-fashioned silence in meetings has a familiar shape: someone has a concern, the room goes quiet, and everyone can sense the tension even if nobody names it. Silence around AI often doesn't feel like tension at all. It feels like smooth agreement, because deferring to a system doesn't carry the same visible social weight as deferring to a person. Nobody watches a colleague quietly accept an AI's number and reads it as backing down the way they'd read someone caving to a loud voice in the room. That's what makes it dangerous: a team can look perfectly safe, full of open debate about everything people disagree with each other on, while running an unexamined blind spot around anything the AI touches.
Edmondson and Seth's 2026 research names the mechanism directly: sustained AI use can quietly erode a person's confidence in their own ability to challenge an AI's output, even when they have the exact expertise needed to catch what's wrong. That's the opposite of how confidence normally works, since a domain expert usually gets more assertive with experience, not less. Watching a fluent, professional-looking answer get produced over and over teaches a different lesson: that hesitating to accept it is what needs justifying, not the acceptance itself. Trust when your teammate is an AI covers why AI output is so much harder to read for confidence than a person's; this is the safety consequence of that same mechanic, playing out as quiet self-doubt instead of an open disagreement anyone could point to and fix.
The cost compounds because it's invisible as it happens. Someone who swallows a doubt about a colleague's plan at least knows they did it. Someone who accepts an AI's output without voicing a doubt often doesn't experience it as caving at all; it just feels like moving on. Multiply that across a team over months and you get a group that has quietly stopped exercising the exact judgment it was hired for, with no single moment anyone could point to as where things went wrong.
What Leaders Do to Build Psychological Safety in AI Teams
The teams that avoid this failure mode aren't the ones with the most cautious AI policy. They're the ones where a handful of specific leader behaviors make questioning, admitting, and doubting a normal part of how work gets done, not an exception someone has to justify.

Model Questioning AI Yourself, Visibly
The single highest-leverage move is a leader publicly second-guessing an AI output in front of the team, including the times it turns out to have been right. "I want to check this number before we use it," said by the most senior person in the room, does more to license the same behavior in everyone else than any written policy. Microsoft's research on managers who visibly model their own AI use found real, measurable effects on the whole team's willingness to engage critically with AI, not just their raw adoption of it. The behavior that spreads is the visible act of checking, not a memo saying checking is encouraged.
Reward Catching the Error, Not Just Shipping Fast
If the only recognized outcome on a team is speed, whoever catches a flawed AI output and slows things down to fix it looks like the person who missed the deadline, not the one who saved the project. Recognition systems built entirely around throughput quietly train people to stop looking too closely. Naming and crediting whoever caught an AI error before it shipped, the way a good team credits whoever spotted a bug before release, makes catching mistakes a status-building move instead of a status-risking one.
Treat AI Mistakes as Blameless, Every Time
How a team responds to a mistake trains everyone watching, whether the mistake came from a person or a machine. Blame culture vs. learning culture is exactly as relevant to an agent's error as to a person's. A team that quietly blames whoever "should have caught" the AI's mistake teaches everyone the safest move next time is silence, not scrutiny. A team that treats the mistake as information, what did the agent miss and what changes so it doesn't recur, keeps the door open for the next person to speak up honestly instead of covering for themselves.
Make "I Don't Trust This" a Normal Sentence
Leaders can name this directly rather than hoping it emerges. Saying, in a team meeting, "if you ever look at what an agent produced and think something's off, that's exactly the sentence I want to hear, even if you can't fully explain why yet" does more work than it sounds like it should. It converts a vague, hard-to-justify feeling into an explicitly welcomed contribution, precisely the gap Chan's research found people fall into: avoiding a question not because they lack the tools to ask it, but because nothing in the room told them asking was worth the friction.
Build the Habit of Asking "Why," Not Just "What"
Teams that ask "what did the AI recommend" and stop there are structurally set up for the over-deference this article describes. Teams that build "why did it recommend that, and does the reasoning hold up" into the workflow itself, as a standard step above a defined stakes threshold rather than an occasional gut check, do the opposite of what Chan's participants did when a bonus made the explanation inconvenient to look at. It's the same discipline human-agent teams need for accountability generally, applied to the moment before output ships rather than after something has already gone wrong.
A Quick Audit: Does Your Team Actually Have This?
Most teams can't tell whether they have psychological safety around AI until they look for specific signals, not a general vibe. The table below separates what the presence of safety actually looks like from what its absence looks like, since both can appear calm on the surface.
| Signal | Present | Missing |
|---|---|---|
| Questioning AI output | People say "I want to check this" without hedging or apologizing first | Doubts get raised only privately, after the output has already shipped |
| Admitting AI use | Disclosure is routine, mentioned the same way someone would mention a source | AI use is mentioned only when directly asked, or not at all |
| Flagging or overriding AI errors | The person who caught the error gets named and credited | The error surfaces later with no clear record of who almost caught it |
| Saying "I don't trust this" | Said plainly in team settings, treated as useful information | Said only one-on-one, if said at all, and treated as a confession |
| Leader behavior | Leaders visibly question AI output themselves, including their own | Leaders treat AI output as settled once a senior person has approved it |
A team that recognizes itself mostly in the right-hand column isn't broken; it's simply running the exact setup Chan's research and Skitka's original automation-bias framework both predict will quietly accumulate errors of commission. The fix isn't a new tool. It's making the left-hand column the normal, unremarkable way the team already talks.
Where to Go Next
This article extends the foundational research on psychological safety into a specific new context. From here:
- Psychological safety at work, for Edmondson's original four-stage model and Google's Project Aristotle research this article builds on
- Trust when your teammate is an AI, for the mechanics of calibrated trust and why AI's confident tone can't be read the way a person's can
- Human-agent teams, for the ownership and accountability questions that come up once an agent is a working part of the team
- Blame culture vs. learning culture, for how a team's response to any mistake, human or AI, shapes whether the next person speaks up
- AI cultural debt, for what accumulates when trust and safety gaps around AI go unaddressed
- AI etiquette and workplace norms, for the broader set of unwritten rules around disclosure this article's second risk category sits inside
- The Frontier Firm and the rise of the agent boss, for how the orchestrator role changes what leaders need to model
- Building an AI-ready culture, for the wider set of practices this fits inside
- What is AI-native culture?, for the foundational concept this article's questions sit inside
- Leading culture in the age of AI, for the full leadership picture across all of Section 7
The health of an AI team shows up in a single moment: when someone says "I don't trust this result," does the room treat it as doing the job, or as slowing everyone down? The disclosure norms, the error reporting, the willingness to override a confident machine all follow from the answer. Build for that moment and the rest gets easier. Punish it once and people quietly stop saying it, which is exactly when the automation bias starts to cost you.

Co-Founder, Rework.com
On this page
- Extending Edmondson's Model: Four New Things People Need to Feel Safe Saying
- Safe to Question AI Output
- Safe to Admit You Used AI
- Safe to Flag an AI Error or Override Its Recommendation
- Safe to Say "I Don't Trust This Result"
- Automation Bias and Over-Deference: The New Failure Mode
- What Automation Bias Actually Is
- Why Psychological Safety Is the Missing Variable
- The Loan-Officer Study: When Not Asking "Why" Is the Point
- Key Facts
- Why This Kind of Silence Is Easy to Miss
- What Leaders Do to Build Psychological Safety in AI Teams
- Model Questioning AI Yourself, Visibly
- Reward Catching the Error, Not Just Shipping Fast
- Treat AI Mistakes as Blameless, Every Time
- Make "I Don't Trust This" a Normal Sentence
- Build the Habit of Asking "Why," Not Just "What"
- A Quick Audit: Does Your Team Actually Have This?
- Where to Go Next