AI Groupthink: When Everyone Uses the Same Model

I spend most of my time encouraging teachers to use AI, so a paper warning that AI use can erode how groups think is exactly the kind of argument I want to take seriously. Sunny (2026) opens with a scene that should feel familiar. Five executives walk into a strategy meeting, each having consulted the same AI assistant on their own beforehand.

They arrive in striking agreement, and they read that agreement as rigor. Sunny’s claim is that it isn’t rigor at all. It’s conformity, routed through a shared algorithm that shaped everyone’s thinking before anyone spoke. He calls it Invisible Groupthink, and the name earns its keep.

The empirical hook is hard to wave away. Sunny points to McKinsey’s 2025 survey (cited in Sunny, 2026) of nearly 2,000 organizations, where 88 percent use AI somewhere yet only 6 percent qualify as genuine high performers. He argues part of that gap is collective, not just individual.

How AI Groupthink Differs From the Old Groupthink

Janis described classical groupthink back in 1982 as a product of cohesive groups, directive leaders, and members who feel the pressure to conform and swallow their doubts. Sunny argues this new version needs none of those things. There’s no dominant personality, no social pressure, no leader steering the room. The only requirement is that team members use the same AI tool, and because nobody feels pushed, the convergence reads as validation.

That’s why the standard fixes miss it. Devil’s advocacy, anonymous voting, and the rest of the anti-groupthink toolkit all target social dynamics. Sunny’s mechanism isn’t social, so those tools have nothing to grab.

AI Groupthink

How the Convergence Happens

The mechanism, which Sunny labels the Informational Quiet Homogenization Effect (IQHE), runs in four steps. Team members start with overlapping assumptions from a shared workplace culture. The same model, trained on the same data, hands them correlated answers across what feel like separate queries. Each person refines their prompts, and the AI agrees more warmly with every pass. Then the team pools the results and treats the overlap as independent corroboration, never noticing the algorithm as the common source.

This isn’t a wild guess. Sunny leans on a 2011 experiment by Lorenz and colleagues (cited in Sunny, 2026), where showing people their peers’ estimates made their judgments converge and their accuracy drop, even though they believed they were reasoning independently.

AI outputs act like those peer estimates, except everyone stares at the same one. Sycophancy deepens it. Sunny cites Zhu et al. (2024, cited in Sunny, 2026), who found the tested models conform to majority opinion regardless of accuracy, and fastest when uncertain, the exact moment a hard decision needs a dissenting voice.

Here’s the prediction that gives the paper its edge. Sunny calls it the Amplification Paradox: as AI adoption grows more uniform and consultation more intense, the epistemic diversity a group needs to catch its own errors falls, so collective decisions get less resilient even as individual work looks sharper. He borrows Page’s diversity formula: collective error equals average individual error minus the group’s diversity bonus. Flatten the diversity and collective error climbs, no matter how capable each person gets.

The scale is what unsettles me. Sunny warns that “an organization that deploys the same LLM-based tool across 50,000 knowledge workers has created structural preconditions for IQHE in every team, at every decision level, simultaneously and without any member being aware of the risk” (pp. 5- 6). Read that with a school district in mind.

When every teacher plans with the same assistant, every department drafts with the same model, and every leadership team consults the same tool, the conditions Sunny describes are already in the building. This is the collective cousin of what Sourati and colleagues (2026) found about LLMs homogenizing individual thought, and it extends the team-level effects Riedl and colleagues (2025) traced in their study of AI’s social forcefield in groups.

What Teachers and Leaders Can Do

Sunny’s remediation is refreshingly concrete. His framework, Diversity-Preserving AI Governance (DPAIG), asks organizations to use at least two different model families for important decisions, record each person’s independent take before anyone consults AI, and keep an AI-free red team for high-stakes calls. None of that means abandoning AI. It means using it on purpose.

That lines up with what I keep arguing for teachers. Intentional AI use means building difference back into the process. Have staff draft their own take before the group prompts anything. A single question can go through two different model families, which gives the team a real comparison.

A caveat keeps me from overselling this. Sunny is upfront that the IQHE is a proposed mechanism, not a demonstrated one, and that the Amplification Paradox is a derived prediction still waiting on evidence. The next step he names is the right one: experiments comparing team accuracy under single-AI, multi-AI, and human-only deliberation on problems with known answers. Until that work lands, this is a sharp theory, not a settled finding.

The warning still holds its shape. Sunny closes with this interesting line: “the organizations investing most in AI may be building the most efficient individual workers and the most homogenized collective thinkers” (p. 9). The same caution belongs in education.

We talk constantly about helping each teacher and student work faster with AI. We rarely ask what happens to a school’s collective judgment when everyone leans on the same model. It pairs with what Abdulhai and colleagues (2026) showed about these tools flattening the language we use. Sameness is the hidden cost of convenience.

References

  • Abdulhai, M., White, I., Wan, Y., Qureshi, I., Leibo, J., Kleiman-Weiner, M., & Jaques, N. (2026). How LLMs distort our written language. arXiv preprint arXiv:2603.18161. https://arxiv.org/abs/2603.18161
  • Janis, I. L. (1982). Groupthink: Psychological studies of policy decisions and fiascoes (2nd ed.). Houghton Mifflin.
  • Lorenz, J., Rauhut, H., Schweitzer, F., & Helbing, D. (2011). How social influence can undermine the wisdom of crowds. Proceedings of the National Academy of Sciences, 108(22), 9020–9025. https://doi.org/10.1073/pnas.1008636108
  • Riedl, C., Savage, S., & Zvelebilova, J. (2025). AI’s social forcefield: Reshaping distributed cognition in human-AI teams. arXiv preprint. https://arxiv.org/html/2407.17489v2 
  • Sourati, Z., Ziabari, A. S., & Dehghani, M. (2026). The homogenizing effect of large language models on human expression and thought. Trends in Cognitive Sciences. Advance online publication. https://doi.org/10.1016/j.tics.2026.01.003 
  • Sunny, Md Rabius, Invisible Groupthink: AI Homogenization, Epistemic Diversity, and Organizational Decision Resilience (May 06, 2026). Available at SSRN: https://ssrn.com/abstract=6724619
  • Zhu, X., Zhang, C., Stafford, T., Collier, N., & Vlachos, A. (2024). Conformity in large language models. arXiv:2410.12428. https://arxiv.org/abs/2410.12428

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top