What makes a team smart?
Hiring measures one person at a time. In 2010, a paper in Science argued that teams have a measurable intelligence of their own — one that the talent of individual members barely predicts. What followed was a decade of replications, a serious statistical fight, and a set of sturdier neighboring findings that every organization can act on.
The finding: Across small groups working through deliberately varied tasks, a general collective-intelligence factor — c — emerged and predicted performance on held-out tasks. It correlated only weakly with members’ average or maximum individual intelligence. What did track it: members’ social sensitivity, the evenness of conversational turn-taking, and the proportion of women in the group.
The mechanism: Team intelligence behaves like a property of the interaction, not a sum of the members. The adjacent, better-established literatures agree: psychological safety, disciplined information sharing, and shared mental models all predict team learning and performance — and all of them are trainable structure rather than casting luck.
The product: Future Proof™ builds the shared-cognition layer directly — team knowledge maps that show what each member knows, targeted cross-training that closes overlap gaps, and cohort learning that turns individual mastery into shared mental models.
In this article
- 01The credential assumption
- 02Two studies and a factor called c
- 03It survived the move online
- 04The critique, honestly told
- 05The sturdier neighbor: psychological safety
- 06What teams share, and what they know together
- 07Team intelligence is infrastructure
- 08What the evidence doesn’t show
- 09What this means for practice
Organizations hire individuals and hope for teams. Nearly every formal instrument of talent management measures one person at a time: the résumé screen, the structured interview, the assessment battery, the performance review. The staffing model built on those instruments is quietly additive. Put enough individually excellent people in a room, and collective excellence should follow.
Anyone who has worked in more than one team knows the failure mode. Two groups drawn from the same talent pool, with matching credentials, identical tools, and the same deadline, routinely produce work of entirely different quality. And the difference persists from task to task — as if the team itself, rather than any member of it, had an ability level.
In 2010, a short paper in Science took that intuition seriously enough to measure it. It reported that groups do have a general, testable intelligence of their own — one that individual smarts barely predict (Woolley et al., 2010). The paper launched a research program, two major waves of replication, and one of the sharper statistical fights in modern organizational psychology. This article walks through all of it.
The original finding; the online and large-scale replications; the critique that keeps the construct honest. Then the neighboring literatures — psychological safety, information sharing, shared team cognition — which are older, larger, and directly actionable. The destination is practical. Whatever collective intelligence finally turns out to be, the things that predict it are properties of interaction and shared knowledge, not of résumés. And properties of interaction can be trained.
The credential assumption
Start with why the question needed asking at all. A century of selection research has made individual ability one of the most dependable predictors of individual job performance. Organizations have responded rationally: they test for it, interview for it, and pay for it. The unexamined step is the leap from person to group.
Staffing practice treats team capability as a portfolio sum — average the talent, perhaps overweight the star. Project plans are drawn as if ten strong individual contributors mechanically compose a strong team of ten. That assumption is not a small operational detail. It decides who gets hired, how teams are assembled, and what organizations think they are buying when they buy talent.
Psychometrics offered a template for testing the leap. Since the early twentieth century, researchers have known that a person’s performance on one cognitive task correlates positively with performance on nearly any other. A single general factor — g — summarizes a large share of the variation.
The question Woolley and colleagues posed was structural: run the same logic one level up. Give many groups many different tasks. If a group’s performance on one task says little about its performance on the next, then “team intelligence” is a category error and the additive staffing arithmetic is roughly fine. If a general factor emerges, then teams have an ability of their own. And then it becomes an empirical question, not a matter of taste, what that ability is made of (Woolley et al., 2010).
Two studies and a factor called c
Woolley, Chabris, Pentland, Hashmi and Malone ran exactly that test. Across two studies, they assembled several hundred people into small working groups. Each group went through a deliberately varied battery — brainstorming problems, judgment problems, planning problems, negotiation-flavored problems. The session finished with more complex criterion tasks, held out as the real test. Group performance proved consistent across very different kinds of work: groups that did well on one task tended to do well on the others. A single statistical factor, which the authors called c for collective intelligence, captured a substantial share of the variance — and predicted how groups performed on the held-out tasks (Woolley et al., 2010).
The result that made the paper famous, though, was what c turned out not to be. It was only weakly related to the average individual intelligence of the group’s members. And it was only weakly related to the maximum — the smartest person in the room. Assembling a team of the highest-scoring individuals available is, on this evidence, a surprisingly unreliable way to produce a smart team. Whatever c measures, it is not a simple sum of individual cognitive horsepower. Neither of the two variables staffing decisions optimize hardest carried much of the signal.
Three things did predict it, and they point in one direction. First, members’ average social sensitivity, measured with the Reading the Mind in the Eyes test. That task asks people to infer what someone is thinking or feeling from photographs cropped to the eye region. Groups whose members read others well scored markedly higher on c. Second, the evenness of conversational turn-taking. Groups in which speaking time was spread relatively equally beat groups dominated by one or two voices.
Third, the proportion of women in the group. The authors linked that effect statistically to social sensitivity, on which women scored higher on average — not to gender itself (Woolley et al., 2010). None of these relationships is anywhere near deterministic. The associations are meaningful rather than overwhelming, and c itself leaves plenty of variance unexplained. But the pattern is coherent. Collective intelligence behaved like a property of how the group processed each other — who spoke, who listened, who noticed — not of what each member privately knew.
It survived the move online
A finding this convenient invites suspicion. The first thing skeptics reached for was the setting: face-to-face laboratory groups, rich in nonverbal signal. So the most interesting early replication removed the faces. Engel and colleagues re-ran the collective-intelligence battery with groups collaborating online, through text, with no visual contact at all. The Reading the Mind in the Eyes test predicted group performance roughly as well online as it did in person (Engel et al., 2014).
A test built from photographs of eyes predicted the performance of people who could not see each other’s eyes. The authors’ interpretation sits in their title: the instrument is really “reading between the lines.” It indexes a general capacity to model other minds from whatever evidence is available — word choice, timing, silence — rather than literal face-reading. For distributed teams, the practical implication is direct and a little comforting. The raw material of collective intelligence travels through text.
The second question was scale, and it took a decade to answer properly. Riedl and colleagues pooled the accumulated evidence — hundreds of groups across many samples, tasks, and settings. The basic structure replicated: a general factor emerges, it is stable, and it predicts how a group will perform on new work (Riedl et al., 2021).
The large-sample picture also refined the story in a way the headline version usually skips. Collective intelligence is not a choice between composition and process; both contribute. Who is on the team matters — skill composition carries real weight. So does how the team works together: social perceptiveness and interaction process add predictive power beyond what member ability explains. Ten years on, the defensible version of the finding is not that talent is irrelevant. It is that talent is one input into a system whose output depends heavily on collaboration structure.
The critique, honestly told
Replication at scale did not end the argument — it sharpened it. The argument deserves to be told straight, because the airport-book version of this literature usually leaves it out. In 2017, Credé and Howardson published a detailed statistical comment arguing that the evidence for c is weaker than advertised (Credé & Howardson, 2017). Their central objection is subtle but serious. Showing that group performances correlate across tasks does not, by itself, establish a distinct group ability.
Groups that are good at tasks are — almost by definition — good at tasks. A general performance factor could reflect member ability, task similarity, shared method variance, or mundane confounds like effort and coordination time. None of that would make it an “intelligence” in any deep sense. On their reanalysis, the factor structure was less clean than the original claim. And c’s incremental validity — what it predicts beyond what you already knew from the members individually — was contested.
How should a practical reader hold this? Treat c as a productive research program, not settled law. The dispute concerns what the correlations mean — one underlying ability versus a looser bundle of performance influences. It does not concern whether the correlations exist: the large-scale replication is real, and so is the critique of its interpretation (Riedl et al., 2021). That is a normal, healthy condition for a young construct, and it sharpens the practical question rather than dissolving it. Because the strongest case for building smarter teams never rested on c alone; it rests on a set of adjacent literatures that are older, broader, and less fragile — and that keep pointing at the same levers.
Evidence for a collective intelligence factor in the performance of human groups.Woolley et al., Science, 2010 — the title of the paper
The sturdier neighbor: psychological safety
The best-established of those adjacent literatures began in 1999, when Amy Edmondson gave a name and a measure to something operating beneath the level of skill. She called it psychological safety: a shared belief that the team is safe for interpersonal risk-taking (Edmondson, 1999). In a psychologically safe team, a member can admit a mistake, ask a naive question, flag a concern, or float a half-formed idea — without expecting to be embarrassed or punished for it.
Studying intact work teams inside a real company, Edmondson found that safety predicted learning behavior — seeking feedback, discussing errors, experimenting, asking for help. Through learning behavior, it predicted team performance. The causal grammar matters. Safety is not niceness, and it is not low standards. It is the climate condition under which the information a team needs actually surfaces.
Two decades later, the construct has the kind of evidence base that c is still accumulating. Frazier and colleagues meta-analyzed the psychological-safety literature across a large number of independent samples. Safety was reliably associated with learning behaviors, engagement, and task performance, across industries and team types (Frazier et al., 2017). The associations are consistent rather than enormous — this is organizational psychology, not physics. But they survive pooling across studies, settings, and measures, which is precisely the test most workplace folk theories fail.
The finding also has a famous corporate echo, and it is worth citing carefully. Google studied hundreds of its own teams in an internal effort it called Project Aristotle. Its analysts reportedly reached the same headline: who was on the team mattered less than how the team treated each other, and psychological safety was the standout differentiator. That result reached the public through journalism — Charles Duhigg’s New York Times Magazine account — not through peer review. It should be weighed exactly as what it is: an unpublished internal analysis, unauditable from outside, that happens to converge with the published evidence (Duhigg, 2016). Convergence at corporate scale is worth something; it is not worth confusing with replication.
The most quoted exhibit in this whole literature — Project Aristotle — reached the public through a magazine story, not a journal. Its methods cannot be audited from outside. Cite it as a corporate anecdote that happens to agree with the published evidence, never as the evidence itself.
What teams share, and what they know together
Safety describes the climate; the next two literatures describe the currency. Teams exist, in large part, to pool information that no single member holds. A meta-analysis by Mesmer-Magnus and DeChurch confirmed two things: the pooling pays, and teams are systematically bad at it. Across studies, information sharing reliably predicted team performance.
But teams share the wrong information. Discussion gravitates toward what everyone already knows, while uniquely held information — the private fact, the unshared doubt, the one member’s contradictory data point — surfaces far too rarely (Mesmer-Magnus & DeChurch, 2009). The failure is structural rather than moral. The information a team most needs to hear is precisely the information least likely to come up on its own. No one else can prompt for what only one person knows.
The fix is a prompt, not a virtue. Uniquely held information surfaces when someone explicitly asks the room for it — what do you know that the rest of us do not? — and structures like round-robins and written first drafts exist to force the pooling that free discussion reliably skips.
The second currency is shared understanding. DeChurch and Mesmer-Magnus meta-analyzed the research on team cognition — team mental models, the overlapping pictures members hold of the task, the tools, and each other. Shared cognition predicted both team process and team performance (DeChurch & Mesmer-Magnus, 2010). Teams whose members know what each other know coordinate implicitly: they route questions to the right person, anticipate handoffs, and notice gaps before the gaps become failures. Teams without shared models coordinate by meeting, which is slow, or by assumption, which is worse. In this literature, “being on the same page” stops being a figure of speech and becomes a measurable property with a performance payoff.
Team intelligence is infrastructure
Put the strands side by side and a consistent picture emerges. The c program says group performance is general across tasks and only weakly inherited from member IQ; its predictors live in the interaction — social perception, even turn-taking (Woolley et al., 2010). The safety literature says teams perform when members can afford to say what they actually know (Frazier et al., 2017). The information-sharing literature says performance tracks whether uniquely held knowledge gets pooled (Mesmer-Magnus & DeChurch, 2009). The team-cognition literature says performance tracks whether members carry a shared map of the work and of each other (DeChurch & Mesmer-Magnus, 2010). Four literatures, four methods, one moral.
4 literatures Collective intelligence, psychological safety, information sharing, and shared team cognition — four independent research programs, four methods, one moral: the signal lives in the interaction, not the parts list (Woolley et al., 2010), (Frazier et al., 2017), (Mesmer-Magnus & DeChurch, 2009), (DeChurch & Mesmer-Magnus, 2010).
Every one of those properties belongs to the team’s operating system, not to its parts list. And every one can be changed: turn-taking can be structured, safety can be led, sharing can be prompted, mental models can be built deliberately. That is the deepest practical revision this evidence demands. Organizations treat team quality as a casting problem — find the right people — because casting is what their instruments can see. The research says team quality behaves more like infrastructure: built, maintained, degraded by neglect, and largely invisible to measurement systems that only ever look at individuals.
What the evidence doesn’t show
Every literature in this article is alive, which means the honest caveats are load-bearing. Six matter most:
- c’s factor structure is genuinely contested. The reanalysis by Credé and Howardson is not a fringe objection; whether c is a distinct latent ability or a redescription of general task performance remains unresolved (Credé & Howardson, 2017).
- Lab tasks dominate the c literature. Most of the evidence comes from short-lived groups of strangers on structured tasks. Intact workplace teams — with history, hierarchy, politics, and stakes — are underrepresented, and effects may differ there.
- Direction of causality for safety is unsettled. Most psychological-safety evidence is correlational; success plausibly breeds safety as well as the reverse, and designs that cleanly separate the two directions are still scarce (Frazier et al., 2017).
- Project Aristotle is journalism, not peer review. Google’s internal analysis was never published; its methods cannot be audited, and it should be cited as a convergent corporate anecdote, nothing stronger (Duhigg, 2016).
- Effect sizes are modest. Social sensitivity, turn-taking, safety, and shared cognition each explain a minority of performance variance. They move averages; they do not guarantee outcomes for any particular team.
- Interventions to raise c are barely tested. Measuring what correlates with collective intelligence has run far ahead of trials showing how to increase it. Training programs built on this literature are extrapolations — reasonable ones, but extrapolations.
Where the evidence stops
- 1c’s factor structure is genuinely contested
- 2Lab tasks dominate the c literature
- 3Direction of causality for safety is unsettled
- 4Project Aristotle is journalism, not peer review
- 5Effect sizes are modest
- 6Interventions to raise c are barely tested
What this means for practice
Caveats noted, the synthesis converts into a short operating agenda. First, staff for the team, not only the role. When composing a group, look at the knowledge the team will hold collectively: where members overlap, where a critical area is one person deep, and where nobody covers it at all. A living map of who knows what is the cheapest team-intelligence instrument an organization can own. It also doubles as the raw material of shared mental models — the variable the team-cognition meta-analysis puts closest to performance (DeChurch & Mesmer-Magnus, 2010).
Second, train the interaction, because the interaction is where the evidence says the signal lives. Protect turn-taking in meetings — facilitation, round-robins, written first drafts before discussion — so decisions draw on the whole group rather than its loudest fifth (Woolley et al., 2010). Ask explicitly for uniquely held information — what does the room not know that you know? — because the sharing meta-analysis says it will not volunteer itself (Mesmer-Magnus & DeChurch, 2009). And lead for safety in the concrete, behavioral sense Edmondson described. Respond to flagged errors as data rather than as offenses, model fallibility from the front, and make questions cheap to ask (Edmondson, 1999).
Third, onboard people into the team’s shared cognition, not just the job description. A new member becomes productive when they learn what the team knows, who holds which piece of it, and how work actually moves. That learning can be explicit instead of osmotic. None of this requires believing every claim ever made for c. It requires noticing that four independent literatures, including the contested one, keep pointing at the same place: a team gets smart through the structure of what it shares. And structure, unlike talent, is something an organization can build on purpose.
How Future Proof™ applies this.
Future Proof builds the shared-cognition layer this evidence points at. Team knowledge maps show what each member knows and where coverage is one person deep — the who-knows-what backbone of shared mental models. Targeted cross-training closes the overlap gaps the map exposes, and cohort learning runs practice through the team as a group, so individual mastery becomes shared vocabulary and shared models. Assessments and analytics track the collective picture over time — because a team’s intelligence should be managed like the infrastructure it is, not left to casting luck.
See the platform →Selected papers.
This is not an exhaustive bibliography — these are the studies cited above.
The evidence, by year
- 1999Edmondson
- 2009Mesmer-Magnus
- 2010Woolley
- 2010DeChurch
- 2014Engel
- 2016Duhigg
- 2017Credé
- 2017Frazier
- 2021Riedl
- Woolley, A.W., Chabris, C.F., Pentland, A., Hashmi, N., & Malone, T.W. (2010). Evidence for a Collective Intelligence Factor in the Performance of Human Groups. Science 330(6004): 686–688. DOI
- Engel, D., Woolley, A.W., Jing, L.X., Chabris, C.F., & Malone, T.W. (2014). Reading the Mind in the Eyes or Reading between the Lines? Theory of Mind Predicts Collective Intelligence Equally Well Online and Face-to-Face. PLoS ONE 9(12): e115212. PDF
- Riedl, C., Kim, Y.J., Gupta, P., Malone, T.W., & Woolley, A.W. (2021). Quantifying collective intelligence in human groups. PNAS 118(21): e2005737118. PDF
- Credé, M., & Howardson, G. (2017). The structure of group task performance — A second look at “collective intelligence”: Comment on Woolley et al. (2010). Journal of Applied Psychology 102(10): 1483–1492. PDF
- Edmondson, A. (1999). Psychological Safety and Learning Behavior in Work Teams. Administrative Science Quarterly 44(2): 350–383. DOI
- Frazier, M.L., Fainshmidt, S., Klinger, R.L., Pezeshkan, A., & Vracheva, V. (2017). Psychological safety: A meta-analytic review and extension. Personnel Psychology 70(1): 113–165. PDF
- Mesmer-Magnus, J.R., & DeChurch, L.A. (2009). Information sharing and team performance: A meta-analysis. Journal of Applied Psychology 94(2): 535–546. PDF
- DeChurch, L.A., & Mesmer-Magnus, J.R. (2010). The cognitive underpinnings of effective teamwork: A meta-analysis. Journal of Applied Psychology 95(1): 32–53. PDF
- Duhigg, C. (2016). What Google Learned From Its Quest to Build the Perfect Team. The New York Times Magazine, February 25, 2016. PDF
See what your teams know — together.
Book a 20-minute demo. We’ll show you the team knowledge map, the cross-training engine, and the cohort analytics that turn individual learning into collective capability.