Team compatibility test: will these two people work well together?

A team compatibility test measures each person on a validated personality model, then reads how those profiles interact: who shares a working rhythm, which pairs complement each other, and where friction is plausible. It cannot issue a verdict on a pair, and this page is honest about why. Name your team below and the map builds itself.

No signup, no card. You get one link to share, and the map builds itself as each person finishes.

60 questions each, about 8 minutes Built on the Big Five (BFI-2 tradition) Nobody sees anyone else’s answers

Cannot ask them? Read the pair on your own, from behaviour you have already watched.

On this page
  1. Read the pair
  2. What it can predict
  3. A team, mapped
  4. The three surfaces
  5. Personality, or structure?
  6. Five pairings
  7. The agreement
  8. What the evidence says
  9. What it must never decide
  10. Choosing a framework
  11. Questions
  12. Methods & sources

Read the pair

You probably have one person in mind. Start there.

Almost every compatibility tool needs the other person to sign up. That is the one thing you often cannot arrange, because they are busy, or unconvinced, or they are your manager. So this reads the pair from what you can see on your own.

The one-sided pair readNothing is stored, nobody is emailed, and they never have to know you did this.
1 / 6

About you

A decision is being made and you disagree. What do you usually do?

Two questions about you, four about behaviour you have already watched. It runs the same rule order as the real map: two natural drivers first, then two candid styles, then a wide gap on who leads, then a gap on pace, and finally the case where two people are simply very alike. What comes back is a hypothesis on one named channel, which is roughly the shape of the measured answer at much lower resolution.

The last question is the important one, and it is not about personality at all. It asks what was actually unresolved the last time things went badly. A large share of what gets called a personality clash turns out to be an undefined decision or an unstated standard, and those feel identical from the inside. Section 05 is about telling them apart.

What it can predict

The compatibility score you are imagining cannot exist

This is the part every page in this category skips, so it goes near the top rather than in the small print.

In 2017, researchers ran the cleanest test anyone has managed of whether two people’s fit can be predicted in advance. Before meeting, participants completed more than a hundred self-report measures: traits, values, preferences, what they said they wanted. Then they met at speed-dating events, and the researchers pointed machine learning at the results[1].

The models worked, up to a point. They predicted a slice of how much a given person tended to like others, and a slice of how much others tended to like them. Then they reached the thing everyone actually wanted, the unique chemistry between these two specific people, and predicted none of it.

What traits predicted, and what they did not0%10%20%30%Share of the variance predicted, before the two people had metHow much you like peoplein general (actor variance)4% to 18%How much people like youin general (partner variance)7% to 27%How you two get onthis pair specifically (relationship variance)Nothing. 0%.
What traits measured beforehand could and could not predict. The first two bars are real, useful signal about individuals. The third is the thing a compatibility score claims to be, and it came back empty. Ranges span the study’s two samples[1].

That was romantic attraction rather than work, and speed dating is a first meeting rather than a standing team. Both caveats are fair. But the finding points the same way as the workplace evidence, and it sets the honest ceiling for everything below: a personality profile describes two individuals placed side by side. It does not describe the pair.

Which raises the obvious question. If pair fit cannot be predicted, what is a compatibility read for? Two things, and both are worth having. It tells you how each of you tends to be experienced, which is real and measurable. And it names the channel where your two default settings are most likely to collide, which is the difference between "we just clash" and "we have never agreed who makes this call".

Similarity is a first-impression effect

The other half of the folklore is that similar people get on. In first meetings, that is strongly true: across 460 effect sizes, actual similarity correlated .47 with attraction. In existing relationships, it was not significant[2]. Similarity opens a door. It does not keep one open.

At work, the split is sharper still, and it is the most practically useful finding on this page. A recent meta-analysis separated deep-level similarity, meaning values, attitudes and working style, from surface-level similarity, meaning age, gender, tenure and background[3].

Which kind of similarity actually matters at workDeep-level (values, attitudes, working style)Surface-level (age, gender, tenure)0.00.20.40.6Correlation with the outcome (r)Getting on with your manager.46not significantJob satisfaction.37not significantJob performance.32not significant
Deep-level similarity tracks the outcomes people care about. Surface-level similarity does not: the associations were not significant for job performance, job satisfaction, the relationship with your manager, citizenship behaviour or commitment[3]. It is not plotted as a small bar, because the study reports an absence rather than a small effect.

Time sorts the two out. In one landmark study, working together for longer weakened the effects of surface-level difference and strengthened the effects of the deep-level kind[4]. What you notice about a colleague in week one stops mattering. What you find out about them in month nine starts to.

Go deeper: why "opposites attract" is half right

It depends entirely on which channel you mean, and the distinction is unusually clean. Warmth is a corresponding signal: people tend to meet warmth with warmth and coolness with coolness, so similar levels sit comfortably together. Dominance is complementary: it invites its opposite.

In the lab, people who complemented a partner’s dominance, meeting leading with steadying, liked that partner more and felt more comfortable than people who mirrored it[5]. The field version is sharper. Restaurants with extraverted leaders were more profitable when staff were passive, and less profitable when staff were proactive[6]. The same trait in the same leader flipped from asset to liability depending on who was standing next to them. That is what people are reaching for when they say chemistry, and it is why a trait score on its own can never carry the answer.

A team, mapped

Here is a real five-person map. Open Dynamics.

This is a complete team map, computed live in your browser by the same engine that builds real ones. The Dynamics chapter is the compatibility chapter: it is where the pair reads live.

Sample team · five members · Big Five Live, computed in your browser

Sample team · Team Map

The Drivers

Energetic and results-focused, they set the pace and keep the target in view.

A high-energy, outward-facing team

This team brings energy and pushes work forward, and no work dimension here is left without an anchor.

This reading rests on outward energy and disciplined execution, the two dimensions the group covers most strongly. A current pattern in the data, not a fixed label.

5people
Team Map

One dot per person, at their percentile against a general adult sample. Social energy runs up the side, warmth across the bottom.

Scatter plot of interpersonal style. The horizontal axis runs from candid on the left to warm on the right; the vertical axis runs from reserved at the bottom to outgoing at the top. Both are percentiles against a general adult sample, so the 50th percentile is typical. 5 people are plotted. Each position below is given the same way the chart labels it: the percentile toward warm, then the percentile toward outgoing. The team's centre sits at warm 44th, outgoing 66th. A moderate spread of styles. Priya Menon: warm 28th, outgoing 84th, driving & challenging; Diego Alvarez: warm 14th, outgoing 80th, driving & challenging; Sofia Haddad: warm 58th, outgoing 74th, energising & mobilising; Marcus Bell: warm 44th, outgoing 46th, focused & independent; Aisha Okafor: warm 76th, outgoing 48th, steady & supportive. The plot also shows a two-subgroup split: Warm group (3) and Candid group (2).

ENERGISING & MOBILISINGDRIVING & CHALLENGINGSTEADY & SUPPORTIVEFOCUSED & INDEPENDENTPriyaDiegoSofiaMarcusAisha↑ OUTGOING↓ RESERVED← CANDIDWARM →
Warm group (3)Candid group (2)Team centreThis team’s spreadMiddle 50% of adults
Every position
Each team member’s position on the team map, as percentiles against a general adult sample.
PersonWarmOutgoingWorking style
Priya Menon28th84thDriving & challenging
Diego Alvarez14th80thDriving & challenging
Sofia Haddad58th74thEnergising & mobilising
Marcus Bell44th46thFocused & independent
Aisha Okafor76th48thSteady & supportive
Team centre44th66thA moderate spread of styles

A lens on interpersonal style: how much drive a person brings into a room, and how much warmth. It covers two of the five traits. Openness, conscientiousness and emotional steadiness sit in Team Shape. Each dot is one person's percentile against a general adult sample, the centre cross marks the 50th, and the lighter inner panel holds the middle 50% of adults. This team's styles sit at a moderate distance from one another.

Coverage

What this team covers, and where the gaps are

For each way of working, does anyone here anchor it? A dimension is covered when at least one member is genuinely strong; a gap is where no one is.

5 of 5 covered
Pressure handlingCovered
70team best / 100
Anchored by Priya Menon
Influence & energyCovered
84team best / 100
Anchored by Priya Menon
Learning & adaptabilityCovered
64team best / 100
Anchored by Sofia Haddad
Collaboration styleCovered
76team best / 100
Anchored by Aisha Okafor
Execution & disciplineCovered
77team best / 100
Anchored by Priya Menon

Collective strengths

What this mix of people can be counted on for, together.

  1. No single tendency dominates: the team adapts across a range of work, with no one dimension standing out as its anchor.

Worth watching

Where this composition could bite

Modest, directional patterns from the team-composition research. Each one is framed as something to manage, not a flaw in anyone.

A cooperation floor set by one member

Watch

On a weakest-link trait, a team performs closer to its lowest member than to its average. Here that is Diego at 14. The same pattern shows in their own profile as "Direct communication style under pressure". This is not a flaw to fix in one person. It is the place where process has to carry what disposition does not.

A potential subgroup split on collaboration

Watch

This group divides into two clusters, most visibly on collaboration: Sofia Haddad, Marcus Bell, Aisha Okafor on one side, Priya Menon, Diego Alvarez on the other. Differences that line up like this can harden into an us-and-them if no one names them. Mixing the two groups across projects, and saying out loud that the split exists, is usually enough to stop it setting.

3 exercises are prescribed for these patterns, starting with cross-cut projects.

Team-level reads use trait-appropriate operationalizations from Bell (2007), Barrick et al. (1998) and Peeters et al. (2006); faultlines follow Lau & Murnighan (1998); complementary-fit framing follows Kristof-Brown et al. (2005). Effects are modest and directional: a lens for judgement, not a verdict. Team personality is one input into how a group works: it informs, it doesn’t define.

Notice what the pair reads do and do not say. Each one names a channel, describes the pattern, and suggests what to agree. None of them scores a pair, ranks the pairs against each other, or uses the word incompatible. That restraint is deliberate, and section 02 is the reason for it.

Notice also how few pairs appear. A five-person team has ten possible pairs, and the map surfaces only the notable ones, friction first. An all-against-all grid would look more thorough and would mostly be noise, because most pairs in most teams have no strong pull in either direction.

The three surfaces

Colleagues differ in a hundred ways. Three of them cost you.

Most differences between two people never matter. These three reliably do, because teams hit them daily. Each has a pole you can recognise in someone within a week of working with them.

Surface one

Decision-making: who reaches for the wheel

Assertiveness decides who moves first when a call needs making. Two people who both reach get a recurring contest, and it is the most misread pattern of the three, because it looks like ego and is usually ambiguity. Neither of them is wrong that the decision matters. Nobody ever told them whose it was.

A gap on this channel is the good news. Someone who leads paired with someone who steadies is easier for both than two people competing to drive.

The rule the map uses: friction when both sit at or above the 65th percentile. Complementary when they are 25 or more points apart and one of them is above 60.

Surface two

Candour: how disagreement is supposed to feel

Some people say the frank thing at the cost of warmth. Others keep the warmth at the cost of saying the hard thing. Neither is the good one, and a team made entirely of either has a predictable problem: one politely agrees its way into a bad decision, the other keeps every decision honest and every meeting bruising.

The pairing that actually costs is two candid styles together, because there is nothing in the middle to absorb a hard moment. That is when a disagreement about the work becomes a disagreement about each other, and that slide is the part the conflict research is unambiguous about. Knowing your own pattern helps here: the free conflict styles assessment shows each of you what you do when a conversation stops being easy.

The rule the map uses: friction when both sit at or below the 35th percentile on cooperation. A gap on its own is not flagged, for reasons worth reading in section 06.

Surface three

Pace and structure: what finished means

Self-discipline and activity level set a person’s native relationship to deadlines and thoroughness. Put a planner beside an improviser and you have scheduled a recurring argument, in which each experiences the other as a character flaw rather than a setting.

This one responds to the most boring fix on the page, and it works almost every time: a written definition of done, with a named owner.

The rule the map uses: complementary when pace or self-discipline differ by 30 or more points. The team-level version of this gap is worth watching too, because spread on conscientiousness correlates negatively with performance[8].

A fourth surface, if you work remotely

Response time. Distributed teams turn an unstated expectation about replies into a personality judgement with remarkable speed: one person reads a four-hour gap as normal focus, the other reads it as being ignored. Nothing in the Big Five measures this, and no personality map will surface it. It is worth naming anyway, because it produces more day-to-day friction in remote pairs than anything a trait score covers, and it is fixed by one sentence about what counts as urgent.

The rules, in full

Here is the whole decision table our map applies to a pair. We publish it because a compatibility tool that will not show you its rules is asking for more trust than it has earned.

Signal Condition Read Channel, and why
Assertiveness Both at or above the 65th percentile Friction Decision-making. Dominance is a complementary signal, not a matching one. Two highs compete for the same wheel. [5]
Cooperation Both at or below the 35th percentile Friction Candour. The one pairing where a disagreement about the work reliably becomes a disagreement about each other. [10]
Assertiveness A gap of 25 points or more, with one of them at 60 or above Complementary Lead and support. People report complementary partners as easier to be around than mirrored ones. [5]
Activity level or self-discipline A gap of 30 points or more on either Complementary Pace and structure. Covers more ground than either tempo alone, once the difference is named rather than resented. [8]
All four signals Every gap inside 15 to 18 points Rapport Similar styles. Easy rapport, with one cost: two people who agree quickly test each other less. [2]

Rules are checked in that order, and a pair gets at most one read. Percentiles are against a general adult sample, so the 65th percentile means more assertive than about two thirds of adults, not more assertive than your colleagues.

Personality, or structure?

Most personality clashes are not about personality

Before you spend a week understanding a colleague better, spend ten minutes checking whether the problem is one of these instead. Each produces exactly the same feeling from the inside, and none of them is improved by a personality conversation.

Undefined decision rights. Two people who both believe this call is theirs will fight every time it comes up, however well matched they are. The tell: the same argument recurs with different content, and it ends when someone gets tired rather than when someone gets convinced.

An unstated definition of done. One of you thinks shipped means live, the other thinks it means reviewed. Neither has ever said so. The tell: the disagreement always arrives at the end of a piece of work, never at the start.

Mismatched cadence. One works to a weekly rhythm, the other to a quarterly one. The tell: you keep experiencing each other as either frantic or unresponsive, and both of you are right about the timescale you are using.

A reviewer with no authority. Someone is expected to have opinions but not to be able to block, and nobody said which. The tell: their feedback is either ignored or treated as a veto, and it alternates unpredictably.

This is not a technicality. A team-level meta-analysis of the three conflict types found that process conflict, the argument about who does what, had the strongest relationship with team performance of the three[11]. The thing that most looks like a personality problem is the thing most likely to be an org-design problem.

The ten-minute check

Write down the last three times it went badly. For each, finish this sentence: "what was genuinely unresolved was ..." If two of the three answers name a decision, a standard, or a deadline, fix that first. If two of them name how something was said, or how it landed, you have a style difference, and the rest of this page is about those.

There is a version of this that is neither. Some friction is simply load: too much work, too little slack, and a pair who would be fine in an ordinary quarter grinding through a bad one. Personality did not change. Capacity did. If the clash started when the workload did, treat the workload.

Five pairings

Five pairings you will recognise

Each one shows the two people on the channel that matters, the rule that flagged them, what it feels like from inside, and the agreement that helps. The percentiles are illustrative. The rules are the real ones.

Pairing one

Two hands on one wheel

Decision-making (assertiveness percentile): Dan and MarcoDecision-making (assertiveness percentile)11 points apartFollows the roomTakes the wheelDanMarcoFlagged because both sit at or above 65.

Dan and Marco are both natural drivers, and every call gets relitigated. From inside it feels like a personality clash. It is almost always a jurisdiction problem: two capable people who each believe, reasonably, that this decision is theirs. Neither is asked to become less decisive. [5]

What to do: Draw the boundary once, in writing, while nobody is annoyed. Strategy calls are Marco’s, channel and budget are Dan’s, and ties go to whoever owns the metric that quarter. Two sentences of jurisdiction, and the pair that relitigated everything now argues only where arguing helps.

Pairing two

Two candid styles, no shock absorber

Candour (cooperation percentile): Nadia and TomCandour (cooperation percentile)8 points apartSays the frank thingKeeps things smoothNadiaTomFlagged because both sit at or below 35.

This is the one pairing the research genuinely warns you about. Two people who both lead with directness have nothing in the middle to absorb a hard moment, so a disagreement about the work slides into a disagreement about each other. That slide is the part that costs, and the conflict meta-analyses point the same way. [10][11]

What to do: Not politeness. Agree in advance how a call between the two of you gets settled, and put objections on the idea with a fixable reason attached: "the hook is derivative, here is the tired part" rather than a verdict. The bluntness is not the problem. Unaimed bluntness is.

Pairing three

One drives, one steadies

Decision-making (assertiveness percentile): Ben and PriyaDecision-making (assertiveness percentile)37 points apartFollows the roomTakes the wheelBenPriyaFlagged as complementary: a 37-point gap, with one above 60.

Priya leads, Ben steadies, and this pairing usually works. In the lab, people who complemented a partner rather than mirroring them liked that partner more and were more comfortable. It is the closest thing to good chemistry that the evidence actually supports. [5]

What to do: Watch the quiet failure rather than the loud one. Ben stops saying the thing he can see, and the room reads his silence as agreement. Ask him first, by name, on anything that matters. A steadier who has stopped speaking is the most expensive person on a team.

Pairing four

Different tempos, same task

Pace and structure (self-discipline percentile): Ravi and JoPace and structure (self-discipline percentile)53 points apartTidies up laterPlans backwardsRaviJoFlagged as complementary: a 53-point gap, well past 30.

Ravi runs on momentum, Jo runs on structure, and each has quietly written a story about the other. Chaotic. Rigid. Both stories are wrong and both describe something real. The team-level version of this gap is worth watching too: spread on conscientiousness correlates negatively with performance, so this difference is one to manage rather than admire. [8]

What to do: One boring artefact fixes most of it: a written definition of done, with a named owner and a date. The friction was never about character. It was two unstated standards sharing one task.

Pairing five

The mirror pair

All four signals (assertiveness shown): Sam and AlexAll four signals (assertiveness shown)7 points apartFollows the roomTakes the wheelSamAlexFlagged as rapport: every channel inside 15 points.

Sam and Alex get on immediately, agree quickly, and enjoy working together. Nothing here needs fixing, which is exactly why it is worth naming. Two people in the same register reinforce each other’s instincts more readily than they test them, and neither of them can see the thing they both miss. [2]

What to do: Put someone unlike them in the room for decisions that matter, and give that person the first word. Their shared blind spot is the risk in this pairing, not each other.

Go deeper: the pattern the map deliberately does not flag

A large gap on candour, one very accommodating person and one very direct one, is the pattern people describe most often and the one our map stays quiet about. Here is why, and why it still matters.

Picture Priya at the 82nd percentile on cooperation and Dan at the 14th. In Monday’s review, Dan calls her concept derivative, because to him a blunt objection is a form of respect: you take an idea seriously by attacking it. Priya hears something personal, concedes to end the moment, and spends the afternoon rewriting a brief that did not need it. By Thursday, Dan has concluded she cannot defend her thinking, and Priya has concluded he is a bully. Neither story is true.

The map does not flag this, because warmth is a corresponding signal and the evidence for harm sits with two low scorers together, not with a gap. That is the honest reading of the research, and we would rather leave a pattern unflagged than invent a threshold for it. What the gap actually needs is one protocol: Priya’s concessions count only when she says "changing my mind", so yielding to end a moment stops reading as agreement. Notice that neither of them is asked to change. Dan is not asked to become warm, and Priya is not asked to grow armour. The protocol between the styles changes instead.

The agreement

The whole thing fits on one card

Every page on this subject ends at "understand each other better", which is not a thing anyone can do on Monday. This is the artefact instead: four lines, twenty minutes, both people present. Print it and fill it in together.

A working agreement, for two 20 minutes · both of you · one page · revisit in six weeks

1

The call that keeps coming back

Name one recurring decision and say who owns it. Not who is more senior. Who decides, when you two disagree. Add the tiebreaker for when it is genuinely unclear.

Decisions about ............................ are ............................’s. Ties go to ............................
2

What done means

Write the standard you have both been assuming and never said. One sentence. If you cannot agree on it in five minutes, you have found the actual problem.

A piece of work is finished when ............................................................
3

How we disagree

One rule each, aimed at your own channel. A direct person commits to naming a fixable reason rather than giving a verdict. An accommodating person commits to a phrase that means genuine agreement, so that yielding to end a moment stops counting as one.

I will ............................................................   You will ............................................................
4

The date we check

Six weeks out, in both calendars, now. This line is the difference between an agreement and a nice conversation, and it is the one people skip.

We look at this again on ............................
Two rules for the conversation itself. Bring an example each, from the last month, described as an event rather than a pattern: "the Tuesday review" rather than "you always". And if one of you disagrees with the read, that person is right, because the read is a hypothesis and they are the primary source on themselves.

What makes this work is what it does not ask for. Nobody has to become less direct or more assertive, and nothing here depends on either of you changing a trait, which is fortunate, because traits are stable and people resent being asked. You are changing the interface between two styles, and that is a much smaller thing to move. For the wider set of conversations that follow, managing different personalities has the scripts.

What the evidence says

Everything above, graded

Direction, the number worth remembering, and a grade for how much to trust it. A is meta-analytic or replicated for decades, B is strong but narrower, C is famous and contested. There is no C row on this table, because the C-grade material did not make the page.

Factor Direction Evidence The number to remember
Predicting one specific pair from traits Cannot A Two studies, 100+ pre-measured traits Machine learning predicted 4 to 18% of how much you like people and 7 to 27% of how much people like you, and none at all of the fit between two specific people [1]
Similarity, at a first meeting Helps A Meta-analysis, 460 effect sizes Actual similarity r = .47 with attraction, but not significant in existing relationships [2]
Deep-level similarity at work Helps A Meta-analysis, 82 samples r = .46 with getting on with your manager, .37 with job satisfaction, .32 with job performance [3]
Surface-level similarity Does nothing A Same meta-analysis Not significant for job performance, job satisfaction, manager relationship, citizenship or commitment [3]
Complementary dominance Helps B Lab studies, small samples People who complemented a partner liked them more than people who mirrored them [5]
Team floor on agreeableness Helps A Meta-analysis In field teams, the least agreeable member predicted performance where the average did not [7]
Spread in conscientiousness Hurts A Meta-analysis ρ = −.24. Mixing very diligent with very lax beats neither [8]
Team personality effects overall Weak A Most recent meta-analysis, 2024 Team personality traits were only weakly related to team performance, with moderators doing much of the work [9]
Task conflict, the supposedly good kind Hurts A Meta-analysis, direction only Negative for both performance and satisfaction, against what the textbooks say. Worse on complex work [10]
Task conflict under psychological safety Helps B 117 project teams The benefits of task conflict showed up only where psychological safety was high [12]
Arguing about who does what Hurts most A Meta-analysis, 89 studies Process conflict had the strongest relationship with team performance of the three conflict types [11]
Psychological safety Helps A Meta-analysis, 136 samples Over 22,000 people and around 5,000 groups; strongest links to engagement, satisfaction and commitment [13]
Having worked together before Helps B Large archival field study Team familiarity predicted performance, while years of individual experience did not [16]
Fitting the group you are in Helps attitudes A Meta-analysis, 172 studies Moderate links to job satisfaction and to staying, but only a modest link to job performance [17]
Team diversity, all kinds Helps a little A Registered report, 2,638 effect sizes Modest positive correlations, larger on complex or creative work. The authors caution against overselling it [18]
Dormant faultlines Depends A Meta-analysis, 24,953 teams Contrary to widespread belief, not directly related to performance or satisfaction. Only activated ones are [14]
Type labels as a basis for matching Unstable A The publisher’s own manual Only 65% of people get the same four-letter MBTI type when retested four weeks later [19]

Composition sets the odds. Conditions decide the game.

Read the table as a whole and a shape appears. The composition effects are real and small. The condition effects are real and larger: psychological safety, having worked together before, whether disagreement stays aimed at the work. Even the one row that rescues task conflict is a condition rather than a trait, because the benefits appeared only where safety was high[12].

So a compatibility read is a map of where the listening will be hardest. It is not a prediction of whether you will listen. That second part is not measured by any instrument on this page, and it is the part that decides how the pair actually goes.

Numbers we deliberately did not use

This subject runs on statistics that do not survive contact with their sources. The faultline meta-analysis everyone quotes was retracted. Thatcher and Patel (2011) was withdrawn in 2016 for errors that "may affect the overall conclusions"[15], and its figures still circulate, so our faultline claim cites the 2024 replacement, which reverses the conclusion[14]. "93% of communication is nonverbal" is a misreading of two 1967 studies that used single words and about 30 participants each, and whose author spent years asking people to stop. "86% of workplace failures come from poor communication" traces to a 2011 vendor survey of conference attendees, now widely misattributed to Salesforce. "70% of teams fail" has no traceable source at all. And "89 of the Fortune 100 use the MBTI" is publisher marketing, not evidence about whether it works.

One more, about our own table. The conflict meta-analysis in row nine is usually quoted with precise coefficients. We could not verify those against the paper’s own results tables, so this page reports the direction and the moderators, which are stated in the abstract, and leaves the decimals out.

What it must never decide

Change the interface between two people, never whether they get one

This part matters enough to be blunt. A compatibility read describes patterns between working styles. It is support for conversations, structure and development. It is never grounds for removing someone from a project, declining to staff a pair, or quietly writing a person off. The moment a personality read shapes who gets the work, it has become an employment decision made on an instrument that was never designed to carry one, and that is a different activity with real legal weight.

Three reasons beyond decency. The effects are modest, and the most recent meta-analysis on team personality puts them lower than the older ones did[9]. The readings are pair-specific, so the colleague who grinds against one teammate is often the ideal counterweight for another, which means a friction flag says almost nothing about a person’s value to the team. And managed friction is frequently productive: the pairs who argue well are often the ones who catch what everyone else missed.

There is a quieter misuse worth naming too, because it is the common one. A manager reads a pair flag, says nothing, and simply stops putting those two on the same work. Nobody has been told anything, no decision has been recorded, and a person’s range has narrowed for a reason they will never learn. That is the failure mode this category should be judged on.

If what you actually need is to evaluate candidates for a role, that is a different job with different rules, and it wants a structured, job-related process built for selection rather than a team mirror. Our hiring assessment is built for that, and deliberately kept separate from this.

The line, in one sentence

Use a compatibility read to change how two people work together. Never to change whether they get to.

Choosing a framework

Which frameworks can actually read a pair?

Most personality frameworks describe individuals and were retrofitted to pairs afterwards. Here is what each one does at the pair level, and the honest caution that goes with it.

Framework What it does with a pair The honest caution
Big Five traits Continuous percentiles per person, compared channel by channel The model nearly all the team research is written in, so published findings actually apply. It is what the free tools on this page use.
DISC Four styles, with popular pairing guides A fast shared vocabulary a pair can learn in an afternoon. Almost no published evidence links DISC pairings to outcomes, so treat it as language rather than measurement.
MBTI type matching Sixteen types, with published compatibility charts The charts rest on the type being stable, and the publisher’s own manual reports 65% agreement on a four-week retest. A synthesis of 25 years of research found no test-retest or structural-validity studies at all.
FIRO Genuinely pair-level, across inclusion, control and openness The oldest properly dyadic model here, and still the best idea in the category. Modern validation is thin, and it is mostly delivered through consultants.
Enneagram Nine types, with detailed pairing lore Rich, engaging and popular in coaching. The pairing claims are not derived from outcome data, so keep it for conversation rather than decisions.

These combine better than they compete. Plenty of teams measure with traits and still enjoy type or style vocabulary as shared slang, which is fine as long as everyone remembers which one is the measurement. The deeper comparisons live in Belbin team roles vs the Big Five and the DISC assessment for team building guide.

One thing to watch across all of them. A published pairing chart is only as stable as the labels underneath it, and type labels move: 65% of people get the same four-letter MBTI type when they retake it four weeks later, on the publisher’s own figures[19]. A systematic review of 25 years of research found no test-retest or structural-validity studies in the published literature at all[20]. An agreement built on a label that moves will move with it.

Questions

Team compatibility and team chemistry: FAQ

What is a team compatibility test?

A team compatibility test measures each person on a personality model and then reads how those profiles interact: who shares a working rhythm, which pairs complement each other, and where friction is plausible. A good one reports patterns on named channels such as decision-making, candour, and pace. It never issues a verdict on a pair, because no honest instrument can. The useful output is a short list of things to agree on, not a score.

Do compatibility tests actually work?

Partly, and it is worth being precise about which part. Traits predict how a person tends to land on people in general, and how people tend to land on them. They do not predict the unique fit between two specific people. The cleanest test of this used more than 100 self-report measures taken before two people met, then machine learning: it predicted 4 to 18% of how much someone liked others, 7 to 27% of how much others liked them, and none of the relationship variance at all. So treat any compatibility score as a description of two individuals placed side by side, never as a prediction about the pair.

What is team dynamics?

Team dynamics is the set of patterns in how a group actually works together: who speaks and who defers, how decisions get made and by whom, how disagreement is voiced, what happens when a deadline slips. Personality is one input into those patterns and not the largest one. Structure matters too, and so does psychological safety, which is the best-evidenced team-level factor there is. A compatibility read describes the personality layer of your team dynamics, which is a real layer and a partial one.

Can two people simply be incompatible?

Not in the way the word implies. Some pairings carry predictable friction: two natural drivers contest decisions, and two very candid people can let a disagreement turn personal. Those are the patterns worth naming early. But they respond well to explicit working agreements, and knowing the pattern is most of the fix. What the evidence does not support is a fixed sentence on a pair. It supports the opposite, since traits could not predict pair fit at all in the best test of it.

What personality types work best together?

The honest answer is that it depends on the channel, and that type language is the wrong tool for the question. On dominance, difference helps: someone who leads paired with someone who steadies is easier for both than two people reaching for the wheel. On candour, two people who both lead with directness are the one pairing the conflict research warns about. On pace, a gap is workable and needs an agreement rather than a fix. Beyond that, published type-pairing charts rest on types being stable, and only 65% of people get the same four-letter MBTI type four weeks later.

How do you improve team dynamics?

Start with the structural things, because they are faster and they are more often the actual cause: who owns which decision, what finished means, and how disagreement is supposed to be voiced. Then work on psychological safety, which has the strongest evidence of any team-level factor and is mostly built by how leaders respond to bad news. Personality work sits after both. It gives a team language for differences that were already there, which makes those differences discussable instead of personal.

Can I use this if my colleague will not take a test?

Yes, and that is the most common situation. The pair read on this page asks two questions about you and four about behaviour you have already watched, so it works when the other person is unavailable, uninterested, or your manager. It is a hypothesis rather than a measurement and it says so. If you can get everyone to answer for themselves, the team map at the top of the page measures it properly instead.

How does team chemistry affect performance?

Less than the phrase suggests, and not in the direction most people assume. The most recent meta-analysis found team personality traits only weakly related to team performance. What does show up reliably is narrower: teams score on some traits by their weakest member rather than their average, spread in conscientiousness correlates negatively with performance, and having worked together before predicts performance where individual years of experience does not. Composition sets the odds. How people treat each other decides the game.

Do teammates see each other’s answers?

No. Each person keeps their own report, and nobody’s individual answers are shown to anyone else. The team view shows patterns: the style map, pair reads, shared strengths and watch-outs. Pair-level information is more sensitive than team-level information, so the rule is stricter here than it looks: a pair read belongs to the two people in it, and should be read with both of them present or not at all.

How many people do I need?

Two is enough for pair reads. From three the map adds the interpersonal style plot, and from four it can detect subgroup faultlines, where several differences line up into camps. Most teams using it are between three and fifteen people. Past about twenty, map pods separately, because a large map becomes a poster rather than a mirror.

Can I use a compatibility read to decide who works on what?

No. Use it to change the interface between two people, never to decide whether they get to work together. Three reasons beyond decency: the effects are modest, the readings are pair-specific, so someone who grinds against one colleague is often the ideal counterweight for another, and managed friction is frequently productive. The moment a compatibility read quietly shapes staffing, it has become an employment decision made on a personality instrument, which is a different activity with real legal weight.

What if the clash is with my manager?

Then the structural questions matter more, not less, and you have less room to negotiate them. Start by separating the decision-rights question from the style question, because raising a style difference with a manager who has not agreed the decision rights tends to go badly. Ask for the boundary first, in ordinary language: which of these calls is mine to make, and which are yours. If the friction survives a clear answer, it is worth reading as a style difference. And a pair read about your manager is for you, not for circulation.

Keep reading

The rest of the team library

That is everything we know about reading a pair honestly. When you are ready, the free map at the top of the page takes a team name and gives you a link to share. Eight minutes a person, and the pair reads write themselves.

Methods & sources

Who wrote this, and what the tools are based on

Michael Hodge

Founder, SeeMyPersonality · author and reviewer of this page

  • Bachelor of Science (Psychology), University of Wollongong, with coursework in psychometrics and research methods.
  • Designs and reviews the questionnaires on this site: item quality, scoring design, norm referencing, and interpretation frameworks.
  • Wrote the pair-read rules published in section 04 against the studies cited below, and this page against those rules.

How content is written and reviewed on SeeMyPersonality →

What the tools on this page are based on

The individual instrument is a 60-item Big Five inventory in the tradition of the BFI-2, the peer-reviewed modern standard, whose domain scales show strong reliability (alpha .83 to .91) and 8-week retest stability (.76 to .84)[21]. Scores are computed against adult norms, in your browser. The pair layer then compares members on four facet-level signals: assertiveness and activity level from extraversion, cooperation from agreeableness, and self-discipline from conscientiousness. Dominance is treated as a complementary signal and warmth as a corresponding one, following the interpersonal research[5], and the thresholds are the ones published in section 04.

The pair read in section 01 is a different thing and should be held more loosely. It applies the same rule order to six self-reported answers rather than to two measured profiles, so it is a structured hunch. It runs entirely in your browser, stores nothing, and sends no data anywhere.

Limits, stated plainly. Team and pair composition effects are modest correlations, reported here as decision support rather than destiny, and the most recent meta-analysis grades them lower than the older ones did[9]. The unique fit between two specific people has not been successfully predicted from traits by anyone, including us[1]. Personality is one input. It informs how two people are likely to run together. It never determines what either of them can do. Where the evidence is contested, the grade column says so rather than rounding up.

Reviewed by: Michael Hodge Content last reviewed: 16 August 2026 Disclosure: the pair read and the team map on this page are built and run by SeeMyPersonality. There is nothing for sale on this page.
  1. 1. Joel, S., Eastwick, P. W., & Finkel, E. J. (2017). Is romantic desire predictable? Machine learning applied to initial romantic attraction. Psychological Science, 28(10), 1478-1489.
  2. 2. Montoya, R. M., Horton, R. S., & Kirchner, J. (2008). Is actual similarity necessary for attraction? A meta-analysis of actual and perceived similarity. Journal of Social and Personal Relationships, 25(6), 889-922.
  3. 3. Condrea, S., & Iliescu, D. (2026). Similarity at work: a meta-analysis of deep-level and surface-level similarity. Current Psychology.
  4. 4. Harrison, D. A., Price, K. H., & Bell, M. P. (1998). Beyond relational demography: time and the effects of surface- and deep-level diversity on work group cohesion. Academy of Management Journal, 41(1), 96-107.
  5. 5. Tiedens, L. Z., & Fragale, A. R. (2003). Power moves: complementarity in dominant and submissive nonverbal behavior. Journal of Personality and Social Psychology, 84(3), 558-568.
  6. 6. Grant, A. M., Gino, F., & Hofmann, D. A. (2011). Reversing the extraverted leadership advantage: the role of employee proactivity. Academy of Management Journal, 54(3), 528-550.
  7. 7. Bell, S. T. (2007). Deep-level composition variables as predictors of team performance: a meta-analysis. Journal of Applied Psychology, 92(3), 595-615.
  8. 8. Peeters, M. A. G., van Tuijl, H. F. J. M., Rutte, C. G., & Reymen, I. M. M. J. (2006). Personality and team performance: a meta-analysis. European Journal of Personality, 20(5), 377-396.
  9. 9. Han, A., Krieger, F., Kim, S., Nixon, N., & Greiff, S. (2024). Team personality and team performance: a meta-analytic revisit. Journal of Research in Personality, 112, 104526.
  10. 10. De Dreu, C. K. W., & Weingart, L. R. (2003). Task versus relationship conflict, team performance, and team member satisfaction: a meta-analysis. Journal of Applied Psychology, 88(4), 741-749.
  11. 11. O’Neill, T. A., Allen, N. J., & Hastings, S. E. (2013). Examining the "pros" and "cons" of team conflict: a team-level meta-analysis of task, relationship, and process conflict. Human Performance, 26(3), 236-260.
  12. 12. Bradley, B. H., Postlethwaite, B. E., Klotz, A. C., Hamdani, M. R., & Brown, K. G. (2012). Reaping the benefits of task conflict in teams: the critical role of team psychological safety climate. Journal of Applied Psychology, 97(1), 151-158.
  13. 13. Frazier, M. L., Fainshmidt, S., Klinger, R. L., Pezeshkan, A., & Vracheva, V. (2017). Psychological safety: a meta-analytic review and extension. Personnel Psychology, 70(1), 113-165.
  14. 14. Thatcher, S. M. B., Meyer, B., Kim, Y., & Patel, P. C. (2024). Team faultlines: a meta-analytic review and integration. Organizational Psychology Review, 14(2).
  15. 15. Retraction of Thatcher & Patel (2011), "Demographic faultlines: a meta-analysis of the literature." Journal of Applied Psychology, 101(8), 1150 (2016).
  16. 16. Huckman, R. S., Staats, B. R., & Upton, D. M. (2009). Team familiarity, role experience, and performance: evidence from Indian software services. Management Science, 55(1), 85-100.
  17. 17. Kristof-Brown, A. L., Zimmerman, R. D., & Johnson, E. C. (2005). Consequences of individuals’ fit at work: a meta-analysis of person-job, person-organization, person-group, and person-supervisor fit. Personnel Psychology, 58(2), 281-342.
  18. 18. Wallrich, L., Opara, V., Wesołowska, M., Barnoth, D., & Yousefi, S. (2024). The relationship between team diversity and team performance: reconciling promise and reality through a comprehensive meta-analysis registered report. Journal of Business and Psychology, 39, 1303-1354.
  19. 19. Schaubhut, N. A., & Herk, N. A. MBTI Form M Manual Supplement. The Myers-Briggs Company.
  20. 20. Erford, B. T., et al. (2025). A systematic review of Myers-Briggs Type Indicator psychometric research, 1999-2024. Journal of Counseling & Development.
  21. 21. Soto, C. J., & John, O. P. (2017). The next Big Five Inventory (BFI-2): developing and assessing a hierarchical model with 15 facets. Journal of Personality and Social Psychology, 113(1), 117-143.