Personality tests for team building: which one, and what to do with the results

A personality test for team building measures every member on the same instrument, then combines the results into one reading of the group: which working styles are covered, where the team runs thin, and which patterns will show up under pressure. The test is the easy part. What you run afterwards is the part that changes anything.

What are you assessing?
Team name

Free · No signup · Live link in about 10 seconds

60 questions each, about 8 minutes Everyone keeps their own result Nothing to install, nothing to buy

On this page
  1. Which one to run
  2. The eight exercises
  3. The 90-minute session
  4. Choosing a test
  5. Reading the results
  6. A team, mapped
  7. What actually works
  8. When it goes wrong
  9. Ground rules
  10. The other Big Five
  11. Questions
  12. Methods & sources

Which one to run

Start here: which exercise does your team actually need?

Most team building fails at the first step, by prescribing the same activity to every team. Seven questions about behaviour you have already watched, and you get two exercises worth your Tuesday instead of eight worth reading about.

Which two should we run?Seven questions about behaviour you can already observe. Nothing to sign up for, nothing stored.
1 / 7

In your team’s meetings, who does the talking?

You can run any of these on hunch alone, without measuring anything, and it will still beat the escape room. But hunches about personality travel through the lens of your own. A lead high in extraversion reliably underestimates how loud the room already is, and the quietest person on a team is usually invisible to exactly the person who most needs to notice them. Eight minutes a head is a small price for aim.

The eight exercises

Eight exercises, each keyed to a pattern rather than a mood

Every one names the pattern it addresses, what that pattern looks like from inside the room, how long it takes, what you need, and how you will know it worked. Pick the two your team needs.

Exercise You will recognise it if Time Fixes
1Quiet-first rounds The same two or three people speak first in every meeting, and you could name the person who has not finished a sentence in a month. 25 to 45 minutes Airtime
2The pre-mortem People are not arguing against the change out loud. They are just very quiet about it, and the questions arrive one at a time, late, in private messages. 60 minutes Change resistance
3The disagreement contract Disagreements do not settle. They either turn personal or go silent, and afterwards two people are careful with each other for a fortnight. 75 minutes Conflict that lands badly
4The reliability handshake Slips keep happening at the same handoff, and the same person keeps catching them on a Sunday. 60 minutes to set up Delivery
5Stress signatures One hard week moves everybody's mood, and you can tell what kind of week it is from the tone of the first message each morning. 45 minutes once Pressure
6Cross-cut projects You can predict who will side with whom before the meeting starts, and you have been right for months. A 30-minute kickoff A split
7The rotating spokesperson When one specific person is on holiday, demos get postponed. 30 minutes of setup A single point of failure
8The role handoff One person is the answer to too many questions, and they have started apologising for response times. 60 minutes Overload

The eight exercises in full Steps, scripts, materials and a remote variant for each

1

Quiet-first rounds

When a team mixes people who think out loud with people who think before they speak, ordinary discussion becomes a quiet tax on the second group. The fix is not asking quiet people to speak up. It is changing the format so the first word is written, not spoken.

You will recognise it if The same two or three people speak first in every meeting, and you could name the person who has not finished a sentence in a month.
In the results A wide spread on extraversion: members sitting at both the outgoing and the reserved end.
Time 25 to 45 minutes, inside a meeting you already hold
Group size 4 to 12 people
You will need The question sent a day ahead, paper or a shared doc, a timer
  1. Post the question in the invite, a day ahead. Reserved members arrive having already thought, which is their home advantage.
  2. Open with four minutes of silent writing. Everyone drafts an answer before anyone hears anyone else.
  3. Go around once: ninety seconds each, no interruptions, no reactions until the round is complete.
  4. Give the most talkative member a real job that is not talking. Timekeeper and scribe, framed honestly as the hardest role in the room.
  5. Only then open free discussion, working from the written material rather than from whoever spoke last.
"We are writing first today. Four minutes, then one round, then we argue. I am scribing, so you will not hear much from me until the round is done."

What good looks like: Within three or four meetings, ideas from the quietest member start surviving into decisions, and the loudest member can name what the team gained by waiting. If airtime still concentrates, shorten the free-discussion phase rather than the round.

Running it with a remote team

This one is better remote than in person. Use the chat for the silent writing phase and have everyone post at the same moment, so nobody anchors on the first answer. Run the round in a fixed order posted in advance, because the video-call habit of waiting for a gap hands the floor to whoever tolerates interruption best. Keep cameras optional; the round already guarantees presence.

2

The pre-mortem

Teams that lean practical rather than novelty-seeking are often told to be more open to change, which is advice nobody has ever successfully followed. A pre-mortem does something smarter. It gives scepticism a sanctioned outlet and turns it into a risk register. There is a real caveat, and it is worth knowing: the technique itself has never been trialled. What has been tested is the ingredient underneath it. Asking people to explain an outcome that has already happened, rather than one that might, raised the number of correctly identified reasons by about 30% in a laboratory study of individuals[19].

You will recognise it if People are not arguing against the change out loud. They are just very quiet about it, and the questions arrive one at a time, late, in private messages.
In the results A lower team average on openness, with a real change on the calendar.
Time 60 minutes, once per major change
Group size 3 to 10 people
You will need Sticky notes, a wall, dot stickers, one owner per risk
  1. Frame it in one sentence, then stop talking. The technique comes from Gary Klein<a class="thb-cite" href="#ref-20">[20]</a>.
  2. Eight minutes of silent writing. Each person drafts the failure story alone, as specifically as they can.
  3. Collect causes verbatim on a wall, one per note. No debating while collecting.
  4. Cluster the notes, then dot-vote the three most plausible failure paths.
  5. For each of the three, assign one owner and one tripwire metric that would say the failure has started.
  6. Close by asking what the plan should change this week. A pre-mortem that changes nothing was a séance, not a meeting.
"It is twelve months from now. The change failed completely, and we are writing the story of why. Eight minutes, on your own, be specific."

What good looks like: The members most doubtful about the change become the authors of its risk register, which turns resistance into vigilance. You know it worked when a tripwire fires months later and nobody is surprised, because the team already wrote that chapter.

Running it with a remote team

Use a shared board and have everyone type in silence with the board hidden, or the first person to post sets the frame for everyone else. Reveal all notes at once. Dot-voting works natively in most whiteboard tools. Keep the eight minutes strict; remote silence feels longer than it is and people start filling it.

3

The disagreement contract

A candid member is an asset most teams mislabel as a problem. The largest meta-analysis of intragroup conflict, covering 116 studies and 8,880 groups, separates the two kinds cleanly. Arguing about the work is roughly neutral for performance and slightly positive in senior teams. Arguing about each other is not[11]. The point of a contract is to keep the first kind and starve the second, without asking anyone to pretend to be softer than they are. It helps if each person knows their own pattern first: the free conflict style test gives everyone that read in eight minutes.

You will recognise it if Disagreements do not settle. They either turn personal or go silent, and afterwards two people are careful with each other for a fortnight.
In the results A low floor on agreeableness: at least one member sitting well below the rest.
Time 75 minutes, then 15 minutes again after 30 days
Group size The whole team, ideally 4 to 12
You will need Something to write five rules on that stays visible afterwards
  1. Read the distinction aloud: arguing about the work, versus arguing about each other. Agree that the first is welcome here.
  2. Each person writes two sentences. What I sound like when I disagree, and what I need from others when they disagree with me. Share in pairs, then with the room.
  3. Draft five house rules together. Strong candidates:
    • Critique the work in the room, not in messages afterwards.
    • Steelman the other position before countering it.
    • Heated threads move to a call within 24 hours.
    • Every debate ends with an explicit "decided, and here is who decides next time".
  4. Write the five down and have everyone sign, literally. The mild ceremony matters, because it turns a mood into a norm.
  5. Diarise a 30-day review. Which rule earned its place, which one nobody used, what needs adding.
"Nobody here is being asked to be nicer. We are agreeing how to argue, so the arguing stays about the work."

What good looks like: Directness keeps flowing but stops leaving bruises. The blunt member hears, often for the first time, that their candour is wanted. The warmer members get a protocol that makes disagreement feel survivable.

Running it with a remote team

Written disagreement is where distributed teams actually struggle, so weight the rules toward text. Two that earn their place remotely: no disagreement lands in a public channel without a named next step, and any thread over six replies moves to a call within the day. Pin the signed rules in the channel where the arguing happens, not in a wiki nobody opens.

4

The reliability handshake

Teams do not miss deadlines on average. They miss them one handoff at a time, and usually at the same handoff. In the composition research the mix matters more than the mean here: variability in conscientiousness sits at about ρ = −.24 with team performance, a larger cost than the benefit of a high average[7]. The dignified reading is that improvisers are your best firefighters, and this structure exists so the most organised person stops being the team's unpaid safety net.

You will recognise it if Slips keep happening at the same handoff, and the same person keeps catching them on a Sunday.
In the results A low floor on conscientiousness, or a wide spread on it. Both cost, and the spread costs more.
Time 60 minutes to set up, 10 minutes weekly after that
Group size 3 to 10 people
You will need A whiteboard, the last three slipped deliverables, one checklist per recurring handoff
  1. Map the recurring handoffs on a whiteboard. Who hands what to whom, and where the last three slips happened.
  2. For the three most common deliverables, write a "done means" checklist of five lines or fewer. Tested, documented, next person notified, or whatever your work's honest version is.
  3. Give every deliverable exactly one owner. Shared ownership is how a low floor hides.
  4. Set a visible work-in-progress limit for whoever is juggling most. Fewer plates, fewer drops.
  5. Add a ten-minute Friday check: what slipped, what caught it, what we change. No blame. The interesting question is always what caught it.
"We are not fixing anyone's attitude. We are fixing one handoff, because that is where the last three slips happened."

What good looks like: Slips stop clustering at one desk, and the checklist absorbs the vigilance work one person used to do from memory. Re-measure after a quarter. A rising floor here is one of the few numbers that reliably follows a process change.

Running it with a remote team

Distributed teams already write things down, so the win is smaller but easier. Put the "done means" checklist in the ticket template rather than a document, so it is impossible to skip. Replace the Friday meeting with a written thread if your team spans more than four hours of time zones; the ritual matters more than the format.

5

Stress signatures

A team that runs warm under pressure usually has excellent radar. It notices risk early and it cares about outcomes. The cost is that pressure travels through it fast. Swapping stress signatures turns private weather into shared information, which is most of what a team can actually do about stress.

You will recognise it if One hard week moves everybody's mood, and you can tell what kind of week it is from the tone of the first message each morning.
In the results A warmer-running team average on emotional stability: the group feels pressure keenly.
Time 45 minutes once, then five minutes whenever crunch starts
Group size 3 to 10 people
You will need Three written prompts each, and somewhere permanent to put the result
  1. Each person completes three prompts in writing. My early warning sign is. What helps me in the moment is. What reliably makes it worse is.
  2. Share in pairs first, then read the three lines to the room. No commentary, no fixing. The exercise is the telling.
  3. Compile everyone's lines into a one-page pressure playbook, and put it where the team actually looks. Next to the on-call rota, not in a slide deck.
  4. Agree one collective release valve: a daily stand-down time, a no-messages window, whatever fits the work.
  5. When a crunch begins, open with the playbook rather than the plan.
"Watch for each other's signals this week. Here is the release valve, and I am going first: mine is that I go quiet and start rewriting things that were already fine."

What good looks like: Pressure gets spotted a day earlier and named without drama, because "I am seeing my warning sign" is easier to say than "I am struggling". Sensitivity starts operating as the early-warning system it always was.

Running it with a remote team

Remote teams lose the ambient signal entirely, which makes this the highest-value exercise on the list for them. Add a fourth prompt: "what my early warning sign looks like in writing", because tone is the only channel left. Some teams add a one-word status in the team channel each morning. Keep it optional, and let people skip it without being asked why.

6

Cross-cut projects

Faultlines are dividing lines that stay dormant until stress arrives, and their defining feature is alignment: the same few people together on axis after axis[13]. One thing worth knowing before you panic about yours. The 2024 meta-analysis, covering 168 studies and 24,953 teams, found that dormant faultlines are not directly related to team performance or satisfaction. Only activated ones are[14]. So this is preventive work, and it is cheap.

You will recognise it if You can predict who will side with whom before the meeting starts, and you have been right for months.
In the results An alignment flag: the same members grouping together on trait after trait.
Time A 30-minute kickoff, then two to six weeks of ordinary work
Group size Teams of 4 or more, since a split needs two people a side
You will need Two real deliverables. Not an invented exercise.
  1. Show the team the pattern and name it in composition language, not blame language. "We cluster into two styles, and the split runs along social energy."
  2. For the next two real deliverables, form pairs or trios drawn deliberately from both sides. Real work only. A mixed committee for an invented task fools nobody.
  3. Rotate who presents the joint output, so credit and visibility cross the line along with the work.
  4. Keep it up for at least two delivery cycles. Familiarity accumulates through repetition, not through one afternoon.
  5. At 30 days, ask whether cross-line collaboration is happening yet without being scheduled.
"There is a line in how we group up, and it has not cost us anything yet. I would like to keep it that way, so the next two projects are going to cross it."

What good looks like: The tell is in the language. When "the loud half" and "the careful half" become names for tendencies rather than teams, the line has stopped organising the group. New joiners are your canary: they should not be able to guess the old subgroups.

Running it with a remote team

Distributed teams get faultlines for free along time zones and office locations, and those lines are usually aligned with everything else. Cross them deliberately: pair across locations on the real work, and rotate meeting times so the same group is not always the one joining at an awkward hour. A rota that always inconveniences the same half is a faultline generator.

7

The rotating spokesperson

When exactly one member anchors the team's social energy, every demo, pitch and difficult call routes through them. That is efficient right up until they are away. This builds a second and third voice, with scaffolding that respects the fact that fronting a room costs reserved people more.

You will recognise it if When one specific person is on holiday, demos get postponed.
In the results A coverage gap on the outward-facing role: one person holds it, or nobody does.
Time 30 minutes of setup, then every client-facing meeting
Group size Any team with external-facing work
You will need A one-page meeting frame and a two-minute debrief form
  1. Build the scaffold once, together. An opening line, an agenda, the three questions you always ask, and the close.
  2. Rotate the spokesperson seat through willing members, one meeting at a time. The natural anchor goes last.
  3. Recast the anchor as coach. Ten minutes of prep beforehand, ten of debrief after, and notes rather than the lead in the room.
  4. Let reserved members choose their format. Some will front a live demo, others will own the written follow-up that wins the deal. Both are client-facing work.
  5. After a full rotation, ask which meetings felt strongest, and staff future ones by fit now that fit is something you have observed.
"Nobody is being turned into an extravert. We are making sure this does not stop when one person is on leave."

What good looks like: At least two people can credibly front the work, the anchor stops being a single point of failure, and someone usually surprises you. The goal is coverage, not conversion.

Running it with a remote team

Remote makes this easier, because the scaffold can sit on the second screen where nobody can see it. Let the spokesperson keep the frame open and the anchor stay on mute in the same call as a safety net, with an agreed signal for handing back. Record the calls if your clients allow it; watching yourself once is worth more than three debriefs.

8

The role handoff

Balanced role coverage was Belbin's core insight, and it fails in a predictable way. One versatile person quietly ends up owning three roles, does them all at seventy percent, and burns out politely. Worth saying plainly, since this page is not selling you a role instrument: Belbin's idea survives scrutiny better than his questionnaire does, whose forced-choice format makes its factor structure hard to evaluate at all[28]. Use roles as a way to talk about load, not as a measurement.

You will recognise it if One person is the answer to too many questions, and they have started apologising for response times.
In the results A thin bench: one member is the best fit for three or more of the team roles.
Time 60 minutes, then a four-week apprenticeship
Group size 3 to 12 people
You will need The role list on a screen, and one named successor
  1. Put the role list on the screen and let the team react before you do. The overloaded member usually names themselves, with relief.
  2. They pick one role to hand off, their least energising of the three. Handing off a role you secretly love does not stick.
  3. Identify the next-best-fit member, and read the evidence honestly. Some trait-to-role links are well supported and some are merely plausible.
  4. Run a four-week apprenticeship. The new owner does the role, the old owner is consultable twice a week, and the new owner has explicit permission to do it differently.
  5. Review at four weeks against one question: is the role happening without the original owner checking?
"You are doing three of these. I would like you to keep the one you enjoy and give away the one you do at eleven at night."

What good looks like: The versatile member gets ten hours of their month back, someone else gets a growth edge that fits them, and the team learns that roles are assignments rather than identities. If the handoff fails, you have learned something real about whether the honest answer is a next hire.

Running it with a remote team

Overload hides better remotely, because nobody sees the hours. Make the handoff visible in whatever tool holds the work, so the new owner is publicly the owner. Twice-weekly consulting slots should be booked, not offered, or the old owner will simply keep doing it faster than they can explain it.

If your team would rather work in four styles than five traits, the DISC team building activities guide is the same idea in that vocabulary, with twelve activities of its own. Most of what is below transfers between the two with only the labels changed.

The 90-minute session

The session that makes one exercise stick

Results, then one exercise, then a commitment. This is the shape we would run, and the debrief at the end is the part with the best evidence behind it, so protect it when you run late.

Team building session 90 minutes · everyone has taken the test · one exercise chosen in advance

0:00 – 0:08

Set the frame

Name the purpose and the limits before anyone looks at anything. This is the sentence that decides whether the next eighty minutes feel safe.

"This describes how we work, not how good we are. There is no best result here, nothing touches reviews, and each of us is the final authority on our own."

0:08 – 0:20

Own results first

Each person names one thing their own result got right and one thing it overstates. Starting with disagreement establishes that the instrument is fallible and that people outrank charts. A minute each, briskly.

"What did yours nail, and what did it get a bit wrong?"

0:20 – 0:35

The team reading, together

Show the group-level view and let people find themselves before you interpret anything. Name patterns as facts about the roster, never as facts about a person in the room, even when everyone can do the arithmetic.

"Where did this put you, and does your neighbour agree?"

0:35 – 1:15

Run the exercise

One exercise, chosen in advance from the eight above, run properly rather than two run hurriedly. Say which pattern it is aimed at before you start, so the room knows this was a diagnosis and not a party game.

1:15 – 1:27

Debrief it

The highest-value twelve minutes of the session. Three questions: what happened, what surprised us, and what we will do differently. A structured debrief of about this length came out roughly 25% better than no debrief across 46 samples[4], which makes it the best-evidenced part of the whole day.

"What did we just learn that we would not have learned by talking about it?"

1:27 – 1:30

One change, one owner, one date

Goal setting and role clarification are the two components of team building that reliably move outcomes[2], and this is where both happen. Two kept changes beat five aspirations.

If someone's result feels wrong to them, agree with them out loud. If one voice dominates, switch to rounds. If the reading shows a low floor, discuss the norm it suggests and never the name it implies. And book the review date before the meeting ends, because that date is the whole difference between a workshop and a decision. For a shorter version that skips the exercise, the 45-minute debrief agenda on the hub does results and norms only.

Choosing a test

Which personality test to use for team building

Five families you will meet, described by what they are actually for. Any of them beats nothing. They are not interchangeable.

Family What it measures Use it when One honest caution
Big Five Five continuous traits, scored against adult norms Use it when you want the results to mean something. Nearly all published team-composition research is written in these five traits, so findings actually transfer. The assessment on this page uses it. Trait language is less memorable than type language, and some teams find "moderately high on conscientiousness" harder to rally around than a four-letter code.
DISC Four working styles Use it when the goal is a shared vocabulary for friction, fast. A team can learn it in an afternoon and still be using the words a year later. Very little published evidence links DISC composition to team performance. Treat it as language, not as evidence about people.
MBTI and 16-type tests Sixteen types from four either-or letters Use it as an on-ramp when a team is sceptical about the whole idea. People enjoy it, which is not nothing. Around half of takers get a different type on retest a few weeks later[26]. Do not build role assignments on a letter that moves.
Belbin team roles Nine preferred roles Use it when the conversation your team needs is about who does what, rather than who is like what. The idea is sound and the questionnaire is weak: its forced-choice format makes the factor structure difficult to evaluate[28].
Strengths inventories Ranked talent themes per person Use it when morale is the problem and you want a positive vocabulary that nobody can be insulted by. Themes do not map onto the five traits the composition research measures, so team-level findings do not transfer to them.

These combine better than they compete. Plenty of teams measure with the Big Five and still enjoy type vocabulary as shared slang, which is fine as long as everyone remembers which one is the measurement. The rule worth holding is narrower than "use a good test": whatever you use, do not let a label decide who gets which work. For the deeper comparisons, see Belbin team roles vs the Big Five.

The instrument on this page is a 60-item Big Five inventory in the tradition of the BFI-2, the modern research standard[1]. It takes about eight minutes and every person keeps their own result. If you want your own profile before you ask anyone else for theirs, which is the right order, the individual version is the same instrument.

Reading the results

Your team is not the average of its members

Here is the part almost every tool gets wrong. Averaging every trait assumes every trait works the same way in a group, and the research says it does not. Four careful planners plus one chronic improviser do not make an above-average team. They make a team whose delivery floor is the improviser, and everyone already knows it.

So each trait gets read its own way. Conscientiousness and agreeableness by the lowest member, because deadlines are missed and arguments are set by the floor rather than the mean[6][8]. Openness and emotional stability by the average, because curiosity and composure roughly add up. Extraversion by the spread, because a room of equally loud voices competes for the floor instead of building on it. One caveat worth carrying: the floor and spread readings matter most when work actually flows between people, and much less when members mostly work alone[9].

That is also what routes you to an exercise. Each kind of reading points somewhere different:

WHAT THE MAP READS WHAT TO RUN FloorConscientiousness, agreeableness3Disagreement contract4Reliability handshakeAverageOpenness, emotional stability2The pre-mortem5Stress signaturesSpreadExtraversion1Quiet-first roundsAlignmentThe same people, every axis6Cross-cut projectsCoverageWho holds which role7Rotating spokesperson8The role handoff
Five kinds of team-level reading, and the exercises each one points to. A floor problem and a spread problem look similar from inside the room and need opposite responses, which is the main reason generic team building underperforms.

The spread problem, in one picture

Extraversion is the clearest example, because you can watch it happen. In an unstructured hour, speaking time does not divide evenly, and it never has. The reason to care is that in studies of group problem-solving, variance in speaking turns correlated −.41 with a group's measured collective intelligence, while the average intelligence of members barely registered[17]. That finding is contested and worth holding loosely[18], but the practical move it suggests costs nothing.

A typical unstructured hour, six people share of speaking time 34%26%18%12%7% The loudest member speaks eleven times as long as the quietest The same six people, with turns taken in a round 17%17%17%17%16%16% Nobody is silenced; the floor is simply handed on
Illustrative, not measured: the shape of an unstructured discussion against the same six people taking turns in a round. The point is not that everyone must speak equally. It is that the default format decides who does, and the default is not neutral.

A team, mapped

Here is a real five-person team. Walk around it.

This is a complete team result, computed live in your browser by the same engine that builds real ones. Click through the chapters, and on Goal Fit change the objective and watch the reading change. It is the clearest way to see what "reading the floor" actually looks like.

Sample team · five members · Big Five Live, computed in your browser

Sample team · Team Map

The Drivers

Energetic and results-focused, they set the pace and keep the target in view.

A high-energy, outward-facing team

This team brings energy and pushes work forward, and no work dimension here is left without an anchor.

This reading rests on outward energy and disciplined execution, the two dimensions the group covers most strongly. A current pattern in the data, not a fixed label.

5people
Team Map

One dot per person, at their percentile against a general adult sample. Social energy runs up the side, warmth across the bottom.

Scatter plot of interpersonal style. The horizontal axis runs from candid on the left to warm on the right; the vertical axis runs from reserved at the bottom to outgoing at the top. Both are percentiles against a general adult sample, so the 50th percentile is typical. 5 people are plotted. Each position below is given the same way the chart labels it: the percentile toward warm, then the percentile toward outgoing. The team's centre sits at warm 44th, outgoing 66th. A moderate spread of styles. Priya Menon: warm 28th, outgoing 84th, driving & challenging; Diego Alvarez: warm 14th, outgoing 80th, driving & challenging; Sofia Haddad: warm 58th, outgoing 74th, energising & mobilising; Marcus Bell: warm 44th, outgoing 46th, focused & independent; Aisha Okafor: warm 76th, outgoing 48th, steady & supportive. The plot also shows a two-subgroup split: Warm group (3) and Candid group (2).

ENERGISING & MOBILISINGDRIVING & CHALLENGINGSTEADY & SUPPORTIVEFOCUSED & INDEPENDENTPriyaDiegoSofiaMarcusAisha↑ OUTGOING↓ RESERVED← CANDIDWARM →
Warm group (3)Candid group (2)Team centreThis team’s spreadMiddle 50% of adults
Every position
Each team member’s position on the team map, as percentiles against a general adult sample.
PersonWarmOutgoingWorking style
Priya Menon28th84thDriving & challenging
Diego Alvarez14th80thDriving & challenging
Sofia Haddad58th74thEnergising & mobilising
Marcus Bell44th46thFocused & independent
Aisha Okafor76th48thSteady & supportive
Team centre44th66thA moderate spread of styles

A lens on interpersonal style: how much drive a person brings into a room, and how much warmth. It covers two of the five traits. Openness, conscientiousness and emotional steadiness sit in Team Shape. Each dot is one person's percentile against a general adult sample, the centre cross marks the 50th, and the lighter inner panel holds the middle 50% of adults. This team's styles sit at a moderate distance from one another.

Coverage

What this team covers, and where the gaps are

For each way of working, does anyone here anchor it? A dimension is covered when at least one member is genuinely strong; a gap is where no one is.

5 of 5 covered
Pressure handlingCovered
70team best / 100
Anchored by Priya Menon
Influence & energyCovered
84team best / 100
Anchored by Priya Menon
Learning & adaptabilityCovered
64team best / 100
Anchored by Sofia Haddad
Collaboration styleCovered
76team best / 100
Anchored by Aisha Okafor
Execution & disciplineCovered
77team best / 100
Anchored by Priya Menon

Collective strengths

What this mix of people can be counted on for, together.

  1. No single tendency dominates: the team adapts across a range of work, with no one dimension standing out as its anchor.

Worth watching

Where this composition could bite

Modest, directional patterns from the team-composition research. Each one is framed as something to manage, not a flaw in anyone.

A cooperation floor set by one member

Watch

On a weakest-link trait, a team performs closer to its lowest member than to its average. Here that is Diego at 14. The same pattern shows in their own profile as "Direct communication style under pressure". This is not a flaw to fix in one person. It is the place where process has to carry what disposition does not.

A potential subgroup split on collaboration

Watch

This group divides into two clusters, most visibly on collaboration: Sofia Haddad, Marcus Bell, Aisha Okafor on one side, Priya Menon, Diego Alvarez on the other. Differences that line up like this can harden into an us-and-them if no one names them. Mixing the two groups across projects, and saying out loud that the split exists, is usually enough to stop it setting.

3 exercises are prescribed for these patterns, starting with cross-cut projects.

Team-level reads use trait-appropriate operationalizations from Bell (2007), Barrick et al. (1998) and Peeters et al. (2006); faultlines follow Lau & Murnighan (1998); complementary-fit framing follows Kristof-Brown et al. (2005). Effects are modest and directional: a lens for judgement, not a verdict. Team personality is one input into how a group works: it informs, it doesn’t define.

Two things to look for while you are in there. The Dynamics chapter is where the floor and spread readings live, so that is what routes you to an exercise above. The Playbook chapter is the bridge from an interesting chart to a different Tuesday, and the session agenda in the 90 minutes is built to run with it open.

What actually works

What the evidence says about team building

The honest table, including the rows that are bad for anyone selling an away day. Direction, the number worth remembering, and a grade for how much to trust it: A for meta-analytic, B for strong but narrower, C for famous and contested.

Factor Direction Evidence The number to remember
A structured debrief Helps A Meta-analysis, 46 samples About 25% better than control (d = .67), across 2,136 people. Average length: 18 minutes. [4]
Goal setting and role clarification Helps A Meta-analysis, 60 effect sizes The two components of team building that reliably move outcomes. The other two do less. [2][25]
Team building, aimed at performance No effect A Two meta-analyses agree No significant direct effect on performance. The 1999 analysis found a non-significant decrease on objective measures. [2][3]
Workshops and simulations Helps A 72 interventions, 8,439 people Teamwork behaviours d = .68, team performance d = .92 [5]
Lecturing a team about teamwork No effect A Same meta-analysis Didactic education on its own produced no significant improvement in teamwork. [5]
A one-off away day Unproven B Sports teams, few cases Programmes under two weeks: ES 0.31, with a confidence interval running from −0.97 to 1.60. [30]
The team floor on agreeableness Helps A Meta-analysis, field teams The least agreeable member predicts performance better than the team average does. [6]
Mixing very diligent with very lax Hurts A Meta-analysis Variability in conscientiousness ρ = −.24, a bigger cost than a high average is a benefit (ρ = .20). [7]
Arguing about the work Neutral A 116 studies, 8,880 groups ρ = −.01 with performance, and +.09 in senior teams. Task conflict is roughly free. [11]
Arguing about each other Hurts A Same meta-analysis ρ = −.16 with performance and −.54 with satisfaction. [11]
A dormant split in the team Depends A 168 studies, 24,953 teams Not directly related to performance or satisfaction. Only splits that get activated are. [14]
Even turn-taking in meetings Helps B Lab groups, and contested Variance in speaking turns correlated −.41 with a group's measured collective intelligence. [17][18]
Being less dependable, after a shock Helps B Single study, 73 teams Before an unexpected change, traits explained nothing. After it, higher openness and lower dependability adapted best. [10]

The exercise is the delivery mechanism. The debrief is the drug.

Read the table from the top and an uncomfortable pattern appears. Team building aimed straight at performance does not reach it[2][3]. What does reach it is unglamorous and specific: setting goals, clarifying who does what, and sitting down afterwards to ask what happened. The last of those, a structured debrief averaging about eighteen minutes, is the single best-evidenced intervention in the whole category[4].

We are an assessment company, so it is worth saying plainly what that implies about our own product. Measuring a team is aim, not treatment. It tells you which exercise is worth running and it makes the debrief specific, which is a real contribution and a modest one. Any page that tells you a personality test will improve your team's performance is describing an effect the literature has looked for and not found.

Numbers we deliberately did not use

This niche runs on statistics that do not survive contact with their sources. "70% of teams fail": a citation loop with no underlying study, which the authors most often blamed for it described as an unscientific estimate. "Task conflict hurts performance, ρ = −.23": superseded. With eighty more studies the effect is ρ = −.01[11]. The most-cited faultline meta-analysis: retracted by its publisher in 2016 for table errors, which is why our splits claim cites the 2024 replacement[14]. "93% of communication is non-verbal": two small 1967 studies about liking, which their own author says do not apply here. Forming, storming, norming, performing: a 1965 literature review, not a validated sequence[29]. Good vocabulary, bad model. Any "$X billion lost to poor teamwork" figure: every one we traced was commissioned by a company selling the cure.

When it goes wrong

Four ways these sessions go wrong, and the sentence that saves each one

Most of the damage this category has done comes from four failures. All are avoidable, and all are common enough that a facilitator should have the words ready rather than inventing them under pressure.

Someone becomes their result

The reading says a person anchors the finishing work, and within a month they are not invited to brainstorms. A description of tendency hardens into a job description. Say it out loud early, and again the moment you hear it happen.

"That is a pattern, not a permission. Nobody here just got assigned a lane."

The room agrees with everything

People rate vague, flattering feedback as uncannily accurate, and a session where nobody pushes back on anything has measured nothing except politeness. Disagreement is the quality control.

"Someone tell me what this got wrong about them. I will start: mine says I am organised, and my inbox is a crime scene."

A result gets used as ammunition

"Well, the test says you are low on agreeableness" is the fastest way to guarantee nobody answers honestly again. One use of a result against a person breaks the tool permanently for the whole team, and you will not be told that it happened.

"We do not quote this at each other. If it comes up in an argument, the argument stops and we start again."

Development data drifts into decisions

A team-building result quietly consulted during a promotion moves the exercise into territory with real legal rules, where instruments face fairness and validity standards this use was never designed to meet. The wall has to be solid and stated, not assumed.

"Nothing from today goes near a review, a promotion or a hiring decision. If we ever need assessment for those, it will be a different process, and you will know about it."

For the one-to-one conversations that follow a session like this, managing different personalities has the scripts. If the same gap keeps appearing and no exercise closes it, that is the honest signal the answer is a change to the roster rather than a change to the meetings, and hiring is a separate discipline with separate guardrails.

Ground rules

Run these without making anyone a specimen

The four ground rules

Consent. People take it knowingly, told what is shared and why, and it stays voluntary. Ownership. Each member keeps their own result, and sharing detail is their choice. Dignity. No trait is bad, so there is no bottom of the class. Purpose. The reading informs how the team runs. The moment it becomes a lever against an individual, stop.

The dignity rule is the one that does the most work, and it is worth spelling out. A low score on agreeableness is a candour engine. A high score on neuroticism is early-warning radar. An exercise that treats either as a defect will earn the resistance it gets, and deserve it. Patterns get named at the team level, never pinned to a person in front of the room, even when everyone present can do the arithmetic.

The legal position, briefly and not as legal advice. A standing team reflecting on how it works together is not making an employment decision, so the demanding rules that govern selection testing do not attach. That stops being true the moment a result touches hiring, promotion or termination. The ground rules above are what keep the ethical bar higher than the legal one, and on why the differences are worth the friction at all, personality diversity in teams makes the case with the costs included.

The other Big Five

Two different things are called "the Big Five" in teamwork

If you have been sent conflicting definitions, this is why. Same nickname, different authors, different subject. Nobody on the first page of results tells you this, and half of them answer the wrong one.

About people

The Big Five personality traits

Five continuous dimensions on which individuals differ. This is the model behind the assessment on this page and behind nearly all published team-composition research.

  • Openness
  • Conscientiousness
  • Extraversion
  • Agreeableness
  • Emotional stability

About behaviour

The "Big Five" of teamwork

Five behaviours a team performs, proposed by Salas and colleagues in 2005[27]. Nothing to do with personality, and not measured by a personality test.

  • Team leadership
  • Mutual performance monitoring
  • Backup behaviour
  • Adaptability
  • Team orientation

They are complementary rather than rival. Traits describe the material you have; the teamwork behaviours describe what a team does with it. The exercises above are aimed at the first and mostly operate through the second: a disagreement contract is backup behaviour written down, and a rotating spokesperson is team leadership spread thinner on purpose. When an exercise produces rules worth keeping, write them into a team charter so they survive the week the workshop fades.

Questions

Personality tests for team building: common questions

How do you use personality tests in team building?

In three steps. Everyone takes the same validated test and keeps their own result. You combine the results into a team-level reading, which is not simply an average, because different traits work differently in a group. Then you run one exercise aimed at what the reading shows, and review it a month later. The mistake almost everyone makes is stopping after step one, which turns a measurement into entertainment.

Which personality test is best for team building?

For measurement, a Big Five instrument, because that is the model nearly all team research is written in, so published findings actually apply to your team. For a shared vocabulary people will still use in six months, DISC is genuinely good. For sorting people into roles, nothing currently available is strong enough to bear that weight. Many teams sensibly measure with the Big Five and chat in whatever labels they enjoy.

Do personality-based team building exercises actually work?

Partly, and less than the industry implies. Two meta-analyses find that team building has no significant direct effect on team performance, though it does improve trust, coordination and how a team feels about itself. What does show a performance effect is narrower: goal setting, role clarification, and a short structured debrief, which came out about 25% better than control across 46 samples. So the honest version is that the exercise is the delivery mechanism and the debrief is the active ingredient.

How long should a team building session be?

Shorter than you think, and repeated. The best-evidenced single intervention is a structured debrief averaging about 18 minutes. In the sports literature, programmes running two weeks or longer showed a clear effect while one-off events under two weeks did not reach significance. Ninety minutes once, followed by ten minutes a week for a month, beats a full day followed by nothing.

What are good team building activities for different personality types?

Match the exercise to the pattern rather than to individuals. A team split between people who think out loud and people who think first needs a written first round, not an escape room. A team where one person catches every mistake needs a checklist at the handoff. A team that is about to go through a big change needs a pre-mortem. The eight exercises on this page each name the pattern they address, so you can pick two and skip the rest.

Can I make my team take a personality test?

You can, and it is usually a mistake. A required test produces careless answers and quiet resentment, which corrupts both the result and the conversation it was meant to start. Make it voluntary, take it yourself first and say so, and show people what the output looks like before they agree. In a team of eight, six honest results beat eight resentful ones.

Can I see my team members' individual results?

Only if they choose to show you, and you should say which it will be before anyone starts. The arrangement that works: each person owns their own full report, and the team-level view is the only shared artefact. Read the team view together rather than reading it about people in private. If you want individual results as a manager, you need a much better reason than curiosity, and you need to say so up front.

Should personality test results be used in performance reviews or hiring?

Not these ones. A team development exercise and a selection decision have different stakes and different rules, and an instrument taken voluntarily for a workshop was not designed to carry an employment decision. The moment a development result is consulted in a promotion, the exercise moves into territory with real legal standards. Keep the wall solid, and use a process built for selection if you need one.

How do you run these exercises with a remote team?

Most of them work better remotely, because the format changes that help distributed teams are the same ones that help mixed-personality teams: write first, take turns in a fixed order, put the outcome somewhere permanent. Each exercise on this page has a remote variant with the specific adjustments. The one genuine loss is ambient signal, which is why the stress-signatures exercise matters more for a distributed team than a co-located one.

What if someone says their result is wrong?

Agree with them, out loud, in front of everyone. Self-report instruments misread people, and the person is the primary source about themselves. This is not just good manners; it is a safeguard. People rate vague, flattering descriptions as uncannily accurate, so a session where nobody disagrees with anything is a session that has learned nothing. Ask what the result got wrong and what it got right, and treat the answer as better data than the score.

Is the Big Five in team building the same as the "Big Five of teamwork"?

No, and the confusion is common enough to be worth naming. The Big Five personality traits are openness, conscientiousness, extraversion, agreeableness and emotional stability, and they describe individual people. The "Big Five of teamwork" is a separate model describing five team behaviours: leadership, mutual performance monitoring, backup behaviour, adaptability and team orientation. Same nickname, different authors, different thing. This page is about the first one.

How often should we do this?

Measure rarely, debrief often. Personality is stable enough that re-testing a team every quarter measures nothing except patience: rank ordering holds at around .64 to .74 over roughly seven years. So re-measure when the roster changes, and refresh about yearly. The debrief habit is the part worth running weekly, and it is free.

Keep reading

The rest of the team library

That is everything we know about doing this well. When you are ready, the assessment at the top of the page takes a team name and gives you a link to share. Eight minutes a person, and you will know which two of the eight to run.

Methods & sources

Who wrote this, and what the assessment is based on

Michael Hodge

Founder, SeeMyPersonality · author and reviewer of this page

  • Bachelor of Science (Psychology), University of Wollongong, with coursework in psychometrics and research methods.
  • Designs and reviews the questionnaires on this site: item quality, scoring design, norm referencing, and interpretation frameworks.
  • Wrote the eight exercises against the composition and team-development literature cited below, and ran each of them with real teams before publishing.

How content is written and reviewed on SeeMyPersonality →

What the assessment on this page is based on

The individual instrument is a 60-item Big Five inventory in the tradition of the BFI-2, the peer-reviewed modern standard whose domain scales show strong reliability and retest stability[1]. Scores are computed against adult norms. The team layer combines members per trait rather than averaging everything: conscientiousness and agreeableness by team minimum, openness and emotional stability by team mean, extraversion by spread, following the composition literature that introduced and tested those operationalisations[6][7][8]. Each exercise is mapped to the reading it addresses, and to the intervention component with the best supporting evidence for that problem[2][25].

Limits, stated plainly. Team-composition effects are modest correlations, offered here as aim rather than as destiny. Team building has no demonstrated direct effect on team performance, and this page says so in its own evidence table rather than in a footnote. Personality is one input among several. It informs how a team is likely to run; it never determines what a team or a person can do. Where the evidence is contested, the grade column says so rather than rounding up.

Reviewed by: Michael Hodge Content last reviewed: 17 August 2026 Disclosure: the assessment and the sample report on this page are built and run by SeeMyPersonality. There is nothing for sale on this page.
  1. 1. Soto, C. J. & John, O. P. (2017). The next Big Five Inventory (BFI-2). Journal of Personality and Social Psychology.
  2. 2. Klein, C., DiazGranados, D., Salas, E., Le, H., Burke, C. S., Lyons, R. & Goodwin, G. F. (2009). Does team building work? Small Group Research.
  3. 3. Salas, E., Rozell, D., Mullen, B. & Driskell, J. E. (1999). The effect of team building on performance: an integration. Small Group Research.
  4. 4. Tannenbaum, S. I. & Cerasoli, C. P. (2013). Do team and individual debriefs enhance performance? A meta-analysis. Human Factors.
  5. 5. McEwan, D., Ruissen, G. R., Eys, M. A., Zumbo, B. D. & Beauchamp, M. R. (2017). The effectiveness of teamwork training on teamwork behaviors and team performance. PLOS ONE.
  6. 6. Bell, S. T. (2007). Deep-level composition variables as predictors of team performance: a meta-analysis. Journal of Applied Psychology.
  7. 7. Peeters, M. A. G., van Tuijl, H. F. J. M., Rutte, C. G. & Reymen, I. M. M. J. (2006). Personality and team performance: a meta-analysis. European Journal of Personality.
  8. 8. Barrick, M. R., Stewart, G. L., Neubert, M. J. & Mount, M. K. (1998). Relating member ability and personality to work-team processes and team effectiveness. Journal of Applied Psychology.
  9. 9. Prewett, M. S., Walvoord, A. A. G., Stilson, F. R. B., Rossi, M. E. & Brannick, M. T. (2009). The team personality-team performance relationship revisited. Human Performance.
  10. 10. LePine, J. A. (2003). Team adaptation and postchange performance: effects of team composition. Journal of Applied Psychology.
  11. 11. de Wit, F. R. C., Greer, L. L. & Jehn, K. A. (2012). The paradox of intragroup conflict: a meta-analysis. Journal of Applied Psychology.
  12. 12. Jehn, K. A. (1995). A multimethod examination of the benefits and detriments of intragroup conflict. Administrative Science Quarterly.
  13. 13. Lau, D. C. & Murnighan, J. K. (1998). Demographic diversity and faultlines: the compositional dynamics of organizational groups. Academy of Management Review.
  14. 14. Thatcher, S. M. B., Meyer, B., Kim, S. & Patel, P. C. (2024). A meta-analytic integration of the faultlines literature. Organizational Psychology Review.
  15. 15. Thatcher, S. M. B. & Patel, P. C. (2012). Group faultlines: a review, integration, and guide to future research. Journal of Management.
  16. 16. Felps, W., Mitchell, T. R. & Byington, E. (2006). How, when, and why bad apples spoil the barrel. Research in Organizational Behavior.
  17. 17. Woolley, A. W., Chabris, C. F., Pentland, A., Hashmi, N. & Malone, T. W. (2010). Evidence for a collective intelligence factor in the performance of human groups. Science.
  18. 18. Credé, M. & Howardson, G. (2017). The structure of group task performance: a second look at collective intelligence. Journal of Applied Psychology.
  19. 19. Mitchell, D. J., Russo, J. E. & Pennington, N. (1989). Back to the future: temporal perspective in the explanation of events. Journal of Behavioral Decision Making.
  20. 20. Klein, G. (2008). Performing a project premortem. IEEE Engineering Management Review (reprint of Harvard Business Review, 2007).
  21. 21. Edmondson, A. C. (1999). Psychological safety and learning behavior in work teams. Administrative Science Quarterly.
  22. 22. Frazier, M. L., Fainshmidt, S., Klinger, R. L., Pezeshkan, A. & Vracheva, V. (2017). Psychological safety: a meta-analytic review and extension. Personnel Psychology.
  23. 23. Roberts, B. W. & DelVecchio, W. F. (2000). The rank-order consistency of personality traits from childhood to old age. Psychological Bulletin.
  24. 24. Bell, S. T., Brown, S. G., Colaneri, A. & Outland, N. (2018). Team composition and the ABCs of teamwork. American Psychologist.
  25. 25. Lacerenza, C. N., Marlow, S. L., Tannenbaum, S. I. & Salas, E. (2018). Team development interventions: evidence-based approaches for improving teamwork. American Psychologist.
  26. 26. Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research.
  27. 27. Salas, E., Sims, D. E. & Burke, C. S. (2005). Is there a "big five" in teamwork? Small Group Research.
  28. 28. Furnham, A., Steele, H. & Pendleton, D. (1993). A psychometric assessment of the Belbin Team-Role Self-Perception Inventory. Journal of Occupational and Organizational Psychology.
  29. 29. Tuckman, B. W. (1965). Developmental sequence in small groups. Psychological Bulletin.
  30. 30. Kwon, S. H. (2024). Analyzing the impact of team-building interventions on team cohesion in sports teams: a meta-analysis. Frontiers in Psychology.