Measurement, interpretation and evidence
Questionnaire responses describe tendencies. Team summaries describe the participants represented. Recommendations are hypotheses to test against real work.
Implementation: ipip-60, version 3 · Norms: emp-2026-09-02 · Report methodology: Atlas v1, September 2026
The current 60-item English assessment uses wording adapted from the Big Five Inventory–2 (BFI-2; Soto and John, 2017). The first-person presentation and some phrasing differ from the published form. The historical URL and internal identifier remain ipip-60; that identifier does not mean the current questions are the original IPIP inventory.
The implementation retains five domain and fifteen facet scoring slots. Historical internal facet identifiers map to the closest BFI-2 constructs. Those legacy names and the adaptation should not be mistaken for an independent validation of the implemented facet structure. The Atlas team report uses the five domain scores consistently.
Participants answer all 60 items on a five-point scale. Reverse-keyed items are scored as six minus the selected response. Each facet sums four keyed items, giving a raw range of 4–20; each domain sums its three facets, giving 12–60. Missing answers are not replaced with a fictional teammate or an average profile.
Each raw score is converted using the versioned empirical percentile table. A percentile is the proportion of reference scores below the score plus half the proportion equal to it, multiplied by 100, rounded and limited to 1–99. Emotional stability is displayed as 100 minus the neuroticism percentile. Higher means more of the trait, not better performance.
The current tables use 58,401 distinct, valid completed result codes recorded on English result pages between 23 June and 1 September 2026. They were generated on 2026-09-02. Malformed codes and results with all fifteen facet sums identical were excluded. This is a sample of people who took this test, not a representative sample of adults, employees or the general population.
Result codes encode scores: identical score profiles were deduplicated. The count therefore describes distinct score vectors, not verified unique people. Excluding identical facet sums is a response-quality heuristic and can also exclude valid responses. Self-selection, language, repeat participation and deduplication can affect the reference distribution. No demographic weighting or population representativeness is established.
The versioned source is the inventory’s percentile tables. Other inventories and extended forms may use different references or model-based estimates; their percentiles should not be treated as interchangeable.
The original BFI-2 validation studied five domains and fifteen facets. In Study 3, domain internal-consistency coefficients were .83–.90 in the Internet sample and .85–.90 in the student sample; eight-week retest correlations were .76–.84 in 110 students. These figures describe the published inventory and study samples. Soto and John, 2017, Tables 2–3.
We have not published equivalent internal-consistency, retest, factor-structure or predictive-validation evidence for this exact adapted implementation. The reference dataset contains score summaries rather than item-level answers, so its size alone cannot establish item reliability. Published BFI-2 evidence does not validate this report’s goal prompts, roles, motivation interpretations, hiring profiles or team-performance predictions.
| Output | Derivation and interpretation |
|---|---|
| Named traits | Scored self-reports, shared with the participant’s permission. They do not establish ability, intent or colleagues’ opinions. |
| Team charts | Arithmetic means and observed ranges of the completed, shared profiles. An average percentile is not a percentile rank for a team. The participation count states how much of the intended team is represented. |
| Individual language | Scores from the 35th through 65th percentile receive a middle-range description. This is an editorial choice to avoid categorical storytelling, not a clinical threshold or a validated boundary. |
| Goal fit | The selected objective chooses relevant trait summaries and discussion prompts. These are authored experimental rules. Exact delivery, innovation and role-fit scores are not presented as measured capability. |
| Working groups | An experimental search uses authored goal profiles to suggest arrangements, while respecting kept groups. Internal weights favour the member-weighted average model match (65%) and lowest group match (35%). Targets, weights and search criteria are product assumptions, not validated performance measures. Members can change the arrangement and should check skills, availability and preferences. |
| Motivations and actions | Participants describe motivations and preferences directly. Teams choose an action, owner, baseline, success measure and review date, then record whether the practice helped. Adoption does not prove effectiveness. |
A percentile of 59 is not evidence of missing capability because it falls below 60. This report does not apply that cutoff or a top-fifth rule to diagnose operational problems. The older “roughly 78% of teams have a gap” argument was a probability illustration under assumed definitions and independence, not an observed rate of dysfunctional teams.
Team-personality research examines different summaries and contexts. Peeters and colleagues analysed average levels and variation, with some relationships differing between student and professional teams. This does not justify a universal rule that a team’s delivery is determined by its least conscientious member. Peeters et al., 2006.
Check recent events, responsibilities, dependencies, capacity, changing priorities and skills before attributing a problem to personality. Two profiles are enough to start a conversation, not enough to establish a reliable performance prediction.
Participants preview their chosen name, five domain scores and fifteen detailed facet scores before joining the shared map. These scores become accessible to anyone with the team map link; the full individual report stays private. Atlas summaries use the five domains, while the original report can also use facets. Participants can respond “Fits me,” “Depends on the situation” or “Doesn’t fit,” add an example, and choose whether each reflection is shared. Reflections are private by default. Only completed, shared profiles enter team summaries; demonstrations use a separate, explicitly fictional roster.
Shared context and actions are visible to report viewers. Personal reflections can be made private again; participants can withdraw their shared profile using the participation controls. The sharing preview explains access before a profile is added to the map.